One-stop multi-mode video generation system and method based on artificial intelligence
Through a one-stop multimodal video generation system, the problems of insufficient multimodal data support and low degree of automation in the video generation process in the existing technology are solved, and high-quality, personalized and diversified video generation is achieved, which improves the efficiency and creativity of video production.
Patent Information
- Application Number
- CN202510540439.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the video generation process has a single function, lacks multimodal data support, limited generation quality, fails to effectively combine user creative needs, has low degree of automation and poor adaptability, resulting in uneven quality and templated videos, which is difficult to meet diversified creative needs.
A one-stop multi-modal video generation system based on artificial intelligence is adopted. By obtaining multiple types of data from multiple input nodes, pre-processing and multi-modal data fusion, video is generated using customized video generation models, and the generation quality is improved through quality evaluation and optimization algorithms, supporting users to personalize customization and automated editing.
It realizes seamless integration and efficient generation of multimodal data, improves the precision and creativity of video generation, supports users to customize creative elements, improves the degree of automation and video quality, and meets diverse creative needs.
Smart Images

Figure CN120343358A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular, to a one-stop multi-modal video generation system and method based on artificial intelligence. Background Art
[0002] With the popularization of the Internet and mobile devices, video, as a powerful information carrier, has seen an exponential growth in demand in fields such as entertainment, education, news, and commerce. Users hope to quickly and conveniently obtain high-quality video content. At the same time, enterprises and content creators also need more efficient video production tools to meet market demands. The traditional video production process is often time-consuming and laborious, requiring professional equipment and technical personnel, and cannot meet the rapidly growing video demands.
[0003] Currently, based on deep learning technology, it is possible to automatically learn the complex patterns and structures of videos, providing strong technical support for video generation. This enables the model to automatically extract features from a large amount of data, learn, and optimize, thereby generating more realistic and smooth video sequences. Additionally, videos can also be generated based on intelligent video editing systems.
[0004] However, there are the following deficiencies in generating videos based on deep learning technology:
[0005] ① Single function: This technology mainly focuses on generating videos from text, lacking support for multi-modal data such as pictures and music. This single data input method limits its application scope and the diversity of generated videos.
[0006] ② Limited generation quality: Since the model mainly relies on text descriptions, the generated videos are often not fine enough in terms of detail processing and creative expression, resulting in uneven video quality.
[0007] ③ Lack of creative integration: This technology fails to effectively combine users' personalized needs and creative elements. The generated video content is relatively template-based. The deep learning model has limited generalization ability in processing complex multi-modal data, resulting in relatively rigid generated video content and being difficult to meet diverse creative needs.
[0008] In addition, there are the following deficiencies in generating videos based on intelligent video editing systems:
[0009] ① Low degree of automation: Many parameters need to be preset by users, and the degree of intelligence is limited, failing to achieve true "one-key generation".
[0010] ② Poor adaptability: It has high requirements for the type and quality of input materials and has poor adaptability.
[0011] ③ Algorithm limitations: Existing algorithms are difficult to accurately identify and integrate diverse materials, resulting in insufficient coherence and logic in the generated videos.
[0012] In summary, the following defects exist in the prior art:
[0013] 1. In the prior art, during the data processing of a single modality (such as generating a video only from text), there is a lack of the ability to comprehensively process multiple data types (text, pictures, music, narration, etc.).
[0014] 2. In the prior art, there are deficiencies in the detail processing and creative expression of video generation, resulting in uneven quality of the generated videos.
[0015] 3. The video content generated in the prior art fails to effectively combine the user's personalized needs and creative elements, lacks innovation and diversity, and the generated video content is relatively templatized.
[0016] 4. In the prior art, there are limitations in terms of automation and adaptability, requiring users to preset many parameters and having a low level of intelligence.
[0017] Therefore, the technical problem that needs to be solved urgently at present is: how to provide an artificial intelligence-based one-stop multi-modal video generation system and method to overcome the problem of single function in the video generation process of the prior art, improve the quality of the generated videos, enhance user interactivity and personalization, increase the level of automation and intelligence, and achieve creative integration and diverse output. Summary of the Invention
[0018] The purpose of this application is to provide an artificial intelligence-based one-stop multi-modal video generation system and method to overcome the problem of single function in the video generation process of the prior art, improve the quality of the generated videos, enhance user interactivity and personalization, increase the level of automation and intelligence, and achieve creative integration and diverse output.
[0019] To achieve the above object, this application provides an artificial intelligence-based one-stop multi-modal video generation method, which includes: in response to a video generation task, obtaining multiple types of data from multiple input nodes; preprocessing the obtained multiple types of data; performing multi-modal data fusion processing on the preprocessed multiple types of data; and generating corresponding videos according to the data after multi-modal data fusion processing by using customized video generation models corresponding to different types of video generation tasks pre-constructed.
[0020] The one-stop multi-modal video generation method based on artificial intelligence as described above, wherein the method further includes: collecting quality defect feature data in the generated video and obtaining user satisfaction, comprehensively evaluating the quality of the generated video according to the quality defect feature data and user satisfaction, and calculating the comprehensive expected loss value of the generated video quality; comparing the size of the comprehensive expected loss value of the generated video quality with a preset threshold, if the comprehensive expected loss value of the generated video quality is greater than the preset threshold, optimizing the customized video generation model for generating the video, otherwise, there is no need to optimize the customized video generation model for generating the video.
[0021] The one-stop multi-modal video generation method based on artificial intelligence as described above, wherein the quality defect feature data includes: video frame anomaly feature data, audio stuttering feature data, and video editing trace feature data.
[0022] The one-stop multi-modal video generation method based on artificial intelligence as described above, wherein the method further includes: automatically identifying exciting segments in the generated video based on a pre-constructed exciting segment automatic recognition model, and performing editing and slicing processing on the exciting segments to obtain an edited video.
[0023] The one-stop multi-modal video generation method based on artificial intelligence as described above, wherein the multi-modal data fusion processing of various types of data after preprocessing includes: splicing pictures and video segments in the order of the script, adding transition effects to make the picture transition natural; adding text to the video in the form of subtitles or voiceovers; inserting music or sound effects according to the video plot and rhythm.
[0024] The one-stop multi-modal video generation method based on artificial intelligence as described above, wherein the method for pre-constructing a customized video generation model includes: obtaining training data; preprocessing the training data; performing secondary processing on the preprocessed training data; inputting the secondary processed training data into an automated training platform for training to obtain a customized video generation model, and storing the obtained model in a user-specific model repository.
[0025] The AI - based one - stop multi - modal video generation method as described above, wherein the pre - constructed highlight automatic recognition model includes: obtaining a training data set and a verification and evaluation data set; inputting the training data set into a basic convolutional neural network model for training to obtain the highlight automatic recognition model; inputting the verification and evaluation data set into the trained highlight automatic recognition model for recognition to obtain verification and evaluation feature data; calculating the recognition reliability evaluation value of the highlight automatic recognition model according to the verification and evaluation data; comparing the recognition reliability evaluation value of the highlight automatic recognition model with a preset reliability threshold. If the recognition reliability evaluation value of the highlight automatic recognition model is less than the preset reliability threshold, then optimize the highlight automatic recognition model; otherwise, there is no need to optimize the highlight automatic recognition model.
[0026] The AI - based one - stop multi - modal video generation method as described above, wherein the verification and evaluation data includes: the time when the highlight automatic recognition model recognizes the result of the object to be recognized, the accuracy rate of the result recognized by the highlight automatic recognition model, and the response time slot of the highlight automatic recognition model.
[0027] As a second aspect of the present application, the present application provides an AI - based one - stop multi - modal video generation system that executes the above - described AI - based one - stop multi - modal video generation method. The system includes:
[0028] A data acquisition module, configured to obtain various types of data from multiple input nodes in response to a video generation task;
[0029] A data pre - processing module, configured to pre - process the obtained various types of data;
[0030] A multi - modal processing module, configured to perform multi - modal data fusion processing on the pre - processed various types of data;
[0031] A video generation module, configured to generate a corresponding video according to the data after multi - modal data fusion processing, using customized video generation models corresponding to different types of video generation tasks pre - constructed.
[0032] Among them, the video generation model can perform customized pre - training of its own model and perform regeneration based on this trained model.
[0033] The AI - based one - stop multi - modal video generation system as described above, wherein the system further includes:
[0034] A data collection module, configured to collect quality defect feature data in the generated video and obtain user satisfaction;
[0035] A data processor for comprehensively evaluating the quality of the generated video based on quality defect feature data and user satisfaction, and calculating the comprehensive expected loss value of the generated video quality.
[0036] A data comparator for comparing the size of the comprehensive expected loss value of the generated video quality with a preset threshold. If the comprehensive expected loss value of the generated video quality is greater than the preset threshold, optimize the customized video generation model for generating this video; otherwise, there is no need to optimize the customized video generation model for generating this video.
[0037] The one-stop multi-modal video generation system based on artificial intelligence as described above, wherein the system further includes:
[0038] An automatic recognition module for automatically recognizing the wonderful segments in the generated video based on a pre-constructed wonderful segment automatic recognition model, and performing editing and slicing processing on the wonderful segments to obtain an edited video.
[0039] The beneficial effects achieved by this application are as follows:
[0040] (1) This application uses artificial intelligence algorithms to automatically generate high-quality video content from input data such as text, pictures, videos, music, sound effects, and editing. It also supports the model personalization customization function. This system is widely used in fields such as film and television previews, advertising, education and training, content creation, and e-commerce videos, and can greatly improve the efficiency and creativity of video production.
[0041] (2) This application realizes the seamless integration and efficient generation of multi-modal data, and provides a full range of video production services.
[0042] (3) This application optimizes artificial intelligence algorithms and models to improve the fineness and creativity of video generation, and ensure the high quality of the output video.
[0043] (4) This application adopts a friendly user interface and interaction mechanism, allowing users to customize generation parameters and creative elements to achieve the customization of personalized video content.
[0044] (5) This application optimizes through intelligent algorithms to achieve a highly automated video generation process, reduce the complexity of user operations, and improve the user experience.
[0045] (6) This application integrates a variety of creative generation modules (such as voice cloning, intelligent editing, etc.), supports diverse video content creation, supports users to upload popular videos, identify the content, generate sliced video segments, and then analyze the keywords of the corresponding segments for secondary processing and generation to meet different scenarios and application requirements. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those skilled in the art, other drawings can also be obtained based on these drawings.
[0047] Figure 1 It is a flowchart of a one-stop multi-modal video generation method based on artificial intelligence according to an embodiment of the present application.
[0048] Figure 2 It is a schematic structural diagram of a one-stop multi-modal video generation system based on artificial intelligence according to an embodiment of the present application.
[0049] Figure 3 It is an architecture diagram of a one-stop multi-modal video generation system based on artificial intelligence according to an embodiment of the present application.
[0050] Figure 4 It is a schematic diagram of the generation method of a customized video generation model according to an embodiment of the present application.
[0051] Figure 5 It is a schematic diagram of the architecture of a workflow generation engine according to an embodiment of the present application. Detailed implementation manners
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0053] Embodiment 1
[0054] As Figure 2 shown, the present application provides a one-stop multi-modal video generation system 1000 based on artificial intelligence, including:
[0055] A data input module 10, used to input various types of data such as text, pictures, music, narration, and videos. The data input module includes multiple input nodes, and each input node can input different types of data.
[0056] A data acquisition module 20, used to obtain various types of data from multiple input nodes in response to a video generation task.
[0057] A data preprocessing module 30, used to preprocess the obtained various types of data.
[0058] The multimodal processing module 40 performs multimodal data fusion processing on the input data. Through a fusion algorithm, it seamlessly integrates multimodal data such as text, pictures, music, narration, and videos to generate high-quality videos.
[0059] Among them, after performing multimodal data fusion processing on the input data, users can customize editing, modification, and clip synthesis operations.
[0060] Among them, the multimodal processing module 40 adopts parallel computing and distributed processing technologies to improve the processing efficiency of multimodal data. It significantly improves the response speed and throughput of the system, ensuring high efficiency and stability in scenarios of large-scale data processing.
[0061] Among them, through weight allocation and optimization strategies, it ensures the coordination and consistency of each modal data in video generation, significantly enhancing the comprehensive expressiveness of the video.
[0062] The video generation module 50 is used to generate corresponding videos according to the data after multimodal data fusion processing, using customized video generation models corresponding to different types of video generation tasks pre-constructed.
[0063] Among them, a customized model is set in the video generation module 50. By performing customized training on the customized model and combining user feedback, the customized model is dynamically optimized to enhance the personalized experience.
[0064] It should be noted that the customized model can meet the user's self-defined model training, and the video effects generated according to this trained model can meet commercial requirements, which is different from the existing video large models that can only generate video clips (with a duration of about 4s - 8s). Secondly, the video generation system of this application can provide local modification of video clips.
[0065] Among them, the video generation model 50 can perform customized pre-training of its own model and regenerate based on this trained model.
[0066] An intelligent editing and slicing model is set in the video generation module 50 for editing and slicing video content. By optimizing the editing algorithm, the coherence and logic of the video content are enhanced.
[0067] The video generation module 50 introduces a voice cloning algorithm to generate narration similar to the user's voice.
[0068] Among them, the video generation module 50 parses the video to generate sliced segments: performs AI recognition for video regeneration, etc.
[0069] The user interaction module is used to receive the user's feedback data.
[0070] This application sets up an operation interface for a one-stop process from text to picture, to video, to music and sound effects, to narration, to editing, and provides rich customization options to enhance the user experience. Through friendly interaction design, it reduces the user's usage threshold and enhances the usability and user stickiness of the system.
[0071] The video generation module 50 integrates a variety of creative generation modules, supporting diverse video content creation. It realizes the deep integration of creative elements, supports diverse video outputs for users in different scenarios and application requirements, and enhances the innovation and attractiveness of the content.
[0072] The output module 60 regenerates the video in combination with the user's feedback data.
[0073] The one-stop multi-modal video generation system 1000 based on artificial intelligence of this application further includes:
[0074] The data acquisition module 70 is used to collect the quality defect feature data in the generated video and obtain the user satisfaction.
[0075] The data processor 80 is used to comprehensively evaluate the quality of the generated video according to the quality defect feature data and the user satisfaction, and calculate the comprehensive expected loss value of the generated video quality.
[0076] The data comparator 90 is used to compare the size of the comprehensive expected loss value of the generated video quality with a preset threshold. If the comprehensive expected loss value of the generated video quality is greater than the preset threshold, the customized video generation model for generating this video is optimized; otherwise, there is no need to optimize the customized video generation model for generating this video.
[0077] The automatic recognition module 100 is used to automatically recognize the wonderful segments in the generated video based on a pre-constructed wonderful segment automatic recognition model, and perform editing and slicing processing on the wonderful segments to obtain a clipped video.
[0078] It should be noted that for a one-stop multi-modal video generation system based on artificial intelligence of this application, where "one-stop" means that the system provides all-round and full-process services, and users do not need to use multiple tools or platforms. "Multi-modal" means covering multiple data types (text, pictures, music, videos, etc.), reflecting the multi-functionality of the system.
[0079] Such as Figure 3As shown in the figure, it is a schematic diagram of the system architecture of a one-stop multi-modal video generation system based on artificial intelligence according to the present application. The architecture of this system includes: an application layer, a presentation layer, a business layer, a model layer, and a data storage layer. Among them, the application layer includes story expansion, graphic explanation, scene analysis, cover creation, intelligent editing, and brand videos, etc. The presentation layer is implemented by front-end technologies (HTML, CSS, JavaScript) and interacts with the background through the HTTP protocol. The business layer includes business modules and technology modules. The model layer includes multiple models. The data storage layer includes relational databases, non-relational databases, object storage services, and in-memory data storage systems, etc.
[0080] As Figure 5 shown in the figure, in response to a video generation task, the present application automatically generates a workflow through a workflow generation engine and generates a video according to the workflow. The architecture of the workflow generation engine of the present application includes: a support system, an AI model library, an AI model generation engine, a core workflow, and a user operation module.
[0081] As Figure 5 shown in the figure, the support system includes: a media processing library, a logging system, a data management system, and a permission control system. The AI model library includes various types of models. The AI model generation engine is linked to the AI model library, and the AI model generation engine is used to perform user requirement analysis, customized model generation, and model deployment and integration in sequence. The AI model generation engine is used for AI content generation. The core workflow includes video input, a preprocessing module, a content analyzer, a timeline builder, and a product extractor, and links video output, graphic explanation, scene analysis, cover image, data analysis, and a user interface through the product extractor. The user interface is linked to the user operation module. The user operation module includes: AI content generation, workflow editing processing, product selection, result preview, and material library management.
[0082] Embodiment 2
[0083] As Figure 1 shown in the figure, the present application provides a one-stop multi-modal video generation method based on artificial intelligence, and this method includes:
[0084] Step S1, in response to a video generation task, obtain various types of data from multiple input nodes.
[0085] Among them, the data input by the input nodes includes various types of data such as text, pictures, music, narration, or videos.
[0086] Among them, there are multiple input nodes, different input nodes are used to input different types of data, and the same input node is used to input the same type of data.
[0087] Step S2, preprocess the obtained various types of data.
[0088] Among them, preprocessing includes: preliminarily processing various types of acquired data. For example, cropping and color - adjusting pictures to make them more unified in style; editing videos to remove unnecessary parts and adjusting the resolution, frame rate, etc. of the videos; editing and converting the format of music to make it meet the rhythm and format requirements of the videos.
[0089] Step S3, perform multi - modal data fusion processing on various types of data after preprocessing.
[0090] Step S3 includes:
[0091] Step S310, splice pictures and video clips in the order of the script and add transition effects to make the picture transition natural.
[0092] Step S320, add text to the video in the form of subtitles or voice - overs.
[0093] Step S330, insert music or sound effects according to the video plot and rhythm.
[0094] As a specific embodiment of the present invention, data of different modalities are fused in chronological order or logical order.
[0095] For example, according to the content development of the text, matching pictures, voice - overs are inserted at specific positions, and music is added, so that these modal data are fused in time series.
[0096] As another specific embodiment of the present invention, features of each modal data are extracted from multi - modal data, and these features are spliced into a joint feature vector. For example, after splicing text features, picture features, music features, voice - over features, etc., a joint feature vector is formed for generating videos.
[0097] Step S4, according to the data after multi - modal data fusion processing, use a customized video generation model corresponding to different types of video generation tasks pre - constructed to generate corresponding videos.
[0098] Among them, different types of video generation tasks correspond to different types of customized video generation models. Inputting the data after multi - modal data fusion processing into the customized video generation model can generate videos corresponding to this type of video generation task, so that the generated videos can be applied to fields such as film and television previews, advertising, education and training, content creation, e - commerce videos, etc.
[0099] Among them, the method for pre - constructing a customized video generation model includes:
[0100] Step T1, obtain training data.
[0101] Among them, the training data includes various types of data such as text, pictures, music, narration, or videos.
[0102] Step T2, preprocess the training data.
[0103] Among them, preprocessing the training data includes: receiving the training data and validating it, checking the quality of the training data. If the check passes, then proceed to the next step and perform secondary processing on the training data. If the check fails, then feedback the quality problem of the training data.
[0104] Step T3, perform secondary processing on the preprocessed training data.
[0105] Such as Figure 4 As shown, performing secondary processing on the preprocessed training data includes:
[0106] Step T310, automatic data processing: classify, clean, and label the training data.
[0107] Step T320, perform automatic data analysis and enhancement processing on the training data.
[0108] Step T330, input the training data after automatic data analysis and enhancement processing into the intelligent bottom mold selector to select a suitable bottom mold.
[0109] Among them, the intelligent bottom mold selector is a tool used to select a suitable bottom mold in specific content creation (such as video creation, 3D modeling, design, etc.).
[0110] Step T4, input the training data after secondary processing into the automated training platform for training to obtain a customized video generation model, and store the obtained model in the user's exclusive model repository.
[0111] Among them, the automated training platform has: a model training system, a model performance evaluation system, a hyperparameter optimizer, an adaptive learning rate adjuster, and a training accelerator, etc.
[0112] Among them, the model training system is used to train the model; the model performance evaluation system is used to evaluate the performance of the model; the hyperparameter optimizer is used for setting hyperparameters to improve the model recognition accuracy; the adaptive learning rate adjuster is used to adjust the learning rate of model training, and the training accelerator is used to accelerate the model training speed.
[0113] Among them, the method for the model training system to train the model is: input the training data into the existing convolutional neural network model for training.
[0114] The automated training platform can dynamically adjust and optimize the customized video generation model according to user feedback and sample data, improving the pertinence and quality of the generated videos and enhancing user satisfaction.
[0115] Step S5, collect the quality defect feature data in the generated video and obtain the user satisfaction. According to the quality defect feature data and the user satisfaction, conduct a comprehensive quality evaluation of the generated video, and calculate the comprehensive expected loss value of the generated video quality.
[0116] Specifically, conduct a quality inspection on the preliminarily produced video. Check whether the video conforms to the theme, style, and purpose set in the creative stage, and check whether the picture is clear, the subtitles are accurate, the audio is normal, etc. At the same time, collect feedback from all parties, such as the modification opinions of customers and the problems found by team members. Modify and optimize the video according to the feedback from the review. It may involve operations such as re-editing some pictures, replacing music, and adjusting subtitles until the video meets the expected quality standards. User satisfaction refers to the degree of satisfaction of users with the generated video, and the user satisfaction is, for example, 80%, 90%, 99%, 99.9%, etc.
[0117] Among them, the quality defect feature data includes: video picture abnormal feature data (for example: unclear picture, low picture resolution, picture distortion, picture jitter, wrong picture subtitles, etc.), audio stuttering feature data (number of audio stuttering times, duration of audio stuttering, etc.), video editing trace feature data (number of video editing trace occurrences, area of video editing traces), etc.
[0118] Among them, the calculation formula for the comprehensive expected loss value of the generated video quality is as follows:
[0119]
[0120] Among them, 1 - My ≠ 0;
[0121] Among them, AC represents the comprehensive expected loss value of the generated video quality; My represents the user satisfaction; μ represents the influence factor of the user satisfaction; q1 represents the influence weight of the video picture abnormal feature data on the comprehensive expected loss value of the generated video quality; M represents the total number of types of video picture abnormal feature data; represents the weight factor of the i-th type of video picture abnormal feature data; C i represents the occurrence times of the i-th type of video picture abnormal feature data; THch il represents the duration of the l-th occurrence of the i-th type of video picture abnormal feature data; Tz represents the total duration of the video; q2 represents the influence weight of the audio stuttering feature data on the comprehensive expected loss value of the generated video quality; K represents the number of audio stuttering occurrences; represents the density value of the audio stuttering; TYch pdenotes the duration of the p-th audio freeze; q3 represents the influence weight of video clip trace feature data on the comprehensive expected loss value of the generated video quality; H represents the number of occurrences of video clip traces; denotes the area of the h-th video clip trace; TJch h denotes the duration of the h-th video clip trace; denotes the density value of the occurrence of video clip traces.
[0122] Among them, the density value of audio freeze The calculation formula is:
[0123]
[0124] Among them, t p(P+1) denotes the interval duration between the p-th audio freeze and the (p + 1)-th audio freeze.
[0125] Among them, the density value of the occurrence of video clip traces The calculation formula is:
[0126]
[0127] Among them, t h(h+1) denotes the interval duration between the h-th video clip trace and the (h + 1)-th video clip trace.
[0128] Step S6, compare the comprehensive expected loss value of the generated video quality with the preset threshold. If the comprehensive expected loss value of the generated video quality is greater than the preset threshold, optimize the customized video generation model for generating this video; otherwise, there is no need to optimize the customized video generation model for generating this video.
[0129] Step S7, based on the pre-constructed wonderful clip automatic recognition model, automatically recognize the wonderful clips in the generated video, and perform editing and slicing processing on the wonderful clips to obtain an edited video.
[0130] As a specific embodiment of the present invention, during the process of editing and slicing the wonderful clips, it also supports users to upload material content for random editing, and can match the music rhythm beat for editing.
[0131] Among them, the method for pre-constructing the wonderful clip automatic recognition model is:
[0132] Step S710, obtain a training data set and a validation and evaluation data set.
[0133] Among them, the training data set is a video with wonderful clips already marked.
[0134] Step S720: Input the training data set into the basic convolutional neural network model for training to obtain the automatic highlight recognition model.
[0135] Among them, the automatic highlight recognition model is obtained by training the basic convolutional neural network model with the highlight training data set.
[0136] Step S730: Input the verification and evaluation data set into the trained automatic highlight recognition model for recognition to obtain the verification and evaluation feature data.
[0137] The verification and evaluation data include: the time when the automatic highlight recognition model recognizes the result of the object to be recognized, the accuracy rate of the result recognized by the automatic highlight recognition model, the proportion of correctly recognized highlights in the actual highlights, and the response time slot of the automatic highlight recognition model. The response time slot is the response recognition duration of the automatic highlight recognition model for the next input object to be recognized after recognizing the current object to be recognized.
[0138] Step S740: Calculate the recognition reliability evaluation value of the automatic highlight recognition model according to the verification and evaluation data.
[0139] Specifically, the calculation formula for the recognition reliability evaluation value of the automatic highlight recognition model is:
[0140]
[0141] Among them, JS represents the recognition reliability evaluation value of the automatic highlight recognition model; R1 represents the influence weight of the time when the automatic highlight recognition model recognizes the result of the object to be recognized; A represents the total number of objects to be recognized in the verification and evaluation data set; Ge a represents the time when the automatic highlight recognition model recognizes the result of the a-th object to be recognized; Az represents the number of objects with accurate recognition results by the automatic highlight recognition model in the verification and evaluation data set; R2 represents the influence weight of the accuracy rate of the result recognized by the automatic highlight recognition model; e = 2.718; ZBs represents the proportion of correctly recognized highlights in the actual highlights; R3 represents the influence weight of the response time slot of the automatic highlight recognition model; Ml a represents the response time slot of the automatic highlight recognition model after recognizing the a-th object to be recognized.
[0142] Step S750: Compare the recognition reliability evaluation value of the automatic highlight recognition model with the preset reliability threshold. If the recognition reliability evaluation value of the automatic highlight recognition model is less than the preset reliability threshold, optimize the automatic highlight recognition model; otherwise, there is no need to optimize the automatic highlight recognition model.
[0143] The present application also provides a computer storage medium. The computer storage medium stores computer instructions, which, when called, are used to execute the address mapping method of the large-capacity solid-state drive. The computer storage medium contains one or more program instructions, and the one or more program instructions are used to be executed by a processor to perform an artificial intelligence-based one-stop multi-modal video generation method.
[0144] An embodiment disclosed by the present invention provides a computer-readable storage medium. Computer program instructions are stored in the computer-readable storage medium. When the computer program instructions run on a computer, the computer is caused to execute the above-mentioned artificial intelligence-based one-stop multi-modal video generation method.
[0145] An embodiment of the present invention provides a processor for processing the above-mentioned artificial intelligence-based one-stop multi-modal video generation method.
[0146] In an embodiment of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0147] It is possible to implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.
[0148] The storage medium may be a memory, for example, it may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory.
[0149] Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).
[0150] The beneficial effects achieved by this application are as follows:
[0151] (1) Through the artificial intelligence algorithm, this application realizes the automatic generation of high-quality video content from input data such as text, pictures, videos, music, sound effects, and clips. At the same time, it supports the model personalization customization function. This system is widely used in fields such as film trailers, advertising, education and training, content creation, and e-commerce videos, and can greatly improve the efficiency and creativity of video production.
[0152] (2) This application realizes the seamless integration and efficient generation of multimodal data, providing comprehensive video production services.
[0153] (3) By optimizing the artificial intelligence algorithm and model, this application improves the fineness and creativity of video generation, ensuring the high quality of the output video.
[0154] (4) This application adopts a friendly user interface and interaction mechanism, allowing users to customize generation parameters and creative elements to achieve the customization of personalized video content.
[0155] (5) Through intelligent algorithm optimization, this application realizes a highly automated video generation process, reduces the complexity of user operations, and improves the user experience.
[0156] (6) This application supports diverse video content creation by integrating multiple creative generation modules (such as timbre cloning, intelligent editing, etc.), enables users to upload popular videos, recognize the content, generate sliced video segments, then analyze the keywords of the corresponding segments, and perform secondary processing to meet different scenarios and application requirements.
[0157] In the description of this application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of this application, "a plurality of" means two or more, unless otherwise specifically defined.
[0158] In the description of this application, the phrase "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in this application is not necessarily to be construed as more preferred or more advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. In the following description, details are set forth for the purpose of explanation. It should be understood that those skilled in the art can recognize that the invention can be practiced without these specific details. In other instances, well-known structures and processes are not described in detail to avoid unnecessary details from obscuring the description of the invention. Therefore, the invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed in this application.
[0159] The above description is only for the embodiments of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A one-stop multi-modal video generation method based on artificial intelligence, characterized in that, The method includes: In response to a video generation task, obtaining various types of data from multiple input nodes; Preprocessing the obtained various types of data; Performing multi-modal data fusion processing on the preprocessed various types of data; According to the data after multi-modal data fusion processing, using a customized video generation model corresponding to different types of video generation tasks pre-constructed to generate a corresponding video; Collecting quality defect feature data in the generated video and obtaining user satisfaction, and based on the quality defect feature data and user satisfaction, comprehensively evaluating the quality of the generated video, and calculating the comprehensive expected loss value of the generated video quality; Comparing the size of the comprehensive expected loss value of the generated video quality with a preset threshold. If the comprehensive expected loss value of the generated video quality is greater than the preset threshold, optimizing the customized video generation model for generating the video, otherwise, there is no need to optimize the customized video generation model for generating the video.
2. The one-stop multi-modal video generation method based on artificial intelligence according to claim 1, wherein Wherein, The quality defect feature data includes: video frame anomaly feature data, audio freeze feature data, and video editing trace feature data.
3. The one-stop multi-modal video generation method based on artificial intelligence according to claim 1, wherein The method further includes: Based on a pre-constructed highlight segment automatic recognition model, automatically recognizing highlight segments in the generated video, and performing editing and slicing processing on the highlight segments to obtain an edited video.
4. The one-stop multi-modal video generation method based on artificial intelligence according to claim 1, wherein Performing multi-modal data fusion processing on the preprocessed various types of data includes: Stitching pictures and video clips in the order of the script, adding transition effects to make the picture transition natural; Adding text to the video in the form of subtitles or narration; Inserting music or sound effects according to the video plot and rhythm.
5. The one-stop multi-modal video generation method based on artificial intelligence according to claim 1, wherein The method for pre-constructing a customized video generation model includes: Obtaining training data; Preprocessing the training data; Performing secondary processing on the preprocessed training data; Inputting the secondary processed training data into an automated training platform for training to obtain a customized video generation model, and storing the obtained model in a user-exclusive model repository.
6. The one-stop multi-modal video generation method based on artificial intelligence according to claim 3, characterized in that The pre-constructed highlight segment automatic recognition model includes: Obtaining a training data set and a validation and evaluation data set; Inputting the training data set into a basic convolutional neural network model for training to obtain a highlight segment automatic recognition model; Inputting the validation and evaluation data set into the trained highlight segment automatic recognition model for recognition to obtain validation and evaluation feature data; Calculating the recognition reliability evaluation value of the highlight segment automatic recognition model according to the validation and evaluation data; Comparing the size of the recognition reliability evaluation value of the highlight segment automatic recognition model with a preset reliability threshold. If the recognition reliability evaluation value of the highlight segment automatic recognition model is less than the preset reliability threshold, optimizing the highlight segment automatic recognition model, otherwise, there is no need to optimize the highlight segment automatic recognition model.
7. The one-stop multi-modal video generation method based on artificial intelligence according to claim 6, wherein, The validation and evaluation data includes: the time when the highlight segment automatic recognition model recognizes the result of the object to be recognized, the accuracy of the result recognized by the highlight segment automatic recognition model, and the response time slot of the highlight segment automatic recognition model.
8. A one-stop multi-modal video generation system based on artificial intelligence, characterized in that, Executing the method according to any one of claims 1-7, the system includes: A data acquisition module, configured to obtain various types of data from multiple input nodes in response to a video generation task; A data preprocessing module for preprocessing various types of acquired data; A multi-modal processing module for performing multi-modal data fusion processing on various types of data after preprocessing; A video generation module for generating corresponding videos according to the data after multi-modal data fusion processing by using customized video generation models corresponding to different types of video generation tasks pre-constructed; 9. The one-stop multi-modal video generation system based on artificial intelligence according to claim 8, characterized in that, The system further includes: A data acquisition module for collecting quality defect feature data in the generated videos and obtaining user satisfaction; A data processor for comprehensively evaluating the quality of the generated videos according to the quality defect feature data and user satisfaction, and calculating the comprehensive expected loss value of the generated video quality; A data comparator for comparing the size of the comprehensive expected loss value of the generated video quality with a preset threshold. If the comprehensive expected loss value of the generated video quality is greater than the preset threshold, the customized video generation model for generating this video is optimized. Otherwise, there is no need to optimize the customized video generation model for generating this video.
10. The one-stop multi-modal video generation system based on artificial intelligence according to claim 8, characterized in that, The system further includes: An automatic recognition module for automatically recognizing exciting segments in the generated videos based on a pre-constructed exciting segment automatic recognition model, and performing editing and slicing processing on the exciting segments to obtain edited videos.
Citation Information
Cited By
AIGC intelligent agent based on fusion of multiple models
CN121660110A