Video content generation method and system based on artificial intelligence, terminal and medium
Through the video content generation method based on artificial intelligence, the time-consuming and cost-effective video production of sports event highlights is solved, and the video content that is automatically generated with emotional and visual coherent is realized, personalized customization is supported, and production efficiency is improved.
Patent Information
- Application Number
- CN202510171138.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems such as time-consuming, costly, lack of emotional capture capabilities, unnatural content integration and lack of personalization in the production of high-light moments of sports events.
Using an artificial intelligence-based video content generation method, by obtaining sports event videos, feature extraction and event recognition are performed, video content that meets the target comic style is generated, and style matching and video synthesis are performed.
It realizes automatic generation of emotional and visually coherent video content, supports personalized customization, improves the production efficiency of high-light videos in sports events, and meets users' personalized needs.
Smart Images

Figure CN120075525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and computer vision, and particularly to a method, system, terminal and medium for generating video content based on artificial intelligence. Background Art
[0002] In the field of video production for highlight moments of sports events, the existing technology mainly relies on manual editing. This process involves screening out exciting clips from a large number of event videos, and then professional video editors perform editing and post-production. This method has several obvious defects:
[0003] 1. Time-consuming and laborious: Manual editing requires editors to watch videos for a long time, identify and select highlight moments, and then perform editing and adjustment. This process not only consumes a large amount of human resources but also has low efficiency, especially when dealing with a large amount of video material.
[0004] 2. High cost: Since manual editing requires professional skills and experience, the cost is relatively high. In addition, additional special effects and animations may be required during the post-production process, which will further increase the cost.
[0005] 3. Lack of emotional capture ability: Existing automated tools often cannot accurately capture and express the emotional changes in sports events. For example, emotional elements such as the excitement of the audience and the joy of the players at the moment of scoring are difficult to be effectively expressed by automated tools.
[0006] 4. Unnatural content integration: Due to the lack of in-depth understanding of the content of sports events, it is difficult for the automatically generated B-Roll video to be seamlessly integrated with the original video content, resulting in the final video appearing abrupt in terms of emotion and vision and lacking attractiveness.
[0007] 5. Lack of personalization: Existing automated tools often cannot provide personalized customization services, resulting in the generated video content being stereotyped and unable to meet the personalized needs of different users for video content and style.
[0008] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide a method, system, terminal and medium for generating video content based on artificial intelligence in view of the above-mentioned defects of the existing technology. The technical solutions adopted by the present invention are as follows:
[0010] In the first aspect, the present invention provides a method for generating video content based on artificial intelligence, wherein the method includes:
[0011] Obtain a sports event video, extract features from the sports event video to obtain key features, and perform event recognition based on the key features to obtain the key events corresponding to the sports event video;
[0012] Perform sentiment analysis on the sports event video to obtain a sentiment analysis result, and associate the sentiment analysis result with the key events;
[0013] Determine a target comic style based on the sentiment analysis result, and generate video content that conforms to the target comic style based on the associated key events, the sentiment analysis result, and the target comic style;
[0014] Perform style matching and video synthesis processing on the generated video content that conforms to the target comic style and the sports event video, and output the final video.
[0015] In one implementation, the obtaining a sports event video, extracting features from the sports event video to obtain key features, and performing event recognition based on the key features to obtain the key events corresponding to the sports event video includes:
[0016] Obtain a sports event video, convert the video frame rate of the sports event video into a standard frame rate, and perform denoising and sharpening processing on the sports event video;
[0017] Perform feature processing on each frame image in the processed sports event video to obtain key features, and mark the moving objects in the sports event video based on an object detection algorithm;
[0018] Based on the key features and the marked moving objects, and in combination with the audio information in the sports event video, obtain the key events.
[0019] In one implementation, the performing sentiment analysis on the sports event video to obtain a sentiment analysis result includes:
[0020] Obtain the commentary in the sports event video, and perform sentiment analysis processing on the commentary to obtain the sentiment tendency of the commentary;
[0021] Obtain the video effects and audio effects in the sports event video, and combine the sentiment tendency to obtain the sentiment analysis result.
[0022] In one implementation, the determining a target comic style based on the sentiment analysis result, and generating video content that conforms to the target comic style based on the associated key events, the sentiment analysis result, and the target comic style includes:
[0023] Obtain a preset comic style template, and match the sentiment analysis result with the comic style template to obtain the target comic style;
[0024] Use a conditional generative adversarial network to generate video content that conforms to the target comic style based on the associated key event, the sentiment analysis result, and the target comic style;
[0025] Post-process the video content that conforms to the target comic style to ensure the visual consistency between the video content that conforms to the target comic style and the sports event video;
[0026] Use key-frame animation technology to generate a dynamic sequence for the video content that conforms to the target comic style.
[0027] In one implementation, the style matching and video synthesis processing of the generated video content that conforms to the target comic style and the sports event video, and outputting the final video, includes:
[0028] Compare the style of the generated video content that conforms to the target comic style with the style of the sports event video, and use image processing technology to fine-tune the generated video content that conforms to the target comic style to ensure the consistency of the two styles;
[0029] Perform sentiment enhancement processing on the generated video content that conforms to the target comic style;
[0030] Align the time axes of the generated video content that conforms to the target comic style and the sports event video, and perform transition effect design, color adjustment, and brightness adjustment on the generated video content that conforms to the target comic style to enhance visual coherence.
[0031] In one implementation, the style matching and video synthesis processing of the generated video content that conforms to the target comic style and the sports event video, and outputting the final video, further includes:
[0032] Adjust the generated video content that conforms to the target comic style based on a preset parameter setting interface, and receive the preview effect in real time.
[0033] In one implementation, the style matching and video synthesis processing of the generated video content that conforms to the target comic style and the sports event video, and outputting the final video, further includes:
[0034] Perform video synthesis processing on the generated video content that conforms to the target comic style and the sports event video;
[0035] Perform quality inspection and optimization on the synthesized video to obtain the final video, and share and publish the final video.
[0036] In a second aspect, an embodiment of the present invention further provides an artificial intelligence-based video content generation system, where the system includes:
[0037] A feature extraction and event recognition module, configured to obtain a sports event video, perform feature extraction on the sports event video to obtain key features, and perform event recognition based on the key features to obtain the key events corresponding to the sports event video;
[0038] An emotion analysis and association module, configured to perform emotion analysis on the sports event video to obtain an emotion analysis result, and associate the emotion analysis result with the key events;
[0039] A comic-style video generation module, configured to determine a target comic style based on the emotion analysis result, and generate video content that conforms to the target comic style based on the associated key events, the emotion analysis result, and the target comic style;
[0040] A style matching and video synthesis module, configured to perform style matching and video synthesis processing on the generated video content that conforms to the target comic style and the sports event video, and output the final video.
[0041] In a third aspect, an embodiment of the present invention further provides a terminal, where the terminal includes a memory, a processor, and an artificial intelligence-based video content generation program stored in the memory and executable on the processor. When the processor executes the artificial intelligence-based video content generation program, the steps of the artificial intelligence-based video content generation method in any one of the above solutions are implemented.
[0042] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where an artificial intelligence-based video content generation program is stored on the computer-readable storage medium. When the artificial intelligence-based video content generation program is executed by a processor, the steps of the artificial intelligence-based video content generation method in any one of the above solutions are implemented.
[0043] Beneficial effects: Compared with the prior art, the present invention provides a method for generating video content based on artificial intelligence. First, the present invention acquires a sports event video, extracts features from the sports event video to obtain key features, and performs event recognition based on the key features to obtain the key events corresponding to the sports event video. Then, perform sentiment analysis on the sports event video to obtain a sentiment analysis result, and associate the sentiment analysis result with the key event. Next, determine the target comic style based on the sentiment analysis result, and generate video content that conforms to the target comic style based on the associated key event, the sentiment analysis result, and the target comic style. Finally, perform style matching and video synthesis processing on the generated video content that conforms to the target comic style and the sports event video, and output the final video. The present invention can automatically generate video content with coherent emotions and visuals and support personalized customization, so as to improve the production efficiency of highlight videos of sports events and meet the personalized needs of users. Description of the Drawings
[0044] Figure 1 It is a flowchart of a preferred embodiment of the method for generating video content based on artificial intelligence provided by an embodiment of the present invention.
[0045] Figure 2 It is a schematic diagram of the architecture of the system for generating video content based on artificial intelligence provided by an embodiment of the present invention.
[0046] Figure 3 It is a schematic block diagram of the terminal provided by an embodiment of the present invention. Detailed Embodiments
[0047] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] The flowchart shown in the drawings is only an example illustration and does not necessarily include all the content and operations or steps, nor does it necessarily execute in the described order. For example, some operations or steps can be decomposed, combined or partially merged, so the actual execution order may change according to the actual situation.
[0049] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0050] It should be understood that, for the convenience of clearly describing the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first control information and the second control information are only used to distinguish different control information, and do not limit their sequence.
[0051] Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit being different.
[0052] It should also be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0053] Aiming at the deficiencies of the prior art, the present invention aims to solve the following technical problems:
[0054] 1. Automatically generating a comic-style B-Roll video: The present invention aims to develop a method and system that can automatically identify highlight moments in a sports event and automatically generate a comic-style B-Roll (auxiliary shots, such as establishing shots, behind-the-scenes footage, or associated content highly relevant to the video content, and any video segment that can set the mood, assist in storytelling, or explain a concept can be called a B-Roll) video that matches these moments. Therefore, it is necessary to have the ability to deeply understand the video content and capture emotions, and be able to automatically create a comic-style video that conforms to the emotions of the scene based on this information.
[0055] 2. Ensuring emotional and visual coherence: The present invention needs to solve the problem of how to ensure the emotional and visual coherence between the automatically generated B-Roll video and the original video. This includes inserting the B-Roll at the correct time point, and ensuring that the style, color, and animation of the B-Roll match the original video, so as to provide a smooth and emotional viewing experience.
[0056] 3. Providing a personalized customization function: The present invention also aims to provide users with the function of personalized customization of B-Roll videos. This means that users can adjust parameters such as the style, content, and duration of the B-Roll according to their own preferences and needs, and even add their own elements, such as custom comic characters or special effects, to create a unique viewing experience.
[0057] To this end, this embodiment provides a method for generating video content based on artificial intelligence. First, the present invention acquires a sports event video, extracts features from the sports event video to obtain key features, and performs event recognition based on the key features to obtain the key events corresponding to the sports event video. Then, perform sentiment analysis on the sports event video to obtain a sentiment analysis result, and associate the sentiment analysis result with the key event. Next, determine the target comic style based on the sentiment analysis result, and generate video content that conforms to the target comic style, that is, B-Roll, based on the associated key event, the sentiment analysis result, and the target comic style. Finally, perform style matching and video synthesis processing on the generated video content that conforms to the target comic style and the sports event video, and output the final video. The present invention can automatically generate video content with coherent emotions and visuals and support personalized customization in a comic style to improve the production efficiency of highlight videos of sports events and meet the personalized needs of users.
[0058] The method for generating video content based on artificial intelligence in this embodiment can be applied to a terminal, and the terminal is an intelligent product terminal such as a computer, a television, and a smart phone. For example, Figure 1 As shown in, the method for generating video content based on artificial intelligence includes the following steps:
[0059] Step S100: Acquire a sports event video, extract features from the sports event video to obtain key features, and perform event recognition based on the key features to obtain the key events corresponding to the sports event video.
[0060] In this embodiment, a sports event video is first input, and then frame extraction and preprocessing are performed on the input sports event video, which specifically includes automatically detecting the video frame rate of the sports event video, then converting the video frame rate of the sports event video into a standard frame rate, and performing denoising and sharpening processing on the sports event video, so as to improve the accuracy of subsequent feature recognition. Next, in this embodiment, a deep learning model, such as ResNet or InceptionNet, is used to perform feature processing on each frame image in the processed sports event video to obtain key features, such as identifying key features such as athletes, balls, and venues. And the moving objects in the sports event video can also be marked based on an object detection algorithm (such as YOLO or SSD), such as marking athletes and balls. Further, in this embodiment, based on the key features and the marked moving objects, combined with the audio information in the sports event video (such as applause and intonation changes in the commentary), a machine learning model (such as a random forest or SVM) is used to identify key events in the video, such as scores and mistakes.
[0061] In specific applications, for example, the sports event video is a 90-minute football game video with a resolution of 1920x1080 and a video frame rate of 25fps. In this embodiment, one frame can be extracted from the football game video every 10 seconds, and a total of 270 key frames are extracted. Each frame is denoised using a 3x3 Gaussian kernel, and the kernel function is:
[0062]
[0063] Among them, x: represents the horizontal coordinate offset of the pixel point in the image, and its value range is from -1 to 1, which is used to determine the relative position of the current pixel point with respect to the center point of the Gaussian kernel in the horizontal direction. When x = 0, it means that the current pixel point is located at the center of the Gaussian kernel; when x = -1 or x = 1, it means that the current pixel point is located at the edge of the Gaussian kernel. y: represents the vertical coordinate offset of the pixel point in the image, and its value range is from -1 to 1, which is used to determine the relative position of the current pixel point with respect to the center point of the Gaussian kernel in the vertical direction. When y = 0, it means that the current pixel point is located at the center of the Gaussian kernel; when y = -1 or y = 1, it means that the current pixel point is located at the edge of the Gaussian kernel. σ: represents the standard deviation of the Gaussian kernel, which is used to control the width and shape of the Gaussian kernel. The larger σ is, the wider the Gaussian kernel is, and the stronger the denoising effect is, but it may cause blurring of image details; the smaller σ is, the narrower the Gaussian kernel is, and the weaker the denoising effect is, but it can better retain image details.
[0064] Among them, the ranges of x and y are from -1 to 1. Then, the Laplace operator is used for sharpening processing, and the operator is:
[0065]
[0066] The sharpened image I sharp is calculated as:
[0067]
[0068] where I is the input image.
[0069] Next, for feature extraction, in this embodiment, the athletes and the ball in the football game can be recognized. Specifically, a pre-trained VGG-16 model can be used, and the key frames are input, and features are extracted through 13 convolutional layers and 3 fully connected layers.
[0070] The core formula for feature extraction is the convolution operation:
[0071]
[0072] I: represents the input image, which is a three-dimensional matrix and usually contains the height, width, and color channels of the image (such as RGB). In the convolution operation, I is used to provide the original image data.
[0073] K: represents the convolutional kernel, which is a small two-dimensional matrix used to slide over the input image and extract local features. The size of the convolutional kernel (such as 3x3, 5x5) determines the range of the image area covered by each convolution operation.
[0074] F: represents the feature map, which is the output result of the convolution operation and is also a three-dimensional matrix. The size of F depends on the size of the input image I, the size of the convolutional kernel K, as well as the stride and padding of the convolution operation. The feature map F contains the local feature information extracted from the input image by the convolutional kernel and is used for subsequent image analysis and processing.
[0075] In this step, this embodiment focuses on the feature changes in the goal area, especially the two corners of the goal, and uses the coordinates of the goal corners as the points of interest.
[0076] When performing event recognition, this embodiment can recognize goal events. Specifically, it can combine visual features (such as changes in the goal area) and audio features (such as changes in the volume of applause), and use the Hidden Markov Model (HMM) for event recognition.
[0077] The state transition probability and observation probability of the HMM are:
[0078] P(X t+1 =j|X t =i)=a ij
[0079] P(O t =k|X t =i)=b ik
[0080] a ij : represents the state transition probability, that is, the probability of transitioning from state i to state j. In the context of event recognition, the state can represent different events in the video (such as no goal, possible goal, confirmed goal), and a ij describes the possibility of conversion between these events. For example, a12 represents the probability of transitioning from the no-goal state to the possible-goal state.
[0081] b ik : represents the observation probability, that is, the probability of observing the observation value k in state i. The observation value can be certain features or signals in the video (such as changes in the goal area, changes in the volume of applause), and b ik describes the probability of the appearance of these features or signals in a specific event state. For example, the probability of observing an obvious change in the goal area in the possible-goal state can be expressed as b2k, where k corresponds to the feature of the change in the goal area.
[0082] J: Represents the number of states. In this embodiment, J = 3, corresponding to the three states of no goal, possible goal, and confirmed goal.
[0083] t: Represents the time step, which is used to describe the time series in the event recognition process. In video analysis, t can correspond to the serial number or timestamp of the video frame.
[0084] Xt+1: Represents the state at time step t + 1, that is, the event state at the next time point of the video. For example, if the state at the current time step t is possible goal Xt = 2, then Xt+1 can be no goal (Xt+1 = 1), possible goal Xt+1 = 2, or confirmed goal Xt+1 = 3, depending on the state transition probability aij.
[0085] Ot: Represents the observation value at time step t, that is, the feature or signal observed at the current time point of the video. For example, Ot can be the change situation in the goal area of the current video frame or the intensity of the audience's applause, etc.
[0086] Among them, a ij is the state transition probability, and b ik is the probability of observing k in state i.
[0087] This embodiment sets 3 states: no goal, possible goal, and confirmed goal. The state transition probability matrix \(A\) and the observation probability matrix B are set according to historical data. For example:
[0088]
[0089]
[0090] Step S200: Perform sentiment analysis on the sports event video to obtain the sentiment analysis result, and associate the sentiment analysis result with the key event.
[0091] In this embodiment, the commentary in the sports event video can be obtained, and sentiment analysis processing is performed on the commentary. A sentiment analysis model (such as a sentiment analysis model based on LSTM) is used to determine the sentiment tendency of the commentary. Moreover, the video effects and audio effects in the sports event video are analyzed to assist in sentiment analysis. Combining the sentiment tendency, the sentiment analysis result can be obtained. Then, the identified key event is associated with the sentiment analysis result to determine the sentiment value of each key event, providing a sentiment background for the generation of the anime-style B-Roll.
[0092] In specific applications, taking the above football game video as an example, when analyzing the emotional tendency in the commentary, based on the emotion dictionary, the occurrence frequencies of positive and negative words in the commentary can be counted, the emotion score can be calculated, and the emotion analysis result can be obtained. The formula for the emotion score is:
[0093]
[0094] where p i is the occurrence frequency of word i, and s i is the emotion value of word i. For example, the emotion value of the word "goal" is +1, and the emotion value of "mistake" is -1. If "goal" appears 5 times and "mistake" appears 2 times, then the emotion score is:
[0095]
[0096] Associate the goal event with the emotion analysis result. According to the occurrence of the goal event, combined with the emotion analysis result, use the logistic regression model to adjust the emotional expression intensity of the subsequent B-Roll:
[0097]
[0098] where W and b are model parameters, and x is the feature vector, which includes the intensity of the goal event and the emotion score of the commentary. For example, if the intensity of the goal event is 0.8 and the emotion score is 0.4, then:
[0099]
[0100] Step S300: Determine the target comic style based on the emotion analysis result, and generate video content that conforms to the target comic style based on the associated key event, the emotion analysis result, and the target comic style.
[0101] Specifically, in this embodiment, a preset comic style template is obtained, and the emotion analysis result is matched with the comic style template to obtain the target comic style, such as exciting, humorous, or sad, and the key visual elements in the target comic style are defined, such as special effect words, speed lines, motion blur, etc. Then, use the conditional generative adversarial network (cGAN) to generate video content that conforms to the target comic style based on the associated key event, the emotion analysis result, and the target comic style, that is, obtain personalized comic B-Roll images. Then, post-process the video content that conforms to the target comic style, such as color correction and detail enhancement, to ensure the visual consistency of the video content that conforms to the target comic style with the sports event video; and use key frame animation technology to generate a dynamic sequence for the video content that conforms to the target comic style to ensure the smoothness and coherence of the actions.
[0102] In specific applications, for example, the passionate style is selected as the target comic style for the goal event. According to the sentiment analysis results, a template of the passionate style is selected, and the color histogram and line thickness are defined to match the sentiment intensity. The color histogram matching formula is:
[0103]
[0104] where H(I) is the color histogram, N is the total number of pixels in the image, and δ is the Dirac δ function. For example, if the color histogram of the goal moment in the original football game video is H original , then the color histogram H B-Roll of the video content B-Roll generated to conform to the target comic style should be adjusted to:
[0105] H B-Roll = H original ·α
[0106] where α is the adjustment coefficient, which is determined according to the sentiment intensity.
[0107] Next, cGAN is used to generate video content that conforms to the target comic style. The key frames and sentiment labels are input, and stylized B-Roll images are output. The loss function of cGAN is:
[0108] L = L GAN + L cGAN
[0109] where L GAN is the loss function of the traditional GAN, and L cGAN is the additional loss term of the conditional GAN to ensure that the generated video content conforms to the passionate style. For example, if the sentiment label is "exciting", a penalty term corresponding to the "exciting" style will be added to the loss function of the conditional GAN.
[0110] To ensure that the generated video content B-Roll is consistent with the original video style, color histogram matching technology is used to adjust the color distribution of the B-Roll to make it close to the original video. The optimization problem of color histogram matching can be expressed as:
[0111] minλ∫|H(I B-Roll ) - H(I Original )| λ dI
[0112] where H(I B-Roll ) and H(I Original ) are the color histograms of the B-Roll and the original video respectively, and λ is the adjustment parameter. For example, if the mean of H(I Original ) is μ original , H(IB-Roll ) has a mean of μ B-Roll , then the adjustment parameter λ can be calculated as:
[0113]
[0114] where σ original is the standard deviation of the original video color histogram.
[0115] Step S400: Match the style and perform video synthesis processing on the generated video content that conforms to the target comic style and the sports event video, and output the final video.
[0116] In this embodiment, the style of the generated video content that conforms to the target comic style is compared with the style of the sports event video, and image processing technology is used to fine-tune the generated video content that conforms to the target comic style to ensure the consistency of the two styles. In practical applications, the style matching algorithm can be optimized through a user feedback loop to improve the accuracy of matching. Then, emotional enhancement processing is performed on the generated video content that conforms to the target comic style, such as adding exaggerated emotional expressions to the generated video content, such as magnifying the impact of key actions or highlighting the expressions of athletes, to enhance the emotional impact of the video. Then, align the time axes of the generated video content that conforms to the target comic style and the sports event video, and design transition effects, color adjustments, and brightness adjustments for the generated video content that conforms to the target comic style to enhance visual coherence. Specifically, in this embodiment, timestamps and event markers can be used to align the generated video content B-Roll with the time axis of the original video to ensure that the B-Roll is inserted at the correct moment in the video. Edit the generated video content B-Roll and adjust its duration to match the video rhythm and narrative requirements. Then, design smooth transition effects, such as dissolve, wipe, or slide effects, to achieve seamless transition between the B-Roll and the original video. Adjust the parameters of the transition effects to adapt to different video content and styles. Then adjust the color and brightness of the generated video content B-Roll to match the color style of the original video and enhance visual coherence.
[0117] Furthermore, in this embodiment, the generated video content conforming to the target comic style is adjusted based on a preset parameter setting interface, and the preview effect is received in real time. Specifically, a parameter setting interface can be provided to allow users to select different B-Roll styles, adjust the position and duration of the B-Roll. Preset B-Roll templates and customization options are provided to meet the needs of different users. And it can also allow users to preview the effect of the B-Roll in real time and provide instant feedback. And according to the user's feedback, the system automatically adjusts the generation parameters of the B-Roll to optimize the final effect. Users can also upload custom images or animations, and the system integrates these elements into the B-Roll to achieve personalized customization.
[0118] Furthermore, in this embodiment, the generated video content conforming to the target comic style is subjected to video synthesis processing with the sports event video, and a high-performance rendering engine is used to ensure the video quality, and video encoding and compression are performed to adapt to different playback platforms and devices. Quality inspection and optimization are performed on the synthesized video, including verification of resolution, frame rate, and encoding format, and according to the quality inspection results, necessary optimizations are performed on the video, such as adjusting the bit rate or fixing synchronization problems, so as to obtain the final video, and the final video is shared and published. The final video can be shared on social media or published on a video sharing platform.
[0119] In specific applications, in this embodiment, timestamps can be used to align the generated video content B-Roll with the goal-scoring moment, and non-linear editing techniques (such as Adobe Premiere Pro) are used for editing. The time alignment can be achieved through the following formula:
[0120] t B-Roll =t Goal +Δt
[0121] where t B-Roll is the insertion time of the B-Roll, t Goal is the timestamp of the goal, and Δt is the delay time, such as 0.5 seconds. If the goal occurs at the 45th minute, that is, 2700 seconds, then:
[0122] t B-Roll =2700 + 0.5=2700.5 seconds
[0123] Then, design smooth transition effects, such as dissolve, wipe, or slide effects, to achieve seamless transitions between the B-Roll and the original video. Use the dissolve effect and set the dissolve time to 1 second to ensure a natural and smooth transition between the B-Roll and the original video. If the duration of the B-Roll is 3 seconds, it should start at 2700.5 seconds and end at 2703.5 seconds. Then, adjust the color and brightness of the B-Roll to match the color style of the original video, enhancing visual coherence. Use color correction techniques to adjust the color balance of the B-Roll to match the color style of the original video. For example, if the original video has a warm color tone, the color temperature of the B-Roll should also be adjusted to a warm color tone.
[0124] Furthermore, this embodiment can provide a user-friendly interface that allows users to select different B-Roll styles, adjust the position and duration of the B-Roll. Users can choose styles such as "passionate" or "humorous", as well as the start time and duration of the B-Roll in the video. For example, users can choose to set the start time of the B-Roll to 1 second after the goal and the duration to 5 seconds. And allow users to preview the effect of the B-Roll in real time and provide instant feedback. Users can watch the effect after the B-Roll is combined with the original video and make adjustments as needed. If users feel that the transition effect of the B-Roll is too abrupt, they can increase the dissolve time from 1 second to 2 seconds. Users can also upload custom images or animations, and the system will integrate these elements into the B-Roll to achieve personalized customization. For example, if a user uploads a comic-style avatar of their favorite football star, the system will integrate this avatar into the B-Roll, especially at the moment of the goal, to enhance the personalized experience.
[0125] Finally, perform the final synthesis of the adjusted B-Roll and the original video, and use a high-performance rendering engine to ensure video quality. Use the H.264 encoding format for video rendering, set the bitrate to 5Mbps, to ensure that the video has a small file size while maintaining high image quality. Conduct a quality check on the output video, including verification of resolution, frame rate, and encoding format. Ensure that the resolution of the video is 1920x1080, the frame rate is 25fps, and the encoding format is H.264. If synchronization issues are found in the video, use frame re-timing techniques to adjust. This embodiment provides sharing and publishing functions, and users can share the final video to social media or publish it to a video sharing platform. Users can directly upload the final video to a video sharing platform to share it with others. The system can also provide a one-click sharing function to facilitate users to quickly share the video.
[0126] This embodiment can automatically generate a comic-style B-Roll video that matches the highlight moments of a sports event. Through deep learning and natural language processing technologies, in-depth understanding and emotion capture of sports video content are achieved, and then a comic-style B-Roll that matches the emotion and visual style of the original video is automatically created, including the following advantages:
[0127] (1) Automated content generation: Use a deep learning model to automatically extract key frames from a sports event video and identify highlight moments, and automatically generate a comic-style B-Roll, significantly improving production efficiency.
[0128] (2) Emotion and visual coherence: Through emotion analysis and video content understanding, ensure the coherence of the B-Roll with the original video in terms of emotion and vision, providing a smooth viewing experience.
[0129] (3) Personalized customization function: Users can adjust the style and content of the B-Roll according to their personal preferences to achieve personalized video enhancement.
[0130] Based on the above embodiment, the present invention also provides an artificial intelligence-based video content generation system, as Figure 2 shown, the system includes: a feature extraction and event recognition module 10, an emotion analysis and association module 20, a comic-style video generation module 30, and a style matching and video synthesis module 40. Specifically, the feature extraction and event recognition module 10 is used to obtain a sports event video, extract features from the sports event video to obtain key features, and perform event recognition based on the key features to obtain the key events corresponding to the sports event video. The emotion analysis and association module 20 is used to perform emotion analysis on the sports event video to obtain an emotion analysis result, and associate the emotion analysis result with the key events. The comic-style video generation module 30 is used to determine a target comic style based on the emotion analysis result, and generate video content that conforms to the target comic style based on the associated key events, the emotion analysis result, and the target comic style. The style matching and video synthesis module 40 is used to perform style matching and video synthesis processing on the generated video content that conforms to the target comic style and the sports event video, and output the final video.
[0131] The working principles of the various modules in the artificial intelligence-based video content generation system of this embodiment are the same as those of the various steps in the above method embodiment, and will not be elaborated here.
[0132] Each module in the above artificial intelligence-based video content generation system can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor in the terminal in hardware form or be independent of the processor, or can be stored in the memory in the terminal in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0133] Based on the above embodiments, the present invention further provides a terminal, and the principle block diagram of the terminal can be as Figure 3 shown. The terminal may include one or more processors 100 ( Figure 3 only one is shown in the figure), a memory 101, and a computer program 102 stored in the memory 101 and executable on one or more processors 100.
[0134] In one embodiment, the so-called processor 100 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0135] In one embodiment, the memory 101 may be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 101 may also include both the internal storage unit and the external storage device of the electronic device. The memory 101 is used to store the computer program and other programs and data required by the terminal. The memory 101 may also be used to temporarily store the data that has been output or will be output.
[0136] Those skilled in the art can understand that Figure 3 the principle block diagram shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, operational database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A video content generation method based on artificial intelligence, characterized in that: The method comprises: Acquire a sports event video, perform feature extraction on the sports event video to obtain key features, and perform event recognition based on the key features to obtain key events corresponding to the sports event video; Performing sentiment analysis on the sports event video to obtain sentiment analysis results, and associating the sentiment analysis results with the key events; Determining a target comic style based on the sentiment analysis result, and generating video content that conforms to the target comic style based on the associated key events, the sentiment analysis result and the target comic style; The generated video content that conforms to the target comic style is style-matched and video-synthesized with the sports event video to output a final video.
2. The method for generating video content based on artificial intelligence according to claim 1, characterized in that: The step of acquiring a sports event video, extracting features from the sports event video to obtain key features, and performing event recognition based on the key features to obtain key events corresponding to the sports event video includes: Acquire a sports event video, convert a video frame rate of the sports event video into a standard frame rate, and perform denoising and sharpening processing on the sports event video; Performing feature processing on each frame of the processed sports event video to obtain key features, and marking the moving objects in the sports event video based on an object detection algorithm; The key event is obtained based on the key features and the marked moving objects and combined with the audio information in the sports event video.
3. The method for generating video content based on artificial intelligence according to claim 1, characterized in that: The performing sentiment analysis on the sports event video to obtain a sentiment analysis result includes: Obtaining commentary in the sports event video, and performing sentiment analysis on the commentary to obtain the sentiment tendency of the commentary; The video effect and the audio effect in the sports event video are obtained, and combined with the emotional tendency to obtain the emotional analysis result.
4. The method for generating video content based on artificial intelligence according to claim 1, characterized in that: The determining a target comic style based on the sentiment analysis result, and generating video content that conforms to the target comic style based on the associated key events, the sentiment analysis result and the target comic style, includes: Obtaining a preset comic style template, and matching the sentiment analysis result with the comic style template to obtain the target comic style; Generate video content that conforms to the target comic style using a conditional generative adversarial network based on the associated key events, the sentiment analysis results, and the target comic style; Post-processing the video content that conforms to the target comic style to ensure visual consistency between the video content that conforms to the target comic style and the sports event video; A keyframe animation technique is used to generate a dynamic sequence for the video content that meets the target comic style.
5. The method for generating video content based on artificial intelligence according to claim 1, characterized in that: The step of performing style matching and video synthesis processing on the generated video content conforming to the target comic style and the sports event video to output a final video includes: Comparing the style of the generated video content that conforms to the target comic style with the style of the sports event video, and using image processing technology to fine-tune the generated video content that conforms to the target comic style to ensure consistency between the two styles; Performing emotion enhancement processing on the generated video content that conforms to the target comic style; The generated video content conforming to the target comic style is aligned with the timeline of the sports event video, and transition effects are designed, colors are adjusted, and brightness is adjusted for the generated video content conforming to the target comic style to enhance visual coherence.
6. The method for generating video content based on artificial intelligence according to claim 5, characterized in that: The step of performing style matching and video synthesis processing on the generated video content conforming to the target comic style and the sports event video to output a final video also includes: The generated video content conforming to the target comic style is adjusted based on a preset parameter setting interface, and a preview effect is received in real time.
7. The method for generating video content based on artificial intelligence according to claim 6, characterized in that: The step of performing style matching and video synthesis processing on the generated video content conforming to the target comic style and the sports event video to output a final video also includes: Performing video synthesis processing on the generated video content conforming to the target comic style and the sports event video; The synthesized video is quality checked and optimized to obtain a final video, and the final video is shared and published.
8. A video content generation system based on artificial intelligence, characterized in that: The system comprises: A feature extraction and event recognition module, which is used to obtain sports event videos, perform feature extraction on the sports event videos to obtain key features, and perform event recognition based on the key features to obtain key events corresponding to the sports event videos; A sentiment analysis and association module, used to perform sentiment analysis on the sports event video, obtain sentiment analysis results, and associate the sentiment analysis results with the key events; A comic-style video generation module, used to determine a target comic style based on the sentiment analysis result, and to generate video content that conforms to the target comic style based on the associated key events, the sentiment analysis result and the target comic style; The style matching and video synthesis module is used to perform style matching and video synthesis processing on the generated video content that conforms to the target comic style and the sports event video, and output the final video.
9. A terminal, characterized in that: The terminal includes a memory, a processor, and an artificial intelligence-based video content generation program stored in the memory and executable on the processor. When the processor executes the artificial intelligence-based video content generation program, the steps of the artificial intelligence-based video content generation method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an artificial intelligence-based video content generation program. When the artificial intelligence-based video content generation program is executed by the processor, the steps of the artificial intelligence-based video content generation method as described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Video editing method and apparatus, device, and storage medium
WO2026157444A1