Fast video editing method based on artificial intelligence
By collecting, identifying and managing video data features and optimizing editing strategies based on user feedback, the problem of manual reliance and low efficiency in existing editing technologies has been resolved. Multi-dimensional quantitative evaluation and personalized adaptation have been achieved, improving the intelligence and fluency of editing.
Patent Information
- Application Number
- CN202510741543.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing intelligent video editing technology relies on manual operation, lacks multi-dimensional feature analysis, has a single editing effect, lacks quantitative standards, has low editing efficiency, and cannot achieve automation and personalized adaptation.
Through the intelligent collection, identification and management of video data features, combined with visual, audio and text feature scoring, the initial editing sequence is generated, and a preference model is built through user feedback to optimize the editing strategy, achieving multi-dimensional quantitative evaluation and personalized adaptation.
The intelligence and efficiency of video editing have been improved, the editing results are more in line with user habits, and the fluency and scientificity of the editing effects have been improved.
Smart Images

Figure CN120711230A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and more specifically, to an artificial intelligence-based fast video editing method. Background Art
[0002] With the development of science and technology, the production and consumption of digital media have become a part of daily life. As an important means of information dissemination, the demand for video is increasing. The core goal of intelligent video editing is to automatically achieve high-quality video editing. It combines advanced technologies in multiple fields such as computer vision, image processing, and machine learning. By analyzing and processing elements such as pictures and audio in the video, it automatically identifies important information such as key frame shot switching points, and automatically generates edited videos according to preset rules or user needs.
[0003] Artificial intelligence technology can analyze and learn large amounts of video data through deep learning models, thereby automatically identifying important information in the video, such as key frames and motion changes. This information can serve as the basis for video editing, helping users quickly generate attractive video works. It is widely used in various fields, such as news reporting and social media, film and television production, etc.
[0004] However, in actual use, it still has some shortcomings. For example, the existing video intelligent editing technology relies on manual screening, sorting and optimization of a large number of video frames in actual application, which cannot achieve automated editing and lacks analysis of multi-dimensional features in the video, resulting in a single editing effect. Existing video editing effect evaluation lacks quantitative standards, insufficient video fluency, and a lack of automated learning mechanism for user preferences. Existing methods are not flexible enough, resulting in low editing efficiency. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a video fast editing method based on artificial intelligence, which is used to solve the problems raised in the above-mentioned background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a method for rapid video editing based on artificial intelligence, comprising the following steps: Step S01: Intelligent collection of video data features: used to obtain the video to be edited, and collect visual data, audio data, and text data in the video frame.
[0007] Step S02: Intelligent recognition of video data features: including a visual feature recognition sub-step, an audio feature recognition sub-step, and a text feature recognition sub-step. The visual feature recognition sub-step is used to analyze the visual feature recognition score of the video frame based on the visual data, the audio feature recognition sub-step is used to analyze the audio feature recognition score of the video frame based on the audio data, and the text feature recognition sub-step is used to analyze the text feature recognition score of the video frame based on the text data.
[0008] Step S03: Intelligent generation of video clips: used to obtain the visual feature recognition score, audio feature recognition score and text feature recognition score of the video frame, score each frame of the video to be edited, generate a video frame sequence arranged according to the score, and generate the initial video editing sequence according to the video editing rules input by the user.
[0009] Step S04: Intelligent management of video editing effects: Based on the initial video editing sequence, the editing quality evaluation data of the video frames is obtained, and the editing effect smoothness evaluation coefficients of adjacent video frames are analyzed.
[0010] Step S05: Video editing strategy evaluation: used to obtain the editing effect smoothness evaluation coefficient between adjacent video frames, compare it with the preset editing effect smoothness evaluation coefficient, and process it.
[0011] Step S06: User preference learning: Through the interactive interface, collect user feedback data on the initial video editing sequence, build a user preference model, and when editing a new video, adjust the video frame selection features of the initial video editing sequence based on the user preference model.
[0012] Preferably, the step S01: intelligent collection of video data features is specifically as follows: S21: Obtain a video to be edited and pre-process the video to be edited, including data cleaning and unifying the video frame data format, wherein the unified video frame data format includes resolution standardization, audio sampling rate standardization, and text encoding standardization; S22: Receive the pre-processed video to be edited, divide the video to be edited into video frames, collect features, and number the video frames in sequence as 1, 2, ..., i, ..., n; S23: The visual data collected includes color features, clarity features, and content features; S24: The audio data collection content includes time features and audio features; S25: The text data collection content is text recognition features.
[0013] Preferably, the visual feature recognition sub-step: S31: extracting the number of HSV spatial histogram intervals, the number of non-empty intervals, the hue of each interval, and the saturation of each interval according to the color features in the video frame, and calculating the color richness in the video frame; S32: extracting a grayscale image according to the clarity feature in the video frame, calculating the variance value in the video frame using a Laplace operator, and obtaining the clarity in the video frame; S33: Based on the content features in the video frame, use the pre-trained object detection model to extract the detected confidence level and the number of detected objects, and calculate the target object detection score in the video frame; S34: Based on the color richness, clarity and target object detection score in the video frame, obtain a visual feature recognition score of the video frame.
[0014] Preferably, the audio feature recognition sub-step: S41: extracting the total duration of the speech segment and the total duration of the audio according to the time features in the video frame, and calculating the effective audio detection score in the video frame; S42: Divide the audio in the video frame into multiple segments, extract the feature vector of each segment, and mark the audio segments in the video frame as 1, 2, ...j, ...J in sequence, and calculate the difference between the feature vectors of the audio segments in the video frame; S43: obtaining the audio layer richness in the video frame by calculating the average value of the feature vector differences between the audio segments in the video frame; S44: Based on the effective audio detection score and audio layer richness in the video frame, obtain an audio feature recognition score for the video frame.
[0015] Preferably, the text feature recognition sub-step: According to the text recognition features in the video frame, the text in the video frame is recognized and output as a character sequence , confidence , and obtain the text feature recognition score of the video frame.
[0016] Preferably, the step S03: intelligent generation of video clips is specifically as follows: S61: scoring each frame of the video to be edited based on the visual feature recognition score, the audio feature recognition score, and the text feature recognition score of the video frame, and sorting the video frames according to the scores to obtain a sorted video frame sequence; S62: According to the video editing rule input by the user, video frames that meet the rule are selected from the sorted video frame sequence to generate an initial video editing sequence; S63: Smoothing the initial video clip sequence to remove redundant video frames to obtain a final video clip result.
[0017] Preferably, the step S04: intelligent management of video editing effects is specifically as follows: S71: The clipping quality assessment data is the text area coverage and video fluency of each video frame in the initial video clip sequence; S72: For adjacent video frames in the initial video clip sequence, calculating a text area coverage difference based on the text area coverage, and calculating a video fluency difference based on the video fluency; S73: Based on the text area coverage difference and video smoothness difference of adjacent video frames in the initial video editing sequence, calculate the editing effect smoothness evaluation coefficient between adjacent video frames; wherein the text area coverage difference is used to evaluate the visibility and coherence of text information between adjacent video frames, and the video smoothness difference is used to evaluate the similarity between adjacent video frames. The smaller the difference between adjacent video frames, the higher the video smoothness.
[0018] Preferably, the step S05: video editing strategy evaluation is specifically as follows: Obtain the editing effect smoothness evaluation coefficient between adjacent video frames and compare it with the preset editing effect smoothness evaluation coefficient. If the editing effect smoothness evaluation coefficient between adjacent video frames is less than the preset editing effect smoothness evaluation coefficient, it indicates that the editing effect between adjacent video frames is not smooth and needs further adjustment. Otherwise, it indicates that the editing effect between adjacent video frames is smooth, and then the video editing strategy is continuously iterated and optimized for video frames with poor editing results.
[0019] Preferably, the step S06: user preference learning is specifically as follows: Step S91: collecting user feedback data on the initial editing sequence through an interactive interface, wherein the user feedback data includes but is not limited to adjusting the editing order, deleting segments, and modifying transition effects; Step S92: extracting user preference features based on user feedback data: visual feature preference recognition, audio feature preference recognition, and text feature preference recognition; Step S93: Build a user preference model based on the user preference features. When editing a new video, predict the user's preference probability for each video frame of the video to be edited based on the user preference model, and adjust the video frame selection features of the initial video editing sequence.
[0020] Technical effects and advantages of the present invention: 1. The present invention provides an artificial intelligence-based fast video editing method. The method obtains a video to be edited, obtains a visual feature recognition score of the video frame based on the color richness, clarity and target object detection score in the video frame, obtains an audio feature recognition score of the video frame based on the effective audio detection score and audio layer richness in the video frame, and obtains a text feature recognition score of the video frame based on the text output in the video frame as a character sequence and a confidence level. The visual feature score can automatically identify video frames with high visual quality and key content. The audio feature score can identify high-quality audio segments. The text feature score can evaluate the quality of the text in the video frame, which is conducive to multi-dimensional quantitative evaluation and improves the scientific nature of editing decisions. By obtaining the visual feature recognition score, audio feature recognition score and text feature recognition score of the video frame, each frame of the video to be edited is scored, and a video frame sequence arranged according to the score is generated. According to the video editing rules input by the user, an initial video editing sequence is generated. The video frames to be edited are converted into a comparable standard numerical sequence using the visual, audio and text scores, thereby improving the intelligence and efficiency of editing, achieving personalized and rapid adaptation based on user instructions, and realizing efficient collaboration between data intelligence and artificial creativity. 2. The present invention provides a method for rapid video editing based on artificial intelligence. By obtaining the text area coverage and video fluency of each video frame in the initial video editing sequence, the text area coverage difference and the video fluency difference are calculated, thereby calculating the editing effect fluency evaluation coefficient between adjacent video frames, and comparing it with the preset editing effect fluency evaluation coefficient, the video editing strategy is continuously iterated and optimized for video frames with poor editing results. At the same time, through an interactive interface, user feedback data on the initial video editing sequence is collected to construct a user preference model. When editing a new video, the video frame selection features of the initial video editing sequence are adjusted based on the user preference model. The evaluation and optimization of the editing effect are conducive to improving the editing fluency. By learning user preferences, the generated initial editing sequence is more in line with user habits. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 The present invention is a flowchart of a method for rapid video editing based on artificial intelligence.
[0022] Figure 2 This is a structural diagram of step S02 of the present invention: intelligent recognition of video data features. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0024] See Figure 1 As shown, the present invention provides a method for rapid video editing based on artificial intelligence, comprising the following steps: The step S01: intelligent collection of video data features: is used to obtain the video to be edited, and collect visual data, audio data, and text data in the video frame.
[0025] In a possible design, the step S01: intelligent collection of video data features is specifically as follows: S11: Obtain a video to be edited and pre-process the video to be edited, including data cleaning and unifying the video frame data format, wherein the unified video frame data format includes resolution standardization, audio sampling rate standardization, and text encoding standardization; S12: Receive the pre-processed video to be edited, divide the video to be edited into video frames, collect features, and number the video frames in sequence as 1, 2, ..., i, ..., n; S13: The visual data collected includes color features, clarity features, and content features; S14: The audio data collection content includes time features and audio features; S15: The text data collection content is text recognition features.
[0026] See Figure 2 As shown, the step S02: intelligent recognition of video data features: includes a visual feature recognition sub-step, an audio feature recognition sub-step, and a text feature recognition sub-step. The visual feature recognition sub-step is used to analyze the visual feature recognition score of the video frame based on the visual data, the audio feature recognition sub-step is used to analyze the audio feature recognition score of the video frame based on the audio data, and the text feature recognition sub-step is used to analyze the text feature recognition score of the video frame based on the text data.
[0027] In a possible design, the step S02: intelligent identification of video data features is specifically as follows: Visual feature recognition sub-steps: S21: extracting the number of HSV spatial histogram intervals, the number of non-empty intervals, the hue of each interval, and the saturation of each interval based on the color features in the video frame, and calculating the color richness in the video frame; S22: extracting a grayscale image according to the clarity feature in the video frame, calculating the variance value in the video frame using a Laplace operator, and obtaining the clarity in the video frame; S23: Based on the content features in the video frame, use the pre-trained object detection model to extract the detected confidence level and the number of detected objects, and calculate the target object detection score in the video frame; S24: obtaining a visual feature recognition score of the video frame based on the color richness, clarity, and target object detection score in the video frame; Audio feature recognition sub-steps: S25: extracting the total duration of the speech segment and the total duration of the audio according to the time features in the video frame, and calculating the effective audio detection score in the video frame; S26: Divide the audio in the video frame into multiple segments, extract the feature vector of each segment, and mark the audio segments in the video frame as 1, 2, ...j, ...J in sequence, and calculate the difference between the feature vectors of the audio segments in the video frame; S27: obtaining the audio layer richness in the video frame by calculating the average value of the feature vector differences between the audio segments in the video frame; S28: Obtaining an audio feature recognition score for the video frame based on the effective audio detection score and audio layer richness in the video frame; Text feature recognition sub-steps: S29: Based on the text recognition features in the video frame, the text in the video frame is recognized and output as a character sequence , confidence , and obtain the text feature recognition score of the video frame.
[0028] In this embodiment, it should be specifically explained that the visual feature recognition sub-step is specifically as follows: S001: Convert the video frame image into an HSV spatial histogram, quantize the HSV spatial histogram into multiple intervals (H: 8 intervals, S: 4 intervals, V: 4 intervals), count the number of pixels in each interval, and obtain an 8×4×4 feature vector. The number of HSV spatial histogram intervals is 8×4×4=128. The number of intervals occupied by the color distribution is the number of non-empty intervals. The more non-empty intervals there are, the more color types the image in the video frame contains. S002: The color richness calculation formula is:
[0029] in, Represented as the color richness of the i-th video frame, It is expressed as the number of non-empty intervals of the i-th video frame, Expressed as the number of HSV spatial histogram intervals of the i-th video frame, It is represented by the hue of the cth interval of the i-th video frame, It is expressed as the saturation of the cth interval of the i-th video frame, It is represented as the maximum value of the hue of the i-th video frame, It is represented as the maximum saturation value of the i-th video frame; S003: Use the Laplace operator to calculate the variance value in the video frame. The formula is:
[0030] in, Expressed as the variance value of the i-th video frame, Represented as the grayscale image of the i-th video frame, Expressed as the Laplace operator; Specifically, the Laplace operator is a second-order derivative operator that is sensitive to changes in edges and details in an image. In a clear image, the edges of objects are sharp and the grayscale values change dramatically. After being processed by the Laplace operator, a larger response value is generated. However, blurred images have smooth edges and slowly changing grayscale values, resulting in a smaller Laplace response value. By calculating the variance of the image after Laplace transformation, the degree of discreteness of the pixel values in the image can be measured as a whole. The larger the variance, the more areas with large grayscale value changes in the image, that is, the image is rich in details and has high clarity. Conversely, a smaller variance indicates a more blurred image. The clarity calculation formula is:
[0031] in, Represented as the clarity of the i-th video frame, Expressed as the variance value of the i-th video frame, Expressed as the minimum value of the variance, Expressed as the maximum value of the variance; S004: The calculation formula of the target object detection score is:
[0032] in, It is represented as the target object detection score of the i-th video frame, Expressed as the confidence of the yth target object in the i-th video frame, Represented as the number of detected targets in the i-th video frame, Expressed as the target importance weight of the y-th target object; S005: The calculation formula for the visual feature recognition score is: ,in, Denotes the visual feature recognition score of the i-th video frame.
[0033] In this embodiment, it should be specifically explained that the audio feature recognition sub-step is specifically as follows: S001: The calculation formula for the effective audio detection score is:
[0034] in, Denotes the effective audio detection score of the i-th video frame, It is represented as the total duration of the speech segment of the i-th video frame, Represented as the total audio duration of the i-th video frame; S002: The calculation of the difference between the feature vectors of the audio clips in the video frame is as follows:
[0035] in, Expressed as the difference of the feature vector of the i-th video frame, It is represented as the feature vector of the pth audio segment of the i-th video frame, is represented as the feature vector of the kth audio segment of the i-th video frame, and J is the number of audio segments; S003: The audio layer richness in the video frame is obtained by calculating the average value of the feature vector differences between the audio clips in the video frame. The formula is:
[0036] in, It is represented as the audio layer richness of the i-th video frame; Specifically, a larger value for audio layer richness indicates greater differences in the characteristics of different parts of the audio, and a richer layer, indicating that the audio has more variations in content, timbre, frequency, etc. A smaller value for audio layer richness indicates that the audio sounds more monotonous and has simpler layers. S004: Based on the effective audio detection score and audio layer richness in the video frame, the audio feature recognition score of the video frame is obtained, and the formula is: ,in, Denotes the audio feature recognition score of the i-th video frame.
[0037] In this embodiment, it should be specifically explained that the text feature recognition sub-step is specifically as follows: The text in the video frame is recognized and output as a character sequence according to the text recognition features in the video frame , confidence , get the text feature recognition score of the video frame, the formula is:
[0038] in, It is represented as the text feature recognition score of the i-th video frame, It is represented as the confidence of the xth character in the i-th video frame, and b is the total number of characters; Specifically, the higher the recognition confidence of a single character in a video frame, the higher the accuracy of text feature recognition.
[0039] The step S03: intelligent generation of video clips: is used to obtain the visual feature recognition score, audio feature recognition score and text feature recognition score of the video frame, score each frame of the video to be edited, generate a video frame sequence arranged according to the score, and generate an initial video editing sequence according to the video editing rules input by the user.
[0040] In a possible design, the step S03: intelligent generation of video clips is specifically as follows: S31: scoring each frame of the video to be edited based on the visual feature recognition score, audio feature recognition score, and text feature recognition score of the video frame, and sorting the video frames according to the scores to obtain a sorted video frame sequence; S32: According to the video editing rule input by the user, video frames that meet the rule are selected from the sorted video frame sequence to generate an initial video editing sequence; S33: Smoothing the initial video clip sequence to remove redundant video frames to obtain a final video clip result; The smoothing process may include removing repeated video frames, removing overly short video frames, and other operations to improve the quality and fluency of the video clip.
[0041] In this embodiment, it should be specifically explained that, based on the visual feature recognition score, audio feature recognition score and text feature recognition score of the video frame, each frame of the video to be edited is scored, specifically: score = (visual feature recognition score + audio feature recognition score + text feature recognition score) / 3.
[0042] The step S04: intelligent management of video editing effects: based on the initial video editing sequence, obtaining the editing quality evaluation data of the video frames, and analyzing the editing effect smoothness evaluation coefficients of adjacent video frames.
[0043] In a possible design, the step S04: intelligent management of video editing effects is specifically as follows: S41: The clipping quality assessment data is the text area coverage and video fluency of each video frame in the initial video clip sequence; S42: For adjacent video frames in the initial video clip sequence, calculating a text area coverage difference based on the text area coverage, and calculating a video fluency difference based on the video fluency; S43: Based on the text area coverage difference and video smoothness difference of adjacent video frames in the initial video editing sequence, the editing effect smoothness evaluation coefficient between adjacent video frames is calculated; wherein the text area coverage difference is used to evaluate the visibility and coherence of text information between adjacent video frames, and the video smoothness difference is used to evaluate the similarity between adjacent video frames. The smaller the difference between adjacent video frames, the higher the video smoothness.
[0044] In this embodiment, it should be specifically explained that the intelligent management of video editing effects is as follows: For adjacent video frames in the initial video clip sequence, the text area coverage difference is calculated using the text area coverage. The formula is:
[0045] in, It is expressed as the difference in text region coverage of the fth video frame of the initial video clip sequence, It is expressed as the text area coverage of the f+1th video frame of the initial video clip sequence, It is expressed as the text area coverage of the f-th video frame of the initial video clip sequence, Expressed as the maximum value of text area coverage, Expressed as the minimum value of text area coverage; The video fluency difference is calculated by video fluency, and the formula is:
[0046] in, It is expressed as the video smoothness difference of the f-th video frame in the initial video clip sequence, It is represented as the video smoothness of the f+1th video frame in the initial video clip sequence, It is represented as the video smoothness of the f-th video frame in the initial video clip sequence, Expressed as the maximum value of video smoothness; Based on the difference in text area coverage and video smoothness between adjacent video frames in the initial video editing sequence, the editing effect smoothness evaluation coefficient between adjacent video frames is calculated. The formula is:
[0047] in, It is represented as the evaluation coefficient of the editing effect smoothness of the f-th video frame of the initial video editing sequence, and e is represented as a natural constant.
[0048] The step S05: video editing strategy evaluation: is used to obtain the editing effect smoothness evaluation coefficient between adjacent video frames, compare it with the preset editing effect smoothness evaluation coefficient, and process it.
[0049] In a possible design, the step S05: video editing strategy evaluation is specifically as follows: Obtain the editing effect smoothness evaluation coefficient between adjacent video frames and compare it with the preset editing effect smoothness evaluation coefficient. If the editing effect smoothness evaluation coefficient between adjacent video frames is less than the preset editing effect smoothness evaluation coefficient, it indicates that the editing effect between adjacent video frames is not smooth and needs further adjustment. Otherwise, it indicates that the editing effect between adjacent video frames is smooth, and then the video editing strategy is continuously iterated and optimized for video frames with poor editing results.
[0050] Step S06: User preference learning: Collect user feedback data on the initial video editing sequence through an interactive interface, build a user preference model, and adjust the video frame selection features of the initial video editing sequence based on the user preference model when editing a new video.
[0051] In a possible design, the step S06: user preference learning is specifically as follows: Step S61: collecting user feedback data on the initial editing sequence through an interactive interface, wherein the user feedback data includes but is not limited to adjusting the editing order, deleting segments, and modifying transition effects; Step S62: extracting user preference features based on user feedback data: visual feature preference recognition, audio feature preference recognition, and text feature preference recognition; Step S63: Build a user preference model based on the user preference features. When editing a new video, predict the user's preference probability for each video frame of the video to be edited based on the user preference model, and adjust the video frame selection features of the initial video editing sequence.
[0052] In this embodiment, it should be specifically explained that the present invention obtains a video to be edited, obtains a visual feature recognition score of the video frame based on the color richness, clarity and target object detection score in the video frame, obtains an audio feature recognition score of the video frame based on the effective audio detection score and audio layer richness in the video frame, and obtains a text feature recognition score of the video frame based on the text output in the video frame as a character sequence and a confidence level. The visual feature score can automatically identify video frames with high visual quality and key content, the audio feature score can identify high-quality audio segments, and the text feature score can evaluate the quality of the text in the video frame, which is conducive to multi-dimensional quantitative evaluation and improves the scientific nature of editing decisions. By obtaining the visual feature recognition score, audio feature recognition score and text feature recognition score of the video frame, each frame of the video to be edited is scored, and a video frame sequence arranged according to the score is generated. According to the video editing rules input by the user, an initial video editing sequence is generated. The video frames to be edited are converted into a comparable standard numerical sequence using visual, audio and text scores, thereby improving the intelligence and efficiency of editing, achieving personalized and rapid adaptation based on user instructions, and realizing efficient collaboration between data intelligence and human creativity. The present invention obtains the text area coverage and video fluency of each video frame in the initial video editing sequence, calculates the text area coverage difference and the video fluency difference, and thus calculates the editing effect fluency evaluation coefficient between adjacent video frames, and compares it with the preset editing effect fluency evaluation coefficient. Then, the video editing strategy is continuously iterated and optimized for video frames with poor editing results. At the same time, through an interactive interface, user feedback data on the initial video editing sequence is collected to construct a user preference model. When editing a new video, the video frame selection features of the initial video editing sequence are adjusted based on the user preference model. The evaluation and optimization of the editing effect are conducive to improving the editing fluency. By learning from user preferences, the generated initial editing sequence is more in line with user habits.
[0053] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for rapid video editing based on artificial intelligence, characterized in that: The following steps are involved: Step S01: Intelligent collection of video data features: used to obtain the video to be edited, and collect visual data, audio data, and text data in the video frame; Step S02: Intelligent recognition of video data features: including a visual feature recognition sub-step, an audio feature recognition sub-step, and a text feature recognition sub-step. The visual feature recognition sub-step is used to analyze the visual feature recognition score of the video frame based on the visual data, the audio feature recognition sub-step is used to analyze the audio feature recognition score of the video frame based on the audio data, and the text feature recognition sub-step is used to analyze the text feature recognition score of the video frame based on the text data; Step S03: Intelligent generation of video clips: used to obtain visual feature recognition scores, audio feature recognition scores, and text feature recognition scores of video frames, score each frame of the video to be edited, generate a sequence of video frames sorted by score, and generate an initial video editing sequence based on the video editing rules input by the user; Step S04: Intelligent management of video editing effects: Based on the initial video editing sequence, obtaining the editing quality evaluation data of the video frames, and analyzing the editing effect smoothness evaluation coefficients of adjacent video frames; Step S05: Video editing strategy evaluation: used to obtain the editing effect smoothness evaluation coefficient between adjacent video frames, compare it with the preset editing effect smoothness evaluation coefficient, and process it; Step S06: User preference learning: Through the interactive interface, collect user feedback data on the initial video editing sequence, build a user preference model, and when editing a new video, adjust the video frame selection features of the initial video editing sequence based on the user preference model.
2. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The step S01: intelligent collection of video data features is specifically as follows: S21: Obtain a video to be edited and pre-process the video to be edited, including data cleaning and unifying the video frame data format, wherein the unified video frame data format includes resolution standardization, audio sampling rate standardization, and text encoding standardization; S22: Receive the pre-processed video to be edited, divide the video to be edited into video frames, collect features, and number the video frames in sequence as 1, 2, ..., i, ..., n; S23: The visual data collected includes color features, clarity features, and content features; S24: The audio data collection content includes time features and audio features; S25: The text data collection content is text recognition features.
3. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The visual feature recognition sub-step: S31: extracting the number of HSV spatial histogram intervals, the number of non-empty intervals, the hue of each interval, and the saturation of each interval based on the color features in the video frame, and calculating the color richness in the video frame; S32: extracting a grayscale image according to the clarity feature in the video frame, calculating the variance value in the video frame using a Laplace operator, and obtaining the clarity in the video frame; S33: Based on the content features in the video frame, use the pre-trained object detection model to extract the detected confidence level and the number of detected objects, and calculate the target object detection score in the video frame; S34: Based on the color richness, clarity and target object detection score in the video frame, obtain a visual feature recognition score of the video frame.
4. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The audio feature recognition sub-step: S41: extracting the total duration of the speech segment and the total duration of the audio according to the time features in the video frame, and calculating the effective audio detection score in the video frame; S42: Divide the audio in the video frame into multiple segments, extract the feature vector of each segment, and mark the audio segments in the video frame as 1, 2, ...j, ...J in sequence, and calculate the difference between the feature vectors of the audio segments in the video frame; S43: obtaining the audio layer richness in the video frame by calculating the average value of the feature vector differences between the audio segments in the video frame; S44: Based on the effective audio detection score and audio layer richness in the video frame, obtain an audio feature recognition score for the video frame.
5. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The text feature recognition sub-step: According to the text recognition features in the video frame, the text in the video frame is recognized and output as a character sequence , confidence , and obtain the text feature recognition score of the video frame.
6. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The step S03: intelligent generation of video clips is specifically as follows: S61: scoring each frame of the video to be edited based on the visual feature recognition score, the audio feature recognition score, and the text feature recognition score of the video frame, and sorting the video frames according to the scores to obtain a sorted video frame sequence; S62: According to the video editing rule input by the user, video frames that meet the rule are selected from the sorted video frame sequence to generate an initial video editing sequence; S63: Smoothing the initial video clip sequence to remove redundant video frames to obtain a final video clip result.
7. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The step S04: intelligent management of video editing effects is specifically as follows: S71: The clipping quality assessment data is the text area coverage and video fluency of each video frame in the initial video clip sequence; S72: For adjacent video frames in the initial video clip sequence, calculating a text area coverage difference based on the text area coverage, and calculating a video fluency difference based on the video fluency; S73: Based on the text area coverage difference and video smoothness difference of adjacent video frames in the initial video editing sequence, calculate the editing effect smoothness evaluation coefficient between adjacent video frames; wherein the text area coverage difference is used to evaluate the visibility and coherence of text information between adjacent video frames, and the video smoothness difference is used to evaluate the similarity between adjacent video frames. The smaller the difference between adjacent video frames, the higher the video smoothness.
8. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The step S05: video editing strategy evaluation is specifically as follows: Obtain the editing effect smoothness evaluation coefficient between adjacent video frames and compare it with the preset editing effect smoothness evaluation coefficient. If the editing effect smoothness evaluation coefficient between adjacent video frames is less than the preset editing effect smoothness evaluation coefficient, it indicates that the editing effect between adjacent video frames is not smooth and needs further adjustment. Otherwise, it indicates that the editing effect between adjacent video frames is smooth, and then the video editing strategy is continuously iterated and optimized for video frames with poor editing results.
9. The method for rapid video editing based on artificial intelligence according to claim 1, characterized in that: The step S06: user preference learning is specifically as follows: Step S91: collecting user feedback data on the initial editing sequence through an interactive interface, wherein the user feedback data includes but is not limited to adjusting the editing order, deleting segments, and modifying transition effects; Step S92: extracting user preference features based on user feedback data: visual feature preference recognition, audio feature preference recognition, and text feature preference recognition; Step S93: Build a user preference model based on the user preference features. When editing a new video, predict the user's preference probability for each video frame of the video to be edited based on the user preference model, and adjust the video frame selection features of the initial video editing sequence.
Citation Information
Cited By
Intelligent automatic control software system for post-processing of digital video and implementation method
CN121053593A