AI-based short video analysis and processing method, system, and storage medium
By analyzing the complexity entropy value and motion vector change rate of short videos, combining HSV histogram and SIFT feature analysis, dynamic segmentation and key frame screening, the problems of unreasonable short video segmentation and key frame extraction are solved, and the processing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510283465.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-03-11
AI Technical Summary
Existing technologies cannot dynamically adapt to content complexity in short video segmentation processing, resulting in resource waste and inaccurate key frame extraction, affecting the efficiency of anomaly review.
By analyzing the complexity entropy value and motion vector change rate of short videos, an adaptive segmentation operation is performed, and the candidate key frames are identified by combining HSV histogram and SIFT feature analysis. The preset key frame selection redundancy prevention mechanism is used to screen the key frames.
It achieves dynamic adjustment of the fragmentation strategy according to the video content, improves the adaptability and efficiency of fragmentation processing, ensures the conciseness and efficiency of key frame sets, and improves the accuracy of subsequent analysis.
Smart Images

Figure CN120220021B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of short video abnormality review, and relates to a short video analysis and processing method, system and storage medium based on AI intelligence. Background Art
[0002] With the rapid development of short videos, their volume and scale are exploding. Traditional processing methods struggle to cope with this massive volume of content, hindering efficient analysis and precise management. Slicing and keyframe extraction of short videos are crucial for improving the efficiency of anomaly review and ensuring efficient use of review resources. Efficient and reasonable slicing and keyframe extraction are crucial prerequisites for short video anomaly review. Therefore, research on AI-based short video analysis and processing is of great significance.
[0003] Existing technical solutions for short video segmentation usually adopt a method of segmenting long videos based on fixed preset durations. This processing method cannot dynamically adapt to the complexity of the content, ignores the information density and motion intensity of short videos, and easily causes waste of resources.
[0004] Existing technical solutions for key frame extraction of short videos usually use a fixed number of equally spaced frames. This processing method cannot adapt to changes in video content, resulting in a large number of similar frames being repeatedly extracted, causing excessive redundancy of image information. It is impossible to dynamically adjust the extraction strategy according to the complexity of the video. Key frames may fall outside the interval, resulting in the loss of important content, affecting the accuracy and efficiency of subsequent analysis. Summary of the Invention
[0005] In view of this, in order to solve the problems of unreasonable fixed segmentation process and low availability of key frame extraction proposed in the above background technology, a short video analysis and processing method, system and storage medium based on AI intelligence are proposed.
[0006] The purpose of the present invention can be achieved through the following technical solutions: The first aspect of the present invention provides a short video analysis and processing method, system and storage medium based on AI intelligence, including: S1, performing data analysis on the target short video frame image to obtain the complexity entropy value and motion vector change rate of the target short video.
[0007] S2. According to the complexity entropy value and the motion vector change rate of the target short video, a dynamic segmentation operation is performed based on a preset adaptive segmentation mechanism trigger logic to obtain each video segment.
[0008] S3. Based on HSV histogram and SIFT feature analysis methods, the HSV histogram difference and SIFT feature matching degree of each video clip corresponding to each frame image are respectively obtained, and then each candidate key frame corresponding to each video clip is identified.
[0009] S4. Selecting a redundancy prevention mechanism based on a preset key frame to screen candidate key frames to obtain key frames.
[0010] The second aspect of the present invention provides an AI-based short video analysis and processing system, including: a video data analysis module for performing data analysis on a target short video frame image to obtain the complexity entropy value and motion vector change rate of the target short video.
[0011] The adaptive segmentation module is used to trigger logic to perform dynamic segmentation operations based on a preset adaptive segmentation mechanism according to the complexity entropy value and motion vector change rate of the target short video to obtain various video segments.
[0012] The candidate key frame identification module is used to obtain the HSV histogram difference and SIFT feature matching degree of each video clip corresponding to each frame image based on HSV histogram and SIFT feature analysis, and then identify each candidate key frame corresponding to each video clip.
[0013] The key frame extraction module is used to select the candidate key frames based on the preset key frame selection redundancy prevention mechanism to obtain the key frames.
[0014] The third aspect of the present invention provides a storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the above-mentioned AI intelligence-based short video analysis and processing method.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention analyzes the complexity entropy value and motion vector change rate of the target short video, and then triggers the logic to perform dynamic segmentation operations based on a pre-set adaptive segmentation mechanism. It can dynamically adjust the segmentation strategy according to the real-time changes in the video content, thereby improving the versatility and adaptability of segmentation processing.
[0016] (2) The present invention obtains the HSV histogram difference and SIFT feature matching degree of each frame image corresponding to each video clip to identify candidate key frames, and selects each candidate key frame based on the preset key frame selection redundancy prevention mechanism to obtain each key frame, ensuring that the key frame set is streamlined and efficient, improving the efficiency and quality of each link in video processing, and ensuring the accuracy of subsequent data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 Schematic diagram of the implementation of the method steps of the present invention.
[0019] Figure 2 This is a schematic diagram of the connection of various modules of the system of the present invention.
[0020] Figure 3 A schematic diagram of the storage medium structure provided by the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0022] See also Figure 1 As shown, the first aspect of the present invention provides a short video analysis and processing method, system and storage medium based on AI intelligence, including: S1, performing data analysis on the target short video frame image to obtain the complexity entropy value and motion vector change rate of the target short video.
[0023] In a preferred embodiment of the present invention, the specific analysis method of the complexity entropy value and the motion vector change rate is as follows: grayscale processing is performed on each frame image corresponding to the target short video, the grayscale value corresponding to each pixel point is obtained, the number of pixels corresponding to each grayscale value of each frame image is counted, and the ratio of the number of pixels corresponding to each grayscale value to the total number of pixels is calculated to obtain the frequency of each grayscale value corresponding to each frame image.
[0024] Using the formula Analyze and obtain the complexity entropy value of the target short video ,in Indicates the frequency of each grayscale value corresponding to each frame image, Indicates the number corresponding to the frame image, , Indicates the number of frame images. Indicates the number corresponding to the grayscale value, , represents the number of grayscale values, .
[0025] The motion vector variation of each pixel point corresponding to each frame image is obtained based on the optical flow method, where the motion vector variation corresponding to the first frame image is 0.
[0026] It should be explained that the above complexity entropy value is used to calculate the complexity entropy value of the target short video. Its construction logic is based on the concept of information entropy. The complexity of the video content is measured by statistical analysis of the grayscale value frequency of each frame image of the video. The details are as follows: 1. Grayscale value frequency statistics: grayscale processing is performed on each frame image corresponding to the target short video, and the number of pixels corresponding to each grayscale value of each frame image is counted. Then, the ratio of the number of pixels corresponding to each grayscale value to the total number of pixels is calculated to obtain the frequency of each grayscale value corresponding to each frame image. This step provides basic data for subsequent calculations and reflects the distribution of different grayscale values in the image.
[0027] 2. Double summation to calculate information volume: Double summation is used in the formula ,in Iterate over all images, Iterate over all grayscale values. It is used to calculate the amount of information of each grayscale value in the corresponding image. By accumulating the amount of information of all grayscale values of all images, the total amount of information about the grayscale value distribution of the entire short video can be obtained.
[0028] 3. Averaging: Based on the double summation, there is a coefficient in front of the formula This is to average the total amount of information and obtain the average amount of information per frame, so as to more reasonably measure the complexity of the content of each frame in the short video, and finally obtain the complexity entropy value of the target short video. .
[0029] Using the formula Analyze and obtain the motion vector change rate of the target short video ,in Indicates the motion vector of each pixel corresponding to each frame image, Indicates the pixel number. , Indicates the number of pixels.
[0030] It should be explained that the motion vector change rate analysis formula is used to calculate the motion vector change rate of the target short video. , which measures the degree of change of pixel motion in the video as a whole and reflects the dynamic characteristics of the video image. Its construction logic is as follows: 1. Motion vector : Indicates the Frame image The motion vector of each pixel in a video. A motion vector describes the direction and magnitude of a pixel's movement between frames and is essential data for analyzing dynamic changes in a video. For example, in a video clip of a person walking, each pixel on that person has a corresponding motion vector between frames.
[0031] 2. Double Summation :The outer layer here Indicates the target short video Frame image is traversed, the inner Indicates the Pixel points are traversed. Is to take the motion vector The modulus is the magnitude of the pixel's motion vector (regardless of direction). Double summation is the sum of the motion vectors of all pixels in all frames, representing the total motion change of the pixels in the video.
[0032] 3. Coefficient : The first part of the formula It is to average the double summation results. is the total number of frames in the video, Is the number of pixels in each frame, and the two are multiplied to get the total number of all pixels in the video. Divided by It can eliminate the influence of the number of video frames and the number of pixels on the total motion change, and obtain the average motion vector change degree of each pixel, that is, the motion vector change rate. The greater the rate of change of the motion vector, the more intense the overall movement of the pixels in the video, and the more obvious the dynamic effect of the picture; conversely, the smaller the rate of change, the relatively smoother the video picture.
[0033] S2. According to the complexity entropy value and the motion vector change rate of the target short video, a dynamic segmentation operation is performed based on a preset adaptive segmentation mechanism trigger logic to obtain each video segment.
[0034] In a preferred embodiment of the present invention, the specific content of the adaptive segmentation mechanism triggering logic is as follows: extract the complexity entropy value and motion vector change rate of the target short video, and then compare them with the pre-set complexity entropy value threshold and motion vector change rate threshold respectively. If the complexity entropy value of the target short video is greater than the complexity entropy value threshold and the motion vector change rate is greater than the motion vector change rate threshold, the adaptive segmentation mechanism is triggered.
[0035] It should be explained that the above content describes the conditional judgment logic that triggers the adaptive segmentation mechanism in short video processing. The specific explanation is as follows: first, the complexity entropy value (measures the complexity of the video content; the larger the value, the more uniform the grayscale distribution of the picture and the more complex the content) and the motion vector change rate (reflects the intensity of the object movement in the video; the larger the value, the more intense the object movement) of the target short video are calculated through a specific formula.
[0036] The resulting complexity entropy value and motion vector change rate are compared with pre-set complexity entropy thresholds and motion vector change rate thresholds, respectively. The thresholds are reference standards set based on experience or specific needs. The adaptive segmentation mechanism is triggered only when the complexity entropy value of the target short video exceeds the preset complexity entropy threshold and the motion vector change rate also exceeds the corresponding threshold. This means that adaptive segmentation is only activated when the video content is complex and the objects are in intense motion, allowing for more reasonable video processing, such as adjusting encoding strategies and optimizing storage.
[0037] In a preferred embodiment of the present invention, the specific steps of performing the dynamic sharding operation are as follows: using the formula Analyze and obtain the segment length of the target short video ,in Indicates the pre-set basic shard duration. They represent the preset complexity entropy value weight coefficient and motion vector change weight coefficient, respectively, which are used to suppress the fragmentation sensitivity of high-complexity scenes and intense motion scenes.
[0038] It should be noted that the above formula is used to calculate the segment duration of the target short video, taking into account the complexity of the video content and the intensity of the movement. Its construction logic and the meaning of each parameter are as follows: 1. Basic segment duration : It is a fixed value set in advance, which provides a basic reference for the segmentation duration. It can be understood as the expected approximate duration of the video segment under normal circumstances, which is the starting basis for calculating the final segmentation duration. For example, in the early planning of video processing, according to experience or demand, A specific number of seconds.
[0039] 2. Weight coefficient : is the complexity entropy weight coefficient, are motion vector change weight coefficients. Their function is to suppress the segmentation sensitivity of high-complexity scenes and intense motion scenes. When the video is in a high-complexity scene (complexity entropy value Large) or intense motion scenes (motion vector change rate When the number of segments is large, these two weight coefficients are used to adjust their impact on the segmentation duration to avoid excessive segmentation or unreasonable segmentation duration. For example, in a special effects video with complex scenes and fast-moving objects, appropriate weight coefficients can make the segmentation duration more in line with actual processing requirements.
[0040] 3. Complexity entropy and motion vector change rate : Measures the complexity of the video content. A larger value indicates a more uniform distribution of grayscale values and more complex content. Reflects the intensity of the object's motion in the video; larger values indicate more intense motion. These two values are key factors influencing the segment duration. Multiplied by the weight coefficient, they contribute to the final segment duration.
[0041] 4. Calculation logic: In the formula, the denominator Will follow and The whole fraction is Divide by this denominator, so when the complexity of the video content or the intensity of the movement increases, the denominator becomes larger and the segment length will be reduced accordingly, which means that the video will be divided into shorter segments for more detailed processing; on the contrary, when the video content is simple and the motion is smooth, the segmentation time will be shortened. Will increase.
[0042] In one possible embodiment, assuming , , based on the segmentation duration analysis formula of the target short video, data simulation calculation is performed to obtain the corresponding simulation calculation results. Some simulation results can be referred to Table 1.
[0043] Table 1. Simulation data and simulation calculation results corresponding to some shard durations
[0044]
[0045] It should be noted that the present invention analyzes the complexity entropy value and motion vector change rate of the target short video, and then triggers the logic to perform dynamic segmentation operations based on a pre-set adaptive segmentation mechanism. It can dynamically adjust the segmentation strategy according to the real-time changes of the video content, thereby improving the versatility and adaptability of segmentation processing.
[0046] S3. Based on HSV histogram and SIFT feature analysis methods, the HSV histogram difference and SIFT feature matching degree of each video clip corresponding to each frame image are respectively obtained, and then each candidate key frame corresponding to each video clip is identified.
[0047] In a preferred embodiment of the present invention, the specific analysis process of the HSV histogram difference is as follows: convert each frame image corresponding to each video clip into HSV space, and obtain the number of pixels corresponding to each bin of each frame image and the total number of pixels, which are recorded as 、 ,in Indicates the box number, , Indicates the number of bins, .
[0048] Using the formula Analyze and obtain the HSV histogram difference corresponding to each frame image ,in Indicates the Frame image The number of pixels in the bin, Indicates the The total number of pixels in the frame image, Indicates taking the minimum value.
[0049] It should be explained that the above formula is used to calculate the HSV histogram difference corresponding to each frame image , which measures the difference between two adjacent image frames in the HSV color space. The details are as follows: 1. HSV Color Space Conversion: Before calculating the HSV histogram difference, each image frame needs to be converted from its original color space to the HSV (Hue, Saturation, Value) color space. The HSV color space better aligns with human color perception and can better distinguish the characteristics of different colors, facilitating subsequent analysis of color differences.
[0050] 2. Binning statistics: Divide the converted image into multiple bins (the number of bins here is ,and , and count the number of pixels in each bin in each frame image And the total number of pixels in the frame image Binning statistics can quantify the color information of an image. Each bin represents a color feature within a certain range. By counting the number of pixels in each bin, the color distribution of the image can be obtained.
[0051] 3. Calculate the similarity of each bin: In the formula This part is to calculate the two adjacent frames of images (the Frame and For each bin , Indicates the Frame image The ratio of binned pixels to total pixels, Indicates the Frame image The ratio of binned pixels to total pixels. By taking the minimum of these two ratios Two frames of images can be obtained in the The similarity of the color distribution of the bins. By adding up the similarities of all bins, we can get the similarity of the overall color distribution of the two frames.
[0052] 4. Calculate the HSV histogram difference: Subtract the sum of the similarities calculated above from 1 to get the result. HSV histogram difference of frame image The larger the difference value, the greater the difference in color distribution between the two frames; the smaller the difference value, the more similar the color distribution of the two frames. For example, in a video, if the scene between the two frames undergoes a significant color change, such as switching from an outdoor scene in the daytime to an indoor scene at night, then the difference in their HSV histograms will be relatively large; on the other hand, if the two frames only have a slight movement in the picture and the color distribution does not change significantly, then the difference will be relatively small.
[0053] It should be noted that when When , the corresponding histogram difference The other parameters in the present invention are assigned values according to this method during the comparative analysis process. The selection of a single frame image does not affect the key frame recognition process and does not affect other behaviors such as subsequent review.
[0054] The key points of each frame image are obtained by scale-invariant feature transformation, and the descriptors of each key point are generated. The nearest neighbor search algorithm is used to obtain the nearest neighbor key points and the second nearest neighbor key points of each key point descriptor of each adjacent frame. The distance between each key point descriptor and the nearest neighbor key point descriptor is recorded as the nearest neighbor distance of each key point, and the distance between each key point descriptor and the second nearest neighbor key point descriptor is recorded as the second nearest neighbor distance of each key point. Then, the ratio of the nearest neighbor distance to the second nearest neighbor distance of each key point is calculated to obtain the distance ratio of each key point pair.
[0055] The distance ratio of each key point pair is compared with the preset distance ratio threshold. If the distance ratio of a key point pair is less than the distance ratio threshold, the key point pair is identified as a valid matching point pair, and the number of valid matching point pairs in each frame image is obtained by counting.
[0056] The SIFT feature matching degree of each frame image is obtained by calculating the ratio of the number of valid matching point pairs and the total number of key points in each frame image.
[0057] In a preferred embodiment of the present invention, the specific process of identifying each candidate key frame corresponding to each video clip is as follows: using the formula Analyze and obtain the key frame score of each frame image ,in Indicates the SIFT feature matching degree of frame image, Represent the pre-set influence weight factors of color change and structure change respectively.
[0058] It should be explained that the above formula is used to calculate the key frame score of each frame image , taking into account the color change and structural change of the image, to determine the possibility of the frame image becoming a key frame, the specific explanation is as follows: 1, HSV histogram difference With color changes: It is obtained by calculating the difference between two adjacent frames of images in the HSV color space. The HSV color space can better reflect people's perception of color. The larger the The greater the difference in color distribution between a frame and the next frame, the more obvious the color change. For example, scene switching and lighting changes in a video can cause Increase. In the formula, This part reflects the contribution of color change to the keyframe score. is a preset color change influence weight factor, which determines the importance of color change in keyframe score calculation. The larger the value, the greater the impact of color change on the key frame score, and the more emphasis is placed on frames with obvious color changes as key frames.
[0059] 2. SIFT feature matching With structural changes: Indicates the SIFT feature matching degree of the frame image. SIFT features are used to describe the local structural features of the image. The higher the The higher the similarity between the frame image and the next frame in terms of structural features, the smaller the structural changes. For example, when the object in the video moves slowly or has no obvious movement, the SIFT feature matching degree will be higher. It indicates the degree of structural change. The larger the value, the greater the structural change. It reflects the contribution of structural changes to the key frame score. is a pre-set structural change impact weight factor, which controls the importance of structural changes in keyframe score calculation. A larger value means that frames with obvious structural changes are given more importance as key frames.
[0060] 3. Keyframe score calculation : Add the effects of color change and structure change to get the key frame score of each frame image Through this comprehensive score, each frame of the video can be evaluated. In practical applications, a key frame score threshold is pre-set. When the value of the image is greater than the threshold, the image is considered to be a key frame. This allows for a more comprehensive screening of representative key frames containing important information in the video by comprehensively considering both color and structural changes. This can improve efficiency and accuracy in applications such as video summarization and video retrieval.
[0061] In one possible embodiment, assuming ,Based on the key frame score analysis formula, data simulation calculation is performed to obtain the corresponding simulation calculation results.,Part of the simulation results can be referred to Table 2.
[0062] Table 2. Simulation data and simulation calculation results corresponding to some shard durations
[0063]
[0064] It needs to be explained that Impact: Represents the HSV histogram difference between adjacent frame images, reflecting color changes. As can be seen from the table, when other conditions remain unchanged, The larger the keyframe score The higher. For example, comparing Group 1 and 3, From 0.2 to 0.8, It increases from 0.2 to 0.72. This shows that the more obvious the color change is, the greater the improvement effect on the key frame score is, and frames with significant color changes are more likely to become key frames.
[0065] Impact: It is the SIFT feature matching degree, which reflects the degree of structural similarity. It indicates the degree of structural change. Reduce (i.e. Increase), keyframe score For example, groups 2 and 4, From 0.6 to 0.7, From 0.4 to 0.3, It increased from 0.46 to 0.72. This means that the greater the structural change, the greater the contribution to the key frame score, and frames with obvious structural changes are more important in key frame judgment.
[0066] If the complexity entropy value of the target short video is greater than the preset complexity entropy value threshold and the motion vector change rate is less than the motion vector change rate threshold, let .
[0067] If the complexity entropy value of the target short video is less than the preset complexity entropy value threshold and the motion vector change rate is greater than the motion vector change rate threshold, let .
[0068] In other cases ,in They represent the pre-set weight factors of different levels of color change. They respectively represent the impact weight factors of different pre-set level structure changes.
[0069] It should be noted that during the short video keyframe extraction process, this section dynamically adjusts the weights of color and structural changes in the keyframe score calculation based on the video's complexity entropy and motion vector change rate, thereby more accurately selecting keyframes. The details are as follows: 1. The significance of weight adjustment in different scenarios: Short videos have widely varying content characteristics. Some have complex images but gentle motion, others are simple but with intense motion, and there are many other scenarios. By determining the relationship between the complexity entropy and motion vector change rate and their respective thresholds and dynamically adjusting the weights, we can better adapt to different video scenarios and ensure that the keyframe selection is more consistent with the actual content of the video.
[0070] 2. When the complexity entropy value is large and the motion vector change rate is small: When the complexity entropy value of the target short video is greater than the pre-set complexity entropy value threshold and the motion vector change rate is less than the motion vector change rate threshold, it means that the video content is complex, but the object movement is relatively smooth. In this case, the changes in the color and structural details of the picture may be the key information. , which means increasing the weight of color change (Depend on Determine), so that in the key frame score calculation, the impact of color changes on key frame selection is more prominent, making those frames with rich color changes and that can reflect complex picture details more likely to be selected as key frames.
[0071] 3. When the complexity entropy value is small and the motion vector change rate is large: If the complexity entropy value is less than the pre-set complexity entropy value threshold and the motion vector change rate is greater than the motion vector change rate threshold, it indicates that the video content is relatively simple, but the object moves violently. In this case, the structural change caused by motion may be a more important feature. So let , that is, to adjust the weight appropriately so that the weight of the structure changes (Depend on The keyframe score calculation is more prominent, which makes it easier to identify frames with obvious structural changes due to motion as keyframes.
[0072] 4. Other situations: In addition to the two extreme situations mentioned above, This set of weights is designed for scenes with relatively balanced video complexity and motion. It comprehensively considers the impact of color and structural changes on keyframe selection, ensuring that keyframe scores are reasonably calculated in a variety of common video scenarios and that representative keyframes are selected.
[0073] 5. The role of weight factor: and These pre-set weighting factors for different levels of color and structural changes were determined based on extensive experiments and real-world application experience. They provide customized keyframe screening strategies for different video scenarios, ensuring that the keyframe extraction algorithm can adapt to diverse short video content, improving the accuracy and effectiveness of keyframe extraction, and ultimately enhancing the performance of related applications such as video analysis, retrieval, and summarization.
[0074] It should be noted that , .
[0075] The key frame score of each frame image is compared with a preset key frame score threshold. If the key frame score of a frame image is greater than the key frame score threshold, the frame image is identified as a candidate key frame.
[0076] For example, the key frame score threshold is .
[0077] S4. Selecting a redundancy prevention mechanism based on a preset key frame to screen candidate key frames to obtain key frames.
[0078] In a preferred embodiment of the present invention, the specific analysis process of the key frame selection redundancy prevention mechanism is as follows: extract each candidate key frame, obtain the number of each candidate key frame, and then arrange them in ascending order to obtain a candidate key frame set, and calculate the difference between the numbers of adjacent candidate key frames to obtain the monitoring distance of each adjacent candidate key frame.
[0079] The monitoring interval of each adjacent candidate key frame is compared with the preset minimum key frame monitoring interval threshold. If the monitoring interval of an adjacent candidate key frame is less than the preset minimum key frame monitoring interval threshold, it is identified that the adjacent candidate key frame is redundant, and the subsequent candidate key frame corresponding to the adjacent candidate key frame is eliminated.
[0080] It should be noted that in the process of extracting key frames from short videos, the above steps are an important part of the redundancy prevention mechanism for key frame selection, which aims to remove redundant frames that may exist in the candidate key frames, so that the key frames finally determined are more representative and efficient. The specific explanations are as follows: 1. The meaning of the monitoring interval between adjacent candidate key frames: Candidate key frames are image frames that have been preliminarily screened and are considered to be key frames. These candidate key frames are arranged in the order in which they appear in the video, and the interval between adjacent candidate key frames on the video timeline is the monitoring interval. For example, if the video has a total of 100 frames, and the candidate key frames are the 10th frame, the 20th frame, the 30th frame, etc., then the monitoring interval between the 10th frame and the 20th frame is 10 frames.
[0081] 2. The Minimum Keyframe Monitoring Interval Threshold: The Minimum Keyframe Monitoring Interval Threshold is a pre-set reference value used to determine whether there is redundancy between adjacent candidate keyframes. It is determined based on video processing experience or specific requirements, and is used to measure whether the distance between adjacent candidate keyframes is reasonable. If the monitoring interval between adjacent candidate keyframes is less than this threshold, it means that the two adjacent candidate keyframes are too close in time, and the information they contain may be highly similar, resulting in redundancy.
[0082] 3. Redundant Frame Removal: When the monitoring distance between adjacent candidate keyframes is less than the minimum keyframe monitoring distance threshold, these adjacent candidate keyframes are identified as redundant. To ensure the quality and representativeness of the keyframes, one of them needs to be removed. In practice, the last candidate keyframe is typically removed. This is because, in the chronological order of a video, the information of the later frame may already be largely reflected in the previous frame. Removing the later frame avoids duplication of keyframe information, making the keyframe set more streamlined. For example, in a video of a person walking slowly, the slow movement of the person may result in several consecutive frames being initially selected as candidate keyframes, but the differences between these frames are relatively small. By comparing the monitoring distance between adjacent candidate keyframes with the minimum keyframe monitoring distance threshold, if the monitoring distance between them is less than the threshold, removing the later candidate keyframe removes this redundant information, ensuring that the final keyframe more accurately reflects the key content of the video.
[0083] 4. Final Results: This method effectively removes redundant frames from candidate keyframes, optimizing keyframe selection. The resulting filtered keyframe set retains the video's key information while avoiding excessive duplication, improving keyframe quality and utilization efficiency. This is crucial for subsequent video processing tasks such as video summarization, video retrieval, and video compression, reducing processing time and storage space while improving performance.
[0084] In a preferred embodiment of the present invention, the specific steps of screening the candidate key frames to obtain the key frames are as follows: performing a candidate key frame elimination operation based on a key frame selection redundancy prevention mechanism, and then recording the candidate key frames retained after elimination as key frames.
[0085] It should be noted that the present invention obtains the HSV histogram difference and SIFT feature matching degree of each frame image corresponding to each video clip and then identifies the candidate key frames, and selects the redundant prevention mechanism based on the preset key frames to screen the candidate key frames to obtain each key frame, thereby ensuring that the key frame set is streamlined and efficient, improving the efficiency and quality of each link in video processing, and ensuring the accuracy of subsequent data analysis.
[0086] See also Figure 2 As shown, the second aspect of the present invention provides a short video analysis and processing system based on AI intelligence, including a video data analysis module, an adaptive segmentation module, a candidate key frame recognition module and a key frame extraction module, wherein the video data analysis module is connected to the adaptive segmentation module, the adaptive segmentation module is connected to the candidate key frame recognition module, and the candidate key frame recognition module is connected to the key frame extraction module.
[0087] The video data analysis module is used to perform data analysis on the target short video frame image to obtain the complexity entropy value and motion vector change rate of the target short video.
[0088] The adaptive segmentation module is used to trigger logic based on a preset adaptive segmentation mechanism to perform dynamic segmentation operations according to the complexity entropy value and motion vector change rate of the target short video to obtain various video segments.
[0089] The candidate key frame identification module is used to obtain the HSV histogram difference and SIFT feature matching degree of each video clip corresponding to each frame image based on HSV histogram and SIFT feature analysis, and then identify each candidate key frame corresponding to each video clip.
[0090] The key frame extraction module is used to screen candidate key frames based on a preset key frame selection redundancy prevention mechanism to obtain key frames.
[0091] See also Figure 3 As shown, the third aspect of the present invention provides a storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the AI intelligence-based short video analysis and processing method as described in any of the above embodiments.
[0092] The storage medium described here includes RAM (Random Access Memory), internal memory, ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), registers, hard disks, removable disks, or any other form of storage medium known in the technical field.
[0093] The above contents are merely examples and explanations of the concept of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, they should all fall within the scope of protection of the present invention.
Claims
1. The short video analysis and processing method based on AI intelligence is characterized by: include: S1. Analyze the frame image data of the target short video to obtain the complexity entropy value and motion vector change rate of the target short video; S2. According to the complexity entropy value and motion vector change rate of the target short video, a dynamic segmentation operation is performed based on a preset adaptive segmentation mechanism trigger logic to obtain each video segment; S3, based on HSV histogram and SIFT feature analysis methods, respectively obtaining the HSV histogram difference and SIFT feature matching degree of each video clip corresponding to each frame image, and then identifying each candidate key frame corresponding to each video clip; S4, screening each candidate key frame based on a preset key frame selection redundancy prevention mechanism to obtain each key frame; The specific analysis method of the complexity entropy value and the motion vector change rate is as follows: Grayscale each frame of the target short video to obtain the grayscale value corresponding to each pixel. Count the number of pixels corresponding to each grayscale value in each frame. Ratio the number of pixels corresponding to each grayscale value to the total number of pixels to obtain the frequency of each grayscale value in each frame. Using the formula Analyze and obtain the complexity entropy value of the target short video ,in Indicates the frequency of each grayscale value corresponding to each frame image, Indicates the number corresponding to the frame image, , Indicates the number of frame images. Indicates the number corresponding to the grayscale value, , represents the number of grayscale values, ; The motion vector change of each pixel point corresponding to each frame image is obtained based on the optical flow method, where the motion vector change corresponding to the first frame image is 0; Using the formula Analyze the motion vector change rate of the target short video ,in Indicates the motion vector of each pixel corresponding to each frame image, Indicates the pixel number. , Indicates the number of pixels; The specific content of the adaptive sharding mechanism triggering logic is as follows: The complexity entropy value and motion vector change rate of the target short video are extracted, and then compared with the pre-set complexity entropy value threshold and motion vector change rate threshold respectively. If the complexity entropy value of the target short video is greater than the complexity entropy value threshold and the motion vector change rate is greater than the motion vector change rate threshold, the adaptive segmentation mechanism is triggered.
2. The AI-based short video analysis and processing method according to claim 1, wherein: The specific steps of performing the dynamic sharding operation are as follows: Using the formula Analyze and obtain the segment length of the target short video ,in Indicates the pre-set basic shard duration. They represent the preset complexity entropy weight coefficient and motion vector change weight coefficient, respectively, which are used to suppress the fragmentation sensitivity of high-complexity scenes and intense motion scenes; The target short video is dynamically segmented based on the segment duration to obtain several video clips.
3. The AI-based short video analysis and processing method according to claim 1, wherein: The specific analysis process of the HSV histogram difference is as follows: Convert each frame image corresponding to each video clip into HSV space, and count the number of pixels corresponding to each bin and the total number of pixels of each frame image, which are recorded as 、 ,in Indicates the box number, , Indicates the number of bins, ; Using the formula Analyze and obtain the HSV histogram difference corresponding to each frame image ,in Indicates the Frame image The number of pixels in the bin, Indicates the The total number of pixels in the frame image, Indicates taking the minimum value; The key points of each frame image are obtained by scale-invariant feature transformation, and the descriptors of each key point are generated. The nearest neighbor search algorithm is used to obtain the nearest neighbor key points and the next nearest neighbor key points of each key point descriptor of each adjacent frame. The distance between each key point descriptor and the nearest neighbor key point descriptor is recorded as the nearest neighbor distance of each key point, and the distance between each key point descriptor and the next nearest neighbor key point descriptor is recorded as the next nearest neighbor distance of each key point. Then, the ratio of the nearest neighbor distance to the next nearest neighbor distance of each key point is calculated to obtain the distance ratio of each key point pair. Compare the distance ratio of each key point pair with the preset distance ratio threshold. If the distance ratio of a key point pair is less than the distance ratio threshold, the key point pair is identified as a valid matching point pair, and the number of valid matching point pairs in each frame image is obtained by counting; The SIFT feature matching degree of each frame image is obtained by calculating the ratio of the number of valid matching point pairs and the total number of key points in each frame image.
4. The AI-based short video analysis and processing method according to claim 3, wherein: The specific process of identifying each candidate key frame corresponding to each video clip is as follows: Using the formula Analyze and obtain the key frame score of each frame image ,in Indicates the SIFT feature matching degree of frame image, Represent the pre-set influence weight factors of color change and structure change respectively; If the complexity entropy value of the target short video is greater than the preset complexity entropy value threshold and the motion vector change rate is less than the motion vector change rate threshold, let ; If the complexity entropy value of the target short video is less than the preset complexity entropy value threshold and the motion vector change rate is greater than the motion vector change rate threshold, let ; In other cases ,in They represent the pre-set weight factors of different levels of color change. They represent the influence weight factors of different pre-set level structure changes respectively; The key frame score of each frame image is compared with a preset key frame score threshold. If the key frame score of a frame image is greater than the key frame score threshold, the frame image is identified as a candidate key frame.
5. The AI-based short video analysis and processing method according to claim 1, wherein: The specific analysis process of the key frame selection redundancy prevention mechanism is as follows: Extract each candidate key frame, obtain the number of each candidate key frame, and then arrange them in ascending order to obtain a candidate key frame set, and calculate the difference between the numbers of adjacent candidate key frames to obtain the monitoring distance of each adjacent candidate key frame; The monitoring interval of each adjacent candidate key frame is compared with the preset minimum key frame monitoring interval threshold. If the monitoring interval of an adjacent candidate key frame is less than the preset minimum key frame monitoring interval threshold, it is identified that the adjacent candidate key frame is redundant, and the subsequent candidate key frame corresponding to the adjacent candidate key frame is eliminated.
6. The AI-based short video analysis and processing method according to claim 1, wherein: The specific steps of screening the candidate key frames to obtain the key frames are as follows: Based on the key frame selection redundancy prevention mechanism, the candidate key frames are eliminated, and the candidate key frames retained after elimination are recorded as key frames.
7. An AI-based short video analysis and processing system, configured to execute the steps of the AI-based short video analysis and processing method according to any one of claims 1 to 6, characterized in that: include: The video data analysis module is used to analyze the frame image data of the target short video to obtain the complexity entropy value and motion vector change rate of the target short video; An adaptive segmentation module is configured to trigger logic based on a preset adaptive segmentation mechanism to perform dynamic segmentation operations to obtain video segments according to the complexity entropy value and motion vector change rate of the target short video; The candidate key frame identification module is used to obtain the HSV histogram difference and SIFT feature matching degree of each video clip corresponding to each frame image based on HSV histogram and SIFT feature analysis, and then identify the candidate key frames corresponding to each video clip; The key frame extraction module is used to select the candidate key frames based on the preset key frame selection redundancy prevention mechanism to obtain the key frames.
8. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the AI intelligence-based short video analysis and processing method as described in any one of claims 1-6.
Citation Information
Patent Citations
Key frame extraction method based on inter-frame difference and color histogram difference
CN112270247A
Real-time image video compression method for dynamic frame screening
CN119182910A