Video frame alignment method, apparatus, device, and storage medium
By utilizing inter-frame feature matching and residual information calculation during the video frame alignment process, the first frame can be accurately located and aligned, solving the problem of low video frame alignment accuracy in existing technologies and achieving higher alignment accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MIGU VIDEO TECH CO LTD
- Filing Date
- 2022-11-03
- Publication Date
- 2026-05-12
AI Technical Summary
Current technologies have low accuracy in video frame alignment, and existing methods can only provide a vague positioning and alignment of the first frame.
By matching the selected video frames in the reference video with each video frame in the target video, and calculating the inter-frame features based on the residual information of the selected video frame, the target video frame and their adjacent video frames, if the inter-frame features are equal, the reference video and the target video are aligned with the selected video frame and the target video frame as the first alignment frame.
This improves the positioning accuracy of the first frame alignment, thereby improving the accuracy of video frame alignment.
Smart Images

Figure CN115941939B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and in particular to a video frame alignment method, apparatus, device, and storage medium. Background Technology
[0002] Video quality detection and evaluation are crucial for ensuring the quality of video network transmission services. Regardless of the technology used, a fundamental prerequisite for video quality detection and evaluation is aligning the video frames in the target video with their corresponding frames in the reference video. Currently, this alignment method involves tagging both the target and reference videos and aligning based on these tags. However, aligning frames solely based on tags cannot achieve accurate frame alignment. Summary of the Invention
[0003] The main objective of this application is to provide a video frame alignment method, apparatus, device, and storage medium, aiming to solve the technical problem of low video frame alignment accuracy in the prior art.
[0004] To achieve the above objectives, this application provides a video frame alignment method, comprising the following steps:
[0005] The selected video frame in the reference video is matched with each video frame in the target video to obtain the target video frame whose frame information matches the selected video frame.
[0006] Based on the residual information of the selected video frame and the target video frame with their respective neighboring video frames, the inter-frame features of the selected video frame and the target video frame with their respective neighboring video frames are obtained.
[0007] If the inter-frame features of the selected video frame and its neighboring video frames are equal to the inter-frame features of the target video frame and its neighboring video frames, the reference video and the target video are aligned using the selected video frame and the target video frame as the first alignment frame.
[0008] Optionally, the step of obtaining the inter-frame features of the selected video frame and the target video frame and their respective neighboring video frames based on the residual information of the selected video frame and the target video frame and their respective neighboring video frames includes:
[0009] Calculate the motion vectors between the selected video frame / target video frame and its adjacent video frames to obtain the residual image;
[0010] Calculate the total number of pixels and the centroid coordinates of the residual image;
[0011] Based on the centroid coordinates of the residual image, the angle between the centroid coordinates of the residual image and the positive X-axis direction is obtained;
[0012] Based on the centroid coordinates of the residual image, the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the current video frame is obtained;
[0013] The total number of pixels in the residual image, the angle between the centroid coordinates of the residual image and the positive X-axis, and the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / target video frame are used as the inter-frame features between the selected video frame / target video frame and its adjacent video frames.
[0014] Optionally, when the residual image is irregular in shape, the step of calculating the centroid coordinates of the residual image includes:
[0015] The residual image is segmented into multiple regular images;
[0016] Obtain the center coordinates and area of each of the aforementioned regular images;
[0017] Based on the X-coordinate of the center coordinate of each regular image and the area of each regular image, the X-coordinate of the centroid of the residual image is obtained.
[0018] Based on the Y-coordinate of the center coordinate of each regular image and the area of each regular image, the Y-coordinate of the centroid of the residual image is obtained.
[0019] Optionally, when the selected video frame is a scene transition frame, the selected video frame is determined in the following way:
[0020] The variation characteristics of the high-frequency subband coefficients of each video frame in the reference video are extracted by three-dimensional wavelet transform.
[0021] The selected video frame is obtained based on the variation characteristics of the high-frequency subband coefficients.
[0022] Optionally, before the step of matching the selected video frame in the reference video with each video frame in the target video to obtain the target video frame whose frame information matches that of the selected video frame, the method further includes:
[0023] Similarity matching is performed on the reference video and the target video to obtain a reference video sequence and a target video sequence, wherein the reference video sequence and the target video sequence contain the same video frames.
[0024] The step of matching selected video frames in the reference video with each video frame in the target video to obtain a target video frame in the target video whose frame information matches that of the selected video frame includes:
[0025] The selected video frame in the reference video sequence is matched with each video frame in the target video sequence to obtain the target video frame in the target video sequence that matches the frame information of the selected video frame.
[0026] Optionally, the step of performing similarity matching on the reference video and the target video to obtain a reference video sequence and a target video sequence includes:
[0027] Calculate the frame information of each video frame in the reference video to obtain a reference video frame information array;
[0028] Calculate the frame information of each video frame in the target video to obtain a target video frame information array;
[0029] Traverse the reference video frame information array and the target video frame information array to obtain a first video frame and a second video frame with a similarity greater than a preset similarity threshold, wherein the first video frame is located in the reference video frame and the second video frame is located in the target video frame;
[0030] The number of frames in the video sequence starting from the first video frame in the reference video and the number of frames in the video sequence starting from the second video frame in the target video are counted.
[0031] The smaller value between the number of frames in the video sequence starting from the first video frame and the number of frames in the video sequence starting from the second video frame is selected as the number of frames to be extracted.
[0032] Based on the number of frames captured, a reference video sequence is obtained from the reference video, starting from the first video frame;
[0033] Based on the number of frames captured, starting from the second video frame, a target video sequence is obtained from the target video.
[0034] Optionally, after the step of traversing the reference video frame information array and the target video frame information array to obtain a first video frame and a second video frame with a similarity greater than a preset similarity threshold, the method further includes:
[0035] The similarity of the adjacent video frames of the first video frame and the adjacent video frames of the second video frame is compared.
[0036] If, among adjacent video frames undergoing similarity comparison, the proportion of video frames with a similarity greater than a preset similarity threshold is greater than a preset proportion threshold, then the steps of counting the number of frames in the video sequence starting from the first video frame in the reference video and the number of frames in the video sequence starting from the second video frame in the target video are performed.
[0037] In addition, to achieve the above objectives, this application also provides a video frame alignment apparatus, comprising:
[0038] The first matching module is used to perform frame information matching between selected video frames in the reference video and each video frame in the target video to obtain the target video frame that matches the frame information of the selected video frame in the target video.
[0039] The inter-frame feature acquisition module is used to obtain the inter-frame features of the selected video frame and the target video frame and their respective neighboring video frames based on the residual information of the selected video frame and the target video frame and their respective neighboring video frames.
[0040] An alignment module is used to align the reference video and the target video, with the selected video frame and the target video frame as the first alignment frame, if the inter-frame characteristics of the selected video frame and its adjacent video frames are equal to the inter-frame characteristics of the target video frame and its adjacent video frames.
[0041] In addition, to achieve the above objectives, this application also provides a video frame alignment device, the device comprising: a memory, a processor, and a video frame alignment program stored in the memory and executable on the processor, the video frame alignment program being configured to implement the steps of the video frame alignment method described above.
[0042] In addition, to achieve the above objectives, this application also provides a storage medium storing a video frame alignment program, which, when executed by a processor, implements the steps of the video frame alignment method described above.
[0043] This application provides a video frame alignment method, apparatus, device, and storage medium. Compared with existing technologies that align video frames based on tags, this application first matches the frame information of selected video frames in a reference video with each video frame in a target video to obtain a target video frame in the target video whose frame information matches that of the selected video frame. Then, based on the residual information of the selected video frame, the target video frame, and their respective neighboring video frames, the inter-frame features of the selected video frame, the target video frame, and their respective neighboring video frames are obtained. If the inter-frame features of the selected video frame and its neighboring video frames are equal to those of the target video frame and its neighboring video frames, the selected video frame and the target video frame are used as the first alignment frame to align the reference video and the target video. Therefore, this application utilizes the inter-frame features corresponding to video frames to locate the first alignment frame, improving the accuracy of the first alignment frame location and thus improving the accuracy of video frame alignment. Therefore, it overcomes the technical defect of existing technologies that align video frames based on tags, which can only vaguely locate the first alignment frame, and thus solves the technical problem of low accuracy in video frame alignment. Attached Figure Description
[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a schematic diagram of the structure of a video frame alignment device in the hardware operating environment involved in the embodiments of this application;
[0047] Figure 2 This is a flowchart illustrating the first embodiment of the video frame alignment method of this application;
[0048] Figure 3 This is a flowchart illustrating the third embodiment of the video frame alignment method of this application;
[0049] Figure 4 This is a schematic diagram of the functional modules of the first embodiment of the video frame alignment device of this application.
[0050] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0051] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0052] Reference Figure 1 , Figure 1 This is a schematic diagram of the video frame alignment device structure in the hardware operating environment involved in the embodiments of this application.
[0053] like Figure 1 As shown, the video frame alignment device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0054] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the video frame alignment device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0055] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a video frame alignment program.
[0056] exist Figure 1 In the video frame alignment device shown, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the video frame alignment device of this application can be set in the video frame alignment device, and the video frame alignment device calls the video frame alignment program stored in the memory 1005 through the processor 1001 and executes the video frame alignment method provided in the embodiment of this application.
[0057] This application provides a video frame alignment method, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the video frame alignment method of this application.
[0058] In this embodiment, the video frame alignment method includes the following steps:
[0059] Step S10: Match the selected video frames in the reference video with each video frame in the target video to obtain the target video frame in the target video that matches the frame information of the selected video frame.
[0060] In this embodiment, the frame information of the selected video frame in the reference video is a hash value, and the hash value of the selected video frame in the reference video is determined in the following way:
[0061] Step A1: Scale the selected video frame into a grayscale image with a preset resolution;
[0062] It should be noted that the preset resolution can be the resolution of one of the video frames in the reference video, or the resolution of one of the video frames in the target video, or it can be set by those skilled in the art according to the actual application requirements. For example, the preset resolution can be an 8×8 resolution.
[0063] Step A2: Traverse each pixel in the grayscale image and compare the brightness value of each pixel with its next adjacent pixel to obtain the difference value;
[0064] It should be noted that if the brightness value of each pixel is greater than or equal to the brightness value of its next adjacent pixel, the difference value is set to 1; if the brightness value of each pixel is less than the brightness value of its next adjacent pixel, the difference value is set to 0.
[0065] Step A3: Combine all the difference values to obtain a difference value array;
[0066] It should be noted that combining all the difference values to obtain the difference value array is done by combining the difference values obtained in step A2 in the order in which they were obtained.
[0067] Step A4: Extract a preset number of values from the difference value array in sequence, obtain hexadecimal values based on the preset number of values, and concatenate the hexadecimal values to obtain a hexadecimal value array. Convert the hexadecimal value array into a string to obtain the hash value of the selected video frame.
[0068] It should be noted that, in terms of image depth, images can be divided into 8-bit, 16-bit, 24-bit, and 32-bit, etc. Therefore, the preset number in this embodiment can be set according to the depth of the video frames in the reference video. For example, if the depth of the video frames in the reference video is 8 bits, then the preset number is 8.
[0069] Since the difference values in the difference value array are set to 0 or 1, the number of values extracted from the difference value array is a binary value, which needs to be converted into a hexadecimal value.
[0070] It should be noted that, in this embodiment, the frame information of each video frame in the target video is a hash value. The method for determining the hash value of each video frame in the target video is the same as the method for determining the hash value of the selected video frame in the reference video mentioned above, and will not be repeated here.
[0071] It should be noted that, in this embodiment, the frame information of the selected video frame is a hash value, and the frame information of each video frame in the target video is also a hash value. Therefore, the step of matching the frame information of the selected video frame in the reference video with each video frame in the target video to obtain the target video frame in the target video that matches the frame information of the selected video frame includes:
[0072] Step B1: Traverse each video frame in the target video;
[0073] Step B2: If the number of data bits that are different between the hash value of the video frame in the target video and the hash value of the selected video frame is less than or equal to a preset bit threshold, then the video frame in the target video is confirmed as the target video frame that matches the frame information of the selected video frame.
[0074] Step S20: Based on the residual information of the selected video frame and the target video frame and their respective neighboring video frames, obtain the inter-frame features of the selected video frame and the target video frame and their respective neighboring video frames;
[0075] It should be noted that, in this embodiment, the inter-frame features include the total number of pixels in the residual image, the angle between the centroid coordinates of the residual image and the positive X-axis, and the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the current video frame.
[0076] The step of obtaining the inter-frame features of the selected video frame, the target video frame, and their respective neighboring video frames based on the residual information between the selected video frame and the target video frame and their respective neighboring video frames includes:
[0077] Step S21: Calculate the motion vector between the selected video frame / target video frame and its adjacent video frames to obtain the residual image;
[0078] It should be noted that, in this embodiment, the motion vector between the selected video frame / target video frame and its adjacent video frames can be calculated using the optical flow method. The optical flow method is an existing technology and will not be described in detail here.
[0079] It should be noted that, in this embodiment, the step of calculating the motion vector between the selected video frame / target video frame and its adjacent video frames to obtain the residual image includes:
[0080] Step S211: Calculate the motion vector between the selected video frame / target video frame and its adjacent video frames;
[0081] Step S212: Based on the motion vector, calculate the difference between the pixel points in the selected video frame / target video frame and the pixel points in its adjacent video frames;
[0082] Step S213: Construct a residual image based on the difference between the pixels in the selected video frame / target video frame and the pixels in its adjacent video frames.
[0083] Step S22: Calculate the total number of pixels and the centroid coordinates of the residual image;
[0084] It should be noted that, in this embodiment, the step of calculating the total number of pixels in the residual image includes:
[0085] The residual image is binarized to obtain the corresponding mask;
[0086] The area value of the mask is calculated using the Area property of the Regionprops(get the properties of region) function as the total number of pixels in the residual image.
[0087] It should be noted that, in this embodiment, when the residual image is irregular in shape, the step of calculating the centroid coordinates of the residual image includes:
[0088] The residual image is segmented into multiple regular images;
[0089] Obtain the center coordinates and area of each of the aforementioned regular images;
[0090] Based on the X-coordinate of the center coordinate of each regular image and the area of each regular image, the X-coordinate of the centroid of the residual image is obtained.
[0091] Based on the Y-coordinate of the center coordinate of each regular image and the area of each regular image, the Y-coordinate of the centroid of the residual image is obtained.
[0092] It should be noted that, in this embodiment, the formula for calculating the centroid coordinates of the residual image is as follows:
[0093]
[0094] Where x represents the X-coordinate of the centroid of the residual image, y represents the Y-coordinate of the centroid of the residual image, n represents the number of regular images, and S i Represented as the area of the i-th regular image, (G ix G iy ) is represented as the center coordinates of the i-th regular image.
[0095] It should be noted that, in this embodiment, the regular image is one or more of triangles, regular polygons, circles, and ellipses.
[0096] When the residual image is of a regular shape, the calculation of the centroid coordinates of the residual image is the same as the calculation of the center coordinates of the residual image.
[0097] Step S23: Based on the centroid coordinates of the residual image, obtain the angle between the centroid coordinates of the residual image and the positive X-axis direction;
[0098] Step S24: Based on the centroid coordinates of the residual image, obtain the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / the target video frame;
[0099] It should be noted that, in this embodiment, the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / target video frame can be calculated based on the pixel position.
[0100] It should be noted that, in this embodiment, the preset vertex of the selected video frame / the target video frame is preferably the lower left corner vertex of the selected video frame / the target video frame.
[0101] Step S25: The total number of pixels in the residual image, the angle between the centroid coordinates of the residual image and the positive X-axis, and the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / target video frame are used as the inter-frame features between the selected video frame / target video frame and its adjacent video frames.
[0102] Step S30: If the inter-frame features of the selected video frame and its neighboring video frames are equal to the inter-frame features of the target video frame and its neighboring video frames, the reference video and the target video are aligned using the selected video frame and the target video frame as the first alignment frame.
[0103] It should be noted that, in this embodiment, the step of aligning the reference video and the target video using the selected video frame and the target video frame as the first alignment frame includes:
[0104] Based on a preset number of frames, starting from the selected video frame, a video sequence with a preset number of frames is extracted from the reference video;
[0105] Based on a preset number of frames, starting from the target video frame, a video sequence with a preset number of frames is extracted from the target video;
[0106] Align the video sequence with a preset number of frames extracted from the reference video with the video sequence with a preset number of frames extracted from the target video.
[0107] It should be noted that, in this embodiment, the preset number of frames may or may not include the selected video frame / target video frame. Those skilled in the art can set it according to actual application needs, and no specific restrictions are imposed here.
[0108] Compared to existing technologies that align video frames based on tags, this embodiment first matches the frame information of selected video frames in the reference video with each video frame in the target video to obtain the target video frame whose frame information matches that of the selected video frame. Then, based on the residual information of the selected video frame, the target video frame, and their respective neighboring video frames, the inter-frame features of the selected video frame, the target video frame, and their respective neighboring video frames are obtained. If the inter-frame features of the selected video frame and its neighboring video frames are equal to those of the target video frame and its neighboring video frames, the selected video frame and the target video frame are used as the first alignment frame to align the reference video and the target video. Therefore, this embodiment utilizes the inter-frame features corresponding to video frames to locate the first alignment frame, improving the accuracy of the first alignment frame location and thus improving the accuracy of video frame alignment. Therefore, it overcomes the technical defect of existing technologies that align video frames based on tags, which can only vaguely locate the first alignment frame, thus solving the technical problem of low video frame alignment accuracy.
[0109] Furthermore, based on the first embodiment of this application, in another embodiment of this application, the selected video frame is a scene switching frame, and the selected video frame is determined in the following manner:
[0110] The variation characteristics of the high-frequency subband coefficients of each video frame in the reference video are extracted by three-dimensional wavelet transform.
[0111] The selected video frame is obtained based on the variation characteristics of the high-frequency subband coefficients.
[0112] It should be noted that, in this embodiment, the step of obtaining the selected video frame based on the variation characteristics of the high-frequency subband coefficients includes:
[0113] The variation feature of the high-frequency subband coefficient is input into the trained classifier to classify and identify the variation feature of the high-frequency subband coefficient. Based on the classification and identification results, it is determined whether the video frame corresponding to the variation feature of the high-frequency subband coefficient is a scene switching frame.
[0114] It should be noted that in this embodiment, video scene switching frames are used as the first alignment frame. On the one hand, since the number of video scene switching frames in a video is less than that of ordinary video frames, the operation speed of the video frame alignment method can be improved, and the video frames can be aligned quickly. On the other hand, since the change information between video scene switching frames and their adjacent frames in a video is the most drastic, the accuracy of the first alignment frame positioning can be improved, thereby improving the accuracy of video frame alignment.
[0115] Furthermore, referring to Figure 3 Based on the first and second embodiments of this application, in another embodiment of this application, before the step of matching the selected video frame in the reference video with each video frame in the target video to obtain the target video frame whose frame information matches the selected video frame, the method further includes:
[0116] Step S00: Perform similarity matching on the reference video and the target video to obtain a reference video sequence and a target video sequence, wherein the reference video sequence and the target video sequence contain the same video frames.
[0117] Step S10, which involves matching the selected video frames in the reference video with each video frame in the target video to obtain the target video frame whose frame information matches that of the selected video frame, includes:
[0118] The selected video frame in the reference video sequence is matched with each video frame in the target video sequence to obtain the target video frame in the target video sequence that matches the frame information of the selected video frame.
[0119] In this embodiment, before locating and aligning the first frame in the reference video and the target video, a reference video sequence and a target video sequence are extracted from the reference video and the target video, respectively. The extracted reference video sequence and the target video sequence contain the same video frames. On one hand, the fact that the extracted reference video sequence and the target video sequence contain the same video frames can be understood as the existence of overlapping video segments in the reference video and the target video. Determining the existence of overlapping video segments in the reference video and the target video is a prerequisite that must be met when aligning video frames in the reference video and the target video. On the other hand, extracting the reference video sequence and the target video sequence before locating and aligning the first frame in the reference video and the target video, compared to locating and aligning the first frame from the reference video and the target video, improves the computational speed of the video frame alignment method and achieves rapid video frame alignment because the number of frames in the reference video sequence and the target video sequence is smaller than that in the reference video and the target video.
[0120] The step of performing similarity matching between the reference video and the target video to obtain a reference video sequence and a target video sequence includes:
[0121] Step S01: Calculate the frame information of each video frame in the reference video to obtain a reference video frame information array;
[0122] In this embodiment, the frame information of each video frame in the reference video is a hash value, and the hash value of each video frame in the reference video is determined in the following way:
[0123] Step S011: Scale each video frame in the reference video into a grayscale image of a preset resolution;
[0124] It should be noted that the preset resolution can be the resolution of one of the video frames in the reference video, or the resolution of one of the video frames in the target video, or it can be set by those skilled in the art according to the actual application requirements. For example, the preset resolution can be an 8×8 resolution.
[0125] Step S012: Traverse each pixel in the grayscale image, compare the brightness value of each pixel with its next adjacent pixel, and obtain the difference value;
[0126] It should be noted that if the brightness value of each pixel is greater than or equal to the brightness value of its next adjacent pixel, the difference value is set to 1; if the brightness value of each pixel is less than the brightness value of its next adjacent pixel, the difference value is set to 0.
[0127] Step S013: Combine all the difference values to obtain a difference value array;
[0128] It should be noted that combining all the difference values to obtain the difference value array is done by combining the difference values obtained in step S012 in the order in which they were obtained.
[0129] Step S014: Extract a preset number of values from the difference value array in sequence, obtain hexadecimal values based on the preset number of values, and concatenate the hexadecimal values to obtain a hexadecimal value array. Convert the hexadecimal value array into a string to obtain the hash value of each video frame in the reference video.
[0130] It should be noted that, in terms of image depth, images can be divided into 8-bit, 16-bit, 24-bit, and 32-bit, etc. Therefore, the preset number in this embodiment can be set according to the depth of the video frames in the reference video. For example, if the depth of the video frames in the reference video is 8 bits, then the preset number is 8.
[0131] Since the difference values in the difference value array are set to 0 or 1, the number of values extracted from the difference value array is a binary value, which needs to be converted into a hexadecimal value.
[0132] It should be noted that, in this embodiment, after calculating the frame information of each video frame in the reference video, the frame information is sorted according to the encoding of the corresponding video frame in the reference video to obtain a reference video frame information array.
[0133] Step S02: Calculate the frame information of each video frame in the target video to obtain a target video frame information array;
[0134] In this embodiment, the frame information of each video frame in the target video is a hash value, and the hash value of each video frame in the target video is determined in the following way:
[0135] Step S021: Scale each video frame in the target video into a grayscale image of a preset resolution;
[0136] It should be noted that the preset resolution can be the resolution of one of the video frames in the reference video, or the resolution of one of the video frames in the target video, or it can be set by those skilled in the art according to the actual application requirements. For example, the preset resolution can be an 8×8 resolution.
[0137] Step S022: Traverse each pixel in the grayscale image and compare the brightness value of each pixel with its next adjacent pixel to obtain the difference value;
[0138] It should be noted that if the brightness value of each pixel is greater than or equal to the brightness value of its next adjacent pixel, the difference value is set to 1; if the brightness value of each pixel is less than the brightness value of its next adjacent pixel, the difference value is set to 0.
[0139] Step S023: Combine all the difference values to obtain a difference value array;
[0140] It should be noted that combining all the difference values to obtain the difference value array is done by combining the difference values obtained in step S022 in the order in which they were obtained.
[0141] Step S024: Extract a preset number of values from the difference value array in sequence, obtain hexadecimal values based on the preset number of values, and concatenate the hexadecimal values to obtain a hexadecimal value array. Convert the hexadecimal value array into a string to obtain the hash value of each video frame in the target video.
[0142] It should be noted that, in terms of image depth, images can be divided into 8-bit, 16-bit, 24-bit, and 32-bit, etc. Therefore, the preset number in this embodiment can be set according to the depth of the video frames in the target video. For example, if the depth of the video frames in the target video is 8 bits, then the preset number is 8.
[0143] Since the difference values in the difference value array are set to 0 or 1, the number of values extracted from the difference value array is a binary value, which needs to be converted into a hexadecimal value.
[0144] It should be noted that, in this embodiment, after calculating the frame information of each video frame in the target video, the frame information is sorted according to the encoding of the corresponding video frame in the target video to obtain a target video frame information array.
[0145] Step S03: Traverse the reference video frame information array and the target video frame information array to obtain a first video frame and a second video frame with a similarity greater than a preset similarity threshold, wherein the first video frame is located in the reference video frame and the second video frame is located in the target video frame;
[0146] It should be noted that, in this embodiment, the frame information in the reference video frame information array is a hash value, and the frame information in the target video frame information array is also a hash value. Therefore, the step of traversing the reference video frame information array and the target video frame information array to obtain the first video frame and the second video frame with a similarity greater than a preset similarity threshold includes:
[0147] Step S031: Traverse the reference video frame information array and the target video frame information array;
[0148] Step S032: If the number of different data bits between a hash value in the reference video frame information array and a hash value in the target video frame information array is less than or equal to a preset bit threshold, then the video frame corresponding to the hash value in the reference video frame information array is marked as the first video frame, and the video frame corresponding to the hash value in the target video frame information array is marked as the second video frame.
[0149] In practical applications, the number of video frames to be processed is often in the tens of thousands. Traversing the baseline video frame information array and the target video frame information array to obtain the first and second video frames with a similarity greater than a preset similarity threshold is often time-consuming and complex. Therefore, to improve processing speed, in this embodiment, both the baseline video frame information array and the target video frame information array are divided into multiple intervals. The intervals are traversed, and if no first or second video frame with a similarity greater than the preset similarity threshold is obtained in any given interval, the remaining intervals are traversed until a first or second video frame with a similarity greater than the preset similarity threshold is obtained.
[0150] Step S06: Count the number of frames in the video sequence starting from the first video frame in the reference video, and the number of frames in the video sequence starting from the second video frame in the target video;
[0151] Step S07: Select the smaller value between the number of frames in the video sequence starting from the first video frame and the number of frames in the video sequence starting from the second video frame as the number of frames to be extracted;
[0152] Step S08: Based on the number of frames captured, starting from the first video frame, obtain a reference video sequence from the reference video;
[0153] It should be noted that in this embodiment, the number of frames captured may or may not include the first video frame. Those skilled in the art can set it according to actual application needs, and no specific restrictions are imposed here.
[0154] Step S09: Based on the number of frames captured, starting from the second video frame, obtain the target video sequence from the target video.
[0155] It should be noted that in this embodiment, the number of frames to be captured may or may not include the second video frame. Those skilled in the art can set this according to actual application requirements, and no specific limitations are imposed here. However, when obtaining a reference video sequence from the reference video, if the number of frames captured includes the first video frame, then when obtaining a target video sequence from the target video, the number of frames captured also includes the second video frame. Similarly, when obtaining a reference video sequence from the reference video, if the number of frames captured does not include the first video frame, then when obtaining a target video sequence from the target video, the number of frames captured also does not include the second video frame.
[0156] Furthermore, based on the first, second, and third embodiments of this application, in another embodiment of this application, after the step of traversing the reference video frame information array and the target video frame information array to obtain a first video frame and a second video frame with a similarity greater than a preset similarity threshold, the method further includes:
[0157] Step S04: Compare the similarity between the adjacent video frames of the first video frame and the adjacent video frames of the second video frame;
[0158] It should be noted that comparing the similarity of adjacent video frames of the first video frame with that of adjacent video frames of the second video frame can be done in several ways: First, it can be comparing the similarity of the preceding adjacent video frame of the first video frame (the adjacent video frame in the reference video with a smaller encoding than the first video frame) with the preceding adjacent adjacent video frame of the second video frame (the adjacent video frame in the target video with a smaller encoding than the second video frame); second, it can be comparing the similarity of the preceding and following adjacent video frames of the first video frame with the preceding and following adjacent video frames of the second video frame. For example, comparing the similarity of the preceding 5 adjacent video frames and the following 5 adjacent video frames of the first video frame with the preceding 5 adjacent video frames and the following 5 adjacent video frames of the second video frame.
[0159] Step S05: If, among the adjacent video frames for similarity comparison, the proportion of video frames with a similarity greater than a preset similarity threshold is greater than a preset proportion threshold, then proceed to step S06.
[0160] It should be noted that when comparing the similarity of adjacent video frames, the comparison process is the same as that of the first and second video frames mentioned above, which is a comparison of the hash values of the video frames, and will not be repeated here.
[0161] It should be noted that when comparing the similarity of adjacent video frames, the adjacent video frames of the first video frame must correspond to the adjacent video frames of the second video frame. For example, the preceding video frame of the first video frame corresponds to the preceding video frame of the second video frame.
[0162] It should be noted that, in this embodiment, the preset proportion threshold can be set by those skilled in the art according to actual application needs, and no specific restrictions are imposed here. For example, as mentioned above, the similarity of the five adjacent video frames before and after the first video frame with the five adjacent video frames before and after the second video frame is compared. If, among all the adjacent videos that are compared for similarity, the proportion of video frames with a similarity greater than the preset similarity threshold is greater than 0.8, then step S06 is executed. That is, among the 10 adjacent video frames that are compared for similarity, 8 video frames have a similarity greater than the preset similarity threshold.
[0163] Compared to obtaining a first video frame and a second video frame with a similarity greater than a preset similarity threshold, and then acquiring a reference video sequence and a target video sequence starting from the first video frame and the second video frame respectively, in this embodiment, after obtaining a first video frame and a second video frame with a similarity greater than a preset similarity threshold, the adjacent video frames of the first video frame and the adjacent video frames of the second video frame are compared for similarity. If the proportion of video frames with a similarity greater than a preset similarity threshold among the adjacent video frames being compared is greater than a preset proportion threshold, then the reference video sequence and the target video sequence are acquired starting from the first video frame and the second video frame respectively. This improves the accuracy of locating the first video frame and the second video frame, and indirectly improves the accuracy of frame video alignment.
[0164] This application also provides a video frame alignment device, referring to... Figure 4 , Figure 4 This is a schematic diagram of the functional modules of the first embodiment of the video frame alignment device of this application.
[0165] In this embodiment, the video frame alignment device includes:
[0166] The first matching module 10 is used to perform frame information matching between selected video frames in the reference video and each video frame in the target video to obtain the target video frame that matches the frame information of the selected video frame in the target video.
[0167] The inter-frame feature acquisition module 20 is used to obtain the inter-frame features of the selected video frame and the target video frame and their respective neighboring video frames based on the residual information of the selected video frame and the target video frame and their respective neighboring video frames.
[0168] Alignment module 30 is used to align the reference video and the target video, with the selected video frame and the target video frame as the first alignment frame, if the inter-frame features of the selected video frame and its adjacent video frames are equal to the inter-frame features of the target video frame and its adjacent video frames.
[0169] Optionally, the inter-frame feature acquisition module includes:
[0170] The residual image acquisition unit is used to calculate the motion vector between the selected video frame / target video frame and its adjacent video frames to obtain a residual image;
[0171] The first inter-frame feature acquisition unit is used to calculate the total number of pixels in the residual image;
[0172] A centroid coordinate calculation unit is used to calculate the centroid coordinates of the residual image;
[0173] The second inter-frame feature acquisition unit is used to calculate the angle between the centroid coordinates of the residual image and the positive X-axis based on the centroid coordinates of the residual image.
[0174] The third inter-frame feature acquisition unit is used to obtain the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / the target video frame based on the centroid coordinates of the residual image.
[0175] The inter-frame feature determination unit is used to take the total number of pixels in the residual image, the angle between the centroid coordinates of the residual image and the positive X-axis, and the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / the target video frame as the inter-frame features between the selected video frame / the target video frame and its adjacent video frames.
[0176] Optionally, when the residual image is irregular in shape, the centroid coordinate calculation unit is used to:
[0177] The residual image is segmented into multiple regular images;
[0178] Obtain the center coordinates and area of each of the aforementioned regular images;
[0179] Based on the X-coordinate of the center coordinate of each regular image and the area of each regular image, the X-coordinate of the centroid of the residual image is obtained.
[0180] Based on the Y-coordinate of the center coordinate of each regular image and the area of each regular image, the Y-coordinate of the centroid of the residual image is obtained.
[0181] Optionally, when the selected video frame is a scene transition frame, the first matching module includes:
[0182] The determining unit is used to extract the variation characteristics of the high-frequency subband coefficients of each video frame in the reference video through three-dimensional wavelet transformation, and to obtain the selected video frame based on the variation characteristics of the high-frequency subband coefficients.
[0183] Optionally, the video frame alignment device further includes:
[0184] The second matching module is used to perform similarity matching on the reference video and the target video to obtain a reference video sequence and a target video sequence, wherein the reference video sequence and the target video sequence contain the same video frames.
[0185] The first matching module is used to perform frame information matching between selected video frames in the reference video sequence and each video frame in the target video sequence to obtain target video frames in the target video sequence that match the frame information of the selected video frames.
[0186] Optionally, the second matching module includes:
[0187] The first frame information array acquisition unit is used to calculate the frame information of each video frame in the reference video to obtain the reference video frame information array;
[0188] The second frame information array acquisition unit is used to calculate the frame information of each video frame in the target video to obtain the target video frame information array;
[0189] The first comparison unit is used to traverse the reference video frame information array and the target video frame information array to obtain a first video frame and a second video frame with a similarity greater than a preset similarity threshold, wherein the first video frame is located in the reference video frame and the second video frame is located in the target video frame;
[0190] The statistics unit is used to count the number of frames in the video sequence starting from the first video frame in the reference video, and the number of frames in the video sequence starting from the second video frame in the target video.
[0191] The frame number determination unit is used to select the smaller value between the number of frames in the video sequence starting from the first video frame and the number of frames in the video sequence starting from the second video frame as the frame number to be extracted.
[0192] The first segmentation unit is configured to obtain a reference video sequence from the reference video, starting from the first video frame, based on the number of segmented frames.
[0193] The second interception unit is used to obtain a target video sequence from the target video, starting from the second video frame, based on the number of intercepted frames.
[0194] Optionally, the second matching module further includes:
[0195] The second comparison unit is used to compare the similarity between the adjacent video frames of the first video frame and the adjacent video frames of the second video frame.
[0196] In the case where the proportion of video frames with a similarity greater than a preset similarity threshold in adjacent video frames being compared is greater than a preset proportion threshold, the statistical unit is used to count the number of frames in the video sequence starting from the first video frame in the reference video and the number of frames in the video sequence starting from the second video frame in the target video.
[0197] The specific implementation of the video frame alignment device in this application is basically the same as the embodiments of the video frame alignment method described above, and will not be repeated here.
[0198] In addition, to achieve the above objectives, this application also provides a storage medium storing a video frame alignment program, which, when executed by a processor, implements the steps of the video frame alignment method described above.
[0199] The specific implementation of the storage medium in this application is basically the same as the embodiments of the video frame alignment method described above, and will not be repeated here.
[0200] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0201] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0203] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A video frame alignment method, characterized in that, The video frame alignment method includes the following steps: The selected video frame in the reference video is matched with each video frame in the target video to obtain the target video frame whose frame information matches the selected video frame. Based on the residual information of the selected video frame and the target video frame with their respective neighboring video frames, the inter-frame features of the selected video frame and the target video frame with their respective neighboring video frames are obtained. If the inter-frame features of the selected video frame and its neighboring video frames are equal to the inter-frame features of the target video frame and its neighboring video frames, the reference video and the target video are aligned using the selected video frame and the target video frame as the first alignment frame.
2. The video frame alignment method as described in claim 1, characterized in that, The step of obtaining the inter-frame features of the selected video frame and the target video frame and their respective neighboring video frames based on the residual information of the selected video frame and the target video frame and their respective neighboring video frames includes: Calculate the motion vectors between the selected video frame / target video frame and its adjacent video frames to obtain the residual image; Calculate the total number of pixels and the centroid coordinates of the residual image; Based on the centroid coordinates of the residual image, the angle between the centroid coordinates of the residual image and the positive X-axis direction is obtained; Based on the centroid coordinates of the residual image, the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / the target video frame is obtained; The total number of pixels in the residual image, the angle between the centroid coordinates of the residual image and the positive X-axis, and the distance between the centroid coordinates of the residual image and the preset vertex coordinates of the selected video frame / target video frame are used as the inter-frame features between the selected video frame / target video frame and its adjacent video frames.
3. The video frame alignment method as described in claim 2, characterized in that, When the residual image is irregular in shape, the step of calculating the centroid coordinates of the residual image includes: The residual image is segmented into multiple regular images; Obtain the center coordinates and area of each of the aforementioned regular images; Based on the X-coordinate of the center coordinate of each regular image and the area of each regular image, the X-coordinate of the centroid of the residual image is obtained. Based on the Y-coordinate of the center coordinate of each regular image and the area of each regular image, the Y-coordinate of the centroid of the residual image is obtained.
4. The video frame alignment method as described in claim 1, characterized in that, When the selected video frame is a scene transition frame, the selected video frame is determined in the following way: The variation characteristics of the high-frequency subband coefficients of each video frame in the reference video are extracted by three-dimensional wavelet transform. The selected video frame is obtained based on the variation characteristics of the high-frequency subband coefficients.
5. The video frame alignment method as described in claim 1, characterized in that, Before the step of matching the selected video frames in the reference video with each video frame in the target video to obtain the target video frame whose frame information matches that of the selected video frame, the method further includes: Similarity matching is performed on the reference video and the target video to obtain a reference video sequence and a target video sequence, wherein the reference video sequence and the target video sequence contain the same video frames. The step of matching selected video frames in the reference video with each video frame in the target video to obtain a target video frame in the target video whose frame information matches that of the selected video frame includes: The selected video frame in the reference video sequence is matched with each video frame in the target video sequence to obtain the target video frame in the target video sequence that matches the frame information of the selected video frame.
6. The video frame alignment method as described in claim 5, characterized in that, The step of performing similarity matching on the reference video and the target video to obtain a reference video sequence and a target video sequence includes: Calculate the frame information of each video frame in the reference video to obtain a reference video frame information array; Calculate the frame information of each video frame in the target video to obtain a target video frame information array; Traverse the reference video frame information array and the target video frame information array to obtain a first video frame and a second video frame with a similarity greater than a preset similarity threshold, wherein the first video frame is located in the reference video frame and the second video frame is located in the target video frame; The number of frames in the video sequence starting from the first video frame in the reference video and the number of frames in the video sequence starting from the second video frame in the target video are counted. The smaller value between the number of frames in the video sequence starting from the first video frame and the number of frames in the video sequence starting from the second video frame is selected as the number of frames to be extracted. Based on the number of frames captured, a reference video sequence is obtained from the reference video, starting from the first video frame; Based on the number of frames captured, starting from the second video frame, a target video sequence is obtained from the target video.
7. The video frame alignment method as described in claim 6, characterized in that, After the step of traversing the reference video frame information array and the target video frame information array to obtain a first video frame and a second video frame with a similarity greater than a preset similarity threshold, the method further includes: The similarity of the adjacent video frames of the first video frame and the adjacent video frames of the second video frame is compared. If, among adjacent video frames undergoing similarity comparison, the proportion of video frames with a similarity greater than a preset similarity threshold is greater than a preset proportion threshold, then the steps of counting the number of frames in the video sequence starting from the first video frame in the reference video and the number of frames in the video sequence starting from the second video frame in the target video are performed.
8. A video frame alignment device, characterized in that, The video frame alignment device includes: The first matching module is used to perform frame information matching between selected video frames in the reference video and each video frame in the target video to obtain the target video frame that matches the frame information of the selected video frame in the target video. The inter-frame feature acquisition module is used to obtain the inter-frame features of the selected video frame and the target video frame and their respective neighboring video frames based on the residual information of the selected video frame and the target video frame and their respective neighboring video frames. An alignment module is used to align the reference video and the target video, with the selected video frame and the target video frame as the first alignment frame, if the inter-frame characteristics of the selected video frame and its adjacent video frames are equal to the inter-frame characteristics of the target video frame and its adjacent video frames.
9. A video frame alignment device, characterized in that, The device includes: a memory, a processor, and a video frame alignment program stored in the memory and executable on the processor, the video frame alignment program being configured to implement the steps of the video frame alignment method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a video frame alignment program, which, when executed by a processor, implements the steps of the video frame alignment method as described in any one of claims 1 to 7.