Video fluency verification method, device, equipment and media
By determining motion vectors between video frames, the method addresses the inaccuracy of frame rate-based fluency evaluation, enhancing accuracy and enabling adaptive video processing to improve fluency and viewing experience.
Patent Information
- Application Number
- JP2023578986
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-23
- Filing Date
- 2022-08-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-08-17
AI Technical Summary
Current methods for evaluating video fluency using frame rate are inaccurate as they do not consider the influence of human visual perception, leading to low accuracy in assessing video fluency.
A method that determines motion vectors between video frames, including lens and object motion vectors, to verify video fluency based on a pre-established motion fluency mapping relationship, which accounts for human visual perception.
Improves the accuracy of video fluency assessment by considering human visual perception, enabling better classification of moving scenes and quantification of fluency, facilitating adaptive video compression and interpolation for improved viewing experience.
Smart Images

Figure 0007719217000003 
Figure 0007719217000004 
Figure 0007719217000005
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to a Chinese patent application filed on August 23, 2021, bearing application number 202110967767.1 and entitled "Video fluency verification method, device, apparatus and medium," the entire contents of which are incorporated herein by reference. [Technical Field]
[0002] The present disclosure relates to the technical field of video processing, and more particularly to a method, apparatus, device and medium for verifying video fluency. [Background technology]
[0003] 2. Description of the Related Art With the development of science and technology, watching videos over the network has become an important part of life, and the demand for videos is becoming higher and higher.
[0004] Video fluency is one of the important indicators that evaluate video quality and affect the viewing experience. Currently, it is usually evaluated using frame rate, where a higher frame rate indicates higher video fluency and vice versa. However, the human eye's perception of video fluency is closely related to the scenes in the video content, and the method of evaluating video fluency using frame rate does not take into account the influence of the human eye's visual perception and has low accuracy. Summary of the Invention [Means for solving the problem]
[0005] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device and medium for verifying video fluency.
[0006] An embodiment of the present disclosure includes: obtaining a target video; determining motion vectors between different video frames by performing motion estimation on the different video frames in the target video, the motion vectors including a lens motion vector used to indicate a change in position of a lens between the different video frames and an object motion vector used to indicate a change in position of the same photographed object in the different video frames; and verifying the fluency of the target video based on the motion vectors between the different video frames.
[0007] An embodiment of the present disclosure includes: a video acquisition module for acquiring a target video; a motion estimation module for performing motion estimation on different video frames in the target video to determine motion vectors between the different video frames, the motion vectors including a lens motion vector used to indicate a change in position of a lens between the different video frames and an object motion vector used to indicate a change in position of a same photographed object in the different video frames; and a fluency module for verifying the fluency of the target video based on the motion vectors between the different video frames.
[0008] An embodiment of the present disclosure further provides an electronic device including a processor and a memory for storing instructions executable by the processor, the processor being adapted to read the executable instructions from the memory and execute the instructions to implement the video fluency verification method provided by an embodiment of the present disclosure.
[0009] An embodiment of the present disclosure further provides a computer-readable storage medium having stored thereon a computer program for performing the video fluency verification method provided by an embodiment of the present disclosure.
[0010] An embodiment of the present disclosure further provides a computer program product including computer programs / instructions that, when executed by a processor, implement the video fluency verification method provided by an embodiment of the present disclosure. [Effects of the Invention]
[0011] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages: The technical solution for checking video fluency provided by the embodiments of the present disclosure acquires a target video and performs motion estimation on different video frames in the target video to determine motion vectors between the different video frames, where the motion vectors include lens motion vectors and object motion vectors, and checks the fluency of the target video based on the motion vectors between the different video frames. Using the above technical solution, video fluency can be checked using two different types of motion estimation between different video frames in a video, and the results of the two different types of motion estimation are related to the visual characteristics of the human eye, thereby improving the accuracy of checking video fluency.
[0012] These and other features, advantages, and aspects of each embodiment of the present disclosure will become more apparent by reference to the following specific embodiments in conjunction with the drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a schematic flowchart of a video fluency verification method provided by an embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic flowchart of another video fluency verification method provided by an embodiment of the present disclosure. [Figure 3] FIG. 3 is a schematic structural diagram of a video fluency verification device provided according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic structural diagram of an electronic device provided according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the drawings show several embodiments of the present disclosure, it should be understood that the present disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0015] It should be understood that the steps recited in the method embodiments of the present disclosure may be performed in various orders and / or in parallel. Additionally, method embodiments may include additional steps and / or omit steps as shown. The scope of the present disclosure is not limited in this respect.
[0016] As used herein, the term "comprises" and variations thereof mean an open inclusion, including, but not limited to. The term "based on" means "based at least in part on." The term "in one embodiment" refers to "at least one embodiment," the term "in another embodiment" refers to "at least one other embodiment," and the term "in some embodiments" refers to "at least some embodiments." Relevant definitions of other terms are provided below.
[0017] It should be noted that the concepts of "first," "second," etc. referred to in this disclosure are merely intended to distinguish between different devices, modules, or units, and do not limit the order or interdependence of functions performed by these devices, modules, or units.
[0018] It should be understood that the modifications "one" and "multiple" referred to in this disclosure are intended to be exemplary rather than limiting, and that those skilled in the art should understand "one or more" unless the context clearly dictates otherwise.
[0019] The names of messages or information exchanged between devices in the embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0020] Video fluency is one of the important indicators that affect the viewing experience. Currently, it is usually evaluated using frame rate, with a higher frame rate being considered to indicate higher video fluency and vice versa. However, the human eye's perception of video fluency is closely related to the scene content of the video. For example, if the video screen is still, a low frame rate will result in no sense of delay, while in scenes with long periods of screen movement due to moving filming equipment, a low frame rate will create a strong sense of discomfort. Therefore, the frame rate required to achieve fluency varies depending on the scene.
[0021] Currently, there is a lack of research on video fluency, and it is mainly evaluated using frame rate. When using frame rate to express video fluency, it cannot take into account the influence of human visual perception, resulting in low accuracy, and when considering the influence of human visual perception, it cannot be specifically quantified. To solve the above problems, the embodiments of the present disclosure provide a method for checking video fluency, which will be described below in conjunction with specific embodiments.
[0022] 1 is a schematic flowchart of a video fluency verification method provided by an embodiment of the present disclosure, which can be performed by a video fluency verification device, which can be realized in software and / or hardware, and generally can be integrated into an electronic device. As shown in FIG. 1, the method includes steps 101 to 103.
[0023] Step 101: Get the target video.
[0024] The target video may be any one video whose fluency needs to be analyzed and checked, and the specific type and origin are not limited, for example, the target video may be a video shot in real time or a video downloaded from the Internet.
[0025] Step 102: Determine motion vectors between different video frames by performing motion estimation on different video frames in the target video.
[0026] In the embodiments of the present disclosure, video motion can be divided into two dimensions: the overall motion of the scene caused by the lens motion, and the detailed motion of the object in the scene content. Therefore, the motion vector may include two kinds: lens motion vector and object motion vector. The lens motion vector is used to indicate the position change of the lens between different video frames, and the object motion vector is used to indicate the position change of the same captured object in different video frames.
[0027] A video frame may be the smallest unit constituting a video, and can be extracted from a target video. Since the target video may include multiple video frames, the distance between different video frames can be set according to actual conditions when performing motion estimation. In the embodiments of the present disclosure, the distance between different video frames is taken as an example to be one video frame, that is, the different video frames are adjacent video frames.
[0028] Specifically, the step of determining motion vectors between different video frames by performing motion estimation on different video frames in the target video may include the steps of extracting multiple video frames of the target video; determining motion vectors of multiple image blocks between different video frames to obtain multiple motion vectors, where each video frame includes multiple image blocks; clustering the multiple motion vectors to obtain multiple types of vector sets; and determining motion vectors between different video frames based on the multiple types of vector sets.
[0029] After acquiring the target video, multiple video frames included in the target video can be extracted. Different video frames can be characterized as a first video frame and a second video frame. The first video frame and the second video frame can be divided into multiple non-overlapping image blocks. The specific number of image blocks is not limited. For example, in an embodiment of the present disclosure, the video frame can be divided into 16*16 image blocks. Then, for each image block in the first video frame, a block matching three-step search method is used to find the best-matching image block in the second video frame. A motion vector can then be determined based on the two matching image blocks, resulting in a motion vector corresponding to each image block. Clustering and analysis can then be performed on the multiple motion vectors to determine motion vectors for different video frames. If the different video frames are adjacent, the motion vectors for every two adjacent video frames can be determined. The clustering method is not limited. For example, clustering can be performed using a density-based clustering method with noise (DBSCAN).
[0030] Taking the search within the range [-7, 7] as an example, the specific process of the above three-step search method is as follows: centering on the current position of the matching block, eight points are searched vertically, horizontally, and diagonally at an interval of 4. Adding the center point, it constitutes the first step of a "field" character with a side length of 8. Taking the closest point among the search results of the first step as the center, similarly, eight points are searched vertically, horizontally, and diagonally. This time, the interval is halved to search for a "field" character with a side length of 4, which is the second step. Repeating the second step, the interval is further halved to 1. At this time, the most similar point found is the point with the minimum matching error, which is the third step. Optionally, the sum of the absolute differences (SAD) of the most matching can be recorded as the reliability of the motion vector. Optionally, the motion vector calculated by the above three-step search method generally contains a lot of noise, and the motion vector of each image block can be smoothed and filtered by a 5×5 median filtering method to remove the noise.
[0031] Optionally, the step of determining the motion vector between different video frames based on multiple types of vector sets may include: determining the vector set with the largest number of motion vectors among the multiple types of vector sets as the first set, and determining the sets other than the first set as the second set; determining the average value of the motion vectors in the first set as the lens motion vector among the motion vectors between different video frames; and determining the average value of the motion vectors in the second set as the object motion vector among the motion vectors between different video frames.
[0032] The motion vectors of multiple image blocks between different video frames are clustered, and the obtained clustering result is multiple types of vector sets, and each type of vector set may include multiple motion vectors. Then, the number of motion vectors among the multiple types of vector sets can be confirmed, and the vector set with the largest number is determined as the first set, which can be understood as a set representing the overall motion, and the other vector sets other than the first set are determined as the second set, which can be understood as a set representing the detailed motion. Then, the average value of the motion vectors in the first set can be determined as the lens motion vector between different video frames, and the average value of the motion vectors in the second set can be determined as the object motion vector between different video frames.
[0033] Step 103: Check the fluency of the target video based on the motion vectors between different video frames.
[0034] In an embodiment of the present disclosure, the step of confirming the fluency of the target video based on the motion vectors between different video frames may include: confirming the unit fluency of the different video frames based on the motion vectors between the different video frames and a pre-established motion fluency mapping relationship; and confirming the fluency of the target video based on the unit fluency.
[0035] The motion fluency mapping relationship can be obtained by human eye marks, i.e., the motion fluency mapping relationship is subjectively established based on the visual characteristics of the human eye. The motion fluency mapping relationship includes a mapping relationship between a motion vector and a fluency quantification value, where the fluency quantification value is a subjective score for a video, and a larger fluency quantification value means higher fluency.
[0036] In the embodiments of the present disclosure, video motion can be divided into two dimensions: the overall motion of the screen caused by the lens motion, and the detailed motion of the object in the screen content. The motion amplitudes of the two types of motion can be divided into four levels, including still, slight motion, normal amplitude motion, and intense motion. According to the types and motion amplitudes, multiple types of motion scenes can be generated, and the fluency quantification values corresponding to different motion scenes are different. Specifically, see Table 1, which is a motion quantification table reflecting the relationship between motion and specific fluency quantification values. Table 1. Movement quantification table
[0037] [Table 1]
[0038] As shown in Table 1 above, the two types of motion states, overall motion and detailed motion, are each quantified on a scale of 0 to 3 points, with higher scores indicating smaller motion. By simply adding up the scores of the two types of motion states, five types of scene motion fluency quantification values ranging from 2 to 6 points are formed, as shown in Table 2, which is a motion fluency mapping relationship table. That is, video motion scenes can be divided into five types according to the motion situation, with different fluency quantification values corresponding to each type of motion scene. The higher the quantified fluency score, the smaller the overall motion. Table 2. Movement fluency mapping relationship table
[0039] [Table 2]
[0040] In an embodiment of the present disclosure, the step of determining the unit fluency of different video frames based on the motion vectors between different video frames and the pre-established motion fluency mapping relationship may include the steps of determining a first fluency quantification value corresponding to the lens motion vector and a second fluency quantification value corresponding to the object motion vector of the different video frames based on the motion vectors between the different video frames and the motion fluency mapping relationship; and determining the sum of the first fluency quantification value and the second fluency quantification value as a unit fluency quantification value representing the unit fluency.
[0041] After determining the motion vectors between different video frames of the target video, the motion fluency mapping relationship can be searched to determine the first fluency quantification value corresponding to the lens motion vector between different video frames and the second fluency quantification value corresponding to the object motion vector, and the first fluency quantification value and the second fluency quantification value are determined as the unit fluency quantification value, which is the unit fluency.
[0042] In an embodiment of the present disclosure, confirming the fluency of the target video based on the unit fluency may include: taking a weighted average of the quantified values of the unit fluency of different video frames to obtain a quantified value of target fluency; and confirming the fluency of the target video based on the quantified value of target fluency. Optionally, confirming the fluency of the target video based on the quantified value of target fluency may include determining the quantified value of target fluency as the fluency of the target video.
[0043] Since the target video may include multiple video frames, after determining the quantification values of unit fluency between different video frames, the quantification value of the multiple unit fluency can be weighted and averaged to calculate the quantification value of target fluency. If the different video frames are adjacent video frames, the quantification values of unit fluency between every two adjacent video frames can be weighted and averaged to obtain the quantification value of target fluency. The quantification value of target fluency can then be determined as the fluency of the target video. Optionally, the target video can be filtered to remove motion scenes that appear multiple times discontinuously, i.e., to remove the quantification values of unit fluency that appear multiple times discontinuously, thereby improving the accuracy of determining the fluency of the video.
[0044] In an embodiment of the present disclosure, the motion fluency mapping relationship further includes a mapping relationship between a fluency quantification value and a fluent frame rate, and the fluent frame rate is used to indicate the frame rate required for videos with different fluency quantification values to achieve visual fluency. The fluent frame rate can be understood as the frame rate required for videos with different fluency quantification values to achieve visual fluency of the human eye.
[0045] Referring to Table 2, the higher the fluency quantification score, the smaller the overall motion, and the lower the frame rate at which a subjective sense of fluency can be achieved. Conversely, the frame rate requirement is higher. As shown in Table 2, different types of motion scenes require different frame rates to achieve fluent viewing. Specifically, the process of determining the fluent frame rate corresponding to each fluency quantification value may include collecting high-frame-rate videos of the above five types of scenes, generating multiple low-frame-rate videos corresponding to each high-frame-rate video in a uniform frame loss manner, marking the fluency quantification values for the high- and low-frame-rate videos of all different scenes, and establishing a mapping relationship between different categories of motion scenes and fluent frame rates according to the marking results, i.e., establishing a mapping relationship between fluency quantification values and fluent frame rates. The motion fluency mapping relationship table in Table 2 may also include a series of motion vectors, which are not specifically shown in the table.
[0046] Optionally, confirming the fluency of the target video based on the quantified value of the target fluency may include confirming the fluency of the target video based on a comparison result between a fluent frame rate corresponding to the quantified value of the target fluency and a current frame rate of the target video. Optionally, confirming the fluency of the target video based on a comparison result between the fluent frame rate corresponding to the quantified value of the target fluency and a current frame rate of the target video may include confirming the fluency of the target video as not fluent if the current frame rate of the target video is smaller than the fluent frame rate corresponding to the quantified value of the target fluency, and confirming the fluency of the target video as fluent if not.
[0047] After confirming the quantification value of the target fluency, the fluent frame rate corresponding to the quantification value of the target fluency can be confirmed by searching the motion fluency mapping relationship, and the fluent frame rate is compared with the current frame rate of the target video. If the current frame rate of the target video is smaller than the fluent frame rate corresponding to the quantification value of the target fluency, the fluency of the target video is confirmed as not fluent; otherwise, the fluency of the target video is confirmed as fluent.
[0048] In the above technical solution, the fluency of the target video can be a result of fluency or dysfluency, and can also be a specific fluency quantification value, where a larger fluency quantification value means a higher fluency.
[0049] In the technical solution for verifying video fluency provided by the embodiments of the present disclosure, a target video is acquired, and motion estimation is performed on different video frames in the target video to determine motion vectors between the different video frames, where the motion vectors include a lens motion vector and an object motion vector, and the fluency of the target video is verified based on the motion vectors between the different video frames. Using the above technical solution, video fluency can be verified by two different types of motion estimation between different video frames in a video, and the results of the two different types of motion estimation are related to the visual characteristics of the human eye, thereby improving the accuracy of verifying video fluency.
[0050] 2 is a schematic flowchart of another video fluency verification method provided by an embodiment of the present disclosure, which further optimizes the video fluency verification method based on the above embodiment. As shown in FIG. 2, the method includes steps 201 to 206.
[0051] Step 201: Obtain a target video.
[0052] Step 202: Determine motion vectors between different video frames by performing motion estimation on different video frames in the target video.
[0053] The motion vectors include lens motion vectors, which are used to indicate the position change of the lens between different video frames, and object motion vectors, which are used to indicate the position change of the same photographed object in different video frames.
[0054] Optionally, the different video frames are adjacent video frames.
[0055] Optionally, the step of determining motion vectors between different video frames by performing motion estimation on different video frames in the target video may include the steps of extracting a plurality of video frames of the target video, determining motion vectors of a plurality of image blocks between the different video frames to obtain a plurality of motion vectors, where each video frame includes a plurality of image blocks, clustering the plurality of motion vectors to obtain a plurality of types of vector sets, and determining motion vectors between the different video frames based on the plurality of types of vector sets.
[0056] Optionally, the step of determining motion vectors between different video frames based on multiple types of vector sets may include the steps of determining a vector set having the largest number of motion vectors among the multiple types of vector sets as a first set, and determining sets other than the first set as second sets; determining an average value of the motion vectors in the first set as a lens motion vector among the motion vectors between different video frames; and determining an average value of the motion vectors in the second set as an object motion vector among the motion vectors between different video frames.
[0057] Step 203: Determine the unit fluency of different video frames based on the motion vectors between different video frames and the pre-established motion fluency mapping relationship.
[0058] The motion fluency mapping relationship includes a mapping relationship between the motion vector and the fluency quantification value, and a larger fluency quantification value means a higher fluency.
[0059] Optionally, the step of determining the unit fluency of the different video frames based on the motion vectors between the different video frames and the pre-established motion fluency mapping relationship includes the steps of determining a first fluency quantification value corresponding to the lens motion vector and a second fluency quantification value corresponding to the object motion vector of the different video frames based on the motion vectors between the different video frames and the motion fluency mapping relationship; and determining the sum of the first fluency quantification value and the second fluency quantification value as a unit fluency quantification value representing the unit fluency.
[0060] Step 204: weight-averaging the unit fluency quantification values of different video frames to obtain a target fluency quantification value.
[0061] After step 204, step 205 or step 206 can be performed.
[0062] Step 205: Determine the quantified value of the target fluency as the fluency of the target video.
[0063] Step 206: confirm the fluency of the target video based on the comparison result between the fluency frame rate corresponding to the quantified value of the target fluency and the current frame rate of the target video.
[0064] Optionally, the motion fluency mapping relationship further includes a mapping relationship between a fluency quantification value and a fluent frame rate, and the fluent frame rate is used to indicate the frame rate required for videos of different fluency quantification values to achieve visual fluency.
[0065] Optionally, the step of confirming the fluency of the target video based on a comparison result between the fluent frame rate corresponding to the quantification value of the target fluency and the current frame rate of the target video includes the step of confirming the fluency of the target video as not fluent if the current frame rate of the target video is smaller than the fluent frame rate corresponding to the quantification value of the target fluency, and otherwise confirming the fluency of the target video as fluent.
[0066] At present, considering the influence of human eye vision, it cannot be specifically quantified, and it is difficult to provide guidance for improving the fluency of specific videos.The video fluency confirmation method provided in this solution is simple and effective, and can realize the classification of moving scenes and the output of fluency through inter-frame motion estimation, and the output fluency can be quantified.By comparing the current frame rate of the video with the fluent frame rate, it is possible to very conveniently introduce adaptive video compression and frame interpolation technology to improve the video fluency based on the comparison result.For example, in video compression, frame loss operation is performed for parts with low frame rate requirements (still or very small motion parts), and frame interpolation technology performs adaptive frame interpolation for videos of different scenes.
[0067] In a technical solution for verifying video fluency provided by an embodiment of the present disclosure, a target video is acquired, and motion estimation is performed on different video frames in the target video to determine motion vectors between the different video frames; the unit fluency of the different video frames is determined based on the motion vectors between the different video frames and a pre-established motion fluency mapping relationship; the quantified values of the unit fluency of the different video frames are weighted-averaged to obtain a target fluency quantification value; and the quantified value of the target fluency is determined as the fluency of the target video; or the fluency of the target video is verified based on a comparison between a fluent frame rate corresponding to the target fluency quantification value and a current frame rate of the target video. Using the above technical solution, the fluency of the video can be verified using two different types of motion estimation and a pre-marked motion fluency mapping relationship. Because the motion fluency mapping relationship is subjectively established based on the visual perception of the human eye, the visual characteristics of the human eye are fully taken into account, and the accuracy of verifying video fluency can be improved.
[0068] 3 is a schematic structural diagram of a video fluency verification device provided by an embodiment of the present disclosure, which can be realized by software and / or hardware, and generally can be integrated into electronic equipment. As shown in FIG. 3, the device includes: a video acquisition module 301 for acquiring a target video; a motion estimation module 302 for performing motion estimation on different video frames in the target video to determine motion vectors between the different video frames, the motion vectors including: a lens motion vector used to indicate a change in position of a lens between the different video frames; and an object motion vector used to indicate a change in position of the same photographed object in the different video frames; a fluency module 303 for verifying the fluency of the target video based on the motion vectors between the different video frames.
[0069] Optionally, the fluency module 303: a unit fluency unit for determining the unit fluency of the different video frames based on the motion vectors between the different video frames and a pre-established motion fluency mapping relationship; a verifying unit for verifying the fluency of the target video based on the unit fluency.
[0070] Optionally, the motion fluency mapping relationship includes a mapping relationship between a motion vector and a fluency quantification value, where a larger fluency quantification value means a higher fluency.
[0071] Optionally, the unit fluency unit is specifically determining a first fluency quantification value corresponding to a lens motion vector of the different video frames and a second fluency quantification value corresponding to an object motion vector of the different video frames based on the motion vectors between the different video frames and the motion fluency mapping relationship; and determining the sum of the first fluency quantification value and the second fluency quantification value as a unit fluency quantification value representative of the unit fluency.
[0072] Optionally, the unit fluency unit is specifically weighting the unit fluency quantification values of the different video frames to obtain a target fluency quantification value; and verifying the fluency of the target video based on the quantified value of the target fluency.
[0073] Optionally, the motion fluency mapping relationship further includes a mapping relationship between the fluency quantification value and a fluent frame rate, and the fluent frame rate is used to indicate the frame rate required for videos of different fluency quantification values to achieve visual fluency.
[0074] Optionally, the unit fluency unit is specifically determining the target fluency quantification as the fluency of the target video; Alternatively, the fluency of the target video may be confirmed based on a comparison result between a fluent frame rate corresponding to the quantified value of the target fluency and a current frame rate of the target video.
[0075] Optionally, the unit fluency unit is specifically If the current frame rate of the target video is smaller than the fluent frame rate corresponding to the quantification value of the target fluency, it is used to confirm that the fluency of the target video is not fluent; otherwise, it is used to confirm that the fluency of the target video is fluent.
[0076] Optionally, the motion estimation module 302 specifically: extracting a plurality of video frames of the target video; determining motion vectors of a plurality of image blocks between different video frames to obtain a plurality of motion vectors, each of the video frames including a plurality of the image blocks; clustering the plurality of motion vectors to obtain a plurality of types of vector sets; and determining a motion vector between the different video frames based on the plurality of vector sets.
[0077] Optionally, the motion estimation module 302 specifically: determining a vector set having the largest number of motion vectors from among the plurality of types of vector sets as a first set, and determining sets other than the first set as second sets; determining an average value of the motion vectors in the first set as a lens motion vector among the motion vectors between the different video frames; The average value of the motion vectors in the second set is used to determine the object motion vector among the motion vectors between the different video frames.
[0078] Optionally, the different video frames are adjacent video frames.
[0079] The video fluency verification device provided by the embodiments of the present disclosure can implement the video fluency verification method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for implementing the method.
[0080] An embodiment of the present disclosure further provides a computer program product including a computer program / instructions that, when executed by a processor, implements the video fluency verification method provided by any embodiment of the present disclosure.
[0081] When implemented in software, the software may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are uploaded to a computer and executed, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wire (e.g., coaxial cable, fiber optics, digital subscriber line (DSL)) or wireless (e.g., infrared, radio, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device, such as a server or data center, that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)).
[0082] FIG. 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Hereinafter, specifically, referring to FIG. 4, a schematic structural diagram of an electronic device 400 applicable to realizing an embodiment of the present disclosure is shown. The electronic device 400 of the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG. 4 is merely an example and does not impose any limitations on the functions and scope of use of the embodiment of the present disclosure.
[0083] 4, the electronic device 400 may include a processing unit (e.g., a central processor, a graphics processor, etc.) 401, which can perform various appropriate operations and processes based on a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. The RAM 403 further stores various programs and data necessary for the operation of the terminal device 400. The processing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0084] Typically, input devices 406 including, for example, a touch panel, touch pad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc., output devices 407 including, for example, a liquid crystal display (LCD), speaker, oscillator, etc., storage devices 408 including, for example, a magnetic tape, hard disk, etc., and communication devices 409 can be connected to the I / O interface 405. The communication devices 409 can enable the terminal device 400 to communicate with other devices wirelessly or via a wire to exchange data. While FIG. 4 shows the terminal device 400 having various devices, it should be understood that it is not necessary for all of the devices shown to be implemented or included. Instead, more or fewer devices may be implemented or included.
[0085] In particular, according to embodiments of the present disclosure, the processes described with reference to the flowcharts above may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the methods illustrated in the flowcharts. In such embodiments, the computer program may be downloaded and installed from a network via the communication device 409, installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, it performs the above-described functions specific to the video fluency verification method of the embodiments of the present disclosure.
[0086] It should be noted that the computer-readable medium described in this disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may be, but are not limited to, an electrical connection having one or more leads, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact magnetic disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which may be used in or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transmit a program for use in or in connection with an instruction execution system, apparatus, or device. Included in a computer-readable medium.
[0087] In some embodiments, clients and servers may communicate using any network protocol now known or later developed, such as HyperText Transfer Protocol (HTTP), and may be connected to one another by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), the World Wide Web (e.g., the Internet), an end-to-end network (e.g., an ad hoc end-to-end network), and any network now known or later developed.
[0088] The computer-readable medium may be included in the electronic device, or may exist independently and not be assembled within the electronic device.
[0089] The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device acquires a target video and performs motion estimation on different video frames in the target video to determine motion vectors between the different video frames, where the motion vectors include a lens motion vector used to indicate a change in position of a lens between the different video frames and an object motion vector used to indicate a change in position of the same photographed object in the different video frames, and confirms the fluency of the target video based on the motion vectors between the different video frames.
[0090] Computer program code for carrying out the operations of the present disclosure can be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. When referring to a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0091] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program section, or portion of code, which includes one or more executable instructions for implementing a given logical function. It should be noted that in some alternative implementations, the functions shown in the blocks may occur in an order different from that shown in the figures. For example, two blocks shown in succession may actually be executed substantially in parallel or in the reverse order, depending on the functionality involved. It should be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system or a combination of dedicated hardware and computer instructions to perform a given function or operation.
[0092] The units according to the embodiments of the present disclosure may be implemented in the form of software or hardware, and in some cases, the names of the units do not constitute limitations on the units themselves.
[0093] As used herein, the above-described functions may be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that may be used include, but are not limited to, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0094] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program for use in or in connection with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection through one or more cables, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact magnetic disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0095] According to one or more embodiments of the present disclosure, the present disclosure provides a method for manufacturing a semiconductor device, comprising: obtaining a target video; determining motion vectors between different video frames by performing motion estimation on the different video frames in the target video, the motion vectors including: a lens motion vector used to indicate a position change of a lens between the different video frames; and an object motion vector used to indicate a position change of the same photographed object in the different video frames; and verifying the fluency of the target video based on the motion vectors between the different video frames.
[0096] According to one or more embodiments of the present disclosure, in the video fluency checking method provided by the present disclosure, the step of checking the fluency of the target video based on the motion vectors between the different video frames includes: determining unit fluency of the different video frames based on the motion vectors between the different video frames and a pre-established motion fluency mapping relationship; and checking the fluency of the target video based on the unit fluency.
[0097] According to one or more embodiments of the present disclosure, in the video fluency verification method provided by the present disclosure, the motion fluency mapping relationship includes a mapping relationship between a motion vector and a fluency quantification value, and a larger fluency quantification value means higher fluency.
[0098] According to one or more embodiments of the present disclosure, in a video fluency verification method provided by the present disclosure, the step of verifying unit fluency of the different video frames based on the motion vectors between the different video frames and a pre-established motion fluency mapping relationship includes: determining a first fluency quantification value corresponding to a lens motion vector of the different video frames and a second fluency quantification value corresponding to an object motion vector of the different video frames based on the motion vectors between the different video frames and the motion fluency mapping relationship; determining the sum of the first fluency quantification value and the second fluency quantification value as a unit fluency quantification value representative of the unit fluency.
[0099] According to one or more embodiments of the present disclosure, in a video fluency verification method provided by the present disclosure, the step of verifying the fluency of the target video based on the unit fluency includes: weighted averaging the unit fluency quantification values of the different video frames to obtain a target fluency quantification value; and confirming the fluency of the target video based on the quantified value of the target fluency.
[0100] According to one or more embodiments of the present disclosure, in the video fluency verification method provided by the present disclosure, the motion fluency mapping relationship further includes a mapping relationship between the fluency quantification value and a fluent frame rate, and the fluent frame rate is used to indicate the frame rate required for videos of different fluency quantification values to achieve visual fluency.
[0101] According to one or more embodiments of the present disclosure, in a video fluency verification method provided by the present disclosure, the step of verifying the fluency of the target video based on the quantified value of the target fluency includes: determining the target fluency quantification as the fluency of the target video; Alternatively, the method may include a step of confirming the fluency of the target video based on a comparison result between a fluent frame rate corresponding to the quantified value of the target fluency and a current frame rate of the target video.
[0102] According to one or more embodiments of the present disclosure, in a video fluency confirmation method provided by the present disclosure, the step of confirming the fluency of the target video based on a comparison result between a fluency frame rate corresponding to the target fluency quantification value and a current frame rate of the target video includes: The method includes a step of confirming that the fluency of the target video is not fluent if the current frame rate of the target video is smaller than the fluent frame rate corresponding to the quantification value of the target fluency, and otherwise confirming that the fluency of the target video is fluent.
[0103] According to one or more embodiments of the present disclosure, in a method for verifying video fluency provided by the present disclosure, the step of determining motion vectors between different video frames by performing motion estimation on the different video frames in the target video includes: extracting a plurality of video frames of the target video; determining motion vectors of a plurality of image blocks between different video frames to obtain a plurality of motion vectors, each of the video frames including a plurality of the image blocks; clustering the plurality of motion vectors to obtain a plurality of types of vector sets; determining motion vectors between the different video frames based on the plurality of vector sets.
[0104] According to one or more embodiments of the present disclosure, in the video fluency checking method provided by the present disclosure, the step of determining motion vectors between different video frames based on the plurality of vector sets includes: determining a vector set having the largest number of motion vectors from among the plurality of types of vector sets as a first set, and determining sets other than the first set as second sets; determining an average value of the motion vectors in the first set as a lens motion vector among the motion vectors between the different video frames; determining an average value of the motion vectors in the second set as an object motion vector among the motion vectors between the different video frames.
[0105] According to one or more embodiments of the present disclosure, in the video fluency verification method provided by the present disclosure, the different video frames are adjacent video frames.
[0106] According to one or more embodiments of the present disclosure, the present disclosure provides a method for manufacturing a semiconductor device, comprising: a video acquisition module for acquiring a target video; a motion estimation module for performing motion estimation on different video frames in the target video to determine motion vectors between the different video frames, the motion vectors including a lens motion vector used to indicate a change in position of a lens between the different video frames and an object motion vector used to indicate a change in position of a same photographed object in the different video frames; a fluency module for verifying the fluency of the target video based on the motion vectors between the different video frames.
[0107] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the fluency module includes: a unit fluency unit for determining the unit fluency of the different video frames based on the motion vectors between the different video frames and a pre-established motion fluency mapping relationship; a verifying unit for verifying the fluency of the target video based on the unit fluency.
[0108] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the motion fluency mapping relationship includes a mapping relationship between a motion vector and a fluency quantification value, and a larger fluency quantification value means higher fluency.
[0109] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the unit fluency unit specifically includes: determining a first fluency quantification value corresponding to a lens motion vector of the different video frames and a second fluency quantification value corresponding to an object motion vector of the different video frames based on the motion vectors between the different video frames and the motion fluency mapping relationship; and determining the sum of the first fluency quantification value and the second fluency quantification value as a unit fluency quantification value representative of the unit fluency.
[0110] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the unit fluency unit specifically includes: weighting the unit fluency quantification values of the different video frames to obtain a target fluency quantification value; and verifying the fluency of the target video based on the quantified value of the target fluency.
[0111] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the motion fluency mapping relationship further includes a mapping relationship between the fluency quantification value and a fluent frame rate, and the fluent frame rate is used to indicate the frame rate required for videos of different fluency quantification values to achieve visual fluency.
[0112] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the unit fluency unit specifically includes: determining the target fluency quantification as the fluency of the target video; Alternatively, the fluency of the target video may be confirmed based on a comparison result between a fluent frame rate corresponding to the quantified value of the target fluency and a current frame rate of the target video.
[0113] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the unit fluency unit specifically includes: If the current frame rate of the target video is smaller than the fluent frame rate corresponding to the quantification value of the target fluency, it is used to confirm that the fluency of the target video is not fluent; otherwise, it is used to confirm that the fluency of the target video is fluent.
[0114] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the motion estimation module specifically comprises: extracting a plurality of video frames of the target video; determining motion vectors of a plurality of image blocks between different video frames to obtain a plurality of motion vectors, each of the video frames including a plurality of the image blocks; clustering the plurality of motion vectors to obtain a plurality of types of vector sets; and determining a motion vector between the different video frames based on the plurality of vector sets.
[0115] According to one or more embodiments of the present disclosure, in the video fluency verification device provided by the present disclosure, the motion estimation module specifically comprises: determining a vector set having the largest number of motion vectors from among the plurality of types of vector sets as a first set, and determining sets other than the first set as second sets; determining an average value of the motion vectors in the first set as a lens motion vector among the motion vectors between the different video frames; The average value of the motion vectors in the second set is used to determine the object motion vector among the motion vectors between the different video frames.
[0116] According to one or more embodiments of the present disclosure, in the video fluency verification apparatus provided by the present disclosure, the different video frames are adjacent video frames.
[0117] According to one or more embodiments of the present disclosure, the present disclosure provides a method for manufacturing a semiconductor device, comprising: a processor; a memory for storing instructions executable by said processor; The processor provides an electronic device that can be used to read the executable instructions from the memory and execute the instructions to implement any of the video fluency verification methods provided by the present disclosure.
[0118] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having stored thereon a computer program for performing any of the video fluency verification methods provided by the present disclosure.
[0119] According to one or more embodiments of the present disclosure, the present disclosure provides a computer program product including computer programs / instructions that, when executed by a processor, implement any of the video fluency verification methods provided by the present disclosure.
[0120] The above description is merely a description of the preferred embodiments of the present disclosure and the technical principles used. As can be understood by those skilled in the art, the disclosure scope of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also include other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed idea, for example, by replacing the above features with technical features having similar functions disclosed in the present disclosure (but not limited to these).
[0121] In addition, although operations are described in a particular order, this should not be construed as requiring these operations to be performed in the particular order or sequence shown. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although the above discussion includes some specific implementation details, these should not be construed as limitations on the scope of the disclosure. Some features that are described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.
[0122] Although the present subject matter has been described in language specific to structural features and / or logical operations of methods, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the specific features and operations described above are disclosed as example forms for implementing the claims.
Claims
1. A method for checking video fluency, comprising: obtaining a target video; determining motion vectors between different video frames by performing motion estimation on the different video frames in the target video, the motion vectors including a lens motion vector used to indicate a change in position of a lens between the different video frames and an object motion vector used to indicate a change in position of the same photographed object in the different video frames; confirming the fluency of the target video based on the motion vectors between different video frames, wherein the fluency of the target video is used to indicate the continuity and stability of images of the target video during playback of the target video; determining motion vectors between different video frames by performing motion estimation on the different video frames in the target video, extracting a plurality of video frames of the target video; determining motion vectors of a plurality of image blocks between different video frames to obtain a plurality of motion vectors, each of the video frames including a plurality of the image blocks; clustering the plurality of motion vectors to obtain a plurality of types of vector sets; determining a motion vector between the different video frames based on the plurality of vector sets; determining the motion vectors between the different video frames based on the plurality of vector sets, determining a vector set having the largest number of motion vectors from among the plurality of types of vector sets as a first set, and determining sets other than the first set as second sets; determining an average value of the motion vectors in the first set as a lens motion vector among the motion vectors between the different video frames; determining an average value of the motion vectors in the second set as an object motion vector among the motion vectors between the different video frames.
2. The step of checking the fluency of the target video based on the motion vectors between different video frames includes: determining unit fluency of the different video frames based on the motion vectors between the different video frames and a pre-established motion fluency mapping relationship; and verifying the fluency of the target video based on the unit fluency.
3. 3. The method of claim 2, wherein the motion fluency mapping relationship includes a mapping relationship between a motion vector and a fluency quantification value, and a larger fluency quantification value means a higher fluency.
4. determining unit fluency of the different video frames based on the motion vectors between the different video frames and the pre-established motion fluency mapping relationship, determining a first fluency quantification value corresponding to a lens motion vector of the different video frames and a second fluency quantification value corresponding to an object motion vector of the different video frames based on the motion vectors between the different video frames and the motion fluency mapping relationship; determining a sum of the first fluency quantification value and the second fluency quantification value as a unit fluency quantification value representative of the unit fluency.
5. The step of checking the fluency of the target video based on the unit fluency includes: weighted averaging the unit fluency quantification values of the different video frames to obtain a target fluency quantification value; and verifying the fluency of the target video based on the quantified value of the target fluency.
6. 6. The method of claim 5, wherein the motion fluency mapping relationship further includes a mapping relationship between the fluency quantification value and a fluent frame rate, and the fluent frame rate is used to indicate the frame rate required for videos of different fluency quantification values to achieve visual fluency.
7. confirming the fluency of the target video based on the quantification value of the target fluency, determining the target fluency quantification as the fluency of the target video; Alternatively, the method of claim 6 further comprises a step of confirming the fluency of the target video based on a comparison result between a fluent frame rate corresponding to the quantified value of the target fluency and a current frame rate of the target video.
8. confirming the fluency of the target video based on a comparison result between a fluency frame rate corresponding to the quantification value of the target fluency and a current frame rate of the target video, 8. The method of claim 7, further comprising: if the current frame rate of the target video is less than a fluent frame rate corresponding to the quantification value of the target fluency, confirming that the fluency of the target video is not fluent; otherwise, confirming that the fluency of the target video is fluent.
9. 2. The method of claim 1, wherein the different video frames are adjacent video frames.
10. 1. A video fluency verification device, comprising: a video acquisition module for acquiring a target video; a motion estimation module for performing motion estimation on different video frames in the target video to determine motion vectors between the different video frames, the motion vectors including a lens motion vector used to indicate a change in position of a lens between the different video frames and an object motion vector used to indicate a change in position of a same photographed object in the different video frames; a fluency module for checking the fluency of the target video based on the motion vectors between different video frames, wherein the fluency of the target video is used to indicate the continuity and stability of images of the target video during playback of the target video; When the motion estimation module performs the step of determining motion vectors between different video frames in the target video by performing motion estimation on the different video frames, extracting a plurality of video frames of the target video; determining motion vectors of a plurality of image blocks between different video frames to obtain a plurality of motion vectors, each of the video frames including a plurality of the image blocks; clustering the plurality of motion vectors to obtain a plurality of types of vector sets; determining motion vectors between the different video frames based on the plurality of vector sets; When the motion estimation module performs the step of determining the motion vectors between the different video frames based on the plurality of vector sets, determining a vector set having the largest number of motion vectors from among the plurality of types of vector sets as a first set, and determining sets other than the first set as second sets; determining an average value of the motion vectors in the first set as a lens motion vector among the motion vectors between the different video frames; determining an average value of the motion vectors in the second set as an object motion vector among the motion vectors between the different video frames.
11. a processor; a memory for storing instructions executable by said processor; The processor is used to realize the video fluency checking method described in any one of claims 1 to 9 by reading the executable instructions from the memory and executing the instructions.
12. A computer-readable storage medium storing a computer program for executing the method for checking video fluency according to any one of claims 1 to 9.
13. A computer program comprising computer program instructions which, when executed by a processor, implements the method for checking video fluency according to any one of claims 1 to 9.
Citation Information
Patent Citations
Photographic apparatus and photographing method
JP2006157428A
Frame rate conversion method, frame rate conversion device and frame rate conversion program
JP2011049633A
Image processing apparatus, image capturing apparatus, and computer program
JP2013165485A
Image processing device, control method therefor, and image capturing device
JP2020095673A
Image pickup apparatus, control method therefor, program, and storage medium
JP2020150448A