Stabilize video
Through the computing system, the camera movement transformation is identified and modified to generate stable video frames, which solves the problem of video quality degradation caused by jitter by the handheld video recording device, and realizes efficient and stable video streaming without increasing cost and size.
Patent Information
- Application Number
- CN202110655510.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-10-14
- Filing Date
- 2016-09-23
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2036-09-23
Smart Images

Figure CN113344816B_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of Chinese invention patent application No. 201680039070.5, filed on September 23, 2016. Technical Field
[0003] This article generally relates to stabilizing video. Background Art
[0004] Video recording was once the domain of dedicated video recording devices, but it is more common to find everyday devices (such as cell phones and tablet computers) that can record video. The problem with most handheld recording devices is that they suffer from video jitter, where unconscious movements of the user while holding the recording device affect the quality of the video.
[0005] A shaky recording device may result in a shaky video unless, for example, the shakiness is compensated for by a video stabilization mechanism. Optical video stabilization can reduce the shakiness present in the video by mechanically moving components of the recording device, such as the lens or image sensor. However, optical video stabilization may increase the material and manufacturing costs of the recording device. Furthermore, optical video stabilization may increase the size of the recording device, and it is generally desirable to design recording devices smaller. Summary of the Invention
[0006] Described herein are techniques, methods, systems, and other mechanisms for stabilizing video.
[0007] As additional descriptions to the embodiments described below, the present disclosure describes the following embodiments.
[0008] Embodiment 1 is a computer-implemented method. The method includes receiving, by a computing system, a first frame and a second frame of a video captured by a recording device (such as a camera). The method includes identifying, by the computing system, a mathematical transformation using the first and second frames of the video, the mathematical transformation indicating movement of the camera relative to a scene captured by the video (i.e., a scene represented in the video) from the time the first frame was captured to the time the second frame was captured. The method includes generating, by the computing system, a modified mathematical transformation by modifying the mathematical transformation indicating movement of the camera relative to the scene so that the mathematical transformation is less representative of the most recently initiated movement. The method includes generating, by the computing system, a second mathematical transformation that can be applied to the second frame using the mathematical transformation and the modified mathematical transformation to stabilize the second frame. The method includes identifying, by the computing system, an expected distortion that will be present in a stabilized version of the second frame resulting from applying the second mathematical transformation to the second frame based on a difference between: (i) an amount of distortion in a horizontal direction resulting from applying the second mathematical transformation to the second frame, and (ii) an amount of distortion in a vertical direction resulting from applying the second mathematical transformation to the second frame. The method includes determining, by a computing system, an amount to reduce a stabilizing effect of applying a second mathematical transform to a second frame based on a degree to which the expected distortion exceeds an acceptable distortion variation calculated from distortion in a plurality of frames of the video preceding the second frame. The method includes generating, by the computing system, a stabilized version of the second frame by applying the second mathematical transform to the second frame, wherein a stabilizing effect of applying the second mathematical transform to the second frame has been reduced based on the determined amount to reduce the stabilizing effect.
[0009] Embodiment 2 is a method according to embodiment 1, wherein the second frame is a frame of the video that immediately follows the first frame of the video.
[0010] Embodiment 3 is a method according to embodiment 1, wherein the mathematical transformation indicative of movement of the camera comprises a homography transformation matrix.
[0011] Embodiment 4 is a method according to embodiment 3, wherein modifying the mathematical transformation comprises applying a low-pass filter to the homography transformation matrix.
[0012] Embodiment 5 is a method according to embodiment 3, wherein the expected distortion is based on a difference between a horizontal scaling value in the second mathematical transform and a vertical scaling value in the second mathematical transform.
[0013] Embodiment 6 is a method according to embodiment 1, wherein modifying the mathematical transform comprises modifying the mathematical transform so that the modified mathematical transform is more representative of movement that has occurred over a long period of time than the mathematical transform.
[0014] Example 7 is a method according to Example 1, wherein determining the amount of reducing the stabilization effect produced by applying the second mathematical transformation to the second frame is further based on a determined movement speed of the camera from the first frame to the second frame, which exceeds an acceptable change in movement speed of the camera calculated based on the movement speed of the camera between multiple frames of the video preceding the second frame.
[0015] Embodiment 8 is a method according to embodiment 1, wherein generating the stabilized version of the second frame comprises scaling to a version of the second frame generated by applying a second mathematical transform to the second frame.
[0016] Embodiment 9 is a method according to embodiment 1, wherein the operation further comprises: moving the zoomed-in area of the version of the second frame horizontally or vertically to avoid the zoomed-in area of the second frame presenting an invalid area.
[0017] Embodiment 10 is directed to a system including a recordable medium storing instructions that, when executed by one or more processors, cause the operations of the method according to any one of embodiments 1 to 9 to be performed.
[0018] In some cases, certain embodiments may achieve one or more of the following advantages. The video stabilization techniques described herein can, for example, compensate for more than two degrees of freedom (e.g., not just horizontal and vertical) by compensating for eight degrees of freedom (e.g., translation, rotation, scaling, and non-rigid rolling shutter distortion). The video stabilization techniques described herein can operate while video is being captured by a device and may not require information from future frames. In other words, the video stabilization techniques can be able to stabilize the most recently recorded frames using only information from past frames, thereby enabling the system to store a stabilized video stream as it is captured (e.g., without storing multiple unstable video streams, such as not storing more than 1, 100, 500, 1000, or 5000 unstable video frames in a video that is currently being recorded or has already been recorded). Thus, the system may not need to wait for the video to stabilize until the entire video has been recorded. The described video stabilization techniques may have low complexity and therefore can run on devices with moderate processing power (e.g., some smartphones). Furthermore, the video stabilization techniques described herein can operate even when frame-to-frame motion estimation fails in the first step.
[0019] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1A diagram showing a video stream being stabilized by a video stabilization process.
[0021] Figures 2A to 2B A flow chart showing a process for stabilizing video is shown.
[0022] Figure 3 is a block diagram of a computing device (either as a client or as a server or multiple servers) that can be used to implement the systems and methods described herein.
[0023] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION
[0024] Stabilizing video is generally described herein. Video stabilization can be performed by identifying a transformation between a most recently received video frame and a previously received video frame (wherein the transformation indicates frame-to-frame movement of a camera relative to a scene), modifying the transformation based on information from past frames, generating a second transformation based on the transformation and the modified transformation, and applying the second transformation to a currently received frame to generate a stabilized version of the currently received frame. Generally referring to Figure 1 2 for a more detailed description.
[0025] Figure 1 A diagram of a video stream being stabilized by a video stabilization process is shown. The diagram includes three frames 110a through 110c of the video. The frames may be consecutive, such that frame 110b may be the frame captured immediately after frame 110a is captured, and frame 110c may be the frame captured immediately after 110b is captured. Two frames of a video may sometimes be referred to herein as a first frame of the video and a second frame of the video, but the "first" designation does not necessarily mean that the first frame is the initial frame of the entire video.
[0026] Frames 110a through 110c are shown positioned between or near lines 112a through 112b, which indicate the positions of the scenes represented by the frames relative to each other. The lines are provided in this figure to illustrate that the camera was moving when it captured frames 110a through 110c. For example, the camera was pointing further downward when capturing frame 110b than when capturing frame 110a, and was pointing further upward when capturing frame 110c than when capturing frames 110a through 110b.
[0027] The computing system identifies a mathematical transformation (block 120) that indicates the movement of the camera from the first frame 110b to the second frame 110c. This identification can be performed using frames 110b to 110c (as indicated by the arrows in the figure), where frame 110c can be the most recently captured frame. The two frames 110b to 110c can be received from a camera sensor or camera module attached to the computing system, or can be received from a remote device that captures the video frames 110b to 110c. The identification of the mathematical transformation can include generating the mathematical transformation. The mathematical transformation can be a homography transformation matrix, as described in additional detail with reference to block 210 in Figures 2A to 2B.
[0028] The computing system then creates a modified transform (block 125) by modifying the initial transform (block 120) such that the modified transform (block 125) is less representative of recently initiated motion than the initial transform. In other words, generating the modified transform includes modifying the initial transform based on information from one or more video frames preceding the first frame 110b and the second frame 110c. The modified transform can be a low-pass filtered version of the initial transform. Doing so results in a modified mathematical transform (block 125) that is more representative of motion that has occurred over a long period of time, as opposed to recently initiated motion, than the initial transform (block 120).
[0029] As an example, the modified mathematical transform (block 125) may be more representative of a panning motion that has occurred over many seconds than an oscillation that began a fraction of a second ago. Modifying the transform in this way takes into account previous frames of the video, such as Figure 1 110b to box 122. For example, the transformation calculated using the previous frame can be used to identify which movements have occurred over a longer period of time and which movements have started more recently. An example way to use the previous frame to calculate the modified transformation can be to apply a low pass filter to the homography transformation matrix, as shown in FIG. Figures 2A to 2B Additional details of block 220 are described in detail in FIG.
[0030] Block 130 shows a second transform generated from the initial transform (block 120) and the modified transform (block 125). The second transform may be the difference between the initial transform and the modified transform. Figures 2A to 2B Block 230 in FIG. 1 , generating a second transform based on the initial transform and the modified transform is described in additional detail.
[0031] Block 132 illustrates how the computing system identifies expected distortion in the stabilized version of the second frame 110c resulting from applying the second mathematical transform to the second frame 110c based on a difference between (i) the amount of distortion in the horizontal direction resulting from applying the second mathematical transform to the second frame 110c and (ii) the amount of distortion in the vertical direction resulting from applying the second mathematical transform to the second frame 110c. Figures 2A to 2B Block 250 in , calculating the expected distortion is described in additional detail.
[0032] Block 134 illustrates how the computing system reduces the amount of stabilization effect produced by applying the second mathematical transform to the second frame 110c based on the extent to which the expected distortion exceeds an acceptable distortion variation. The acceptable distortion variation may be calculated using multiple frames of the video preceding the second frame 110c, such as Figure 1 2 , the amount of video stabilization to be reduced is described in additional detail. Using the determined amount of video stabilization to reduce the amount of video stabilization by the computing system may include using the determined amount to generate a modified second transform (block 140).
[0033] The computing system generates a stabilized version of the second frame 110 c (block 150) by applying the modified second transform (block 140) to the second frame 110 c. Because the modified second transform (block 140) has been modified based on the determined amount of reduced stabilization, the stabilized version of the second frame generated by the computing system (block 150) is considered to have been reduced based on the determined amount of reduced stabilization.
[0034] In some embodiments, determining the amount to reduce stabilization is further or alternatively based on determining that the speed of movement of the camera relative to the scene from the first frame 110b to the second frame 110c exceeds an acceptable speed change of the camera relative to the scene. The acceptable speed change of the camera can be calculated based on a plurality of frames of the video preceding the second frame 110c, as described with reference to FIG. Figures 2A to 2B 240 and 260 in more detail.
[0035] In some embodiments, generating a stable version of the second frame includes scaling to a version of the second frame generated by applying a second mathematical transformation to the second frame. The computing system may shift the magnified region horizontally, vertically, or both horizontally and vertically to avoid an invalid region of the magnified region that may appear at an edge of the stable second frame. Figures 2A to 2BThis process is described in more detail at blocks 280 and 290 in FIG.
[0036] Figures 2A to 2B A flow chart of a process for stabilizing video is shown. The process is represented by blocks 210 to 290 described below. Figures 2A to 2B The operations described in conjunction with these blocks are performed in the order shown.
[0037] In block 210, the computing system uses the two video frames as input to estimate a matrix representing frame-to-frame motion ("H_interframe"). This frame-to-frame motion matrix can be a homography matrix. A homography matrix can be a matrix that can represent the movement of a scene or a camera that is capturing a scene between two frames of video. As an example, each frame of a video can display a two-dimensional image. Suppose the first frame takes a picture of a square from directly in front of it so that the square has equal-length sides at ninety-degree angles in the video frame (in other words, a square appears). Now suppose the camera is moved to one side (or the square itself is moved) so that the next frame of the video shows the square skewed, with some sides longer than others and the angles of the square not being ninety degrees. The positions of the four corner points of the square in the first frame can be mapped to the positions of the four corner points in the second frame to identify how the camera or scene moves from one frame to the next.
[0038] The mapping of these corner points to each other in the frames can be used to generate a homography that represents the motion of the camera viewpoint relative to the scene being recorded. Given this homography, the first frame and the generated homography can be used to recreate the second frame, for example, by moving pixels in the first frame to different locations according to known homography methods.
[0039] The homography matrix described above can represent not only translational motion, but also rotation, scaling, and non-rigid rolling shutter distortion. In this way, the application of the homography matrix can be used to stabilize video associated with eight degrees of freedom. For comparison purposes, some video stabilization mechanisms only stabilize images to account for translational motion (e.g., up / down and left / right).
[0040] Although other types of homography matrices can be used (and other mathematical representations of frame-to-frame movement can be used, even if not homography matrices or even if not matrices), the above-mentioned homography transformation matrix can be a 3x3 homography transformation matrix. The 3x3 matrix (called H_interfame) can be determined as follows. First, the computing system finds a set of feature points (typically corner points) in the current image, where these points are represented as [x'_i, y'_i], i=1...N (N is the number of feature points). Then, the corresponding feature points in the previous frame are found, where these corresponding feature points are represented as [x_i, y_i]. It should be noted that these points are described as being in the GL coordinate system (i.e., x and y range from -1 to 1 and have the center of the frame as the origin). If these points are in the image pixel coordinate system where x ranges from 0 to the image width and y ranges from 0 to the image height, then these points can be transformed to the GL coordinate system or the resulting matrix can be transformed to compensate.
[0041] The H_interfame matrix above is a 3x3 matrix containing 9 elements:
[0042]
[0043] H_interfame is the transformation matrix that transforms [x_i, y_i] to [x'_i, y'_i], as described below.
[0044] z_i'*[x'_i,y'_i,1]'=H_interframe*[x_i,y_i,1]'. [x'_i,y'_i,1]' is a 3x1 vector that is the transpose of the [x'_i,y'_i,1] vector. [x_i,y_i,1]' is a 3x1 vector that is the transpose of the [x_i,y_i,1] vector. z_i' is a scaling factor.
[0045] Given a set of corresponding feature points, an example algorithm for estimating this matrix is described at Algorithm 4.1 (page 91) and Algorithm 4.6 (page 123) in the following computer vision book: “Hartley, R., Zisserman, A.: Multiple View Geometry in Computer Vision. Cambridge University Press (2000),” available at ftp: / / vista.eng.tau.ac.il / dropbox / aviad / Hartley, %20Zisserman%20-%20Multiple%20View%20Geometry%20in%20Computer%20Vision.pdf.
[0046] In block 220, the computing system estimates a low-pass transformation matrix (H_lowpass). The low-pass transformation matrix can then be combined with the H_interframe matrix to generate a new matrix (H_compensation), which can be used to remove the results of unintentional "high-frequency" movement of the camera. If the system attempted to remove all movement (in other words, without performing the low-pass filtering described herein), the user may not be able to intentionally move the camera and cause the scene depicted by the video to also move. Therefore, the computing system generates a low-pass transform to filter out high-frequency movement. High-frequency movement can be movement that is irregular and not represented by multiple frames, such as movement back and forth over a short period of time. In contrast, low-frequency movement can be movement that is represented by multiple frames, such as a user panning the camera for multiple seconds.
[0047] To perform this filtering, the computing system generates a low-pass transform matrix (H_lowpass) that includes weighted values to emphasize low-frequency motion that has occurred over a long time series. The low-pass transform matrix can be the result of applying a low-pass filter to the H_interframe matrix. Each element in the low-pass transform matrix is generated element by element based on (1) its own time series of low-pass transform matrices from the previous frame, (2) the H_interframe matrix representing the motion between the previous frame and the current frame, and (3) an attenuation ratio specified by the user. In other words, the elements in the matrix that are weighted with significant values can be those that represent motion that has been presented in the H_interframe matrix over multiple frames. The equation for generating H_lowpass can be expressed as follows:
[0048] H_lowpass = H_previous_lowpass * transform_damping_ratio + H_interframe * (1 - transform_damping_ratio) This equation is an example of a two-tap infinite impulse response filter.
[0049] In block 230, the computing system calculates a compensation transformation matrix (H_compensation). The compensation matrix can be a combination of a lowpass matrix (H_lowpass) and a frame-to-frame motion matrix (H_interframe). These two matrices are combined to generate a matrix (H_compensation) that needs to maintain motion from one frame to another, but only motion that has occurred within a reasonable period of time, not including recent "unintentional" motion. The H_compensation matrix can represent the difference in motion between H_interframe and H_lowpass, so that H_lowpass can be applied to the last frame to generate a modified version of the last frame that represents the intentional motion that occurred between the last frame and the current frame, while H_compensation can be applied to the current frame to generate a modified (and stabilized) version of the current frame that represents the intentional motion that occurred between the last frame and the current frame. Roughly speaking, applying H_compensation to the current frame removes unintentional motion from that frame. Specifically, given this calculated H_compensation matrix, the system should be able to take the current frame of the video, apply the H_compensation matrix to the current frame using the transformation process, and obtain a newly generated frame that is similar to the current frame, but excludes any such sudden and small movements. In other words, the system tries to keep the current frame as close as possible to the last frame, but allows for long-term "intentional" movements.
[0050] The compensation transformation matrix can be generated using the following equation:
[0051] ·H_compensation=Normalize(Normalize(Normalize(H_lowpass)*H_previous_compensation)*Inverse(H_interframe))
[0052] The H_previous_compensation matrix is the H_constrained_compensation matrix calculated after this process, but calculated for the previous frame. Inverse() is a matrix inversion operation used to generate the original version of the last frame by inverting the transformation. Combining the original version of the last frame with the low-pass filter matrix enables intentional movement. Combining it with H_previous_compensation compensates for the previous compensation value.
[0053] Normalize() is an operation that normalizes a 3x3 matrix by its second singular value. The normalization process is performed because some steps of the process may result in transformations that do not make much sense in the real world. Therefore, the normalization process can ensure that reasonable results are obtained from each process step. Normalization is performed for each step of the process so that an odd output from one step does not contaminate the remaining steps of the process (for example, if the odd output provides a value close to zero, then this value close to zero will pull the outputs of the remaining steps to close to zero as well). For reasons described below, additional processing can enhance the results of the video stabilization process.
[0054] In block 240, the computing system calculates a speed reduction value. The speed reduction value can be a value used to determine how much to reduce video stabilization when the camera is moving quickly, and because frame-to-frame motion can be unreliable, video stabilization becomes undesirable. To calculate the amount by which video stabilization can be reduced, the inter-frame motion speed is first calculated. In this example, the computing system generates the speed at the center of the frame. The speed in the x-direction is pulled from the element in row 1, column 3 of H_interframe as follows (block 242):
[0055] ·speed_x=H_interframe[1,3]*aspect_ratio
[0056] The velocity in the y direction is extracted from the element at row 2, column 3 in H_interframe as follows (block 242):
[0057] speed_y = H_interframe[2,3]
[0058] The aspect_ratio described above is frame_width divided by frame_height.These identifications of velocity may only consider translational movement between two frames, but in other examples, velocity may consider rotation, scaling, or other types of movement.
[0059] The system can then determine a low-pass motion velocity that accounts for the long-term velocity of the camera (or scene) and excludes sudden and rapid "unintentional" movements. This is done by taking the current velocity and combining it with the previously calculated low-pass velocity, and further applying a falloff ratio that inversely weights the current velocity relative to the previously calculated low-pass velocity, for example, as follows:
[0060] ·lowpass_speed_x=lowpass_speed_x_previous*speed_damping_ratio+speed_x*(1-speed_damping_ratio)
[0061] This equation actually generates the lowpass speed by taking the previously calculated speed and reducing it by the amount specified by the decay ratio. This reduction is compensated by the current speed. In this way, the current speed of the video affects the overall lowpass_speed value, but it is not the only factor that affects the lowpass_speed value. The above equation represents an infinite impulse response filter.
[0062] The same process can be performed for the y velocity to generate a low-pass y velocity, for example, as follows:
[0063] ·lowpass_speed_y=lowpass_speed_y_previous*speed_damping_ratio+speed_y*(1-speed_damping_ratio)
[0064] The decay ratio in this process is set by the user, and an example value is 0.99.
[0065] The process then combines these values to generate a single representation of the low-pass velocity that caused movement in the x and y directions, for example, using the following equations (block 244):
[0066] ·lowpass_speed=sqrt(lowpass_speed_x*lowpass_speed_x+lowpass_speed_y*lowpass_speed_y)
[0067] The lowpass speed calculated essentially represents the long-term speed of movement between frames. In other words, lowpass_speed has less influence on recent speed changes and more influence on longer-term speed trends.
[0068] After calculating the low-pass speed, the system can calculate the speed reduction value. In some examples, the speed reduction value is a value between 0 and 1 (which could also be other boundary values), and the system can generate the speed reduction value based on how the low-pass speed compares to a low threshold and a high threshold. If the low-pass speed is below the low threshold, the speed reduction value can be set to the boundary value 0. If the low-pass speed is above the low threshold, the speed reduction value can be set to the boundary value 1. If the low-pass speed is between the two thresholds, the calculation system can select a speed reduction value between the thresholds that represents the scaled value of the low-pass speed, e.g., the speed reduction value is between the boundary values 0 and 1. The calculation of the speed decay value can be represented by the following algorithm (block 246):
[0069] · If lowpass_speed < low_speed_threshold, then speed_reduction = 0;
[0070] · And if lowpass_speed > high_speed_threshold, then speed_reduction = max_speed_reduction;
[0071] · Otherwise, speed_reduction = max_speed_reduction * (lowpass_speed - low_speed_threshold) / (high_speed_threshold - low_speed_threshold).
[0072] Using this algorithm, low_speed_threshold, high_speed_threshold, and max_speed_reduction are all specified by the user. Example values include: low_speed_threshold = 0.008; high_speed_threshold = 0.016; and max_speed_reduction = 1.0.
[0073] In block 250, the calculation system calculates the distortion reduction value. The calculation system can calculate the distortion reduction value because the compensation transform may produce too much non-rigid distortion when applied to a video frame. In other words, the video stabilization may appear unrealistic, e.g., because the distortion caused by stretching the image more in one direction than in another may occur too quickly and may seem unusual to the user.
[0074] To calculate the distortion reduction value, the computing system may first calculate the compensation scaling factor by looking at the value of the scaling factor in the H_compensation matrix as follows:
[0075] zoom_x = H_compensation[1,1], which is the element in the first row and first column of the H_compensation matrix.
[0076] zoom_y = H_compensation[2,2], which is the element in the 2nd row and 2nd column of the H_compensation matrix.
[0077] Scaling factors may be those that identify how the transformation stretches the image in one dimension.
[0078] The computing system may then determine the difference between the two scaling factors to determine the extent to which the image is distorted by stretching more in one direction than the other, as follows (block 252):
[0079] ·distortion=abs(zoom_x-zoom_y)
[0080] A low-pass filter is applied to the distortion to limit the rate at which the distortion can change and thus ensure that abrupt changes in the distortion are minimized using the following formula (block 254):
[0081] ·lowpass_distortion=previous_lowpass_distortion*distortion_damping_ratio+distortion*(1-distortion_damping_ratio)
[0082] In other words, the algorithm is arranged so that the amount of distortion can change slowly. In the above formula, distortion_damping_ratio is the attenuation ratio of the distortion IIR filter specified by the user. An example value is 0.99.
[0083] After calculating the low-pass distortion, the computing system can calculate a distortion reduction value. In some examples, the distortion reduction value is a value between 0 and 1 (which could also be other boundary values), and the system can generate the distortion reduction value based on how the low-pass distortion compares to a low threshold and a high threshold. If the low-pass distortion is below the low threshold, the distortion reduction value can be set to the boundary value 0. If the low-pass distortion is above the low threshold, the distortion reduction value can be set to the boundary value 1. If the low-pass distortion is between the two thresholds, a scaled value representing the low-pass distortion can be selected between the thresholds (e.g., the resulting distortion reduction value is between the boundary values 0 and 1). The calculation of the distortion reduction value can be represented by the following algorithm (block 256):
[0084] · If lowpass_distortion < low_distortion_threshold, then distortion_reduction = 0;
[0085] · And if lowpass_distortion > high_distortion_threshold, then max_distortion_reduction;
[0086] · Otherwise, distortion_reduction = max_distortion_reduction * (lowpass_distortion - low_distortion_threshold) / (high_distortion_threshold - low_distortion_threshold).
[0087] Using this algorithm, low_distortion_threshold, high_distortion_threshold, and max_distortion_reduction are all specified by the user. Example values include: low_distortion_threshold = 0.001; high_distortion_threshold = 0.01; and max_distortion_reduction = 0.3.
[0088] In block 260, the computing system reduces the intensity of video stabilization based on the determined speed reduction value and the distortion reduction value. To do this, the computing system calculates a reduction value, which in this example is identified as the maximum of the speed reduction value and the distortion reduction value (block 262), as follows:
[0089] ·reduction=max(speed_reduction,distortion_reduction)
[0090] In other examples, the reduction value can be a combination of the two values that results in a fraction of each value (e.g., the values can be added or multiplied together and then possibly multiplied by a predetermined number such as 0.5). The reduction value can be less than or between boundary values 0 and 1, and the closer the reduction value is to 1, the more the computing system can reduce the intensity of image stabilization.
[0091] The computing system can then modify the compensation transformation matrix to generate a reduced compensation transformation matrix (block 264). The computing system can do this by multiplying the compensation transformation matrix by a value obtained by subtracting the reduced value from 1. In other words, if the reduced value is very close to 1 (indicating that image stabilization will be greatly reduced), the values in the compensation matrix may be significantly reduced because they will be multiplied by a number close to zero. The numbers in the modified compensation transformation matrix are then added to the identity matrix, which has been multiplied by the reduced value. Example equations are as follows:
[0092] ·H_reduced_compensation=Identity*reduction+H_compensation*(1-reduction)
[0093] In block 270, the computing system may limit compensation so that the resulting video stabilization does not display invalid regions of the output frame (e.g., those regions outside the frame). As some background, because compensation may distort an image generated by the image stabilization process, the image may display invalid regions substantially outside the image at its boundaries. To ensure that these invalid regions are not displayed, the computing system may zoom in on the image to crop off the edges of the image that may include the invalid regions.
[0094] Returning to the process of box 270, if the camera is moved quickly and significantly, the stabilization can be locked to displaying the old position because the rapid and significant movement can be filtered out, which can introduce the above-mentioned invalid area into the display of the stabilized frame. In this case, if the video is about to display an invalid area, the limiting process described below can ensure that the video stabilization essentially stops having full control over the area in which the frame is displayed. This determination as to whether stabilization needs to give up some control over the frame can be initiated by first setting the corner points of the output image and determining whether these corner points fall outside a pre-specified cropping area. The maximum amount of compensation and scaling can be limited to twice the crop ratio, where the crop ratio can be specified by the user (e.g., crop 15% on each side, or crop 0.15 in the following equation):
[0095] ·max_compensation=cropping_ratio*2
[0096] The computing system can then use the H_reduced_compensation matrix to transform the four corners of the unit square in GL coordinates (x01, y01) = (-1, -1), (x02, y02) = (1, -1), (x03, y03) = (-1, 1), (x04, y04) = (1, 1) into the four corner points (x1, y1), (x2, y2), (x3, y3), (x4, y4). (Note that the video frame does not need to be a unit square, but the size of the unit frame is mapped to the unit square in GL coordinates.) More specifically, we use the following formula to transform (x0i, y0i) to (xi, yi):
[0097] ·dzi*[xi,yi,1]'=H_reduced_compensation*[x0i,y0i,1]'
[0098] In this example, [x0i,y0i,1]' is a 3x1 vector that is the transpose of the [x0i,y0i,1] vector. [xi,yi,1]' is a 3x1 vector that is the transpose of the [xi,yi,1] vector. zi is the scaling factor.
[0099] The computing system can then identify the maximum amount of displacement in each direction (left, right, up, and down) from the corner of each transformed video frame to the side of the unit square as follows:
[0100] ·max_left_displacement=1+max(x1,x3)
[0101] ·max_right_displacement=1-min(x2,x4)
[0102] ·max_top_displacement=1+max(y1,y2)
[0103] ·max_bottom_displacement=1-min(y3,y4)
[0104] If any identified displacement exceeds the maximum compensation amount (the maximum compensation amount is twice the crop ratio described above and would indicate that the invalid area is within the display area of the unit square's magnified area), then the corners of the frame are shifted the same amount away from the sides of the unit square so that the invalid area is not displayed. The equation for shifting the corners is as follows:
[0105] If max_left_displacement>max_compensation, shift the four corner points to the left by max_left_displacement-max_compensation;
[0106] If max_right_displacement>max_compensation, shift the four corner points to the right by max_right_displacement-max_compensation;
[0107] If max_top_displacement>max_compensation, shift the four corner points upward by max_top_displacement-max_compensation;
[0108] If max_bottom_displacement > max_compensation, shift the 4 corner points downward by max_bottom_displacement - max_compensation.
[0109] Shifting these corner points will display the identification of invalid areas even if the display is cropped (block 272).
[0110] After all the above shift operations, the four new corner points can be represented as (x1', y1'), (x2', y2'), (x3', y3'), (x4', y4'). The computing system then calculates the constrained compensation transformation matrix H_constrained_compensation, which maps the four corners of the unit square in GL coordinates (x01, y01) = (-1, -1), (x02, y02) = (1, -1), (x03, y03) = (-1, 1), (x04, y04) = (1, 1) to the four constrained corner points (x1', y1'), (x2', y2'), (x3', y3'), (x4', y4') as follows:
[0111] ·zi'*[xi',yi',1]'=H_constrained_compensation*[x0i,y0i,1]'
[0112] In this example, [x0i,y0i,1]' is a 3x1 vector that is the transpose of the [x0i,y0i,1] vector. [xi',yi',1]' is a 3x1 vector that is the transpose of the [xi',yi',1] vector. zi is a scaling factor. Given 4 pairs of points [x0i,y0i,1]' and [xi',yi',1]', an example algorithm for estimating the matrix is described at Algorithm 4.1 (page 91) in the following computer vision book: "Hartley, R., Zisserman, A.: Multiple View Geometry in Computer Vision. Cambridge University Press (2000)," available at ftp: / / vista.eng.tau.ac.il / dropbox / aviad / Hartley,%20Zisserman%20-%20Multiple%20View%20Geometry%20in%20Computer%20Vision.pdf
[0113] H_constrained_compensation is then saved as H_previous_compensation, which may be used in calculations to stabilize the next frame, as described above with reference to block 230 .
[0114] In block 280, the computing system modifies the constrained compensation matrix so that the stabilized image will be scaled to crop the border. In some examples, the computing system first identifies a scaling factor as follows:
[0115] ·zoom_factor=1 / (1-2*cropping_ratio)
[0116] Thus, the computing system doubles the crop ratio (e.g., by doubling 15% of the value of 0.15 to 0.3), subtracts the resulting value from 1 (e.g., to 0.7), and then divides this result by 1 to obtain the scaling factor (e.g., 1 divided by 0.7 equals a scaling factor of 1.42). The computing system can then divide certain features of the constrained compensation matrix to magnify the display by a certain amount, as follows:
[0117] ·H_constrained_compensation[3,1]=H_constrained_compensation[3,1] / zoom_factor
[0118] ·H_constrained_compensation[3,2]=H_constrained_compensation[3,2] / zoom_factor
[0119] ·H_constrained_compensation[3,3]=H_constrained_compensation[3,3] / zoom_factor
[0120] In block 290, the computing system applies the modified constrained compensation matrix to the current frame to generate a cropped and stabilized version of the current frame. An example of applying the constrained compensation matrix (H_constrained_compensation) to the input frame to generate an output image can be described as follows:
[0121] ·z'*[x',y',1]'=H_constrained_compensation*[x,y,1]'
[0122] [x,y,1]' is a 3x1 vector representing the coordinates in the input frame
[0123] [x',y',1]' is a 3x1 vector representing the coordinates in the output frame
[0124] z' is the scaling factor
[0125] H_constrained_compensation is a 3x3 matrix with 9 elements:
[0126]
[0127] In additional detail, for each pixel [x,y] in the input frame, the above transform is used to find the location [x',y'] in the output frame, and the pixel value from [x,y] in the input frame is copied to [x',y'] in the output frame. Alternatively, for each pixel [x',y'] in the output frame, the inverse transform is used to find the location [x,y] in the input frame, and the pixel value from [x,y] in the input frame is copied to [x',y'] in the output frame. These operations can be performed efficiently in a computing system graphics processing unit (GPT).
[0128] The process described herein with respect to blocks 210 through 290 may then be repeated for the next frame, using some of the values obtained from processing the current frame for the next frame.
[0129] In various embodiments, an operation performed "in response to" or "as a result of" another operation (e.g., determination or identification) is not performed if the previous operation was not successful (e.g., if a determination is not performed). An operation performed "automatically" is an operation performed without user intervention (e.g., intervening user input). Features of the present invention described in conditional language may describe optional embodiments. In some examples, "transmitting" from a first device to a second device includes: the first device placing data on a network for the second device to receive, but may not include the second device receiving the data. Conversely, "receiving" from a first device may include receiving data from a network, but may not include the first device transmitting the data.
[0130] "Determining" by a computing system may include the computing system requesting another device to perform a determination and provide the result to the computing system. Also, "displaying" or "presenting" by a computing system may include the computing system sending data for causing another device to display or present the referenced information.
[0131] In various embodiments, operations described as being performed on a matrix represent operations performed on the matrix or on a version of the matrix that has been modified by the operations described in this disclosure or their equivalents.
[0132] Figure 3 is a block diagram of computing devices 300 and 350 (either as a client or as a server or servers) that can be used to implement the systems and methods described herein. Computing device 300 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Computing device 350 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components shown here, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the embodiments of the invention described and / or claimed herein.
[0133] Computing device 300 includes a processor 302, memory 304, storage device 306, a high-speed interface 310 connected to memory 308 and high-speed expansion port 304, and a low-speed interface 314 connected to a low-speed bus 312 and storage device 306. Each component 302, 304, 306, 308, 310, and 312 is interconnected using a different bus and can be mounted on a common motherboard or interconnected in other ways as needed. Processor 302 can process instructions executed within computing device 300, including instructions stored in memory 304 or storage device 306 for displaying graphical information for a GUI on an external input / output device (such as a display 316 coupled to high-speed interface 308). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and different types of memory, if desired. Similarly, multiple computing devices 300 can be connected, each providing a portion of the necessary operations (e.g., as a server array, a group of blade servers, or a multi-processor system).
[0134] Memory 304 stores information within computing device 300. In one embodiment, memory 304 is one or more volatile memory units. In another embodiment, memory 304 is one or more non-volatile memory units. Memory 304 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0135] Storage device 306 can provide mass storage for computing device 300. In one embodiment, storage device 306 can be or include a computer-readable medium, such as a floppy disk drive, a hard disk drive, an optical disk drive, or a magnetic tape drive, a flash memory or other similar solid-state memory device, or an array of devices (including a storage area network or other configured devices). A computer program product can be tangibly embodied as an information carrier. A computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or machine-readable medium, such as memory 304, storage device 306, or memory on processor 302.
[0136] The high-speed controller 308 manages bandwidth-intensive operations of the computing device 300, while the low-speed controller 312 manages less bandwidth-intensive operations. This allocation of functions is merely an example. In one embodiment, the high-speed controller 308 is coupled to the memory 304, the display 316 (e.g., via a graphics processor or accelerator), and the high-speed expansion port 310, which can accept various expansion cards (not shown). In an embodiment, the low-speed controller 312 is coupled to the storage device 306 and the low-speed expansion port 314. The low-speed expansion port 314 may include various communication ports (e.g., USB, Bluetooth, Ethernet, and wireless Ethernet) and may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, or networking device (such as a switch or router) via a network adapter.
[0137] As shown, computing device 300 can be implemented in a variety of ways. For example, computing device 300 can be implemented as a standard server 320, or multiple times in a group of such servers. Computing device 300 can also be implemented as part of a rack server 324. In addition, computing device 300 can be implemented in a personal computer (such as laptop computer 322). Alternatively, components from computing device 300 can be combined with other components in a mobile device (not shown) (such as device 350). Each such device can contain one or more computing devices 300 and 350, and the entire system can be composed of multiple computing devices 300 and 350 communicating with each other.
[0138] Computing device 350 includes, among other components, a processor 352, a memory 364, an input / output device (such as a display 354), a communication interface 366, and a transceiver 368. Device 350 may also be provided with a storage device, such as a micro hard drive or other device, for providing additional storage. Each of components 350, 352, 364, 354, 366, and 368 are interconnected using various buses, and some components may be mounted on a common motherboard or interconnected in other ways as desired.
[0139] The processor 352 can execute instructions within the computing device 350, including instructions stored in the memory 364. The processor 352 can be implemented as a chipset including individual or multiple analog and digital processor chips. In addition, the processor can be implemented using any number of architectures. For example, the processor can be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimum Instruction Set Computer) processor. The processor can provide, for example, coordination of other components of the device 350, such as control of a user interface, applications executed by the device 350, and wireless communications conducted by the device 350.
[0140] The processor 352 can communicate with the user through a control interface 356 and a display interface 354 coupled to a display 358. The display 354 can be, for example, a TFT (thin film transistor liquid crystal display) display or an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 356 may include suitable circuitry for driving the display 354 to present graphics and other information to the user. The control interface 358 can receive commands from the user and convert the commands for submission to the processor 352. In addition, the external interface 362 can provide communication with the processor 352 so that the device 350 can communicate with other devices in the vicinity. In some embodiments, the external interface 362 can provide, for example, wired communication, or in some embodiments, wireless communication, and multiple interfaces can also be used.
[0141] Memory 364 stores information within computing device 350. Memory 364 can be implemented as one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Expansion memory 374 can also be provided and connected to device 350 via expansion interface 372, which can include, for example, a SIMM (Single Inline Memory Module) card interface. This expansion memory 374 can provide additional storage space for device 350 or can also store applications or other information for device 350. Specifically, expansion memory 374 can include instructions for executing or supplementing the processes described above and can also include security information. Thus, for example, expansion memory 374 can be provided as a security module for device 350 and can be programmed with instructions that allow for secure use of device 350. Furthermore, secure applications can be provided via a SIMM card along with additional information (such as identifying information placed on the SIMM card in a non-invasive manner).
[0142] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly embodied as an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or machine-readable medium, such as memory 364, expansion memory 374, or memory on processor 352, which may be received, for example, via transceiver 368 or external interface 362.
[0143] Device 350 can communicate wirelessly via a communication interface 366, which may include digital signal processing circuitry, if desired. Communication interface 366 can provide communication in various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS text messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication can occur, for example, via a radio frequency transceiver 368. In addition, short-range communication can occur using, for example, Bluetooth, WiFi, or other such transceivers (not shown). In addition, a global positioning system (GPS) receiver module 370 can provide additional navigation- or location-related wireless data to device 350, which can be used by applications running on device 350, if appropriate.
[0144] Device 350 may also communicate audibly using audio codec 360, which may receive spoken information from a user and convert the spoken information into usable digital information. Audio codec 360 may also generate audible sounds for the user, such as through a speaker, for example, a speaker in an earpiece of device 350. Such sounds may include sounds from voice calls, may include recorded sounds (e.g., voice messages, music files, etc.), and may also include sounds generated by applications operating on device 350.
[0145] As shown, computing device 350 can be implemented in a variety of forms. For example, computing device 350 can be implemented as a cellular phone 380. Computing device 350 can also be implemented as part of a smartphone 382, a personal digital assistant, or other similar mobile device.
[0146] In addition, computing device 300 or 350 may include a universal serial bus (USB) flash drive. The USB flash drive may store an operating system and other applications. The USB flash drive may include input / output components, such as a wireless transmitter or a USB connector that may be plugged into a USB port of another computing device.
[0147] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuitry, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs executable and / or interpretable on a programmable system, a storage system, at least one input device, and at least one output device, the programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from the programmable processor and transmit data and instructions to the programmable processor.
[0148] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., a disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network ("LAN"), a wide area network ("WAN"), a peer-to-peer network (with temporary or static members), a grid computing infrastructure, and the Internet.
[0151] A computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0152] Although certain embodiments have been described in detail above, other modifications are possible. Furthermore, other mechanisms for implementing the systems and methods described herein may be used. In addition, the logic flows depicted in the accompanying drawings do not require the specific order or sequential order shown to achieve the desired results. Additional steps may be provided or steps may be deleted from the described flows, and additional components may be added to or removed from the described systems. Therefore, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method comprising: Receiving, by a computing system, a first frame and a second frame of a video captured by a camera; generating, by the computing system and using the first and second frames of the video, a first mathematical transform indicating movement of the camera relative to a scene captured by the video from when the first frame was captured to when the second frame was captured, including recently initiated movement and movement that has occurred over a longer period of time, the computing system generating the first mathematical transform by identifying how feature points present in the first and second frames move between the first and second frames; generating, by the computing system, a modified mathematical transform by modifying the first mathematical transform so that the modified mathematical transform is more representative of movement of the camera relative to the scene that has occurred over a long period of time; generating, by the computing system, a second mathematical transform using the first mathematical transform and the modified mathematical transform by a process that involves removing the identified indication of long-term camera movement from the first mathematical transform, the second mathematical transform being more representative of recently initiated camera movement relative to the scene and less representative of long-term camera movement relative to the scene than the first mathematical transform; A stabilized version of the second frame is generated by the computing system by applying the second mathematical transform to the second frame to remove recently initiated motion from the second frame and retain motion that has occurred long term, without regard to motion of the camera relative to the scene from future frames of the video.
2. The computer-implemented method of claim 1 , wherein: The second frame is a frame of the video that immediately follows the first frame of the video.
3. The computer-implemented method of claim 1 , wherein: The first mathematical transformation indicative of movement of the camera relative to a scene comprises a homography transformation matrix.
4. The computer-implemented method of claim 1 , wherein: Generating the stabilized version of the second frame includes scaling to a version of the second frame generated by applying the second mathematical transform to the second frame.
5. The computer-implemented method of claim 4 , further comprising: The zoomed-in area of the stable version of the second frame is shifted horizontally or vertically to avoid the zoomed-in area of the stable version of the second frame from presenting an invalid area.
6. The computer-implemented method of claim 1 , wherein: Identifying movement of the camera that has occurred over a long period of time includes analyzing frames of the video that the camera has captured prior to the first frame and the second frame of the video to identify movement that has occurred over a long period of time.
7. The computer-implemented method of claim 1 , wherein: Identifying that movement of the camera has occurred over a long period of time includes applying a low-pass filter to the first mathematical transform based on an analysis of frames of the video that the camera has captured prior to the first and second frames of the video.
8. The computer-implemented method of claim 1 , further comprising: An expected distortion to be present in a stabilized version of the second frame resulting from applying the second mathematical transform to the second frame is identified by the computing system based on a difference between: (i) an amount of distortion in the horizontal direction resulting from applying the second mathematical transform to the second frame, and (ii) an amount of distortion in the vertical direction resulting from applying the second mathematical transform to the second frame; determining, by the computing system, an amount to reduce a stabilization effect produced by applying the second mathematical transform to the second frame based on an extent to which the expected distortion exceeds an acceptable distortion variation calculated from distortion in a plurality of frames of the video preceding the second frame, Wherein generating the stabilized version of the second frame comprises reducing a stabilizing effect of applying the second mathematical transform to the second frame based on the determined amount to reduce the stabilizing effect.
9. The computer-implemented method of claim 1 , further comprising: calculating, by the computing system, an acceptable change in movement speed of the camera based on an analysis of the movement speed of the camera between a plurality of frames of the video preceding the second frame; as well as determining, by the computing system, an amount to reduce a stabilization effect produced by applying the second mathematical transform to the second frame based on the determined speed of movement of the camera from the first frame to the second frame exceeding the acceptable change in speed of movement of the camera, Wherein generating the stabilized version of the second frame comprises reducing a stabilization effect produced by applying the second mathematical transformation to the second frame based on the determined amount to reduce the stabilization effect.
10. One or more non-transitory computer-readable devices comprising instructions that, when executed by one or more processors, cause operations to be performed, the operations comprising: Receiving, by a computing system, a first frame and a second frame of a video captured by a camera; generating, by the computing system and using the first and second frames of the video, a first mathematical transform indicating movement of the camera relative to a scene captured by the video from when the first frame was captured to when the second frame was captured, including recently initiated movement and movement that has occurred over a longer period of time, the computing system generating the first mathematical transform by identifying how feature points present in the first and second frames move between the first and second frames; generating, by the computing system, a modified mathematical transform by modifying the first mathematical transform so that the modified mathematical transform is more representative of movement of the camera relative to the scene that has occurred over a long period of time; generating, by the computing system, a second mathematical transform using the first mathematical transform and the modified mathematical transform by a process that involves removing the identified indication of long-term camera movement from the first mathematical transform, the second mathematical transform being more representative of recently initiated camera movement relative to the scene and less representative of long-term camera movement relative to the scene than the first mathematical transform; A stabilized version of the second frame is generated by the computing system by applying the second mathematical transform to the second frame to remove recently initiated motion from the second frame and retain motion that has occurred long term, without regard to motion of the camera relative to the scene from future frames of the video.
11. The one or more non-transitory computer-readable devices of claim 10, wherein: The second frame is a frame of the video that immediately follows the first frame of the video.
12. The one or more non-transitory computer-readable devices of claim 10, wherein: The first mathematical transformation indicative of movement of the camera relative to a scene comprises a homography transformation matrix.
13. The one or more non-transitory computer-readable devices of claim 10, wherein: Generating the stabilized version of the second frame includes scaling to a version of the second frame generated by applying the second mathematical transform to the second frame.
14. The one or more non-transitory computer-readable devices of claim 13, wherein: The operation further includes horizontally or vertically shifting the zoomed-in area of the stable version of the second frame to avoid the zoomed-in area of the stable version of the second frame presenting an invalid area.
15. The one or more non-transitory computer-readable devices of claim 10, wherein: Identifying movement of the camera that has occurred over a long period of time includes analyzing frames of the video that the camera has captured prior to the first frame and the second frame of the video to identify movement that has occurred over a long period of time.
16. The one or more non-transitory computer-readable devices of claim 10, wherein: Identifying that movement of the camera has occurred over a long period of time includes applying a low-pass filter to the first mathematical transform based on an analysis of frames of the video that the camera has captured prior to the first and second frames of the video.
17. The one or more non-transitory computer-readable devices of claim 10, wherein: The operations further include: An expected distortion to be present in a stabilized version of the second frame resulting from applying the second mathematical transform to the second frame is identified by the computing system based on a difference between: (i) an amount of distortion in the horizontal direction resulting from applying the second mathematical transform to the second frame, and (ii) an amount of distortion in the vertical direction resulting from applying the second mathematical transform to the second frame; determining, by the computing system, an amount to reduce a stabilization effect produced by applying the second mathematical transform to the second frame based on an extent to which the expected distortion exceeds an acceptable distortion variation calculated from distortion in a plurality of frames of the video preceding the second frame, Wherein generating the stabilized version of the second frame comprises reducing a stabilizing effect of applying the second mathematical transform to the second frame based on the determined amount to reduce the stabilizing effect.
18. The one or more non-transitory computer-readable devices of claim 10, wherein: The operations further include: calculating, by the computing system, an acceptable change in movement speed of the camera based on an analysis of the movement speed of the camera between a plurality of frames of the video preceding the second frame; and determining, by the computing system, an amount to reduce a stabilization effect produced by applying the second mathematical transform to the second frame based on the determined speed of movement of the camera from the first frame to the second frame exceeding the acceptable change in speed of movement of the camera, Wherein generating the stabilized version of the second frame comprises reducing a stabilization effect produced by applying the second mathematical transformation to the second frame based on the determined amount to reduce the stabilization effect.
Citation Information
Patent Citations
Stable video
CN107851302B