Methods, systems, and devices for stabilizing video to reduce camera and face movement
The computer-implemented video stabilization method uses facial features and sensor information to optimize the virtual camera viewpoint posture, solving the problem of video instability caused by camera and face movement, improving video quality and reducing equipment costs.
Patent Information
- Application Number
- CN202210373495.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-28
- Filing Date
- 2019-04-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2039-04-17
AI Technical Summary
Existing technologies do not provide good video stabilization when the recording device moves unexpectedly, especially when the camera and the person's face move, resulting in a degradation of video quality.
The video stabilization method implemented by computer uses a computing system to receive a video stream and determine the position of facial features of a person's face, combines the information of a motion or orientation sensor, optimizes the posture of the virtual camera viewpoint, generates a stable view, and offsets the movement of the camera and the person's face.
Effectively reduces camera and face movement, improves video stability, reduces the need for mechanical stabilization, reduces manufacturing costs, and improves user experience.
Smart Images

Figure CN114758392B_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of Chinese invention patent application No. 201980026280.4, filed on April 17, 2019. Technical Field
[0003] This document discusses stabilizing video to reduce camera and face movement. Background Art
[0004] Various types of recording devices, such as cameras and smartphones, include image sensors that can be used to record video. Video can be generated by capturing a sequence of images, often referred to as video "frames," and typically captured at a defined frame rate (e.g., thirty frames per second). The captured sequence of frames can be presented by a display device at the same frame rate, and the switching from one frame to another can be largely unnoticeable to humans, making the display appear to be showing actual movement rather than a sequence of rapidly switching images.
[0005] Recording devices sometimes experience unexpected movement during recording, for example, shake caused by being held by a person or attached to a moving vehicle. This movement can have a particularly noticeable effect on the video when the movement is rotational and the scene being captured is far from the recording device.
[0006] Various techniques can stabilize video to limit unintended movement of the recording device. One technique for stabilizing video is mechanical stabilization, in which mechanical actuators counteract external forces. Mechanical stabilization can be achieved by mounting the recording device on an auxiliary device (e.g., a gimbal) that stabilizes the movement of the entire recording device. Mechanical stabilization can also be achieved by integrating actuators within the recording device to stabilize the camera or part of it relative to the movement of the main body of the recording device. Another technique for stabilizing video is digital video stabilization, in which a computer analyzes the recorded video and crops the captured frames in a way that produces a partially magnified version of the stabilized video. Mechanical stabilization and digital video stabilization techniques can be used in combination. Summary of the Invention
[0007] This disclosure describes the following embodiments.
[0008] Embodiment 1 is a computer-implemented video stabilization method. The method includes receiving, by a computing system, a video stream comprising a plurality of frames and captured by a physical camera. The method includes determining, by the computing system, positions of facial features of a human face depicted in a frame of the video stream captured by the physical camera. The method includes determining, by the computing system, a stable position of the facial features taking into account previous positions of the facial features in previous frames of the video stream captured by the physical camera. The method includes determining, by the computing system, a pose of the physical camera in virtual space using information received from a motion or orientation sensor coupled to the physical camera. The method includes mapping, by the computing system, the frames of the video stream captured by the physical camera into the virtual space. The method includes determining, by the computing system, an optimized pose of a virtual camera viewpoint in the virtual space, and generating a stabilized view of the frame from the optimized pose. The optimization process determines a difference between the stabilized position of the facial features and the position of the facial features in the stabilized view of the frame as viewed from a potential pose of the virtual camera viewpoint. The optimization process determines a difference between a potential pose of the virtual camera viewpoint in the virtual space and a previous pose of the virtual camera viewpoint in the virtual space. The optimization process determines a difference between the potential pose of the virtual camera viewpoint in the virtual space and a pose of the physical camera in the virtual space. The method includes generating, by the computing system, a stabilized view of the frame using the optimized pose of the virtual camera viewpoint in the virtual camera space.
[0009] Embodiment 2 is a computer-implemented video stabilization method based on embodiment 1, further comprising presenting, by the computing system, a stabilized view of the frame on a display of the computing system.
[0010] Embodiment 3 is a computer-implemented video stabilization method based on embodiment 1, wherein the motion or orientation sensor comprises a gyroscope.
[0011] Embodiment 4 is a computer-implemented video stabilization method according to embodiment 1, wherein the computing system determines positions of facial features of the face depicted in the frame based on positions of a plurality of corresponding facial landmarks depicted in the frame. Furthermore, the computing system determines the difference between the stabilized positions of the facial features and the positions of the facial features in the stabilized view of the frame by measuring a deviation between the positions of the plurality of facial landmarks in the stabilized view of the frame and the stabilized positions of the facial features.
[0012] Example 5 is a computer-implemented video stabilization method based on Example 1, wherein the optimization process includes: minimizing the value of a pose parameter (hereinafter also referred to as "E_V_0(T)") based on at least one of the following: (i) a deviation between a landmark in the stabilized view of the frame and a stabilized position of the facial feature (hereinafter also referred to as a variable "E_center"); (ii) a difference between a potential pose of the virtual camera viewpoint of the frame and a pose of the virtual camera viewpoint of the previous frame (hereinafter also referred to as a variable "E_rotation_smoothness"); (iii) a difference between a camera rotation in the virtual space of the frame and a real camera rotation in the virtual space of the frame. (iv) the spherical angle between the camera rotation in the virtual space and the real camera rotation in the virtual space (hereinafter also referred to as the variable "E_distortion"); (v) the change in the offset to the virtual principal point between the frame and the previous frame (hereinafter also referred to as the variable "E_offset_smoothness"); and (vi) the number of undefined pixels in the stabilized view of the frame generated using the potential pose of the virtual camera viewpoint in the virtual space (hereinafter also referred to as the variable "E_undefined_pixel").
[0013] Embodiment 6 is a computer-implemented video stabilization method based on embodiment 1, wherein the optimization process includes a nonlinear computational solver that optimizes values of multiple corresponding variables.
[0014] Embodiment 7 is a computer-implemented video stabilization method based on embodiment 1, wherein the optimization process determines the number of undefined pixels in the stabilized view of the frame generated using the potential pose of the virtual camera viewpoint in the virtual space.
[0015] Example 8 is a computer-implemented video stabilization method based on Example 1, wherein the optimization process determines the difference between (a) the offset of the principal point of the stabilized view of the frame generated using the potential pose of the virtual camera viewpoint in the virtual space, and (b) the offset of the previous principal point of the previous stabilized view of the frame generated using the previous pose of the virtual camera viewpoint in the virtual space.
[0016] Example 9 is a computer-implemented video stabilization method based on Example 1, wherein generating a stabilized view of the frame includes: mapping a subset of scan lines of the frame to a viewing angle viewed from an optimized posture of the virtual camera viewpoint and interpolating other scan lines of the frame.
[0017] Example 10 is a computer-implemented video stabilization method based on Example 1. In this embodiment, determining the stable position of the facial feature includes using a position optimization process that (i) determines the difference between the potential stable position of the facial feature and the actual position of the facial feature in the frame; (ii) determines the difference between the potential stable position of the facial feature and the position of the facial feature in the previous frame; and (iii) accounts for constraints on the distance between the potential stable position of the facial feature and the position of the facial feature in the frame. To determine the difference between the potential stable position of the facial feature and the actual position of the facial feature in the frame, for example, how far the potential stable face center deviates from the actual determined face center can be considered to ensure that stabilization does not attempt to stabilize the face so strongly that the video will substantially deviate from the actual depiction of the face in the video. To determine the difference between the potential stable position of the facial feature and the position of the facial feature in the previous frame, for example, how far the potential stable center deviates from the last determined face center can be considered to ensure that stabilization does not attempt to track the current position of the face so strongly that the face moves abruptly in the video.
[0018] Example 11 is a computer-implemented video stabilization method based on Example 1, which includes selecting a face depicted in a frame of the video stream captured by the physical camera as a face to be tracked from multiple faces depicted in the frame of the video stream captured by the physical camera by the following operations: (i) selecting the face based on the size of each face in the multiple faces, (ii) selecting the face based on the distance of each face in the multiple faces to the center of the frame, or (iii) selecting the face based on the distance between the face selected for tracking in a previous frame and each face in the multiple faces.
[0019] Example 12 is a computer-implemented video stabilization method based on Example 1, wherein the optimized pose of the virtual camera viewpoint has a different position and rotation in the virtual space than the pose of the physical camera.
[0020] Embodiment 13 is directed to one or more computer-readable devices having instructions stored thereon, which, when executed by one or more processors, cause the performance of the actions of the method according to any one of embodiments 1 to 10.
[0021] Embodiment 14 is directed to a system comprising one or more processors and one or more computer-readable devices having instructions stored thereon that, when executed by the one or more processors, cause the actions of the method described in any one of embodiments 1 to 10 to be performed. Such a system can determine a stable position of a facial feature in a video frame, taking into account the position of the facial feature in a previous frame. The system can determine a physical camera pose in a virtual space and map the frame into the virtual space. The computing system can determine an optimized virtual camera pose using an optimization process that determines (i) the difference between the stable position of the facial feature and the position of the facial feature when viewed from a potential virtual camera pose, (ii) the difference between the potential virtual camera pose and a previous virtual camera pose; and (ii) the difference between the potential virtual camera pose and the physical camera pose. The computing system can use the optimized virtual camera pose to generate a stable view of the frame.
[0022] This document also describes techniques, methods, systems, and other mechanisms for stabilizing video to reduce camera and face movement. Generally, the mechanisms described herein can generate a stabilized version of a video by determining a virtual recording pose (virtual camera position and orientation) that provides a more stable recording experience than the actual pose of the physical camera. The virtual recording pose can be determined to offset not only the unwanted physical movement of the camera, but also the movement of faces in the scene. A computerized process can distort frames captured by the physical camera so that the frames appear to have been captured from the virtual recording pose rather than from the actual pose of the physical camera, wherein the virtual recording is laterally offset from the pose of the physical camera in the virtual space. This process can be repeated for each frame to produce a stabilized video.
[0023] In addition to camera movement, stabilizing the movement of the face can provide better stabilization results in various situations. For example, suppose a user is using a smartphone's front-facing camera to capture a video (e.g., a "selfie") while riding in a vehicle. The vehicle may cause the camera and the user to bounce around. In this case, a video stabilization mechanism that only stabilizes the physical movement of the camera may actually be counterproductive because the user's face may continue to bounce even if the camera position is stable.
[0024] The techniques described herein can stabilize video to minimize unintended movement of the camera and one or more subjects. This can reduce or eliminate the need for mechanical stabilization, which can reduce manufacturing costs and the space required to house the camera in the recording device. Alternatively, using both mechanical and digital video stabilization can enhance stabilization. Another benefit is that the user of the recording device may not have to focus on stabilizing the recording device or the subject of the video, allowing them to focus on other aspects of the video shooting experience.
[0025] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 Frames of a video stream and movement of objects represented within the frames of the video stream are shown.
[0027] Figure 2 Shows the viewpoints of the physical and virtual cameras in the virtual space and the frames mapped into the virtual space.
[0028] Figures 3A to 3E Flowchart showing a process for stabilizing video to reduce camera and face movement.
[0029] Figure 4 is a conceptual diagram of a system that can be used to implement the systems and methods described in this document.
[0030] Figure 5 is a block diagram of a computing device that can be used to implement the systems and methods described in this document as a client or server or multiple servers.
[0031] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION
[0032] This document generally describes stabilizing video to reduce camera and object motion. The techniques described herein can involve warping frames of a video so that the frames appear to be derived from the pose of a stabilized virtual camera viewpoint rather than the pose of the physical camera that captured the frames. The process of determining the pose of the virtual camera viewpoint can take into account the current pose of the physical camera, the previous pose of the virtual camera viewpoint, and the position of a human face in the scene.
[0033] The position of a person's face in a scene can be important because the person's face can move relative to the camera even if the camera's position is stabilized (e.g., mechanically or using digital stabilization techniques). The movement of the person's face can be particularly noticeable when the person is close to the camera, such as when the person is taking a "selfie," because in this case the face may occupy a large portion of the frame.
[0034] The techniques described herein may involve determining a "stable" position of a person's face in a recently received video frame. This may include calculating the position of the person's face in the video and determining how the face moves from one frame to the next. The stable position of the person's face may be where the person's face would be if the movement of the face were smoothed out (such as if someone were to grab the person by their shoulders to limit any shaking and slow down sudden movements). The video stabilization process may take the stable position of the person's face into account when selecting the pose of the virtual camera viewpoint.
[0035] More generally, the video stabilization process can consider several different factors when selecting the pose of the virtual camera viewpoint. The first factor is the pose of the physical camera, as determined using one or more sensors (e.g., a gyroscope) of the camera that identify the camera's movement and / or rotation. The second factor is the pose determined for the virtual camera in the last frame. In the absence of other factors, the determined pose of the virtual camera viewpoint may be somewhere between the pose of the physical camera and the previous pose of the virtual camera viewpoint.
[0036] One of these other factors is the distance between the stable position of the face (as determined above) and the position of the face in the video frame when the video frame is distorted so that it appears to have been taken from the perspective of the virtual camera viewpoint. As a simple example, a person may have moved his or her face to one side very quickly, which may cause the recording device to determine that the stable position of the face should be between the previous position of the face and the actual position of the face. The effect of sudden face displacement on the recorded video can be reduced by moving the virtual camera viewpoint in the same direction as the face movement, thereby reducing the movement of the face, at least in the stabilized video.
[0037] In this way, the video stabilization techniques described herein can take into account multiple different factors to select an optimized pose for the virtual camera viewpoint, and once that pose is selected, the video frames can be warped so that the video appears to have been obtained from the optimized pose for the virtual camera viewpoint rather than from the pose of the physical camera (in some examples, the position and orientation of the virtual camera viewpoint are different from the position and orientation of the physical camera in virtual space). This process can be repeated for each frame of the video to generate a video that appears to have been captured from a virtual camera that moves in a more stable manner than the physical camera. This process may not require analysis of future frames and can therefore be performed in real time as the video is captured. In this way, the video that appears on the display of the recording device while the recording device is recording the video can be a stabilized video.
[0038] The following description explains this video stabilization technique in more detail with respect to the accompanying drawings. The description will generally follow Figures 3A to 3E This flowchart describes the process of stabilizing a video to reduce camera and face movement. The description of this flowchart will refer to Figure 1 and Figure 2 , to illustrate various aspects of video stabilization technology. Figure 1 Frames of the video are shown, along with the movement of objects from one frame to the next. Figure 2 Shows the hardware that records the video and explains how the hardware and virtual camera viewpoint change over time as frames are captured.
[0039] Now refer to Figure 3AAt block 310, a computing system receives a video stream comprising a plurality of frames. Figure 2 The computing system 220, depicted as a smartphone in FIG, can record video using a front-facing camera 222. The recorded video may include a video stream 110, which is Figure 1 1 is illustrated as including a plurality of component frames 120 , 130 and 140 . Figure 1 An example is shown in which video stream 110 is recorded in real time, and frame 140 represents the most recently captured frame, where frame 130 is the last frame captured and frame 120 is the frame captured before it. Video stream 110 may include hundreds or thousands of frames captured at a predetermined frame rate (such as 30 frames per second). As captured, computing system 220 may store each frame in memory. In some examples, video stream 110 may have been pre-recorded, and frame 140 may represent a frame in the middle of video stream 110.
[0040] At block 312, the computing system selects a target time for the selected frame. The selected frame may be the most recently captured frame (e.g., Figure 1 In some examples, the target time may be defined as the beginning of the exposure duration for the frame, the end of the exposure duration for the frame, or some other time during the exposure duration for the frame. For example, the computing system may define the target time for the selected frame as the middle of the exposure duration for the frame (block 314).
[0041] At block 316, the computing system selects a face depicted in the frame as a face to be tracked from among the multiple faces depicted in the frame. For example, the computing system may analyze the frame to identify each face in the frame and may then determine which of the identified faces to track. Figure 1 In frame 140 , the computing system may choose to track face 150 rather than face 160 due to various factors, such as those described with respect to blocks 318 , 320 , and 322 .
[0042] At block 318, the computing system selects a face based on the size of each face in the plurality of faces. For example, the computing system may determine the size of a bounding box for each face and may select the face with the largest bounding box.
[0043] At block 320, the computing system selects a face from the plurality of faces based on the distance of each face from the center of the frame. For example, the computing system may identify the location of each face (described in more detail below with respect to block 330) and determine the distance between the corresponding face and the center of the frame. The face closest to the center of the frame may be selected.
[0044] At block 322, the computing system may select a face based on the distance between the location of the face selected in the previous frame and the location of each face in the current frame. This operation can help ensure that the system tracks the same face from frame to frame as the face moves around in the video.
[0045] The computing system may weight the operations of blocks 318, 320, and 322 equally or differently to generate an interest value for each face, and may compare the interest values of the faces to identify the face with the highest interest value. A threshold interest value may be defined such that if no face has an interest value exceeding the threshold, no face is selected. The face selection process may consider other factors, such as the orientation of the face, whether the eyes are open in the face, and which faces are smiling. The face selection process may consider any combination of one or more of the above factors.
[0046] At block 330, the computing system may determine the location of a facial feature to be tracked from the selected face in the frame. An example facial feature to be tracked is the center of the face, but the computing system may also track another facial feature (such as the center of the person's nose or mouth). Determining the center of the face may involve the operations of blocks 332, 334, and 336.
[0047] At block 332, the computing system identifies a bounding box of the selected face. The computing system may perform this identification using a face information extraction module.
[0048] At block 334, the computing system uses the center of the face as a facial feature. The computing system can determine the center of the face by identifying the locations of multiple landmarks on the face and determining the average location of those landmarks. Example landmarks include the locations of the eyes, ears, nose, eyebrows, corners of the mouth, and chin.
[0049] Although the operation of block 334 is described with reference to frame 140, this face center determination process may be performed for each frame in the video as that frame is stabilized. Figure 1 The description of frame 140 in includes multiple annotations, but for Figure 1 130 in FIG. 1 presents annotations of the face center determination process, but it should be understood that the face center determination process can be similar when applied to frame 140. Determining the center of face 150 in frame 130 can include the computing system identifying a plurality of landmarks on face 150 using the identified landmarks illustrated within dashed rectangular box 170. The annotation accompanying frame 130 illustrates that the computing system has identified the locations of the eyes and corners of the mouth, but the computing system can identify other facial landmarks. By analyzing these landmark locations 170, the computing system determines that location 180 represents an average location of landmarks 170.
[0050] At block 336, the computing system determines the orientation of the face and uses the face orientation to determine whether the face is in a silhouette mode. In some examples, if it is determined that the face is in a silhouette mode, the computing system may not track the position of the face.
[0051] At block 340, the computing system determines the stable position of the facial features. Continuing with the example above where the face is centered on the facial features being tracked, the computing system determines the stable center of the face (represented as position 182 in Figure 1 frame 140). The stable center of the face can represent the position where the center of the face might be located as the face smoothly moves from a previous frame to the current frame. In this way, the stable position of the face center can be selected to not be too far from the true center of the face (represented as position 184 in Figure 1 frame 140), but to always maintain the stable position of the face center if possible. In this way, identifying the stable position of the face center can involve an optimization process (of positions) that interprets multiple different factors, as described with respect to blocks 342, 344, and 346.
[0052] At block 342, determining the stable position 182 of the facial features interprets the distance between the potential stable position of the facial features and the actual position 184 (e.g., the average of the facial landmarks). For example, the computing system can consider how far the potential stable face center deviates from the actually determined face center to ensure that the stabilization does not strongly attempt to stabilize the face such that the video will substantially deviate from the actual depiction of the face in the video. The term that interprets this factor is called E_follow, which measures how far the stable 2D center H(T) of the current frame is from the true landmark center C(T) determined as the average of all 2D landmarks.
[0053] At block 344, determining the stable position 182 of the facial features interprets the distance between the potential stable position of the facial features and the previously determined position 180 of the facial features. For example, the computing system can consider how far the potential stable center deviates from the position 180 of the last determined center of the face to ensure that the stabilization does not strongly attempt to track the current position of the face such that the face cannot make sudden movements in the video. The term that interprets this factor is called E_smoothness, which measures the change between H(T) and H(T_pre), where H(T_pre) is the estimated 2D head center of the previous frame.
[0054] At block 346, determining the stable position of the facial features interprets the constraint on the distance between the potential stable position of the facial features and the determined position of the facial features. This factor is imposed as a hard constraint |H(T)-C(T)|<CroppedRange such that C(T) moves within the effective range around H(T), which does not create an undefined region. [[ID=]
[18]
[0055] An example process of combining these various terms to determine the stable position of facial features can be summarized as E_H(T) = w_1 * E_smoothness + w_2 * E_follow, such that |H(T) - C(T)| < CroppedRange. The values w_1 and w_2 are used to weight the smoothness and follow factors.
[0056] At box 348, the computing system determines the pose of the physical camera in virtual space. For the first frame of the video, the pose of the physical camera is referred to as R(t) and can be initialized with zero rotation and offset. Thereafter, the computing system can analyze signals received from one or more motion or rotation sensors to determine how the pose of the physical camera changes and the position of subsequently captured frames. For example, the computing system can receive signals from a gyroscope at a high frequency (e.g., 200 Hz) and can use the information in the signals to determine how the pose of the camera changes and the current pose of the camera (where the current pose of the physical camera may have a different orientation and / or position than the initialized pose of the physical camera). In some examples, the computing system can alternatively or additionally use an accelerometer to determine the pose of the physical camera. The camera and one or more motion or rotation sensors can be physically coupled to each other such that the camera and the one or more sensors have the same motion and pose. For example, the camera and one or more sensors can be coupled to the same housing of a smartphone. The pose of the physical camera can be represented as a quaternion representation (4D vector), which defines the pose of the camera in virtual space. The pose of the physical camera in virtual space is Figure 2 represented by camera 214 in. Determining the pose of an object in virtual space can include assigning coordinates and orientation to the object in a common coordinate system and does not require generating a visual representation of the object.
[0057] At box 350, the computing system maps the frame to virtual space. For example, the computing system can apply coordinates to the frame to represent the positions of parts of the frame (such as the corners of the frame) relative to the position of the frame in the virtual space of the physical camera. As Figure 2 illustrated by frame 230 in, this mapping can be performed by constructing a projection matrix that maps the real-world scene to the image. Mapping the frame to virtual space can account for various factors as described with respect to boxes 352 and 354.
[0058] At box 352, mapping the frame to virtual space accounts for the pose of the physical camera in virtual space. For example, the physical camera pose can be used to determine that the position of the frame should be in front of the physical camera, where the principal point (e.g., the center) is aligned with the orientation of the physical camera.
[0059] At block 354, mapping the frames to the virtual space accounts for the physical camera's focus lens and the physical camera's current zoom setting. For example, a non-zoomed frame captured using a camera without a fisheye lens should span a larger portion of the virtual space in front of the physical camera than a zoomed frame captured using a camera without a fisheye lens.
[0060] This process of mapping a frame to a virtual space can be represented by the equation P_(i, j_T) = R_(i, j_T) * K(i, j_T), where i is the frame index and j is the scanline index. R_(i, j_T) represents the pose of the physical camera, also known as the camera's non-intrinsic matrix (rotation matrix) obtained using gyroscope information. K(i, j_T) = [f 0 Pt_x; f 0 Pt_y; 0 0 1] is the camera's intrinsic matrix, where f is the focal length of the current frame and Pt is the 2D principal point set to the center of the image.
[0061] At block 360, the computing system uses an optimization process to determine an optimized pose of the virtual camera viewpoint in virtual space, from which a stabilized view of the frame is generated. This optimization can be performed using a nonlinear motion filtering engine and can select a virtual camera viewpoint that smooths rotations and translations of the virtual camera viewpoint relative to the physical camera. Selecting the optimized pose of the virtual camera viewpoint can involve selecting a position in virtual space of the virtual camera viewpoint and an orientation of the virtual camera viewpoint, one or both of which can differ from the position and orientation in virtual space of the physical camera.
[0062] exist Figure 2 The optimized pose of the virtual camera is illustrated in FIG. 212 and can be determined using an optimization process that accounts for multiple factors, such as the physical camera (in FIG. Figure 2 ), the pose of the virtual camera for the previous frame (illustrated as camera 214 in Figure 2 ), and the actual position of the facial features (e.g., Figure 2 The center of the face 280 in FIG. 2 and the stable position of the facial features (eg, Figure 2 The distance between the stable face center 282 in .
[0063] An example equation for determining the optimized pose of a virtual camera is E_V_0(T) = w_1*E_center+w_2*E_rotation_smoothness+w_3*E_rotaton_following+w_4*E_distortion+w_5*E_undefined_pixel+w_6*E_offset_smoothness. The terms included in this equation (e.g., E_center and E_rotation_smoothness) can take an example virtual camera pose as input and can output a value indicating the suitability of the virtual camera pose for that particular term. The virtual camera pose (also referred to herein as the virtual camera viewpoint) can be expressed as V-0(T) = [R_v(T), O_v(T)], where R_v(T) is the virtual camera extrinsic matrix (rotation matrix) and O_v(T) is the 2D offset of the virtual principal point Pt_v.
[0064] Selecting a given virtual camera pose can affect the value of each term in the equation illustrated above, producing a resulting value of E-V_0(T). Therefore, different virtual camera poses can be input into the above equation to determine different values of E_V_0(T). Instead of inputting many different virtual camera poses into the above equation for E_V_0(T) to identify the virtual camera pose with the best value, a nonlinear solver (such as a Ceres solver) can be used to determine the best virtual camera pose (e.g., the value that minimizes the value of E_V_0(T)). Although E_V_0 is a function of T, it can also be expressed as a function of R_v(T) and O_v(T) because the value of T affects the values of Rv(T) and O_v(T), which affect the value of E_V_0, and can therefore be alternatively illustrated as E-V_0(R_v(T), O_v(T)). Reference boxes 360-380 provide additional details on determining the best virtual camera pose.
[0065] At block 360, the computing system determines whether the optimization process is being performed on the first frame of the video. If so, the computing system uses a virtual camera pose with zero rotation and zero offset as an initialization (block 364). If not, the computing system uses the virtual camera pose from the previous frame during the optimization process (block 366).
[0066] At block 368, the optimization process determines the difference between (1) the stable position of the facial features and (2) the position of the facial features in the stable view of the frame. This is the term of the optimization process that accounts for the movement of the face. The influence and operation of this factor can be referred to Figure 1 and Figure 2 To picture.
[0067] As an illustration, Figure 1Three frames 120, 130, and 140 are shown. Faces 150 and 160 are represented in each of these frames and are positioned at different locations relative to each other in the various frames to illustrate various aspects of the techniques described herein. Frame 120 shows the initial positions of faces 150 and 160. However, between frames 120 and 130, the camera may have moved to the right (e.g., panned), causing the positions of faces 150 and 160 in frame 130 to shift to the left relative to their positions in frame 120. In frame 140, the camera has not moved from its position when frame 130 was captured, but face 150 has moved to the right in frame 140 due to the real-world movement of the face. (Face and camera movement would typically occur simultaneously, but in this illustration the movement is isolated to different frames for ease of description.)
[0068] As previously described with respect to blocks 340-346, the computing system has determined that stable position 182 represents a stable position of face 150, e.g., a desired center of face 150 that stabilizes the movement of face 150 as it moves between frames 130 and 140. Figure 2 Also illustrated in FIG, is a diagram showing frame 140 mapped into virtual space as frame 230. As shown in frame 140 and also in frame 230, because the user moved to the right during the translation from frame 140 to frame 150, the actual center of the user's face in the frame (determined based on facial landmarks) is to the right of stable position 182 of the user's face 150.
[0069] Figure 2 Frame 230 is shown from the perspective of physical camera 214, but if frame 230 is viewed from the perspective of virtual camera 212 rather than physical camera 214, the view and position of certain objects in the frame may change. As an illustration, as virtual camera 212 moves around, the position of stabilizing position 182 may remain fixed, but the position of face 150 may move around in the frame, just as the position of an object within your own field of view may change if you walk to the side and continue facing it. With the ability to move virtual camera 212 around to influence the position of face 150 in the image, optimal stabilization of face 150 may include positioning virtual camera 212 so that the determined center of face 150, as viewed from the viewpoint of virtual camera 212, is aligned with stabilizing position 182. However, optimization algorithms may account for factors, and therefore, the ultimately selected optimal position for virtual camera 212 may not be positioned so that the determined center of face 150 is perfectly aligned with stabilizing position 182.
[0070] At block 370, the locations of the plurality of facial landmarks in the stabilized view of the frame may be used to represent the location of the facial feature in the stabilized view of the frame. For example, the computing system may determine the locations of the plurality of facial landmarks, and these locations may collectively represent the location of the center of the person's face 150. The computation to interpret the location of the facial feature (e.g., the center of the person's face 150) does not require the actual computation of the location of the facial feature and may instead use data indicating the location of the facial feature, such as the locations of the plurality of landmarks.
[0071] At block 372, the difference of block 368 is determined by obtaining an average of the distances between (1) the stable position of the facial feature and (2) the positions of the plurality of facial landmarks in the stable view of the frame. In other words, the computing system may not actually compute the position of the facial feature in the stable view of the frame. Instead, the system may determine how far each landmark is from the stable position of the facial feature in the stable view of the frame, and may identify a virtual camera pose that minimizes this distance among all the landmarks.
[0072] Referring again to the equation for E_V_0(T), the operation in block 368 represents E_center, which measures the average deviation between the position of each projected landmark on the virtual camera plane (e.g., frame 140 viewed from the perspective of the virtual camera) and the estimated 2D head center point H(T), which is the target head center on the virtual camera plane. For each detected landmark 1, the computing system can identify the scan line to which it belongs, calculate the transformation P_v(T) used to map the real image projected by P_(i, j) to that scan line, and map the landmark to the virtual camera plane to obtain its 2D position I_v. The deviation is then calculated as the L2 difference between I_v and H(T), i.e., ||I_v-H(T)||^2. The E_center term ensures that the projected center of the selected face on the stabilized frame follows the estimated 2D head center H(T).
[0073] At block 374, the optimization process determines the difference (eg, difference in position and / or orientation) between the proposed pose of the virtual camera viewpoint in the virtual space and the previous pose of the virtual camera viewpoint in the virtual space. Figure 2 , the optimization process may interpret the distance 216 between the optimized pose 212 of the virtual camera viewpoint and the previous pose 210 of the virtual camera viewpoint as used to generate the stable view of the previous frame 130. (To simplify the explanation, this discussion sometimes refers to a proposed pose of the virtual camera, but it should be understood that this discussion is intended to encompass optimization processes that may not actually test multiple different proposed poses, but rather perform an optimization process, e.g., as described throughout this disclosure.)
[0074] Referring again to the equation for E-V_0(T), the operation of block 374 may represent E_rotation_smoothness, i.e., a rotation smoothness term. This term measures the difference between the virtual camera pose of the current frame and the virtual camera pose of the previous frame. A rotation metric such as the 12-difference between quaternions (4D vectors) may be used. This term may help ensure that changes in the virtual camera pose occur smoothly.
[0075] At block 376, the optimization process determines the difference between the proposed pose of the virtual camera viewpoint in virtual space and the pose of the physical camera in virtual space. Figure 2 , the optimization process may account for the distance 218 between the pose 212 of the optimized virtual camera viewpoint and the pose 214 of the physical camera.
[0076] Referring again to equation E_V_0(T), the operation of block 376 may represent E_rotation_following, i.e., the rotation following term. As mentioned above, this term may measure the difference between the virtual camera rotation of the current frame (another way of referring to the virtual camera "orientation") and the real camera rotation of the current frame, which may ensure that the virtual camera rotation follows the real camera rotation.
[0077] At box 378, the optimization process determines whether the virtual camera has rotated too far away from the real camera. Referring again to the equation for E_V_0(T), the operation of box 378 represents the distortion term E_distortion. This term can measure the weighted spherical angle between the virtual camera rotation and the real camera rotation: E = L(angle) * angle, where the weight L(angle) = 1 / (1+exp(\beta_1*(angle-\beta_0))) is a logistic regression that is close to 0 if the angle is less than a threshold \beta_0 and close to 1 otherwise. The parameter \beta_1 controls the speed of the translation from 0 to 1. This term may only be effective (i.e., output a large value) when the virtual camera is rotated too far away from the real camera. This term can help ensure that the virtual camera is rotated only within a certain range from the real camera and is not rotated far enough to cause visually observable perspective distortion.
[0078] At block 380, the optimization process measures the change in offset from the virtual principal point between the current frame and the previous frame (e.g., the previous frame). Referring again to the equation for E_V_0(T), the operation of block 380 represents the offset smoothness term, E_offset_smoothness. This term measures the change in offset from the virtual principal point between the current frame and the previous frame and helps ensure that the offset changes smoothly across frames.
[0079] Referring again to the equation for E_V_0(T), ( Figures 3A to 3EThe undefined pixel term E_undefined_pixel (not shown in the figure) calculates the transformation used to map the real image projected by P-(i, j) to P_v(T) for each scan line, warps the video frame to the virtual camera plane, and measures the number of undefined pixels in the warped frame. This penalty will result in a virtual pose solution for undefined pixels. A reference quantity r can be used to control the sensitivity of the penalty, for example, such that E = number of undefined pixels / r. By adjusting r to a larger value, this term can output a smaller value to disable the penalty. When r is small, this term can output a larger value to dominate the entire optimization process and avoid undefined pixels.
[0080] The weights of each term in the equation applied to E_V_0(T) can be changed for different frames. For example, when there is no valid face or no face is detected in the frame, the weight w_1 can be set to 0 so that the optimization process does not perform face stabilization when determining the pose of the virtual camera. In addition, the weights w_1 to w_6 can be changed based on the landmark detection confidence. For example, when the landmark detection confidence is low, w_1 can be reduced to avoid fitting with unreliable landmarks, and when the landmarks across frames are unstable, the weights associated with smoothness can be increased to avoid virtual camera movement caused by unstable landmarks. In addition, when the pose is large, the weights associated with smoothness can be increased to avoid virtual camera movement caused by pose changes.
[0081] At block 382, the computing system tests for stability. This test can be performed by generating a test stabilized view of a video stream frame using the optimized pose of the virtual camera viewpoint (block 384). For example, computing system 220 can generate a stabilized view of frame 230 from virtual camera position 212. For example, based on V_0(T), the computing system can calculate a transformation for mapping the real image projected by P_(i, j) to P_v(T) for each scan line.
[0082] At block 386, the computing system determines whether the test stabilized view of the frame includes undefined pixels. If so, the computing system can select a different virtual camera viewpoint (block 387). More specifically, the computing system can determine whether P_v(T) for each scan line will leave any pixels undefined in the output frame. If so, the reference quantity r in the undefined pixel term may be too large. A binary search can be performed for this reference quantity between a preset minimum value and its current value, and a maximum reference quantity that does not result in undefined pixels can be selected, and the optimized result can be used as the final virtual camera pose V(T).
[0083] For example, if the virtual camera viewpoint does not leave any undefined pixels, because the computing system 220 determines that the stabilized view of the frame 230 obtained from the virtual camera's position 212 does not include any undefined pixels, the computing system can use the determined virtual camera viewpoint as the final virtual camera viewpoint (box 388).
[0084] At block 390, the computing system uses the optimized pose of the virtual camera viewpoint to generate a stabilized view of the frame. For example, the image warping engine can load the mapping output from the motion filtering engine and use it to map each pixel in the input frame to the output frame. For example, the computing system can use the final V(T) to generate the final virtual projection matrix P′_v(T) and can calculate the final mapping for each scan line. This task is common in image and graphics processing systems, and different solutions can be implemented based on whether the process is optimized for performance or quality, or some mixture of the two.
[0085] At block 392, the computing system generates a stabilized view of the frame by mapping a subset of the scan lines of the frame to the perspective viewed from the optimized pose of the virtual camera viewpoint, and interpolating the other scan lines in the frame. For example, instead of computing P_(i, j) and mapping for each individual scan line, the computing system may compute only a subset of the scan lines and may use interpolation therebetween to generate a dense map.
[0086] At box 394, the computing system presents a stabilized view of a frame of the video stream on a display and stores the stabilized view of the frame in a memory. In some examples, the presentation of the stabilized view of the frame is performed in real time while the video is being recorded, for example, before the next frame is captured or shortly after the frame is captured (e.g., before another 2, 5, 10, or 15 frames are captured during an ongoing recording). In some examples, storing the stabilized view of the frame in the memory may include deleting or otherwise not persistently storing the original, unstable frame. In this way, after the computing system completes the video recording, the computing system may store a stabilized version of the video and may not store an unstable version of the video.
[0087] At block 396 , the computing system repeats this process for the next frame of the video, for example, by starting the process again at block 312 with the next frame, unless the video includes no more frames.
[0088] The techniques described herein can be used to stabilize video to reduce both camera motion and non-object-oriented motion. For example, a computing system can track the center of another moving object, such as a soccer ball being thrown or a vehicle moving across frames, and stabilize the video to reduce both camera motion and non-object-oriented motion.
[0089] In addition to the foregoing, controls can be provided to the user that allow the user to make choices about whether and when the systems, programs, or features described herein can enable the collection of user information (e.g., location information about the device). Any location information about the place where the location information is obtained (such as a city, zip code, or state level) can be generalized so that the user's specific location cannot be determined. Thus, the user can control what information is collected about the user, how the information is used, and what information is provided to or about the user. Furthermore, any location determination performed with respect to the techniques described herein can only identify the posture of the device relative to the initial posture at the start of video recording, rather than the absolute geographic location of the device.
[0090] Now refer to Figure 4 , illustrates a conceptual diagram of a system that can be used to implement the systems and methods described herein. In this system, a mobile computing device 410 can wirelessly communicate with a base station 440, which can provide the mobile computing device with wireless access to a number of hosted services 460 via a network 450.
[0091] In this illustration, mobile computing device 410 is depicted as a handheld mobile telephone (e.g., a smartphone or app phone) that includes a touch screen display device 412 for presenting content to and receiving touch-based user input from a user of mobile computing device 410. Other visual, tactile, and auditory output components (e.g., LED lights, a vibration mechanism for tactile output, or a speaker for providing tone, sound generation, or recorded output) may also be provided, as well as a variety of different input components (e.g., a keyboard 414, physical buttons, a trackball, an accelerometer, a gyroscope, and a magnetometer).
[0092] An example visual output mechanism in the form of a display device 412 can take the form of a display with resistive or capacitive touch capabilities. The display device can be used to display video, graphics, images, and text, as well as to coordinate the user touch input location with the location of the displayed information so that the device 410 can associate the user contact at the location of the displayed item with the item. Mobile computing device 410 can also take alternative forms, including as a laptop computer, tablet or tablet computer, personal digital assistant, embedded system (e.g., car navigation system), desktop personal computer, or computerized workstation.
[0093] An example mechanism for receiving user input includes a keyboard 414, which can be a full keyboard or a traditional keyboard that includes keys for the numerals "0-9," "*," and "#." Keyboard 414 receives input when a user physically touches or presses a keyboard key. User manipulation of a trackball 416 or interaction with a trackpad enables the user to provide direction and rate of movement information to mobile computing device 410 (e.g., manipulating the position of a cursor on display device 412).
[0094] Mobile computing device 410 is capable of determining the location of physical contact (e.g., the location of contact by a finger or stylus) with touch screen display device 412. Using touch screen 412, various "virtual" input mechanisms can be generated in which a user interacts with graphical user interface elements depicted on touch screen 412 by contacting the graphical user interface elements. An example of a "virtual" input mechanism is a "software keyboard" in which a keyboard is displayed on the touch screen and the user selects keys by pressing an area of touch screen 412 corresponding to each key.
[0095] Mobile computing device 410 may include mechanical or touch-sensitive buttons 418a-d. Additionally, the mobile computing device may include buttons for adjusting the volume of the output of one or more speakers 420, as well as a button for turning the mobile computing device on or off. Microphone 422 allows mobile computing device 410 to convert audible sound into electrical signals that can be digitally encoded and stored in a computer-readable memory or transmitted to another computing device. Mobile computing device 410 may also include a digital compass, an accelerometer, a proximity sensor, and an ambient light sensor.
[0096] An operating system can provide an interface between the hardware (e.g., input / output mechanisms and a processor that executes instructions retrieved from a computer-readable medium) and software of a mobile computing device. Example operating systems include Android, Chrome, iOS, Mac OS X, Windows 7, Windows Phone 7, Symbian, Blackberry, WebOS, various UNIX operating systems, or proprietary operating systems for computing devices. An operating system can provide a platform for executing application programs, which facilitates interaction between the computing device and a user.
[0097] Mobile computing device 410 may present a graphical user interface to touch screen 412. A graphical user interface is a collection of one or more graphical interface elements and may be static (e.g., the display appears to remain unchanged over a period of time) or dynamic (e.g., the graphical user interface includes graphical interface elements that animate without user input).
[0098] A graphical interface element can be text, a line, a shape, an image, or a combination thereof. For example, a graphical interface element can be an icon displayed on a desktop and associated text for the icon. In some examples, a graphical interface element can be selected through user input. For example, a user can select a graphical interface element by pressing an area of a touch screen corresponding to the display of the graphical interface element. In some examples, a user can manipulate a trackball to highlight a single graphical interface element with focus. User selection of a graphical interface element can invoke a predefined action of the mobile computing device. In some examples, a selectable graphical interface element further or alternatively corresponds to a button on keyboard 404. User selection of a button can invoke a predefined action.
[0099] In some examples, the operating system provides a "desktop" graphical user interface that is displayed after turning on the mobile computing device 410, after activating the mobile computing device 410 from a sleep state, after "unlocking" the mobile computing device 410, or after receiving a user selection of the "home" button 418c. The desktop graphical user interface can display several graphical interface elements that, when selected, invoke corresponding applications. The invoked application may present a graphical interface that replaces the desktop graphical user interface until the application is terminated or hidden from view.
[0100] User input can affect the order in which operations of the mobile computing device 410 are performed. For example, a single-action user input (e.g., a single click on the touch screen, a swipe across the touch screen, contact with a button, or a combination of these actions occurring simultaneously) can invoke an operation that changes the display of a user interface. In the absence of user input, the user interface may not change at a particular time. For example, a multi-touch user input using the touch screen 412 can invoke a mapping application to "zoom in" on a location, but the mapping application may default to zooming in after a few seconds.
[0101] The desktop graphical interface may also display "widgets." A widget is one or more graphical interface elements associated with a currently executing application and displays content on the desktop controlled by the currently executing application. The widget's application may be launched when the mobile device is turned on. Furthermore, a widget may not take full screen control. Instead, the widget may "own" only a small portion of the desktop, displaying content and receiving touchscreen user input within that portion of the desktop.
[0102] Mobile computing device 410 may include one or more location identification mechanisms. The location identification mechanism may include a combination of hardware and software that provides an estimate of the geographic location of the mobile device to the operating system and applications. The location identification mechanism may employ satellite-based positioning techniques, identification of base station transmit antennas, triangulation of multiple base stations, determination of Internet access point IP location, inferred identification of the user's location based on search engine queries, and user-provided location identification (e.g., by receiving a user "check-in" to a location).
[0103] Mobile computing device 410 may include other applications, computing subsystems, and hardware. A call processing unit may receive an indication of an incoming phone call and provide the user with the ability to answer the incoming phone call. A media player may allow the user to listen to music or play movies stored in the local memory of mobile computing device 410. Mobile device 410 may include a digital camera sensor and corresponding image and video capture and editing software. An internet browser may enable a user to view content on a web page by typing in an address corresponding to the web page or selecting a link to the web page.
[0104] Mobile computing device 410 may include an antenna for wirelessly communicating information with base station 440. Base station 440 may be one of many base stations in a collection of base stations (e.g., a mobile phone cellular network) that enables mobile computing device 410 to maintain communication with network 450 as mobile computing device 410 is geographically moved. Computing device 410 may alternatively or additionally communicate with network 450 through a Wi-Fi router or a wired connection (e.g., Ethernet, USB, or firewall). Computing device 410 may also wirelessly communicate with other computing devices using the Bluetooth protocol, or may employ an ad hoc wireless network.
[0105] A service provider operating a network of base stations can connect mobile computing device 410 to network 450 to enable communication between mobile computing device 410 and other computing systems providing services 460. Network 450 is illustrated as a single network, although services 460 can be provided over different networks (e.g., the service provider's internal network, the public switched telephone network, and the Internet). The service provider can operate a server system 452 that routes information packets and voice data between mobile computing device 410 and computing systems associated with services 460.
[0106] Network 450 may connect mobile computing device 410 to a public switched telephone network (PSTN) 462 to establish voice or fax communications between mobile computing device 410 and another computing device. For example, service provider server system 452 may receive an indication of an incoming call for mobile computing device 410 from PSTN 462. Conversely, mobile computing device 410 may send a communication to service provider server system 452 to initiate a telephone call using a telephone number associated with a device accessible via PSTN 462.
[0107] Network 450 can connect mobile computing device 410 to a Voice over Internet Protocol (VoIP) service 464, which routes voice communications over an IP network, as opposed to the PSTN. For example, a user of mobile computing device 410 can invoke a VoIP application and use it to initiate a call. Service provider server system 452 can forward the voice data from the call to the VoIP service, which can route the call over the Internet to the corresponding computing device, potentially using the PSTN for the final leg of the connection.
[0108] Application store 466 can provide users of mobile computing device 410 with the ability to browse a list of remotely stored applications that the user can download and install on mobile computing device 410 via network 450. Application store 466 can serve as a repository for applications developed by third-party application developers. Applications installed on mobile computing device 410 can communicate with a server system designated for that application via network 450. For example, a VoIP application can be downloaded from application store 466, thereby enabling the user to communicate with VoIP service 464.
[0109] Mobile computing device 410 can access content on Internet 468 via network 450. For example, a user of mobile computing device 410 can invoke a web browser application that requests data from a remote computing device accessible at a designated universal resource location. In various examples, some services 460 can be accessed via the Internet.
[0110] The mobile computing device can communicate with a personal computer 470. For example, the personal computer 470 can be a home computer for the user of the mobile computing device 410. Thus, the user can stream media from his personal computer 470. The user can also view the file structure of his personal computer 470 and transfer selected documents between computerized devices.
[0111] Speech recognition service 472 can receive voice communication data recorded by microphone 422 of mobile computing device and convert the voice communication into corresponding text data. In some examples, the converted text is provided to a search engine as a web query, and the corresponding search engine search results are transmitted to mobile computing device 410.
[0112] Mobile computing device 410 can communicate with social network 474. The social network may include numerous members, some of whom have been accepted as related to acquaintances. Applications on mobile computing device 410 can access social network 474 to retrieve information based on the acquaintances of the user of the mobile computing device. For example, a "Contacts" application can retrieve phone numbers of acquaintances of the user. In various examples, content can be delivered to mobile computing device 410 based on the social network distance and connection relationships of the user to other members in the member's social network graph. For example, advertisements and news article content can be selected for the user based on the level of interaction with such content by members who are "close" to the user (e.g., members who are "friends" or "friends of friends").
[0113] Mobile computing device 410 can access a personal collection of contacts 476 via network 450. Each contact can identify an individual and include information about the individual (e.g., phone number, email address, and birthday). Because the collection of contacts is hosted remotely to mobile computing device 410, the user can access and maintain contacts 476 across multiple devices as a common collection of contacts.
[0114] Mobile computing device 410 can access cloud-based applications 478. Cloud computing provides applications (e.g., a word processor or email program) that are hosted remotely from mobile computing device 410 and can be accessed through device 410 using a web browser or dedicated programs. Example cloud-based applications include the Google Docs word processor and spreadsheet service, the Google Gmail webmail service, and the Picasa picture manager.
[0115] Mapping service 480 can provide street maps, routing information, and satellite imagery to mobile computing device 410. An example of a mapping service is Google Maps. Mapping service 480 can also receive queries and return location-specific results. For example, mobile computing device 410 can send the estimated location of the mobile computing device and the user-entered query "pizza" to mapping service 480. Mapping service 480 can return a street map with "markers" superimposed on the map identifying the geographic locations of nearby "pizzas."
[0116] Routing service 482 can provide routing instructions to a user-provided destination to mobile computing device 410. For example, routing service 482 can stream to device 410 a street-level view of the device's estimated location, along with data providing audio commands and overlaid arrows to guide the user of device 410 to the destination.
[0117] Mobile computing device 410 can request various forms of streaming media 484. For example, computing device 410 can request a stream of a pre-recorded video file, a live television program, or a live radio program. Example services that provide streaming media include YOUTUBE and PANDORA.
[0118] The microblogging service 486 may receive a user-input post without identifying a recipient of the post from the mobile computing device 410. The microblogging service 486 may broadcast the post to other members of the microblogging service 486 who have agreed to subscribe to the user.
[0119] Search engine 488 can receive a textual or spoken query entered by a user from mobile computing device 410, determine a collection of Internet-accessible documents that respond to the query, and provide information to device 410 to display a list of results of the search for responsive documents. In the example of receiving a spoken query, speech recognition service 472 can convert the received audio into a textual query that is sent to the search engine.
[0120] These and other services can be implemented in a server system 490. A server system can be a combination of hardware and software that provides a service or collection of services. For example, a group of physically separate and networked computerized devices can operate together as a logical server system unit to handle the operations required to provide services to hundreds of computing devices. A server system is also referred to herein as a computing system.
[0121] In various embodiments, an operation performed "in response to" or "as a result of" another operation (e.g., a determination or identification) is not performed if the previous operation was unsuccessful (e.g., if it was determined that it was not performed). An operation performed "automatically" is an operation performed without user intervention (e.g., intervening user input). Functionality in this document described in conditional language may describe optional implementations. In some examples, "transmitting" from a first device to a second device includes the first device placing data into a network for receipt by the second device, but may not include the second device receiving the data. Conversely, "receiving" from a first device may include receiving data from a network, but may not include the first device transmitting the data.
[0122] "Determining" by a computing system may include requesting another device to perform a determination and provide the result to the computing system. In addition, "displaying" or "presenting" by a computing system may include sending data for causing another device to display or present the reference information.
[0123] Figure 5 is a block diagram of a computing device 500, 550 that can be used as a client or a server or multiple servers to implement the systems and methods described in this document. Computing device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown here, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the inventions described and / or claimed in this document.
[0124] Computing device 500 includes a processor 502, memory 504, storage device 506, a high-speed interface 508 connected to memory 504 and a high-speed expansion port 510, and a low-speed interface 514 connected to a low-speed bus 514 and storage device 506. Each of components 502, 504, 506, 508, 510, and 512 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. Processor 502 can process instructions for execution within computing device 500, including instructions stored in memory 504 or on storage device 506 for displaying graphical information for a GUI on an external input / output device such as a display 516 coupled to high-speed interface 508. In other embodiments, multiple processors and / or multiple buses can be used, along with multiple memories and multiple types of memory, as appropriate. In addition, multiple computing devices 500 can be connected, with each device providing a portion of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
[0125] Memory 504 stores information within computing device 500. In one embodiment, memory 504 is one or more volatile memory units. In another embodiment, memory 504 is one or more non-volatile memory units. Memory 504 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0126] The storage device 506 can provide mass storage for the computing device 500. In one embodiment, the storage device 506 can be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory, or other similar solid-state storage device or array of devices, including devices in a storage area network or other configuration. A computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as the memory 504, the storage device 506, or a memory on the processor 502.
[0127] The high-speed controller 508 manages bandwidth-intensive operations of the computing device 500, while the low-speed controller 512 manages less bandwidth-intensive operations. This allocation of functionality is merely an example. In one embodiment, the high-speed controller 508 is coupled to the memory 504, the display 516 (e.g., via a graphics processor or accelerator), and to the high-speed expansion port 510, which can accept various expansion cards (not shown). In this embodiment, the low-speed controller 512 is coupled to the storage device 506 and the low-speed expansion port 514. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, or networking device such as a switch or router, for example, via a network adapter.
[0128] As shown in the figure, computing device 500 can be implemented in many different forms. For example, it can be implemented as a standard server 520, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. In addition, it can be implemented in a personal computer such as laptop computer 522. Alternatively, components from computing device 500 can be combined with other components in a mobile device (not shown) such as device 550. Each of such devices can contain one or more of computing devices 500, 550, and the entire system can include multiple computing devices 500, 550 in communication with each other.
[0129] Computing device 550 includes a processor 552, a memory 564, an input / output device such as a display 554, a communication interface 566, and a transceiver 568, among other components. Device 550 may also be provided with a storage device, such as a microdrive or other device, to provide additional storage. Each of components 550, 552, 564, 554, 566, and 568 is interconnected using various buses, and several of the components may be mounted on a common motherboard or otherwise as appropriate.
[0130] The processor 552 can execute instructions within the computing device 550, including instructions stored in the memory 564. The processor can be implemented as a chipset of chips that include separate and multiple analog and digital processors. In addition, the processor can be implemented using any of a variety of architectures. For example, the processor can be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimum Instruction Set Computer) processor. The processor can, for example, provide coordination for other components of the device 550, such as control of a user interface, applications run by the device 550, and wireless communications performed by the device 550.
[0131] The processor 552 can communicate with the user through a control interface 558 and a display interface 556 coupled to the display 554. The display 554 can be, for example, a TFT (thin film transistor liquid crystal display) display or an OLED (organic light emitting diode) display or other appropriate display technology. The display interface 556 may include appropriate circuits for driving the display 554 to present graphics and other information to the user. The control interface 558 can receive commands from the user and convert them for submission to the processor 552. In addition, an external interface 562 in communication with the processor 552 can be provided to enable near-area communication of the device 550 with other devices. The external interface 562 can be provided for wired communication in some embodiments, for example, or for wireless communication in other embodiments, and multiple interfaces can also be used.
[0132] Memory 564 stores information within computing device 550. Memory 564 can be implemented as one or more of the following: one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Expansion memory 574 can also be provided and connected to device 550 via expansion interface 572, which can include, for example, a SIMM (Single In-line Memory Module) card interface. This expansion memory 574 can provide additional storage space for device 550 or store applications or other information for device 550. Specifically, expansion memory 574 can include instructions for executing or supplementing the above-described processes and can also include security information. Thus, for example, expansion memory 574 can be provided as a security module for device 550 and can be programmed with instructions that allow for secure use of device 550. Furthermore, secure applications and additional information can be provided via a SIMM card, such as placing identification information on the SIMM card in an unhackable manner.
[0133] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine readable medium, such as memory 564, expansion memory 574, or memory on processor 552, which can be received, for example, via transceiver 568 or external interface 562.
[0134] Device 550 can communicate wirelessly via a communication interface 566, which may include digital signal processing circuitry, if necessary. The communication interface 566 may provide for communication in various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication may occur, for example, via a radio frequency transceiver 568. In addition, short-range communication may occur, such as using Bluetooth, WiFi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 570 may provide additional navigation and location-related wireless data to device 550, which may be used as appropriate by applications running on device 550.
[0135] Device 550 may also communicate audibly using audio codec 560, which may receive spoken information from a user and convert it into usable digital information. Audio codec 560 may also generate audible sounds for the user, such as through a speaker, for example, in a headset of device 550. Such sounds may include sounds from voice phone calls, may include recorded sounds (e.g., voice messages, music files, etc.), and may also include sounds generated by applications operating on device 550.
[0136] As shown in the figure, computing device 550 can be implemented in many different forms. For example, it can be implemented as a cellular phone 580. It can also be implemented as part of a smart phone 582, a personal digital assistant, or other similar mobile device.
[0137] The additional computing device 500 or 550 may include a Universal Serial Bus (USB) flash drive. The USB flash drive may store an operating system and other application programs. The USB flash drive may include input / output components, such as a wireless transmitter or a USB connector that may be plugged into a USB port of another computing device.
[0138] Various implementations of the systems and techniques described herein can be implemented using digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable on a programmable system comprising at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0139] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms "machine-readable medium," "computer-readable medium," and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) that the user can use to provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, voice, or tactile input.
[0141] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface or a web browser that a user can use to interact with implementations of the systems and techniques described herein), or includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), a peer-to-peer network (with self-organizing or static members), a grid computing infrastructure, and the Internet.
[0142] A computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0143] Although a few embodiments have been described in detail above, other modifications are possible. Furthermore, other mechanisms for implementing the systems and methods described in this document may be used. In addition, the logic flows depicted in the accompanying drawings do not require the particular order or sequential sequence shown to achieve the desired results. Other steps may be provided, or steps may be eliminated from the described flows, and other components may be added to or removed from the described systems. Therefore, other embodiments are within the scope of the appended claims.
Claims
1. A method for stabilizing a video, comprising: receiving, by a computing system, a video stream comprising a plurality of frames captured by a physical camera; identifying, by the computing system, in a current frame of the plurality of frames of the video stream, a location of a feature of an object depicted in the current frame; The computing system determines a stable position of the feature in the current frame by: determining a first distance between a potential stable position of the feature in the current frame and the identified position of the feature, determining a second distance between the potential stable position of the feature and a previously determined position of the feature, and wherein determining the stable position of the feature comprises determining a weighted average of the first distance and the second distance on the premise that the first distance is within a cropping range of the video stream; and A stabilized view of the current frame is generated by the computing system using the stabilized position of the feature.
2. The method according to claim 1, wherein The cropping range includes a defined area of the current frame of the video stream surrounding the identified location of the feature.
3. The method according to claim 1, wherein The stabilized view of the current frame has a different crop than a crop of a stabilized view of a previous frame in the plurality of frames of the video stream.
4. The method according to claim 1, wherein Determining the stable position of the feature further comprises: determining a compliance value for the feature, the compliance value comprising a first position change between the potential stable position of the feature and the identified position of the feature; determining a smoothness value for the feature, the smoothness value determined based on a second change in position of the feature between the potentially stable position of the feature and a previous position of the feature in a previous frame of the plurality of frames of the video stream; and A sum of the compliance value of the feature and the smoothness value of the feature is optimized, the stable position of the feature corresponding to the optimized sum of the compliance value and the smoothness value.
5. The method according to claim 4, wherein The sum of the compliance value of the feature and the smoothness value of the feature includes the sum of the weighted compliance value of the feature and the weighted smoothness value of the feature.
6. The method according to claim 1, wherein Identifying the locations of features of the object depicted in the current frame includes: identifying a bounding box of the object; and A center of the bounding box is identified as the feature of the object by identifying positions of a plurality of object landmarks within the bounding box depicted in the current frame and determining an average position of the plurality of object landmarks.
7. The method according to claim 1, further comprising: The object depicted in the current frame of the video stream is selected as a tracking object from among a plurality of objects depicted in the current frame of the video stream by: selecting the tracked object based on a size of each of the plurality of objects, selecting the tracked object based on a distance of each of the plurality of objects from a center of the current frame, or The tracked object is selected based on a distance between a tracked object selected in a previous frame and each of the plurality of objects.
8. The method according to claim 1, further comprising: A stabilized view of the current frame is presented by the computing system on a display of the computing system.
9. The method according to claim 1, wherein The stabilized view of the current frame represents a partially zoomed-in version of the video stream.
10. The method according to claim 1, further comprising: Determining the pose of the physical camera in space; as well as Determine the optimal pose of the virtual camera viewpoint in virtual space, and Wherein, generating the stabilized view of the current frame is based on the optimized posture of the virtual camera viewpoint.
11. A computing device comprising: camera; A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed, cause one or more processors of the computing device to: receiving a video stream comprising a plurality of frames captured by the camera; identifying, in a current frame of the plurality of frames of the video stream, a location of a feature of an object depicted in the current frame; A stable position of the feature in the current frame is determined using a process, the process: determining a first distance between a potential stable position of the feature in the current frame and the identified position of the feature, determining a second distance between the potential stable position of the feature and a previously determined position of the feature, and wherein determining the stable position of the feature comprises determining a weighted average of the first distance and the second distance on the premise that the first distance is within a cropping range of the video stream; and A stabilized view of the current frame is generated using the stabilized positions of the features.
12. The computing device of claim 11, wherein: The cropping range includes a defined area of the current frame of the video stream surrounding the identified location of the feature.
13. The computing device of claim 11, wherein: The stabilized view of the current frame has a different crop than a crop of a stabilized view of a previous frame in the plurality of frames of the video stream.
14. The computing device of claim 11, wherein: The process of determining the stable position of the feature further comprises: determining a compliance value for the feature, the compliance value comprising a first position change between the potential stable position of the feature and the identified position of the feature; determining a smoothness value for the feature, the smoothness value determined based on a second change in position of the feature between the potentially stable position of the feature and a previous position of the feature in a previous frame of the plurality of frames of the video stream; and A sum of the compliance value of the feature and the smoothness value of the feature is optimized, the stable position of the feature corresponding to the optimized sum of the compliance value and the smoothness value.
15. The computing device of claim 14, wherein: The sum of the compliance value of the feature and the smoothness value of the feature includes the sum of the weighted compliance value of the feature and the weighted smoothness value of the feature.
16. The computing device of claim 11, wherein: The computer-executable instructions, when executed, cause the one or more processors of the computing device to identify the location of the feature of the object in the following manner: identifying a bounding box of the object; and A center of the bounding box is identified as the feature of the object by identifying positions of a plurality of object landmarks within the bounding box depicted in the current frame and determining an average position of the plurality of object landmarks.
17. The computing device of claim 11 , the non-transitory computer-readable storage medium further comprising computer-executable instructions that, when executed, cause the one or more processors of the computing device to: The object depicted in the current frame of the video stream is selected as a tracking object from among a plurality of objects depicted in the current frame of the video stream by: selecting the tracked object based on a size of each of the plurality of objects, selecting the tracked object based on a distance of each of the plurality of objects from a center of the current frame, or The tracked object is selected based on a distance between a tracked object selected in a previous frame and each of the plurality of objects.
18. The computing device of claim 11 , the non-transitory computer-readable storage medium further comprising computer-executable instructions that, when executed, cause the one or more processors of the computing device to: A stabilized view of the current frame is presented on a display of the computing device.
19. The computing device of claim 11, wherein: The stabilized view of the current frame represents a partially zoomed-in version of the video stream.
20. A computer program product comprising instructions which, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 10.
Citation Information
Patent Citations
Stabilizing video
CN107851302A
Real-time video stabilization for mobile devices based on on-board motion sensing
US20170332018A1