Temporally stable and occlusion-free caption rendering position in the Z plane for stereoscopic video
By using a system to adjust disparity values for captions in stereoscopic video based on optical flow maps and just noticeable difference requirements, the issues of occlusion and eye strain are resolved, ensuring stable and comfortable viewing.
Patent Information
- Application Number
- JP2024135171
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-07-06
- Filing Date
- 2024-08-14
- Publication Date
- 2026-02-02
- Estimated Expiration
- 2044-08-14
AI Technical Summary
Captions in stereoscopic video can cause occlusion, eye strain, and temporal instability due to improper placement and movement in the Z plane, leading to viewer discomfort.
A system determines disparity values for captions using a high-quality frame-by-frame optical flow/disparity map and applies a bidirectional window to smooth these values, adjusting them iteratively to meet a just noticeable difference requirement, ensuring stable placement and minimal movement in the Z plane.
The solution prevents captions from being occluded and minimizes eye strain by maintaining consistent placement and reducing perceptible movement, enhancing the viewing experience in stereoscopic video.
Smart Images

Figure 0007809762000001 
Figure 0007809762000002 
Figure 0007809762000003
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001]
[0001] Pursuant to 35 U.S.C. § 119(e), this application is entitled to and claims the benefit of the filing date of U.S. Provisional Patent Application No. 63 / 627,646, entitled "Automatic Method to Produce Temporarily Stable Occlusion Free Subtitle Rendering Position in Z-plane for Stereoscopic Video," filed January 31, 2024, the contents of which are incorporated herein by reference in their entirety for all purposes. [Background technology]
[0002]
[0002] Stereoscopic video is a type of media presentation that uses a pair of videos to reconstruct disparity, which is the basis of three-dimensional (3D) visual perception in the human eye. Disparity, or Z-plane position, is defined as the horizontal displacement between the left and right videos. In a convergent camera setup, a more negative value of disparity for an object implies that the object is closer to the viewer. In a parallel camera setup, the disparity value is always positive, and a larger disparity value in a parallel camera setup indicates that the object is closer to the viewer.
[0003]
[0003] Captions may be rendered in stereoscopic video. The placement of captions in the Z plane can cause problems for viewers. If captions are placed too far away from the viewer, they may be occluded by objects in front of them. Also, if captions are placed too close to the viewer in the Z plane, the viewer may have to focus on objects that are farther away in the Z plane and on captions that are too close, which can cause eye strain. Another problem is when the Z plane position of the caption changes drastically (e.g., from close to far) across successive frames, which causes eye strain as the viewer searches for captions at different Z plane positions.
[0004]
[0004] The included drawings are for illustrative purposes and serve only to provide examples of possible structure and operation for the disclosed inventive systems, apparatus, methods, and computer program products. These drawings in no way limit any changes in form and detail that may be made by one skilled in the art without departing from the spirit and scope of the disclosed implementations. [Brief explanation of the drawings]
[0005] [Figure 1] FIG. 1 illustrates a system for processing stereoscopic video, according to some embodiments. [Figure 2]
[0006] FIG. 2 illustrates a simplified flowchart for processing disparity values for captions, according to some embodiments. [Figure 3]
[0007] FIG. 3 illustrates a simplified flowchart of backward-pass disparity adjustment, according to some embodiments. [Figure 4]
[0008] FIG. 4 illustrates a simplified flowchart for forward-pass disparity adjustment, according to some embodiments. [Figure 5]
[0009] FIG. 5 illustrates a graph of raw disparity values for stereoscopic video, according to some embodiments. [Figure 6]
[0010] FIG. 6 illustrates a graph with adjustments using a rendering information generator, according to some embodiments. [Figure 7]
[0011] FIG. 7 shows a graph of adjusted disparity values according to some embodiments. [Figure 8]
[0012] FIG. 8 illustrates an example of a computing device, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0006]
[0013] Techniques for video display systems are described herein. In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of some embodiments. Some embodiments, as defined by the claims, may include some or all of the features in these examples, alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0007]
[0014] <System Overview>
[0008]
[0015] In some embodiments, a system determines disparity values for text used to display captions in stereoscopic video. The text may be different types of text. In some embodiments, the text may be referred to as captions, which may be subtitles or closed captions for stereoscopic video. The term "caption" may indicate the function of the text. In some embodiments, the term captions may be either closed captions or subtitles. Closed captions provide a text transcription of the interactive portions of a video. They are designed for use by hard-of-hearing viewers. Subtitles provide a text translation of the interactive portions of a video. Subtitles may assume that the viewer can hear the audio but cannot understand the language. Both closed captions and subtitles are timed text that reflects the content of the interactive portions of a video. In some embodiments, a stereoscopic video may have a first video and a second video, referred to as the left video and the right video. Captions may be displayed on both the left video and the right video. In the following description, the process may be applied to the left video or the right video.
[0009]
[0016] To process where to display captions in stereoscopic video, the system uses a map describing the motion of objects in frames of the stereoscopic video and the difference in position or depth of corresponding objects in a pair of frames from the left and right videos of the stereoscopic video. The map may be a high-quality frame-by-frame optical flow / disparity map. The map may include disparity values for one or more pixels of a pair of frames. The left and right videos may include the same frame numbers. The left and right videos may be the same size or different sizes. For example, the resolution of the left and right views of the stereoscopic video need not be the same for this process to work. The disparity map may reference disparity values for frame numbers. When stereoscopic video is referenced in the processes herein, the process may determine the disparity values from the relationship between the left and right videos.
[0010]
[0017] The system then applies a window, such as a fixed bidirectional window, to filter / smoothe, etc., the noisy disparity values across multiple frames. The bidirectional window may consider disparity values in the forward and backward directions for a selected frame. After the disparity values are smoothed, the system starts with a first value set to x, such as the lowest (most negative) disparity value from the global set of disparity values in the stereoscopic video that is closest to the viewer. The system then finds the frame containing the cell value at x. The system loops through each frame and each cell location that contains a disparity value equal to x. The cell location may be associated with a disparity value. In the loop, the system adjusts the disparity values in both the backward and forward directions to meet a requirement, such as a just noticeable difference (JND) requirement. The just noticeable difference requirement may define a minimum rate of change in the Z plane for the caption to be perceptible. After that processing, the system gradually adjusts the disparity values to be farther away (e.g., increasing) from the viewer, and each time, the system iteratively adjusts the disparity values in both the backward and forward directions to meet the requirements. After finishing the processing, such as when the lowest processed disparity value meets a threshold, the system outputs the disparity value for the frame. The system then uses the adjusted disparity value to display captions along with the stereoscopic video.
[0011]
[0018] <System>
[0012]
[0019] 1 illustrates a system 100 for processing stereoscopic video, according to some embodiments. Stereoscopic video is a type of media presentation that utilizes a pair of videos to reconstruct the parallax that is the basis of the human eye's three-dimensional (3D) visual perception. Stereoscopic video can be video that simulates depth perception using two videos captured using two cameras, one for the viewer's left eye and one for the right eye. These images, when viewed via a client device 104, create the illusion of three-dimensional depth.
[0013]
[0020] The convergent camera system 106 includes multiple cameras, such as left camera 116-1 and right camera 116-2, that capture respective videos (e.g., left video and right video). The convergent camera system 106 may be used in any stereoscopic camera geometry, such as a convergent stereoscopic camera setup, a parallel stereoscopic camera setup, any convergent camera setup that can be calibrated to a parallel stereoscopic camera, etc. A video of an object 114 may be captured (this video may include multiple objects). Parallax, or Z-plane position, is defined as the horizontal displacement between the left and right videos of the captured image of the object 114. In a convergent camera setup, a more negative value of parallax implies that the object is closer to the viewer.
[0014]
[0021] The server system 102 receives the stereoscopic video. Although depicted as receiving video from the converged camera system 106, the server system 102 may receive the stereoscopic video from any source. Captions may be displayed along with the stereoscopic video. As mentioned above, captions may include subtitles, closed captions, or other text that is displayed along with the stereoscopic video.
[0015]
[0022] The server system 102 includes a rendering information generator 108 that analyzes the stereoscopic video and generates caption rendering information for the captions. The caption rendering information may include information used to display the captions along with the stereoscopic video, such as the coordinates of the Z-plane position of the caption. The coordinates of the Z-plane position may be determined based on the adjusted disparity values, as described below.
[0016]
[0023] The client device 104 receives the stereoscopic video and caption rendering information. The client device 104 may include different devices such as a virtual reality headset, a 3D television, a smartphone, a tablet, 3D glasses, etc. The client device 104 includes a media player 112 that plays the stereoscopic video in 3D space. Two videos (e.g., from the left camera 116-1 and the right camera 116-2, respectively) can be viewed in the left view 110-2 and the right view 110-2. Captions can be displayed based on the respective disparity values of the left video in the left view 110-2 and the right video in the right view 110-2.
[0017]
[0024] To render captions in stereoscopic video, for each caption phrase, the caption rendering information may include a specified disparity value to instruct the media player 112 on the client device 104 to render the caption at an appropriate location in the Z plane. The client device 104 may have different processes for using the disparity value to display the caption in the video. The appropriate location is defined as a range, such as a narrow, achievable range of disparity values, so that the caption is not placed too far from the viewer's viewpoint in 3D space, which may cause the caption to appear buried under the video, and also not placed too close to the viewpoint, which may cause occlusion and fatigue for the viewer.
[0018]
[0025] It is common for caption phrases to last longer than a fraction of a second. It can be equally important to ensure that the caption appears to be held in a constant location in 3D space for the duration that it is displayed in the stereoscopic video. If the caption moves with small jitters throughout that time, it is likely to cause discomfort and a subpar viewing experience for the user. And if the caption moves significantly throughout that time, the eye movements required to locate (e.g., find) the moved caption in 3D space are likely to be distracting and uncomfortable for the viewer.
[0019]
[0026] Also, the underlying scene of the stereoscopic video may change while the captions are being rendered, possibly changing the premise of the first problem to be solved: where to place the captions. The rendering information generator 108 uses a process that can solve both the problem of caption placement and constraining the movement of the captions in 3D space.
[0020]
[0027] In some embodiments, the rendering information generator 108 uses maps such as high-quality per-frame optical flow / disparity maps. The optical flow / disparity maps may describe the motion of objects in a frame and the difference in the object's position or depth in a pair of frames in the video. There may be a disparity map for each frame, and a particular disparity map may or may not be for the full set of pixels in the frame. In the full-pixel case, there are the same number of disparity values per frame, as is the resolution of the video frame. In the non-full-pixel case, there may be N*M disparity values. The disparity map may be N*M in size, where N may be >= 1 and <= the video height and M may be >= 1 and <= the video width. When the processing of disparity values is described for a disparity map, it may be for either the first disparity map or the second disparity map.
[0021]
[0028] The map may be generated in different ways, such as by a prediction network (e.g., a neural network) based on an estimation method for generating frame-by-frame disparity maps for stereoscopic video, but may also be generated using other processes. The process for determining the rendering information may be independent of the particular choice of disparity map estimation method, as long as the estimation method is frame-by-frame and of high quality, such as above a threshold. The rendering information generator 108 may then apply a window, such as a fixed bidirectional window, to filter / smooth noisy measurements of the disparity values. The bidirectional window may consider disparity values in the forward and backward directions for the frame. A loop is then used in which the rendering information generator 108 starts from a first value, such as the lowest (most negative) disparity value, and gradually increases that value, each time the subtitle processing system 108 iteratively adjusts the disparity values both backward and forward to meet requirements, such as a just-noticeable-difference requirement, for the disparity value change rate. The just-noticeable-difference requirement may define a threshold, such as a minimum rate of change that may be perceptible, or another desired value (e.g., a set rate of change). Finally, the rendering information generator 108 clamps the processed disparity values to an achievable range as specified.
[0022]
[0029] This process offers many advantages: for example, captions displayed using adjusted disparity values may no longer be occluded by objects in the stereoscopic video, and problems with captions being too close in the Z plane or shifting across multiple frames may be avoided.
[0023]
[0030] The process for generating adjusted disparity values for captions is described in more detail below.
[0024]
[0031] <Overall process>
[0025]
[0032] FIG. 2 illustrates a simplified flowchart 200 for processing disparity values for captions, according to some embodiments. The following process may be performed for either the right or left video. The following process may be based on a setup in which a more negative value of disparity for an object implies that the object is closer to the viewer. In another setup, the disparity value is always positive, with larger disparity values indicating closer to the viewer, and this setup may have a different flow. For example, the minimum disparity may be changed to the maximum disparity, and the value of x may be decremented. At 202, the rendering information generator 108 determines a disparity estimation. The disparity estimation may calculate the difference in position between two corresponding images from two videos, such as the left and right images from the left and right videos of a stereoscopic video. The disparity estimation may generate a disparity map for each frame, such as a disparity map for corresponding frame 1, corresponding frame 2, corresponding frame 3, etc., of the left and right videos. As mentioned above, the disparity estimation may use different processes to generate the disparity map.
[0026]
[0033] At 204, the rendering information generator 108 calculates the minimum disparity for a cell of a predetermined grid. Although a minimum disparity value is described, the determined disparity value may be the disparity value closest to the viewer (in some cases, the most positive). In some implementations, this may be the maximum disparity value. A value other than the minimum for a cell may also be used, such as an average disparity value for the cell. A cell may be a predetermined region of a frame, which may be a portion of a frame or the entire frame. There may be one or more cells in a frame. The minimum (most negative) disparity value may be determined from the disparity values from the cells of the disparity map.
[0027]
[0034] At 206, the rendering information generator 108 performs a smoothing operation, such as a fixed window temporal smoothing operation. The temporal smoothing operation may use a fixed window, which may reduce variation in disparity estimates across multiple frames. This may remove or adjust disparity values that may be outliers. The smoothing operation may or may not be performed.
[0028]
[0035] The rendering information generator 108 may then process the disparity values and adjust some of the disparity values to improve the display of captions in the stereoscopic video. The adjustment may adjust the disparity values in a frame to just a noticeable distance in the backward and forward directions. In some embodiments of the process, at 208, the rendering information generator 108 sets a value x, where x = a global minimum disparity for all frames in the video. For example, the disparity values for all cells are analyzed, and the minimum disparity value from the disparity values of these cells is set as the value of x.
[0029]
[0036] At 210, the rendering information generator 108 determines whether the value x is greater than N, where N may be a threshold value (e.g., a positive or negative number). In some embodiments, the value of N may be predetermined, such as set to "0," although other values may be used. This check is used to determine when to stop adjusting the disparity values. As described below, the process starts with the minimum disparity value and gradually increases the disparity value found from the global disparity value. The process may be stopped when a maximum disparity value is reached. For example, the system may not adjust the disparity value when the disparity value is greater than a value such as zero. Stopping the process when no further checks of the disparity value are necessary may improve the speed of the adjustment.
[0030]
[0037] If the value of x is not greater than N, then at 212, the rendering information generator 108 finds a frame with a disparity value equal to x (disparity == x). There may be frames with more than one disparity value equal to x (e.g., some frames may contain one cell with the minimum disparity value, and some frames may have multiple cells with the minimum disparity value). For each frame found, the following may be performed as described: At 214, the rendering information generator 108 determines the cell location in each frame that has a disparity value equal to x (disparity == x). For example, the minimum disparity value may be -8000.
[0031]
[0038] For each determined cell location, the following may be performed as described at 216 and 218. At 216, the rendering information generator 108 performs a backward pass disparity adjustment for each of the locations containing a disparity value equal to x. This process is described in FIG. 3. Then, at 218, the rendering information generator 108 performs a forward pass disparity adjustment for each of the locations in the frame containing a disparity value equal to x. This process is further described in FIG. 4. The process may analyze locations in previous frames (backward pass) and subsequent frames (forward pass) to adjust the disparity values of these locations, if necessary.
[0032]
[0039] After performing the backward-pass disparity adjustment and the forward-pass disparity adjustment for each location, the following is performed: Each location may be associated with a disparity value that may be adjusted. At 220, the rendering information generator 108 increments the disparity value to the next lowest minimum disparity value (x++) to determine the next disparity value that is farther away from the viewer. If a more negative disparity value is farther away from the viewer, the value of x may be decremented. For example, if the minimum disparity value is -8000, the next highest minimum disparity value may be greater than -8000, such as -7999. This step may move to disparity values farther away from the viewer, which may be smaller positive numbers if the setup has larger disparity values closer to the viewer. The process then repeats to 210, where the process is executed using the new value of x.
[0033]
[0040] If the value x is greater than N, then at 222, the rendering information generator 108 may perform a clamping operation to clamp the disparity values (e.g., adjusted or unadjusted) to a feasible range. Clamping may adjust disparity values that are outside the feasible range to values that fill the feasible range. If the feasible range is not used or clamping is undesired, clamping may not be performed. The process then ends, and an adjusted disparity map may be output. The adjusted disparity map may then be used to generate rendering information for captions. For example, the rendering information may be at the adjusted disparity value or at a set distance from the adjusted disparity value. The rendering information for the cell is then sent to the client 104 for use in displaying captions for the frame. The client 104 may use the rendering information to display captions in the Z plane along with the stereoscopic video.
[0034]
[0041] The following describes backward-pass and forward-pass disparity adjustments. These adjustments may establish a rate of Z change at which the caption's movement (in / out in the Z plane) becomes unnoticeable. A just-noticeable difference value may be determined at which the viewer is no longer able to notice a difference in Z change for the caption. This allows the rendering information generator 108 to adjust the disparity values at a rate that may become undetectable to the viewer. This may improve the temporal stability of the captions. The just-noticeable difference value may be set to a different desired value, exceeding the point that would be noticeable if desired. The adjustments may also satisfy a minimum disparity value to avoid occluded captions. The disparity values are iteratively adjusted in both the forward and backward directions to satisfy the minimum disparity value, changing the disparity values at an undetectable rate.
[0035]
[0042] <Backward pass parallax adjustment>
[0036]
[0043] FIG. 3 illustrates a simplified flowchart 300 of backward-pass disparity adjustment, according to some embodiments. The following process may be based on a setup where a more negative value of disparity for an object implies that the object is closer to the viewer. In another setup, where the disparity value is always positive and a larger disparity value indicates closer to the viewer, this setup may have a different flow. For example, the just-noticeable difference may be negative, and the comparison in 308 may be "greater than" instead of "less than." The following may be performed for each location determined in 214 of FIG. 2: At 302, the rendering information generator 108 sets the current frame index to c. The index may be the frame selected in 212 of FIG. 2 based on its disparity value. At 304, the rendering information generator 108 sets the previous frame index to the value of p, where p=c−1. This may point to a position that is the frame before the current frame, if applicable (e.g., the previous frame index cannot go before the start of the video). The loop is then executed until the value of p is less than 0 (e.g., the value of p cannot go before the first frame of the video).
[0037]
[0044] At 306, the rendering information generator 108 determines whether the value of p is less than zero. This may test whether the first frame of the video has been reached and already analyzed. If so, the process ends. If not, at 308, the rendering information generator 108 determines whether disp[p] is less than disp[c] to determine whether disp[p] is closer to the viewer. Note that more negative numbers may be less than less negative numbers (e.g., −8000<−7800). Similarly, less positive numbers have smaller values (+1000<+2000). This may determine whether the disparity value of the previous index p (e.g., the pixel in the previous frame that is in the same position as the current pixel in the current frame being analyzed) is less than the disparity value of the current index c (e.g., the pixel in the current frame that is in the same position). If so, the process may end. One reason the process may terminate is because the disparity value of the previous frame may be smaller than the disparity value of the current frame and does not need to be adjusted (e.g., the disparity value may have already been adjusted in a previous iteration because the process starts from the minimum disparity value). If not, at 310, the rendering information generator 108 determines whether the disparity value at the previous index is greater than the disparity value of the current index plus just the noticeable difference value. For example, a disparity value for the previous frame (e.g., −7800) that is greater than the disparity value for the current frame (e.g., −8000) plus the minimum noticeable distance value (+100) may be (e.g., −7800 > −8000 + 100 = −7800 > −7900). Also, the disparity value for the previous frame (e.g., −8100) that is less than the disparity value for the current frame (e.g., −8000) plus the minimum noticeable distance value (+100) may be (e.g., −8100<−8000+100=−8100<−7900).If this is true, then at 312, the rendering information generator 108 adjusts the disparity value at the previous index p, such as by setting the disparity value at the previous index p to be equal to the disparity value at the current index plus just the noticeable difference value (e.g., −8000 + 100 = −7900). The adjusted disparity value may also be set to the disparity value at the current index plus a value that is just less than or equal to the noticeable difference value. Adjusting the disparity value may ensure that the change in disparity value between the previous frame and the current frame is just less than or equal to the noticeable difference value for this pixel. That is, a viewer may not notice a change in the position of a caption that changes from −8000 to −7900 instead of −8000 to −7800. Otherwise, the process proceeds directly to 314.
[0038]
[0045] At 314, the rendering information generator 108 decrements the values of c and p by 1. For example, if the frame number was c=1000 and p=999, the new value of c is 999 and the new value of p is 998. The process then repeats to 306, where it is determined whether p is less than zero. If p is less than 0, the process ends. A value of zero means it is the beginning of the video; any value representing the beginning of the video can be used. The process iteratively adjusts the disparity values for locations in the backward direction to just meet the noticeable difference requirement for the disparity value change rate, until the disparity value for the location in the current frame is greater than the disparity value for the location in the previous frame. The adjustment process may stop when the disparity value for the location in the next frame is less than the current disparity value for the location in the current frame, which may indicate that the disparity value for the location in the next frame has already been adjusted or that the beginning of the video has been reached. This process is called for each location determined at 214 in FIG. 2.
[0039]
[0046] <Forward path parallax adjustment>
[0040]
[0047] FIG. 4 illustrates a simplified flowchart 400 for forward-pass disparity adjustment, according to some embodiments. The following process may be based on a setup where a more negative value of disparity for an object implies that the object is closer to the viewer. In another setup, the disparity value is always positive, with larger disparity values indicating closer to the viewer, and this setup may have a different flow. For example, the just-noticeable difference may be negative, and the comparison in 408 may be "greater than" instead of "less than." The following may be performed for each location determined in 214 of FIG. 2. Each location may be associated with a disparity value that may be adjusted. At 402, the rendering information generator 108 sets the current index to c. The index may be the frame selected based on its disparity value in FIG. 2. At 404, the rendering information generator 108 sets the next index to the value of n, where n=c+1. This may point to the next frame, if applicable. A loop is then performed until n is greater than or equal to a value that may be the last frame of the video.
[0041]
[0048] In that case, the following loop can be executed. At 406, the rendering information generator 108 determines whether the value of the next index n is greater than or equal to the maximum number of frames of the stereoscopic video. If so, it may have reached the end of the stereoscopic video and the process ends. If not, at 408, the rendering information generator 108 determines whether the disparity of the next index, disp[n]<disp[c], is less than the disparity of the current index, to determine whether disp[n] is closer to the viewer. This can determine whether the disparity value of the next index n (for example, the pixel in the next frame at the same position as the current pixel in the currently analyzed current frame) is less than the disparity value of the current index c (for example, the pixel in the current frame at the same position). If so, the process may end. One reason the process may end is that the disparity value of the next frame is smaller than the disparity value of the current frame and does not need to be adjusted (for example, since the process starts from the minimum disparity value, the disparity value may already have been adjusted in the previous iteration). If not, at 410, the rendering information generator 108 determines whether the disparity value of n is greater than the disparity value of the current index plus exactly the just noticeable difference value. If this is true, at 412, the rendering information generator 108 adjusts the disparity value at the next index n, such as setting the disparity value at the next index n to be equal to the disparity value of the current index plus exactly the just noticeable difference value. For example, the disparity value for the next frame (for example, -7800) may be greater than the disparity value for the current frame (for example, -8000) plus the minimum just noticeable distance value (+100) (for example, -7800 > -8000 + 100 = -7800 > -7900). Also, the disparity value for the next frame (for example, -8100) may be less than the disparity value for the current frame (for example, -8000) plus the minimum just noticeable distance value (+100) (for example, -8100 < -8000 + 100 = -8100 < -7800). Adjusting the disparity value can ensure that the change in the disparity value between the next frame and the current frame is no more than exactly the just noticeable difference value for this pixel.Alternatively, the adjusted disparity value may be set to the disparity value of the current index plus a value that is just less than or equal to the noticeable difference value, i.e., the viewer may not notice any change in the position of the caption. Otherwise, the process proceeds directly to step 414.
[0042]
[0049] At 414, the rendering information generator 108 increments the value of the current index c and the value of the next index n by 1. For example, if the frame numbers were c=1000 and n=1001, the new value of c is 1001 and the new value of p is 1002. The process then repeats to 406. The process may end when the value of the next index is greater than or equal to the end of the stereoscopic video. The process iteratively adjusts the disparity value in the forward direction to just meet the noticeable difference requirement for the disparity value change rate until the disparity value of the location in the current frame is greater than the disparity value of the location in the previous frame or the end of the video is reached. The adjustment process may stop when the disparity value for the next frame is less than the current disparity value, which may indicate that the disparity value for the next frame has already been adjusted. This process is called for each location determined at 214 in FIG. 2.
[0043]
[0050] <Example>
[0044]
[0051] 5 illustrates a graph 500 of raw disparity values for stereoscopic video according to some embodiments. The Y-axis indicates the minimum disparity value for a frame of the stereoscopic video, and the X-axis indicates the frame number in the stereoscopic video. At 502, an original smoothed raw disparity curve is shown, representing the minimum disparity value for a frame in the video. The disparity values may be unadjusted values.
[0045]
[0052] The following describes problems that can occur when these disparity values are used to display captions: At 504, a problem occurs in that the rate of Z change may be too fast: any captions displayed over this time may visibly move in the video and may not be temporally stable.
[0046]
[0053] At 506, a fixed Z approach for captions solves most of the occlusion problem and resolves temporal stability. Here, a fixed Z value may allow the caption to not be occluded by the object. However, at 508, a fixed Z approach may cause searching and fatigue when the difference between a small (negative) fixed Z value is far from the disparity value for the object in the video. Here, the object may be further away from the caption if the caption is displayed with a fixed Z value.
[0047]
[0054] At 510, similar to 504, but in the opposite direction, another problem occurs: the rate of Z change is too fast. Any captions displayed over this time may visibly move in the video and may not be temporally stable. At 512 and 514, another problem occurs: the rate of Z change may be too fast, and any captions displayed over this time may appear to flicker, even if the disparity values are not moving significantly.
[0048]
[0055] 6 shows a graph 600 with adjustments made using the rendering information generator 108, according to some embodiments. At 601, the processed output disparity curve is shown as a dotted line. At 602, the rendering information generator 108 adjusts the rate of change of the disparity values to just meet the noticeable difference range. This change avoids the problem that the rate of Z change may be so fast that captions displayed over this time may not visibly move in the video from the viewer's perspective. The rate of change at which captions are displayed across frames can help with the temporal stability of the captions by not moving by more than just a noticeable difference in successive frames.
[0049]
[0056] At 604, the fixed Z value prevents the caption from being occluded. The fixed Z value can be closer than any object in the video. At 606, the rendering information generator 108 adjusts the disparity value as quickly as possible to get as close as possible to the measured disparity for the video while adhering to just noticeable differences.
[0050]
[0057] At 608 and 610, the rendering information generator 108 prevents the caption from being occluded by adjusting the caption minimum value. At 612 and 614, the rendering information generator 108 adjusts the rate of change to just within the noticeable difference range to avoid flicker.
[0051]
[0058] 7 shows a graph 700 of adjusted disparity values, according to some embodiments. A portion of the disparity curve 702 has been adjusted as described in FIG. 6, and the portion of the curve that has not been adjusted is the original disparity values from FIG. 5.
[0052]
[0059] <Conclusion>
[0053]
[0060] Therefore, the rendering information generator 108 adjusts the disparity values to avoid problems that could have occurred when displaying captions, and the adjustment improves the display of captions in stereoscopic video when the adjusted disparity values are used.
[0054]
[0061] <System>
[0055]
[0062] FIG. 8 illustrates an example of a computing device, according to some embodiments. According to various embodiments, a system 800 suitable for implementing embodiments described herein includes a processor 801, a memory 803, a storage device 805, an interface 811, and a bus 815 (e.g., a PCI bus or other interconnect fabric). The system 800 may operate as any device or service described herein. While a specific configuration is described, various alternative configurations are possible. The processor 801 may perform operations such as those described herein. Instructions for performing such operations may be embodied in the memory 803, on one or more non-transitory computer-readable media, or in some other storage device. Various specially configured devices may also be used in place of or in addition to the processor 801. The memory 803 may be random access memory (RAM) or other dynamic storage device. The storage device 805 may include a non-transitory computer-readable storage medium that retains information, instructions, or any combination thereof, such as instructions that, when executed by the processor 801, cause the processor 801 to be configured or operable to perform one or more actions of the methods described herein. A bus 815 or other communication component may support communication of information within the system 800. An interface 811 may be connected to the bus 815 and configured to send and receive data packets over a network. Examples of supported interfaces include, but are not limited to, Ethernet, Fast Ethernet, Gigabit Ethernet, Frame Relay, Cable, Digital Subscriber Line (DSL), Token Ring, Asynchronous Transfer Mode (ATM), High Speed Serial Interface (HSSI), and Fiber Distributed Data Interface (FDDI). These interfaces may include ports suitable for communication with an appropriate medium. They may also include an independent processor and / or volatile RAM.The computer system or computing device may include or be in communication with a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.
[0056]
[0063] Any of the disclosed implementations may be embodied in various types of hardware, software, firmware, computer-readable media, and combinations thereof. For example, some techniques disclosed herein may be implemented, at least in part, by a non-transitory computer-readable medium containing program instructions, state information, etc., for configuring a computing system to perform various services and operations described herein. Examples of program instructions include both machine code, such as produced by a compiler, and higher-level code that may be executed via an interpreter. The instructions may be embodied in any suitable language, such as, for example, Java, Python, C++, C, HTML, any other markup language, JavaScript, ActiveX, VBScript, or Perl. Examples of non-transitory computer-readable media include, but are not limited to, magnetic media such as hard disks and magnetic tapes, optical media such as flash memory, compact discs (CDs) or digital versatile discs (DVDs), magneto-optical media, and other hardware devices such as read-only memory ("ROM") devices and random access memory ("RAM") devices. The non-transitory computer-readable medium may be any combination of such storage devices.
[0057]
[0064] In the foregoing specification, various techniques and mechanisms may be described in the singular for clarity. However, it should be noted that some embodiments include multiple iterations of a technique or multiple instantiations of a mechanism unless otherwise stated. For example, a system may use a processor in various contexts while using multiple processors while remaining within the scope of the present disclosure unless otherwise stated. Similarly, various techniques and mechanisms may be described as including a connection between two entities. However, a connection does not necessarily imply a direct, unobstructed connection, as various other entities (e.g., bridges, controllers, gateways, etc.) may exist between the two entities.
[0058]
[0065] Some embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform methods described by some embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured or operable to perform those described in some embodiments.
[0059]
[0066] As used in this description and throughout the claims that follow, the words "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in this description and throughout the claims that follow, the meaning of "in" includes "in" and "on," unless the context clearly dictates otherwise.
[0060]
[0067] The above description illustrates various embodiments, along with examples of how aspects of some embodiments may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of some embodiments as defined by the following claims. Based on the above disclosure and the following claims, other arrangements, embodiments, implementations, and equivalents may be used without departing from the scope of the invention as defined by the claims.
[0061]
[0068] Some embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform methods described by some embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured or operable to perform those described in some embodiments.
[0062]
[0069] As used in this description and throughout the claims that follow, the words "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in this description and throughout the claims that follow, the meaning of "in" includes "in" and "on," unless the context clearly dictates otherwise.
[0063]
[0070] The above description illustrates various embodiments, along with examples of how aspects of some embodiments may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of some embodiments as defined by the following claims. Based on the above disclosure and the following claims, other arrangements, embodiments, implementations, and equivalents may be used without departing from the scope of the invention as defined by the claims. The inventions described in the claims of the present application as originally filed are set forth below. [C1] 1. A method comprising: determining a disparity value from a plurality of disparity values in a current frame of stereoscopic video, wherein the disparity value is based on a difference in value for a pixel between a first video and a second video of the stereoscopic video; determining a location in the current frame that contains the disparity value; analyzing one or more first frames prior to the current frame to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values; analyzing one or more second frames after the current frame to adjust disparity values in the one or more second frames to generate one or more adjusted second disparity values; outputting the one or more adjusted first disparity values and the one or more adjusted second disparity values for use in displaying captions in the first video or the second video; A method for providing [C2] the one or more adjusted first parity values do not change by more than a difference value in successive first frames; the one or more adjusted second parity values do not change by more than the difference value in successive second frames; The method described in C1. [C3] Determining the disparity value comprises: determining a disparity value from the plurality of disparity values that is deemed closest to a viewer of the stereoscopic video; The method of C1, comprising: [C4] Determining the disparity value comprises: determining a disparity value for a location in a plurality of locations that is considered closest to a viewer of the stereoscopic video; determining the location from the plurality of locations; The method of C1, comprising: [C5] The method of C1, wherein the location comprises a cell defined as an area in the current frame, the area being used to display captions in the first video or the second video. [C6] Determining the location includes: determining one or more frames that include the disparity values; determining, for each of the one or more frames, one or more locations that contain the disparity value; The method of C1, comprising: [C7] For each of the one or more locations: analyzing one or more first frames prior to each frame for the location to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values; analyzing one or more second frames subsequent to the respective frame for the location to adjust the disparity values in the one or more second frames to generate one or more adjusted second disparity values; The method of C6, comprising: [C8] Analyzing one or more first frames prior to the current frame includes: determining a first frame; determining whether a disparity value for the first frame is smaller than a disparity value for the current frame, where smaller indicates closer to a viewer of the stereoscopic video; When the disparity value for the first frame is not less than the disparity value for the current frame, determining whether the disparity value for the first frame is greater than the disparity value for the current frame plus a difference value, where greater indicates greater distance from the viewer of the stereoscopic video; and The method of C1, comprising: [C9] Analyzing one or more first frames prior to the current frame includes: adjusting the disparity value for the first frame to an adjusted first disparity value that is less than or equal to the disparity value plus the difference value; The method of claim C8, comprising: [C10] Analyzing one or more first frames prior to the current frame includes: determining another first frame prior to the determined first frame; analyzing the disparity value of the another first frame and the adjusted first disparity value of the first frame to determine whether to adjust the disparity value of the another first frame to generate an adjusted first disparity value for the another first frame; The method of C9, comprising: [C11] Analyzing one or more first frames prior to the current frame includes: terminating the analysis of the one or more first frames prior to the current frame when the disparity value for the first frame is less than the disparity value for the current frame. The method of claim C8, comprising: [C12] Analyzing one or more second frames after the current frame includes: determining a second frame; determining whether a disparity value for the second frame is less than a disparity value for the current frame, where less indicates closer to a viewer of the stereoscopic video; and When the disparity value for the second frame is not less than the disparity value for the current frame, determining whether the disparity value for the second frame is greater than the disparity value for the current frame plus a difference value, where greater indicates greater distance from the viewer of the stereoscopic video; and The method of claim C1, comprising: [C13] Analyzing one or more second frames after the current frame includes: adjusting the disparity value for the second frame to an adjusted second disparity value that is less than or equal to the disparity value plus the difference value. The method of claim C8, comprising: [C14] Analyzing one or more second frames after the current frame includes: determining another second frame prior to the determined second frame; analyzing the disparity value of the another second frame and the adjusted second disparity value of the second frame to determine whether to adjust the disparity value of the another second frame to generate an adjusted second disparity value for the another second frame; The method of C9, comprising: [C15] Analyzing one or more second frames after the current frame includes: terminating the analysis of the one or more second frames prior to the current frame when the disparity value for the second frame is less than the disparity value for the current frame. The method of claim C8, comprising: [C16] the one or more adjusted first disparity values change at a rate that is less than or equal to a difference threshold; the one or more adjusted second disparity values change at the rate that is less than or equal to the difference threshold. The method described in C1. [C17] the disparity values for one or more first frames prior to the current frame are iteratively adjusted to generate the one or more adjusted first disparity values; the disparity values for one or more second frames after the current frame are iteratively adjusted to generate the one or more adjusted second disparity values. The method described in C1. [C18] Displaying captions in the first video or the second video using the one or more adjusted first disparity values and the one or more adjusted second disparity values. The method of C1, further comprising: [C19] 1. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a computing device, cause the computing device to: determining a disparity value from a plurality of disparity values in a current frame of stereoscopic video, wherein the disparity value is based on a difference in value for a pixel between a first video and a second video of the stereoscopic video; determining a location in the current frame that contains the disparity value; analyzing one or more first frames prior to the current frame to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values; analyzing one or more second frames after the current frame to adjust disparity values in the one or more second frames to generate one or more adjusted second disparity values; outputting the one or more adjusted first disparity values and the one or more adjusted second disparity values for use in displaying captions in the first video or the second video; 1. A non-transitory computer-readable storage medium operable to perform the steps of: [C20] 1. An apparatus comprising: one or more computer processors; determining a disparity value from a plurality of disparity values in a current frame of stereoscopic video, wherein the disparity value is based on a difference in value for a pixel between a first video and a second video of the stereoscopic video; determining a location in the current frame that contains the disparity value; analyzing one or more first frames prior to the current frame to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values; analyzing one or more second frames after the current frame to adjust disparity values in the one or more second frames to generate one or more adjusted second disparity values; outputting the one or more adjusted first disparity values and the one or more adjusted second disparity values for use in displaying captions in the first video or the second video; a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable to perform An apparatus comprising:
Claims
1. 1. A method comprising: determining a disparity value from a plurality of disparity values in a current frame of stereoscopic video, wherein the disparity value is based on a difference in value for a pixel between a first video and a second video of the stereoscopic video; determining a location in the current frame that includes the disparity value, the location comprising a cell defined as a region in the current frame; analyzing one or more first frames prior to the current frame to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values, wherein the one or more adjusted first disparity values do not change by more than just a noticeable difference between successive first frames; analyzing one or more second frames after the current frame to adjust disparity values in the one or more second frames to generate one or more adjusted second disparity values, wherein the one or more adjusted second disparity values do not change by more than the just noticeable difference value in successive second frames; outputting the one or more adjusted first disparity values and the one or more adjusted second disparity values for use in displaying captions in the first video or the second video; displaying captions in the first video or the second video using the one or more adjusted first disparity values and the one or more adjusted second disparity values; A method for providing
2. Determining the disparity value comprises: determining a disparity value from the plurality of disparity values that is deemed closest to a viewer of the stereoscopic video; The method of claim 1 , comprising:
3. Determining the disparity value comprises: determining a disparity value for a location in a plurality of locations that is considered closest to a viewer of the stereoscopic video; determining the location from the plurality of locations; The method of claim 1 , comprising:
4. The method of claim 1 , wherein the region is used to display captions in the first video or the second video.
5. Determining the location includes: determining one or more frames that include the disparity values; determining, for each of the one or more frames, one or more locations that contain the disparity value; The method of claim 1 , comprising:
6. For each of the one or more locations: analyzing one or more first frames prior to each frame for the location to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values; analyzing one or more second frames subsequent to the respective frame for the location to adjust the disparity values in the one or more second frames to generate one or more adjusted second disparity values; The method of claim 5 , comprising:
7. Analyzing one or more first frames prior to the current frame includes: determining a first frame; determining whether a disparity value for the first frame is less than a disparity value for the current frame, where less indicates closer to a viewer of the stereoscopic video; When the disparity value for the first frame is not less than the disparity value for the current frame, determining whether the disparity value for the first frame is greater than the disparity value for the current frame plus the just noticeable difference value, where greater indicates greater distance from the viewer of the stereoscopic video; and The method of claim 1 , comprising:
8. Analyzing one or more first frames prior to the current frame includes: adjusting the disparity value for the first frame to the disparity value plus an adjusted first disparity value that is less than or equal to the just noticeable difference value; The method of claim 7, comprising:
9. Analyzing one or more first frames prior to the current frame includes: determining another first frame prior to the determined first frame; analyzing the disparity value of the another first frame and the adjusted first disparity value of the first frame to determine whether to adjust the disparity value of the another first frame to generate an adjusted first disparity value for the another first frame; The method of claim 8 , comprising:
10. Analyzing one or more first frames prior to the current frame includes: terminating the analysis of the one or more first frames prior to the current frame when the disparity value for the first frame is less than the disparity value for the current frame. The method of claim 7, comprising:
11. Analyzing one or more second frames after the current frame includes: determining a second frame; determining whether a disparity value for the second frame is less than a disparity value for the current frame, where less indicates closer to a viewer of the stereoscopic video; and When the disparity value for the second frame is not less than the disparity value for the current frame, determining whether the disparity value for the second frame is greater than the disparity value for the current frame plus the just noticeable difference value, where greater indicates greater distance from the viewer of the stereoscopic video; and The method of claim 1 , comprising:
12. Analyzing one or more second frames after the current frame includes: adjusting the disparity value for the second frame to the disparity value plus an adjusted second disparity value that is less than or equal to the just noticeable difference value. The method of claim 7, comprising:
13. Analyzing one or more second frames after the current frame includes: determining another second frame prior to the determined second frame; analyzing the disparity value of the another second frame and the adjusted second disparity value of the second frame to determine whether to adjust the disparity value of the another second frame to generate an adjusted second disparity value for the another second frame; The method of claim 8 , comprising:
14. Analyzing one or more second frames after the current frame includes: terminating the analysis of the one or more second frames prior to the current frame when the disparity value for the second frame is less than the disparity value for the current frame. The method of claim 7, comprising:
15. the one or more adjusted first disparity values change at a rate that is less than or equal to the just-noticeable-difference threshold; the one or more adjusted second disparity values change at the rate that is less than or equal to the just-noticeable-difference threshold. The method of claim 1.
16. the disparity values for one or more first frames prior to the current frame are iteratively adjusted to generate the one or more adjusted first disparity values; the disparity values for one or more second frames after the current frame are iteratively adjusted to generate the one or more adjusted second disparity values. The method of claim 1.
17. 1. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a computing device, cause the computing device to: determining a disparity value from a plurality of disparity values in a current frame of stereoscopic video, wherein the disparity value is based on a difference in value for a pixel between a first video and a second video of the stereoscopic video; determining a location in the current frame that includes the disparity value, the location comprising a cell defined as a region in the current frame; analyzing one or more first frames prior to the current frame to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values, wherein the one or more adjusted first disparity values do not change by more than just a noticeable difference between successive first frames; analyzing one or more second frames after the current frame to adjust disparity values in the one or more second frames to generate one or more adjusted second disparity values, wherein the one or more adjusted second disparity values do not change by more than the just noticeable difference value in successive second frames; outputting the one or more adjusted first disparity values and the one or more adjusted second disparity values for use in displaying captions in the first video or the second video; displaying captions in the first video or the second video using the one or more adjusted first disparity values and the one or more adjusted second disparity values; 1. A non-transitory computer-readable storage medium operable to perform the steps of:
18. 1. An apparatus comprising: one or more computer processors; determining a disparity value from a plurality of disparity values in a current frame of stereoscopic video, wherein the disparity value is based on a difference in value for a pixel between a first video and a second video of the stereoscopic video; determining a location in the current frame that includes the disparity value, the location comprising a cell defined as a region in the current frame; analyzing one or more first frames prior to the current frame to adjust disparity values in the one or more first frames to generate one or more adjusted first disparity values, wherein the one or more adjusted first disparity values do not change by more than just a noticeable difference between successive first frames; analyzing one or more second frames after the current frame to adjust disparity values in the one or more second frames to generate one or more adjusted second disparity values, wherein the one or more adjusted second disparity values do not change by more than the just noticeable difference value in successive second frames; outputting the one or more adjusted first disparity values and the one or more adjusted second disparity values for use in displaying captions in the first video or the second video; displaying captions in the first video or the second video using the one or more adjusted first disparity values and the one or more adjusted second disparity values; a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable to perform An apparatus comprising:
Citation Information
Patent Citations
Stereo-image-data transmitting apparatus, stereo-image-data transmitting method, stereo-image-data receiving apparatus, and stereo-image-data receiving method
JP2011239169A
Stereoscopic image processing device and stereoscopic image imaging method
JP2012015774A
Stereoscopic image data transmission device, stereoscopic image data transmission method, stereoscopic image data reception device, and stereoscopic image data reception method
JP2012120143A
Combining 3D images and graphical data
JP2012518314A
Image processing apparatus, information processing system, image processing method, and program
JP2013239834A