Unshielded video overlay
Patent Information
- Application Number
- JP2022533180
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-07-29
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2040-07-29
AI Technical Summary
Existing video streaming technologies overlay content on top of the original video stream, often obscuring important content such as faces, text, or fast-moving objects, leading to inefficient use of computing resources and reduced viewer engagement.
A machine learning-based system identifies exclusion zones in video frames containing important content, aggregates these zones over time, and superimposes overlaid content within inclusion zones that avoid these areas, using optical character recognition and convolutional neural networks to detect and track text, human features, and objects of interest.
This approach ensures that important content is not obscured, optimizing bandwidth usage and enhancing viewer experience by providing additional content without interfering with the underlying video, thus improving screen area efficiency and viewer interaction.
Smart Images

Figure 00000018_0000 
Figure 00000019_0000 
Figure 00000020_0000
Abstract
Description
Background Art
[0001] Videos streamed to a user can include additional content overlaid on top of the original video stream. The overlaid content can be provided to the user within a rectangular area that overlays and blocks a portion of the original video screen. In some approaches, the rectangular area for providing the overlaid content is placed at the lower center of the video screen. If important content of the original video stream is placed at the lower center of the video screen, the important content may be blocked or hindered by the overlaid content.
Summary of the Invention
Problems to be Solved by the Invention
[0002] This specification relates to techniques for overlaying content on a video stream while avoiding areas of the video screen that feature useful content in the underlying video stream, such as areas in the original video stream that contain important objects like faces, text, or fast-moving objects.
Means for Solving the Problems
[0003] Generally, a first innovative aspect of the subject matter described herein can be embodied in a method that includes the steps of: identifying a corresponding exclusion zone for each video frame in a sequence of video frames, based on the detection of a specified object within the area of the video frame that is in the corresponding exclusion zone; aggregating the corresponding exclusion zones for video frames in a specified duration or number of frames; defining an inclusion zone in a specified duration or number of frames of the sequence of video frames in which the overlaid content is eligible for inclusion, wherein the inclusion zone is defined as an area of video frames in a specified duration or number outside the aggregated corresponding exclusion zones; and providing the overlaid content to include a specified duration or number of frames of the sequence of video in the inclusion zone while the video is being displayed on a client device. Other implementations of this aspect include corresponding devices, systems, and computer programs encoded on a computer storage device and configured to perform aspects of these methods.
[0004] In some embodiments, the identification of exclusion zones may include the step of identifying one or more areas in the video where text appears, for each video frame in a sequence of frames, and the method further includes the step of generating one or more bounding boxes that define one or more areas from the rest of the video frame. The identification of one or more areas where text appears may include the step of identifying one or more areas using an optical character recognition system.
[0005] In some embodiments, exclusion zone identification may include the step of identifying one or more regions in a video where human features appear, for each video frame in a sequence of frames, and the method further includes the step of generating one or more bounding boxes that define one or more regions from the rest of the video frame. The step of identifying one or more regions where human features appear may include the step of identifying one or more regions using a computer vision system trained to identify human features. The computer vision system may be a convolutional neural network system.
[0006] In some embodiments, exclusion zone identification may include the step of identifying one or more regions in the video in which significant objects appear, for each video frame in a sequence of frames, wherein the identification of regions in which significant objects appear is identification using a computer vision system configured to recognize objects from a selected set of object categories that do not contain text or human features. Exclusion zone identification may also include the step of identifying one or more regions in the video in which significant objects appear, based on the detection of objects moving beyond a selected distance between consecutive frames, or the detection of objects moving between a specified number of consecutive frames.
[0007] In some embodiments, the aggregation of corresponding exclusion zones may include the step of generating a joining of bounding boxes that define the corresponding exclusion zones from other parts of the video. The demarcation of an inclusion zone may include the step of identifying a set of rectangles within a sequence of video frames that do not overlap with the aggregated corresponding exclusion zones over a specified duration or number of frames, and the step of providing overlaid content to be contained within an inclusion zone may include the step of identifying an overlay having dimensions that fit within one or more rectangles in the set of rectangles, and the step of providing the overlay within one or more rectangles over a specified duration or number of frames.
[0008] The subject matter described herein may be implemented in specific embodiments to achieve one or more of the following advantages:
[0009] While a user is watching a video stream that fills a video screen, content of value to the user within that video screen area may not fill the entire area of the video screen. For example, useful content, such as faces, text, or important objects like fast-moving objects, may occupy only a portion of the video screen area. Therefore, there is an opportunity to present additional useful content to the user in the form of overlaid content that does not obstruct the portion of the video screen area containing the useful content below. Aspects of this disclosure provide the advantage of identifying exclusion zones that exclude overlaid content, because overlaying content on these exclusion zones would obstruct or obscure useful content contained in the underlying video stream, resulting in wasted computing resources by delivering video to the user when the user cannot perceive the useful content. In some situations, a machine learning engine (such as a Bayesian classifier, optical character recognition system, or neural network) can identify interesting features in a video stream, such as faces, text, or other important objects, and can identify exclusion zones that encompass these interesting features, and then present the overlaid content outside of these exclusion zones. As a result, users can receive overlaid content without interference to the valuable content of the underlying video stream, thus preventing wasted computing resources required to deliver the video. This results in a more efficient video delivery system that prevents wasted computing system resources (e.g., network bandwidth, memory, processor cycles, and limited client device display space) through video delivery that obscures valuable content or is otherwise imperceptible to the user.
[0010] This has further advantages in improving screen area efficiency with respect to the bandwidth of useful content delivered to the viewer. When a user is watching a video in which the useful content of the video occupies only a small portion of the viewing area, as is typical, the bandwidth available for delivering the useful content to the viewer is not fully utilized. By using a machine learning system to identify that small portion of the viewing area that contains the useful content of the video stream below, aspects of this disclosure provide overlaying additional content outside that small portion of the viewing area, resulting in more efficient use of the screen area for delivering the useful content to the viewer.
[0011] In some methods, the overlaid content includes, for example, a box or other icon that the viewer can click to remove the overlaid content if the overlaid content interferes with the informative content in the video below. A further advantage of this disclosure is that the overlaid content is less likely to interfere with the informative content in the video below, resulting in less disruption to the viewing experience and a higher likelihood that the viewer will not "click away" the presented overlaid content.
[0012] The various features and advantages of the subject matter described above will be explained below with reference to the drawings. Additional features and advantages will be evident from the subject matter and claims described herein. [Brief explanation of the drawing]
[0013] [Figure 1] This diagram shows an overview of how exclusion zones are aggregated and how inclusion zones are defined for videos containing frames. [Figure 2] This figure shows an example of frame-by-frame aggregation of excluded zones for the example in Figure 1. [Figure 3] This figure shows an example of a machine learning system for identifying and aggregating exclusion zones and selecting overlaid content. [Figure 4] This diagram shows a flowchart of the process, including aggregating excluded zones and selecting overlaid content. [Figure 5] This is a block diagram of an exemplary computer system. [Modes for carrying out the invention]
[0014] The same reference number and name in various drawings refer to the same element.
[0015] It is generally desirable to provide overlaid content on an underlying video stream, to provide additional content to viewers of the video stream, and to improve the amount of content delivered within the viewing area with a given video streaming bandwidth. However, there is a technical problem in determining how to position the overlaid content so as not to obstruct the valuable content in the underlying video. This is a particularly difficult problem in situations where content is overlaid on video, as the location of important content in the video can change rapidly over time. Therefore, a particular location in the video that is a good candidate for overlaid content at one point in time may become a bad candidate for overlaid content at a later point in time (for example, due to the movement of characters in the video).
[0016] This specification presents a solution to this technical problem by describing a machine learning method and system that can identify exclusion zones corresponding to areas of a video stream that are more likely to contain valuable content, aggregate these exclusion zones over time, and then position the overlaid content within an inclusion zone outside the aggregated exclusion zones so that the overlaid content is less likely to interfere with the valuable content in the underlying video.
[0017] Figure 1 shows an illustrative example of exclusion zones, inclusion zones, and overlaid content in a video stream. In this example, the original, underlying video stream 100 is shown on the left side of the figure as frames 101, 102, and 103. Each frame may contain areas that could contain valuable content that should not be obscured by overlaid content features. For example, a frame of video may contain one or more areas of text 111, such as closed caption text or text displayed on features within the video, such as product labels, road signs, or text displayed on a whiteboard on a screen in a video of a school lecture. Note that features within the video are part of the video stream itself, and any overlaid content is separate from the video stream itself. As will be further discussed below, machine learning systems such as optical character recognition (OCR) systems can be used to identify areas having frames containing text, and this identification may include identifying a bounding box surrounding the identified text, as shown. As shown in the example in Figure 1, text 111 may be located in different positions within different frames, and therefore, overlaid content that persists over a duration including multiple frames 101, 102, and 103 will not be located in any particular location where text 111 is located across multiple frames.
[0018] Other areas that should not be obscured by overlaid content features (for example, to ensure that important content of the video below is not obscured) are areas 112 that contain people or human features. For example, a frame may contain one or more people, or parts thereof, such as a human face, torso, limbs, or hands. As will be further discussed below, machine learning systems such as convolutional neural network (CNN) systems can be used to identify areas within a frame that contain people (or parts thereof, such as a face, torso, limbs, or hands), and such identification may include identifying bounding boxes surrounding the identified human features, as shown. As shown in the example in Figure 1, human features 112 may be located in different positions within different frames, and therefore, overlaid content that persists over a duration encompassing multiple frames 101, 102, 103 will not be located anywhere across multiple frames where human features 112 are located. In some techniques, human features 112 may be distinguished by limiting human feature detection to larger human features, i.e., features that are in the foreground and closer to the video's viewpoint, as opposed to background human features. For example, larger human faces corresponding to people in the foreground of a video may be included in the detection scheme, while smaller human faces corresponding to people in the background of a video, such as faces in a crowd, may be excluded from the detection scheme.
[0019] Other areas that should not be obscured by overlaid content features are areas 113 containing other potential objects of interest, such as animals, plants, streetlights or other road features, bottles or other containers, furniture, etc. As will be further discussed below, machine learning systems, such as convolutional neural network (CNN) systems, can be used to identify areas within a frame that contain potential objects of interest. For example, a machine learning system can be trained to classify various objects selected from a list of object categories, such as identifying dogs, cats, vases, flowers, or any other category of objects that a viewer might potentially be interested in. Identification can include identifying a bounding box surrounding the detected object, as shown. As shown in the example in Figure 1, the detected object 113 (a cat in this example) may be located in different positions within different frames, and therefore, overlaid content that persists over a duration encompassing multiple frames 101, 102, 103 will not be located in any particular location where the detected object 113 is located across multiple frames.
[0020] In some methods, detected objects 113 can be distinguished by limiting object detection to moving objects. Moving objects are generally more likely to convey content important to the viewer and are therefore potentially less desirable to be obscured by overlaid content. For example, detected objects 113 can be limited to objects that move a minimum distance within a selected time interval (or selected frame interval), or objects that move within a specified number of consecutive frames.
[0021] As shown by the video screen 120 in Figure 1, exclusion zones for overlaid content can be identified based on detected features in the video that are more likely to attract the viewer's attention. For example, an exclusion zone could include an area 121 identified as where text appears within the video frame (to be compared with identified text 111 in the original video frames 101, 102, and 103); an exclusion zone could also include an area 122 identified as where human features appear within the video frame (to be compared with identified human features 112 in the original video frames 101, 102, and 103); and an exclusion zone could also include an area 123 identified as where other objects of interest (such as objects moving faster, or objects identified from a selected list of object categories) appear within the video frame (to be compared with identified objects 113 in the original video frames 101, 102, and 103).
[0022] Since overlaid content may be contained over a selected duration (or a selected number of frames of the underlying video), exclusion zones can be aggregated over a selected duration (or a selected span of frames) to define an aggregated exclusion zone. For example, an aggregated exclusion zone may be the sum of all exclusion zones corresponding to a feature of interest detected in each video frame within a sequence of video frames. The selected duration can be 1 second, 5 seconds, 10 seconds, 1 minute, or any other duration suitable for displaying the overlaid content. Alternatively, the selected span of frames can be 24 frames, 60 frames, 240 frames, or any other span of frames suitable for displaying the overlaid content. The selected duration can correspond to a selected span of frames, as determined by the frame rate of the underlying video, and vice versa. The example in Figure 1 shows aggregation over only three frames 101, 102, and 103, but this is for illustrative purposes only and is not intended to limit the possibilities.
[0023] Figure 2 shows an example of how the aggregation of exclusion zones can proceed for each frame of the example of FIG. 1. First, the minimum number of consecutive frames in which the exclusion zones can be aggregated can be selected (or, depending on the situation, the minimum time interval can be selected and converted into several consecutive frames based on the video frame rate). In this example, for illustrative purposes only, the minimum number of consecutive frames is three frames corresponding to frames 101, 102, and 103 as shown.
[0024] As shown in column 200 of Figure 2, each frame 101, 102, and 103 in the video frames below may contain a feature of interest, such as text 111, human features 112, or other potential objects of interest 113. For each frame, a machine learning system such as an optical character recognition system, a Bayesian classifier, or a convolutional neural network classifier can be used to detect the feature of interest within the frame. As shown in column 210 of Figure 2, detecting a feature within a frame may include determining bounding boxes to enclose the features. Thus, the machine learning system may output a bounding box 211 enclosing the detected text in each frame, a bounding box 212 enclosing the human features 112 in each frame, and / or a bounding box 213 enclosing the other potential objects of interest 113 in each frame. The bounding box 212 enclosing the human features 112 can be selected to enclose the entirety of any human features detected in the frame, or it can be selected to enclose only a portion of any human features detected in the frame (for example, only the face, head and shoulders, torso, hands, etc.). As shown in column 220 of Figure 2, the bounding boxes 211, 212, and 213 of the consecutive frames can correspond to exclusion zones 221, 222, and 223, respectively, which can be accumulated or aggregated per frame. The newly added exclusion zone 230 is accumulated for the exclusion zone of frame 102, and the newly added exclusion zone 240 is accumulated for the exclusion zone of frame 103. As a result, as seen in the bottom right frame of Figure 2, the aggregation includes all bounding boxes for all features detected within all frames in the selected interval of the consecutive frames. Note that in this example, the aggregated exclusion zone seen in the bottom right of Figure 2 is the aggregated exclusion zone of frame 101 for the selected interval of the consecutive frames (three in this example).Thus, for example, in the case of frame 102, the aggregated exclusion zone includes the single-frame exclusion zones of frame 102, frame 103, and a fourth frame not shown, and similarly, the aggregated exclusion zone of frame 102 includes the single-frame exclusion zones of frame 103, a fourth frame not shown, and a fifth frame not shown, and so on for the following.
[0025] By having an aggregated exclusion zone over a selected duration or selected span of frames, an inclusion zone can be defined where the overlaid content is eligible for display. For example, inclusion zone 125 in FIG. 1 corresponds to an area of viewing area 120 where none of text 111, human feature 112, or other notable object 113 is displayed over the span of frames 101, 102, 103.
[0026] In some methods, the containment zone can be defined as a set of containment zone rectangles whose entirety is defined by its combination. The set of containment zone rectangles can be calculated by iterating over all the bounding boxes accumulated to define the aggregate exclusion zone (for example, rectangles 211, 212, and 213 in Figure 2). For a given bounding box within the accumulated bounding boxes, the upper right corner is selected as the starting point (x,y), and then expanded up, left, and right to find the largest box that does not overlap with any other bounding boxes (or edges of the viewing area), and that largest box is added to the list of containment zone rectangles. Next, the lower right corner is selected as the starting point (x,y), and then expanded up, down, and right to find the largest box that does not overlap with any other bounding boxes (or edges of the viewing area), and that largest box is added to the list of containment zone rectangles. Next, the upper left corner is selected as the starting point (x,y), and then expanded upwards, left, and right to find the largest box that does not overlap with any of the other bounding boxes (or edges of the viewing area), and that largest box is added to the list of containment zone rectangles. Next, the lower left corner is selected as the starting point (x,y), and then expanded downwards, left, and right to find the largest box that does not overlap with any of the other bounding boxes (or edges of the viewing area), and that largest box is added to the list of containment zone rectangles. These steps are then repeated for the next bounding box in the accumulation of bounding boxes. Note that these steps can be completed in any order. The containment zone thus defined is the containment zone of frame 101, as it defines the area in frame 101, 102, and 103, i.e., in any frame within the selected interval of consecutive frames, where the overlaid content can be placed in frame 101 and subsequent frames during the selected interval of consecutive frames (three in this example) without occluding any detected features.The inclusion zone of frame 102 can be similarly defined, but includes the supplementary single-frame exclusion zones of frame 102, frame 103, and a fourth frame not shown; similarly, the inclusion zone of frame 103 includes the supplementary single-frame exclusion zones of frame 103, a fourth frame not shown, and a fifth frame not shown, and so on.
[0027] Once the containment zone is defined, appropriate overlaid content can be selected to be displayed within the containment zone. For example, a set of overlaid content candidates may be available, and each item in the set of overlaid content candidates may have specifications that include, for example, the width and height of each item, the minimum duration each item is provided to the user, etc. One or more items from the set of overlaid content candidates may be selected to fit within the defined containment zone. For example, as shown in the viewing area of Figure 1, two items of overlaid content 126 may be selected to fit within the containment zone 105. In various methods, the number of overlaid content features may be limited to one, two, or more features. In some methods, the first overlaid content feature may be provided during a first time span (or frame span), and the second overlaid content feature may be provided during a second time span (or frame span), and so on, and the first, second, etc. time spans (or frame spans) may completely overlap, partially overlap, or not overlap at all.
[0028] After selecting the overlaid content 126 to be provided to the viewer along with the underlying video stream, Figure 1 shows an example of a video stream 130 containing both the underlying video and the overlaid content. As shown in this example, the overlaid content 126 does not obscure or interfere with the features of interest that were detected in the underlying video 100 and used to define the exclusion zone for the overlaid content.
[0029] Referring next to Figure 3, an exemplary example is shown as a block diagram of a system that selects an inclusion zone and provides overlaid content on a video stream. The system may operate as a video pipeline, receiving the original video on which the content to be overlaid as input, and providing the video with the overlaid content as output. System 300 may include a video preprocessor unit 301 that can be used to provide downstream uniformity of video specifications such as frame rate (which can be adjusted by a resampler), video size / quality / resolution (which can be adjusted by a rescaler), and video format (which can be adjusted by a format converter). The output of the video preprocessor is a video stream 302 in a standard format for further processing by the system's downstream components.
[0030] System 300 includes a text detector unit 311 that receives a video stream 302 as input and provides as output a set of regions in the video stream 302 where text appears. The text detector unit may be a machine learning unit, such as an optical character recognition (OCR) module. For efficiency, the OCR module only needs to find regions in the video where text appears without actually recognizing the text present in those regions. Once regions are identified, the text detector unit 311 can generate (or specify) bounding boxes that define (or otherwise define) the regions in each frame determined to contain text, which can be used to identify exclusion zones for overlaid content. The text detector unit 311 may output the detected text bounding boxes as an array (indexed by frame number), for example, where each element of the array is a list of rectangles defining the text bounding boxes detected in that frame. In some methods, the detected bounding boxes may be added to the video stream as metadata information for each frame of the video.
[0031] System 300 also includes a person or human feature detector unit 312 that receives a video stream 302 as input and provides a set of regions of the video containing a person (or a part thereof, such as a face, torso, limbs, or hands) as output. The person detector unit may be a computer vision system such as a machine learning system, for example, a Bayesian image classifier or a convolutional neural network (CNN) image classifier. The person or human feature detector unit 312 may be trained on labeled training samples, for example, labeled with human features illustrated by the training samples. Once trained, the person or human feature detector unit 312 may output labels that identify one or more human features detected in each frame of the video, and / or confidence values that indicate the level of confidence that one or more human features are located within each frame. The person or human feature detector unit 312 may also generate bounding boxes that define the areas where one or more human features are detected, which can be used to identify exclusion zones for overlaid content. For efficiency, the human feature detector unit only needs to find areas in the video where human features appear without actually recognizing the identification information of the people present in those areas (for example, without recognizing the faces of specific people present in those areas). The human feature detector unit 312 can output the detected human feature bounding boxes as an array (indexed by frame number), for example, where each element of the array is a list of rectangles defining the human feature bounding boxes detected in that frame. In some methods, the detected bounding boxes can be added to the video stream as metadata information for each frame of the video.
[0032] System 300 also includes an object detector unit 313 that receives a video stream 302 as input and provides as output a set of video regions containing potential objects of interest. Potential objects of interest can be objects classified as belonging to an object category in a selected list of object categories (e.g., animals, plants, road or terrain features, containers, furniture, etc.). Potential objects of interest can also be limited to identified moving objects, for example, objects that move a minimum distance within a selected time interval (or selected frame interval), or objects that move within a specified number of consecutive frames in the video stream 302. The object detector unit may be a computer vision system such as a machine learning system, for example, a Bayesian image classifier, or a convolutional neural network image classifier. The object detector unit 313 may be trained on labeled training samples labeled with objects classified as belonging to an object category in a selected list of object categories. For example, an object detector can be trained to recognize animals such as cats or dogs, or furniture such as tables and chairs, or terrain or road features such as trees or road signs, or any combination of such selected object categories. The object detector unit 313 can also generate bounding boxes that define (or otherwise specify) the area of the video frame in which the identified object is identified. The object detector unit 313 can output the detected object bounding boxes as an array (indexed by frame number) in which each element of the array is a list of rectangles defining an object bounding box detected within that frame. In some methods, the detected bounding boxes can be added to the video stream as metadata information for each frame of the video. In other exemplary examples, the system 300 may include at least one of a text detector 311, a person detector 312, or an object detector 313.
[0033] System 300 also includes an inclusion zone calculator unit or module 320 that receives input from one or more of the following: a text detector unit 311 (having information about areas where text appears in the video stream 302), a person detector unit 312 (having information about areas where a person or part thereof appears in the video stream 302), and an object detector unit 313 (having information about areas where various potential objects of interest appear in the video stream 302). Each of these areas can define an exclusion zone, and the inclusion zone calculator unit can aggregate these exclusion zones, and then the inclusion zone calculator unit can define an inclusion zone in which the overlaid content is eligible for inclusion.
[0034] The aggregated exclusion zone can be defined as a combination of lists of rectangles, each containing potentially noteworthy features such as text, people, or other objects of interest. This can be represented as an accumulation of bounding boxes generated by detector units 311, 312, and 313 over a selected number of consecutive frames. Firstly, bounding boxes can be aggregated frame by frame. For example, if text detector unit 311 outputs a first array (indexed by frame number) of text bounding boxes within each frame, human feature detector 312 outputs a second array (indexed by frame number) of human feature bounding boxes within each frame, and object detector unit 313 outputs a third array (indexed by frame number) of bounding boxes for objects detected within each frame, these first, second, and third arrays can be merged to define a single array (also indexed by frame number), each element being a single list merging all bounding boxes for all features (text, people, or other objects) detected within that frame. Next, bounding boxes can be aggregated over a selected interval of consecutive frames. For example, a new array (again indexed by frame number) may be defined, where each element is a single list merging all bounding boxes for all features found within frames i, i+1, i+2, ..., i+(N-1), where N is the number of frames within the selected interval of consecutive frames. In some methods, the aggregated exclusion zone data can be added to the video stream as metadata information for each frame of the video.
[0035] The containment zone calculated by the containment zone calculator unit 320 can then be defined as a supplement to the accumulation of bounding boxes generated by the detector units 311, 312, and 313 over a selected number of consecutive frames. The containment zone can be specified, for example, as another list of rectangles whose combination forms the containment zone, or as a polygon having horizontal and vertical sides, which can be described, for example, by a list of polygonal vertices, or as a list of such polygons, if the containment zone includes a cut-off area of the viewing screen. If the accumulated bounding boxes are represented by an array (indexed by frame number) as described above, and each element is a list merging all bounding boxes for all features detected in that frame and in the following N-1 consecutive frames, then the containment zone calculator unit 320 can store the containment zone information as a new array (also indexed by frame number), where each element is a list of containment rectangles for that frame, taking into account all bounding boxes accumulated in that frame and in the following N-1 consecutive frames, by iterating across each accumulated bounding box and across the four corners of each bounding box, as described above in the context of Figure 2. Note that these containment zone rectangles can be overlapping rectangles that collectively define the containment zone. In some methods, this containment zone data can be added to the video stream as metadata information for each frame of the video.
[0036] System 300 also includes an overlaid content matcher unit or module 330 that receives input from an inclusion zone calculator unit or module 320, for example, in the form of inclusion zone specifications. The overlaid content matcher unit can select content suitable for overlaying on video within an inclusion zone. For example, the overlaid content matcher can access a catalog of overlaid content candidates, each item in the catalog of overlaid content candidates having specifications that may include, for example, the width and height of each item, and the minimum duration for which each item should be provided to the user. The overlaid content matcher unit can select one or more items from the set of overlay content candidates to fit within an inclusion zone provided by the inclusion zone calculator unit 320. For example, if containment zone information is stored in an array (indexed by frame number), and each element in the array is a list of containment zone rectangles for that frame (taking into account all bounding boxes accumulated over that frame and the subsequent N-1 consecutive frames), then for each item in the catalog of candidate overlay content, the overlay content matcher can identify containment zone rectangles in the array that are large enough to fit the selected item, and these can be ranked in order of size and / or in order of persistence (for example, if the same rectangle appears in multiple consecutive elements of the array, indicating that the containment zone is available for more than the minimum number of consecutive frames N), and then select a containment zone rectangle from that ranked list for the containment of the selected overlay content.In some techniques, the overlaid content may be scalable, for example, across a range of possible x or y dimensions, or across a range of possible aspect ratios, and in these techniques, the containment zone rectangle matching the overlaid content item may be selected, for example, the largest area containment zone rectangle that can accommodate the scalable overlaid content, or a containment zone rectangle that is large enough to persist for the longest duration of a series of frames.
[0037] System 300 also includes an overlay unit 340 that receives both the underlying video stream 302 and the selected overlay content 332 (and its position) from the overlay content matcher 330 as input. The overlay unit 340 can then provide the video stream 342 containing both the underlying video content 302 and the selected overlay content 332. On the user device, a video visualizer 350 (for example, a video player embedded in a web browser or a video app on a mobile device) displays the video stream with the overlay content to the user. In some methods, the overlay unit 340 may reside on the user device and / or be embedded within the video visualizer 350, in other words, both the underlying video stream 302 and the selected overlay content 332 may be delivered to the user (for example, over the internet) and combined on the user device to display the overlaid video to the user.
[0038] Referring now to Figure 4, an exemplary example is shown as a process flow diagram for a method of providing overlaid content on a video stream. This process involves identifying, in 410, a corresponding exclusion zone from which the overlaid content is excluded for each video frame in a sequence of video frames. For example, the inclusion zones may correspond to areas of the viewing area that contain video features that are more likely to attract the viewer's attention, such as areas containing text (e.g., area 111 in Figure 1), areas containing people or human features (e.g., area 112 in Figure 1), and areas containing specific objects of interest (e.g., area 113 in Figure 1). These areas can be detected using machine learning systems, such as an OCR detector for text (e.g., text detector 311 in Figure 3), a computer vision system for people or human features (e.g., person detector 312 in Figure 3), and a computer vision system for other objects of interest (e.g., object detector 313 in Figure 3).
[0039] The process also includes aggregating corresponding exclusion zones for video frames within a specified duration or number of frames in a sequence of frames. For example, rectangles 121, 122, and 123, which are bounding boxes of potential features of interest, as shown in Figure 1, can be aggregated across a sequence of frames to define an aggregated exclusion zone, which is a combination of exclusion zones for that sequence of frames. This combination of exclusion zone rectangles can be calculated, for example, by the inclusion zone calculator unit 320 in Figure 3.
[0040] The process, in 430, is to define containment zones within a specified duration or number of video frames where overlaid content is eligible for containment, further comprising defining the containment zones as areas of video frames within a specified duration or number outside of the aggregated corresponding exclusion zones. For example, the containment zone 125 in Figure 1 may be defined as a supplement to the aggregated exclusion zone, and the containment zone may be described as a combination of rectangles that collectively fill the containment zones. The containment zones may be calculated, for example, by the containment zone calculator unit 320 in Figure 3.
[0041] The process further includes providing overlaid content to include a specified duration or number of video frames in a containment zone while the video is being displayed on a client device. For example, the overlaid content may be selected from a catalog of overlaid content candidates, for example, based on the dimensions of the items in the catalog of overlaid content candidates. In the example of Figure 1, two overlaid content features 126 are selected to be contained within the containment zone 125. The overlaid content (and its position within the viewing area) may also be selected by the overlaid content matcher 330 in Figure 3, and the overlay unit 340 can overlay the overlaid content onto the underlying video stream 302 to define a video stream 342 having the overlaid content that is provided to the user for viewing on a client device.
[0042] In some methods, the overlaid content selected from the catalog of overlaid content may be selected in the following way: For frame i, the containment zone may be defined as a combination of a list of containment area rectangles, where the containment area rectangle is a rectangle that does not intersect with any of the exclusion zones (bounding boxes of detected objects) in frames i, i+1, ..., i+(N-1), where N is the minimum number of consecutive frames selected. Then, for frame i and a given candidate item from the catalog of overlaid content, the containment area rectangle is selected from a list of containment area rectangles that can fit the candidate item. These are containment area rectangles that can fit the candidate item in the minimum number of consecutive frames N selected. The same process can be performed for frame i+1, and then by taking the intersection of the results for frame i and frame i+1, a list of containment area rectangles that can fit the candidate item for N+1 consecutive frames can be obtained. Again, by taking the intersection with the result for frame i+2, a set of containment areas that can fit the candidate item for N+2 consecutive frames can be obtained. This process can be iterated over any selected span of frames (including the total duration of the video) to obtain a rectangle suitable for containing candidate items with frame durations N, N+1, ..., N+(k-1), where N+k is the longest possible duration. Thus, for example, the position of overlaid content can be selected from a list of containment area rectangles that can persist for the longest duration, i.e., for N+k frames, without occluding the detected features.
[0043] In some methods, two or more content features may be included simultaneously. For example, a first item of overlaid content may be selected, and then a second item of overlaid content may be selected by defining an additional exclusion zone surrounding the first item of overlaid content. In other words, the second item of overlaid content may be positioned by treating the video superimposed with the first item of overlaid content as a new underlying video suitable for overlaying additional content. The exclusion zone for the first item of overlaid content may be made significantly larger than the overlaid content itself to increase the spatial separation between different items of overlaid content within the viewing area.
[0044] In some methods, the selection of overlaid content may include selection that allows a specified level of intrusion on exclusion zones. For example, some area-based intrusion may be tolerated by weighting the containment zone rectangles by the extent to which the overlaid content spatially extends outside each containment zone rectangle. Alternatively or additionally, some time-based intrusion may be tolerated by ignoring temporary exclusion zones that exist for relatively short periods. For example, if an exclusion zone is defined for only a single frame out of 60 frames, it can be given a lower weight than an area where the exclusion zone exists for the entire 60 frames, and is therefore more likely to be occluded. Alternatively or additionally, some content-based intrusion may be tolerated by ranking the relative importance of different types of exclusion zones corresponding to different types of detected features. For example, detected text features may be ranked more important than detected non-text features, and / or detected human features may be ranked more important than detected non-human features, and / or faster-moving features may be ranked more important than slower-moving features.
[0045] Figure 5 is a block diagram of an exemplary computer system 500 that may be used to perform the operations described above. The system 500 includes a processor 510, memory 520, storage device 530, and input / output device 540. Each of the components 510, 520, 530, and 540 may be interconnected, for example, using a system bus 550. The processor 510 is capable of processing instructions to be executed within the system 500. In some implementations, the processor 510 is a single-threaded processor. In other implementations, the processor 510 is a multi-threaded processor. The processor 510 is capable of processing instructions stored in memory 520 or on storage device 530.
[0046] Memory 520 stores information within the system 500. In one implementation, memory 520 is a computer-readable medium. In some implementations, memory 520 is a volatile memory unit. In another implementation, memory 520 is a non-volatile memory unit.
[0047] The storage device 530 may provide a mass storage device to the system 500. In some implementations, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 may include, for example, a hard disk device, an optical disk device, a storage device shared over a network by multiple computing devices (e.g., a cloud storage device), or some other mass storage device.
[0048] The input / output device 540 provides input / output operations for the system 500. In some implementations, the input / output device 540 may include one or more of the following: a network interface device, e.g., an Ethernet card; a serial communication device, e.g., an RS-232 port; and / or a wireless interface device, e.g., an 802.11 card. In another implementation, the input / output device may include a driver device configured to receive input data and send output data to an external device 460, e.g., a keyboard, printer, and display device. However, other implementations may be used, such as a mobile computing device, a mobile communication device, or a set-top box television client device.
[0049] An exemplary processing system is illustrated in Figure 5, but the subject matter and functional operations described herein may be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware, or a combination of one or more of these, including the structures disclosed herein and their structural equivalents.
[0050] The subject matter and embodiments of operation described herein may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, or in one or more combinations thereof, including the structures disclosed herein and their structural equivalents. Embodiments of the subject matter described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on one or more computer storage media for execution by a data processing device or for controlling the operation of a data processing device. The computer storage media may be temporary or non-temporary. Alternatively or additionally, the program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by a data processing device. The computer storage media may be or be included in computer-readable storage devices, computer-readable storage boards, random or serial access memory arrays or devices, or one or more combinations thereof. Furthermore, the computer storage media may be the source or destination of computer program instructions encoded on artificially generated propagating signals, rather than the propagating signals themselves. Computer storage media may also be, or comprise, one or more distinct physical components or media (for example, multiple CDs, disks, or other storage devices).
[0051] The operations described herein may be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.
[0052] The term "data processing device" encompasses all types of devices, machines, and equipment for processing data, including, for example, programmable processors, computers, systems on a chip, or a combination of the above. A device may include dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, a device may also include code that creates an execution environment for the computer program in question, such as processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or code that constitutes one or more of these. Devices and execution environments can realize a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0053] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed as standalone programs or in any form, including modules, components, subroutines, objects, or other units suitable for use in a computing environment. Computer programs may, but do not, correspond to files in a file system. A program may be stored in a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to that program, or in multiple collaborative files (e.g., files that store one or more modules, subprograms, or parts of code). Computer programs may be deployed to run on one computer or on multiple computers located in one site or distributed across multiple sites and interconnected by a communication network.
[0054] The processes and logic flows described herein may be executed by one or more programmable processors that run one or more computer programs to perform actions by operating on input data and producing outputs. The processes and logic flows may also be implemented by dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the devices may also be implemented as such.
[0055] Processors suitable for executing computer programs include, for example, both general-purpose and dedicated microprocessors. Generally, a processor will receive instructions and data from read-only memory, random-access memory, or both. Essential elements of a computer are a processor for performing actions according to instructions, and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more high-capacity devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or will be operablely coupled to them to receive data from them, transfer data to them, or both. However, a computer does not have to have such devices. Furthermore, a computer may be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a Universal Serial Bus (USB) flash drive), to name just a few examples. Devices suitable for storing computer program instructions and data include, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks, encompassing all forms of non-volatile memory, media, and memory devices. Processors and memory may be complemented by or incorporated into dedicated logic circuits.
[0056] To provide user interaction, embodiments of the subject matter described herein may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide user interaction; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic input, voice input, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from that web browser.
[0057] Embodiments of the subject matter described herein may be implemented in a computing system that includes, for example, a data server, a backend component, or a middleware component, such as an application server, or a frontend component, such as a client computer having a graphical user interface or web browser through which a user can interact with the implementation of the subject matter described herein, or any combination of one or more such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks).
[0058] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is established by computer programs running on each computer and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to the client device (for example, to display data to a user interacting with the client device and to receive user input from that user). Data generated on the client device (e.g., the results of user interaction) may be received by the server from the client device.
[0059] This specification includes many specific implementation details, which should not be construed as limitations on the scope of any invention or claimable invention, but rather as descriptions of specific features in specific embodiments of a particular invention. Some features described herein in the context of separate embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately in multiple embodiments or in any preferred partial combination. Furthermore, features may be described above as operating in several combinations, and may even be initially claimed as such, but one or more features from a claimed combination may, in some cases, be removed from the combination, and the claimed combination may relate to a partial combination or a variation of a partial combination.
[0060] Similarly, while operations are shown in a specific order in the drawings, this should not be understood as requiring that such operations be performed in a specific illustrated order or sequence, or that all illustrated operations be performed, in order to achieve the desired result. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be incorporated together in a single software product or packaged into multiple software products.
[0061] The above describes specific embodiments of this subject matter. Other embodiments are within the scope of the following claims. In some cases, the actions enumerated in the claims may be performed in a different order, but the desired results can still be achieved. In addition, the processes illustrated in the accompanying figures do not necessarily require the specific order or sequence shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous. [Explanation of symbols]
[0062] 100 video streams 101 frames 102 frames 103 frames 105 Inclusion Zone 111 Text 112 Human Characteristics 113 Potential objects to focus on 120 video screens 121 area, rectangle 122 area, rectangle 123 area, rectangle 125 Inclusion Zones 126 Overlaid Content 130 video streams 200 columns 210 columns 211 Boundary Box 212 Boundary Box 213 Boundary Box 221 Exclusion Zone 222 Exclusion Zone 223 Exclusion Zone 230 Exclusion Zones 240 Exclusion Zones 300 Systems 301 Video Preprocessor Unit 302 Video Streams 311 Text Detector Unit 312 Human Feature Detector Unit 313 Object Detector Unit 320 Inclusion Zone Computer Units or Modules 330 Overlaid Content Matcher Units or Modules 332 Overlaid Content 340 Overlay Units 342 video streams 350 Video Visualizer 460 External devices 500 Computer Systems 510 Processor 520 memory 530 Storage Devices 540 Input / Output Devices 550 System Bus
Claims
1. for each video frame in a sequence of frames of video, identifying a corresponding exclusion zone that excludes overlaid content based on detection of a specified object in an area of the video frame that is within the corresponding exclusion zone; aggregating the corresponding exclusion zones for the video frames in a specified duration or number of the sequence of frames; defining an inclusion zone within the specified duration or number of the sequence of frames of the video within which overlaid content is eligible for inclusion, the inclusion zone being defined as an area of the video frames within the specified duration or number that is outside the aggregated corresponding exclusion zone; providing overlaid content for including the specified duration or number of the sequence of frames of the video within the inclusion zone during display of the video on a client device; A method comprising:
2. 2. The method of claim 1 , wherein identifying the exclusion zone comprises identifying, for each video frame in the sequence of frames, one or more regions in which text appears in the video, and wherein the method further comprises generating one or more bounding boxes that define the one or more regions from other portions of the video frames.
3. The method of claim 2 , wherein identifying the one or more regions in which text is to be displayed comprises identifying the one or more regions using an optical character recognition system.
4. 4. The method of claim 1, wherein identifying the exclusion zone comprises identifying, for each video frame in the sequence of frames, one or more regions in which human features appear in the video, and wherein the method further comprises generating one or more bounding boxes that define the one or more regions from other portions of the video frames.
5. 5. The method of claim 4, wherein identifying the one or more regions in which human characteristics are displayed comprises identifying the one or more regions using a computer vision system trained to identify human characteristics.
6. The method of claim 5 , wherein the computer vision system is a convolutional neural network system.
7. 7. The method of claim 1, wherein identifying the exclusion zone comprises identifying, for each video frame in the sequence of frames, one or more regions in which significant objects appear in the video, and wherein identifying the regions in which significant objects appear is done using a computer vision system configured to recognize objects from a selected set of object categories that do not include text or human features.
8. 8. The method of claim 7, wherein identifying the exclusion zone comprises identifying the one or more regions in which the object of interest appears in the video based on detecting an object moving more than a selected distance between consecutive frames or detecting an object moving for a specified number of consecutive frames.
9. 9. The method of claim 1, wherein aggregating the corresponding exclusion zones comprises generating a union of bounding boxes that define the corresponding exclusion zones from other parts of the video.
10. wherein defining the inclusion zones comprises identifying a set of rectangles within the sequence of frames of the video that do not overlap with the aggregated corresponding exclusion zones over the specified duration or number; providing overlaid content for inclusion in the inclusion zone, identifying an overlay having dimensions that fit within one or more rectangles in the set of rectangles; providing said overlay within said one or more rectangles for said specified duration or number; Including, The method of claim 9.
11. one or more processors; one or more memories storing computer-readable instructions configured to cause the one or more processors to perform the method of any one of claims 1 to 10; A system including:
12. A computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 10.