System and method for processing video of objects or areas of interest

By adopting deep learning technology in the video processing system, rapid detection and tracking of objects or areas of interest in live videos is achieved, solving the problem of limited user interaction in existing technologies and improving the video experience.

CN115633211BActive Publication Date: 2025-09-16AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210804974.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-12
Filing Date
2022-07-08
Publication Date
2025-09-16
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

Existing technologies have difficulty in quickly and accurately detecting and tracking objects or areas of user interest in live video, especially on low-cost home media players or set-top box units, and user interaction methods are limited, affecting the video experience.

Method used

Using deep learning object detection and tracking technology, the system on chip (SoC) or multi-chip module can identify and display potential objects or areas of interest, and provide enhanced video content through appropriate user interface control selection and adjustment.

Benefits of technology

It enables accurate tracking and enhanced display of objects or areas of interest during video playback, improving the user video experience, especially on low-cost devices, and supports multiple interaction methods such as remote control or voice control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115633211B_ABST
    Figure CN115633211B_ABST
Patent Text Reader

Abstract

The present application relates to systems and methods for processing video of objects or regions of interest. The systems and methods may include a processor. The processor may be configured to perform object detection to detect visual indications of potential objects of interest in a video scene, receive a selection of an object of interest from the potential objects of interest, and provide enhanced video content of the object of interest indicated by the selection within the video scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a system and method for processing video of an object or region of interest. Background Art

[0002] The present disclosure relates to video processing, including but not limited to video processing using machine learning. In a digital video system, including but not limited to a set-top box, tuner, and / or video processor, a user can perform functions, such as slow motion, fast forward, pause, and rewind, on detected and tracked video objects, typically mimicking the visual feedback provided by a digital video recorder (DVR) during these operations. Further actions and information can enhance the user's video experience. Summary of the Invention

[0003] In one aspect, the present application relates to a method comprising: providing a first video stream for display; receiving a user selection of an object of interest; and providing a second video stream directed to the same video content as the first video stream, wherein the second video stream includes enhanced video content of the object of interest indicated by the user selection.

[0004] In another aspect, the present application relates to a video processing system comprising: a processor configured to perform object detection to detect visual indications of potential objects of interest in a video scene, the processor configured to receive a selection of an object of interest from the potential objects of interest, and the processor configured to provide enhanced video content of the object of interest indicated by the selection within the video scene.

[0005] In another aspect, the present application relates to an entertainment system for providing a video for viewing by a user, the entertainment system comprising: an interface configured to receive a selection; and one or more processors, one or more circuits, or any combination thereof, configured to: provide a visual indication of potential objects of interest in a video scene; receive a selection of an object of interest from the potential objects of interest; and provide enhanced video content of the object of interest indicated by the selection within the video scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various visual objects, aspects, features and advantages of the present disclosure will become more apparent and better understood by reference to the detailed description taken in conjunction with the accompanying drawings, in which like reference characters identify corresponding elements throughout. In the accompanying drawings, like reference numerals generally indicate identical, functionally similar and / or structurally similar elements.

[0007] Figure 1 is a general block diagram depicting an example system capable of providing enhanced video content in accordance with some embodiments.

[0008] Figure 2is a diagram depicting a video scene with indications of potential objects of interest, according to some embodiments.

[0009] Figure 3 is a depiction of augmented video content with an object of interest according to some embodiments Figure 2 Schema of the video scene.

[0010] Figure 4 is a depiction of augmented video content with an object of interest using a picture-in-picture mode according to some embodiments Figure 2 Schema of the video scene.

[0011] Figure 5 is a diagram depicting a Figure 1 A flowchart of the operation of the system described in

[0014] is used to provide an example enhanced video.

[0012] Figure 6 According to some embodiments, Figure 1 Block diagram of an example video processing system of the system illustrated in .

[0013] Figure 7 According to some embodiments, Figure 1 Block diagram of an example video processing system of the system illustrated in .

[0014] Figure 8 According to some embodiments, Figure 1 Block diagram of an example video processing system of the system illustrated in .

[0015] Figure 9 is a diagram depicting an example electronic program guide display according to some embodiments.

[0016] Figure 10A -B depicts a diagram comprising a Figure 1 A flow chart of a tracking operation for providing example enhanced video is provided by the system described in .

[0017] Figure 11 is a block diagram depicting an example set-top box system configured for detection and tracking operations in accordance with some embodiments.

[0018] Figure 12 is a block diagram depicting an example set-top box system configured for picture-in-picture operation in accordance with some embodiments.

[0019] Figure 13 is a block diagram depicting an example set-top box system configured for object of interest selection in accordance with some embodiments.

[0020] Figure 14is a block diagram depicting an example set-top box system configured for sharing metadata and tracking information according to some embodiments.

[0021] Figure 15 is a block diagram depicting an example set-top box system and television configured for providing enhanced video content according to some embodiments.

[0022] Figure 16 is a block diagram depicting an example sample index format according to some embodiments.

[0023] Figure 17 is a diagram illustrating the use of Figure 16 Block diagram of an example video processing system that uses the index format described in to facilitate trick play in OOI / ROI mode.

[0024] The details of various embodiments of the methods and systems are set forth in the accompanying drawings and the description below. DETAILED DESCRIPTION

[0025] The following is a more detailed description of various concepts related to methods, devices, and systems for video operations and implementations of the methods, devices, and systems. Before turning to a more detailed description and figures that describe example implementations in detail, it should be understood that the present application is not limited to the details or methods set forth in the description or illustrated in the figures. It should also be understood that the terminology used is for descriptive purposes only and should not be considered limiting.

[0026] The present disclosure generally relates to systems and methods for providing object of interest (OOI) or region of interest (ROI) video features that can enhance a user's video experience. As used herein, the term object of interest is intended to refer to an object, person, animal, region, or any video feature of interest. In some embodiments, a video processing system allows a user to automatically zoom in on a visual object or region of a video scene based on his or her interests. For example, for a sports video, a user can view an athlete of interest in greater detail or with more information; for a movie or TV show, a user can highlight his or her favorite actor; for a travel channel, a user can zoom in on a specific scene area; for a shopping channel, a user can zoom in on a particular product, etc.; for a training video, a user can zoom in on a component or part of a device.

[0027] In certain embodiments, video processing system advantageously overcomes the problem associated with object of interest or district moving rapidly between frames in live video.If live video is played on low-cost home media player or set-top box (STB) unit, these problems may be especially difficult so.In addition, in certain embodiments, video processing system advantageously overcomes the problem associated with selecting object of interest or district by remote controller (with or without voice control capability) of low-cost home media player or set-top box unit.In certain embodiments, video processing system tracks potential object of interest or district accurately and provides indication, therefore can more easily select described potential object of interest or district.

[0028] In some embodiments, the video system and method have an object or region of interest video playback architecture that provides a processing flow to address object and region of interest selection, detection, and tracking. In some embodiments, deep learning object detection and tracking technology is provided in a system on a chip (SoC) or a multi-chip module system. In some embodiments, object detection or metadata is used to identify potential objects or regions of interest and display them on the screen. Selection of the object or region of interest can be controlled via an appropriate user interface, such as a remote control or microphone (i.e., a voice interface) on a player or set-top box unit. In some embodiments, object tracking is used to automatically adjust and indicate the object or region of interest in subsequent frames during video playback.

[0029] Some embodiments relate to systems, methods, and apparatus for processing video including a processor configured to perform object detection to detect visual indications of potential objects of interest in a video scene, receive a selection of an object of interest from the potential objects of interest, and provide enhanced video content of the object of interest indicated by the selection within the video scene.

[0030] Some embodiments relate to an entertainment system for providing a video for viewing by a user. The entertainment system includes an interface configured to receive a selection and one or more processors, one or more circuits, or any combination thereof. The one or more processors, one or more circuits, or any combination thereof, are configured to: provide a visual indication of a potential object of interest in a video scene; receive a selection of an object of interest from among the potential objects of interest; and provide enhanced video content of the object of interest indicated by the selection within the video scene.

[0031] Some embodiments relate to a method. The method includes providing a first video stream for display and receiving a user selection of an object of interest. The method also includes providing a second video stream directed to the same video content as the first video stream, wherein the second video stream includes enhanced video content of the object of interest indicated by the user selection.

[0032] Figure 1 1 is a block diagram depicting an example of an entertainment system 10. Entertainment system 10 is any system for providing video (including but not limited to educational systems, training systems, design systems, simulators, gaming systems, home theaters, televisions, augmented reality systems, remote auction systems, virtual reality systems, live video conferencing systems, etc.). In some embodiments, entertainment system 10 includes a user interface 12, a video processing system 14, and a monitor 16. In some embodiments, entertainment system 10 provides object of interest processing for video playback. In some embodiments, entertainment system 10 uses video processing system 14 to process user input to provide augmented video content of a user-selected object of interest. Video including the augmented video content is provided on monitor 16 for viewing by the user.

[0033] The video processing system 14 receives video frames 32 associated with a video stream from a source. The source is any video source, including but not limited to a media player, a cable television provider, an Internet subscription service, a headend, a camera, a storage media server, a satellite provider, a set-top box, a VCR, a computer, or other source of video material. The video processing system 14 includes a selector 20, a tracker 22, and a video enhancer 24.

[0034] Selector 20 identifies or detects potential objects of interest in video frames 32 received at the input and receives a user selection from user interface 12. Selector 20 uses metadata 36 at the input, sound information 34 at the input, and / or video processing of video frames 32 to identify potential objects or regions of interest. In some embodiments, various video and data processing techniques may be used to detect objects of interest and potential objects of interest. In some embodiments, selector 20 and tracker 22 use a deep learning object detection system-on-chip (SoC). In some embodiments, video object detection or metadata is used to identify potential objects or regions of interest.

[0035] Tracker 22 tracks the selected object of interest and potential objects of interest in video frames 32 and provides data to video enhancer 24 so that video enhancer 24 can provide enhanced video of the selected object of interest. The enhanced video is provided to monitor 16 as a video frame in a stream. In some embodiments, tracker 22 uses frame history and motion vectors to track the objects of interest and potential objects of interest. In some embodiments, tracker 22 uses metadata 36, ​​sound information 34 (e.g., sound cues), and / or video processing of video frames 32 to track the objects of interest and potential objects of interest. Tracker 22 automatically tracks the selected object or region of interest in subsequent frames during video playback.

[0036] The video enhancer 24 uses the tracked potential and selected objects or areas of interest from the tracker 22 and provides enhanced video or instructions in subsequent frames. In some embodiments, the video enhancer 24 automatically provides a zoomed image of the object of interest or local area of ​​the scene selected by the user. The zoom level can be controlled by the user interface 12. In some embodiments, the video enhancer 24 automatically provides a highlighted image, a recolored image, a high-contrast image, a higher-definition image, or a three-dimensional image as video enhancement of the object of interest or local area of ​​the scene selected by the user. In some embodiments, the enhanced video includes text information, graphics, icons, or symbols that provide additional information about the object of interest in a video format. In some embodiments, the video enhancer 24 also provides instructions for potential objects of interest so that the user can select those objects of interest. The instructions and enhanced video are provided in the video signal provided to the monitor 16. The video signal can be a stream or sequence of video frames.

[0037] The user interface 12 may be a smartphone, remote control, microphone, touch screen, tablet computer, mouse, or any device for receiving user input (e.g., which may include selection of an object of interest and a region of interest of video enhancement type). In some embodiments, the user interface 12 receives a command from the user interface 12 to initiate an object of interest or region of interest selection process on a set-top box unit or video recorder. The user interface 12 may include a far-field voice interface or a push-to-talk interface, a game controller, buttons, a touch screen, or other selectors. In some embodiments, the user interface 12 is part of a set-top box unit, a computer, a television, a smartphone, a fire stick, a home control unit, a gaming system, an augmented reality system, a virtual reality system, a computer, or other video system.

[0038] Monitor 16 may be any type of screen or viewing medium for the video signal from video processing system 14. Monitor 16 is a liquid crystal display (LCD), a plasma display, a television, a computer monitor, a smart TV, an eyeglass display, a head-mounted display, a projector, a heads-up display, or any other device for presenting images to a user. In some embodiments, monitor 16 is part of or connected to a simulator, a home theater, a set-top box unit, a computer, a smartphone, a smart TV, a Firestick, a home control unit, a gaming system, an augmented reality system, a virtual reality system, or other video system.

[0039] The video stream processed by the video processing system 14 may be in the form of video frames provided from a media server or client device. Examples of media servers include set-top boxes (STBs) that can perform digital video recorder functions, home or enterprise gateways, servers, computers, workstations, and the like. Examples of client devices include televisions, computer monitors, mobile computers, projectors, tablet computers, or handheld user devices (e.g., smartphones). In some embodiments, the media server or client device is configured to output audio, video, program information, and other data to the video processing system 14. The entertainment system 10 has components interconnected by wired or wireless connections (e.g., wireless networks). For example, connections may include coaxial cables, BNC cables, fiber optic cables, composite cables, s-video, DVI, HDMI, component, VGA, DisplayPort, or other audio and video transmission technologies. The wireless network connection may be a wireless local area network (WLAN) and may use Wi-Fi under any of various Wi-Fi standards. In some embodiments, the video processing system 14 is implemented as a single chip or a system on a chip (SOC). In some embodiments, detection of objects of interest, as well as the provision of indicators and enhanced video, is provided in real time.

[0040] In some embodiments, the video processing system 14 includes one or more decoding units, a display engine, a code converter, a processor, and a storage unit (e.g., a frame buffer, memory, etc.). The video processing system 14 includes one or more microprocessors, digital signal processors (CPUs), application specific integrated circuits (ASICs), programmable logic devices, servers, and / or one or more other integrated circuits. The video processing system 14 may include one or more processors that can execute instructions stored in memory to perform the functions described herein. The storage unit includes, but is not limited to, a disk drive, a server, dynamic random access memory (DRAM), flash memory, memory registers, or other types of volatile or non-volatile fast memory. The video processing system 14 may include Figure 1 For example, video processing system 14 may include additional buffers (e.g., an input buffer for storing compressed video frames before they are decoded by a decoder), network interfaces, controllers, memory, input and output devices, conditional access components, and other components for audio / video / data processing.

[0041] The video processing system 14 may provide the video stream in a number of formats, such as different resolutions (e.g., 1080p, 4K, or 8K), frame rates (e.g., 60 fps versus 30 fps), bit precision (e.g., 10 bits versus 8 bits), or other video characteristics. For example, in some embodiments, the received or provided video stream associated with the video processing system 14 includes a 4K ultra-high-definition (UHD) (e.g., 3,840×2,160 pixels or 2160p) or even an 8K UHD (7680×4320) video stream.

[0042] refer to Figure 2 , video processing system 14 provides a video scene 100 on monitor 16. Although video scene 100 is shown as a track and field meet, video scene 100 may be any type of video scene, including any sporting event, movie, television program, auction, simulation, training video, educational video, etc. In some embodiments, video processing system 14 provides frames 102, 104, 106, 108, 110, and 112 around each athlete as an indication of potential objects of interest.

[0043] In some embodiments, boxes 102 , 104 , 106 , 108 , 110 , and 112 are bounding boxes and include labels or numbers for enabling user selection. Figure 2 A video frame 101 is shown of a video scene 100. Boxes 102, 104, 106, 108, 110, and 112 may also be provided around spectators, coaches, and referees or other officials. Although indicators are shown as boxes 102, 104, 106, 108, 110, and 112, other indicators or symbols (e.g., arrows, labels, icons, highlights, etc.) may be used.

[0044] 102,104,106,108,110 and 112 and include athlete's identification, time, track number, name, current position, athlete's game statistics, speed, etc. (e.g., text information 122). In some embodiments, text information may include the price, current quote, or other information of the product in the relevant home shopping application. In some embodiments, text information may be provided together with the athlete's zoomed image or be provided in a part of the screen (e.g., lower left corner) that is unassociated with the action. Text information may include a digital form #1 to #n for identifying frames 102,104,106,108,110 and 112 and selecting one or more of frames 102,104,106,108,110 and 112.

[0045] The user may select one or more of the potential objects of interest for the enhanced video via the user interface 12. Figure 2In the example of FIG, the athlete in box 108 is selected and provided as an enhanced image in a zoomed image by video enhancer 24. The zoomed image may appear in video scene 100 at its tracked location or may be provided on another portion of video scene 100. Video mixing techniques may be used to provide an enhanced video image within video scene 100 to reduce stark contrast.

[0046] The user can adjust the size and position of the object of interest, such as zooming in, out, moving left / right / up / down, or enlarging or reducing the image of the object of interest through the user interface 12. The area of ​​interest can be selected using one object or multiple objects as a group.

[0047] refer to Figure 3 , video processing system 14 provides frame 200 of video scene 100 on monitor 16. Frame 200 is a future frame from frame 101 of video scene 100 and includes box 108 ( Figure 2 ) as a larger zoomed image compared to the other players in scene 100. Text information may be provided with boxes 102, 106, 104, 108, 110, and 112, including current position and speed (e.g., text information 214 ( Figure 2 )). In some embodiments, scene 100 is cropped in frame 200 to provide proportionality of the scaled image.

[0048] refer to Figure 4 , the video processing system 14 provides a frame 300 in the video scene 100 on the monitor 16. Frame 300 is a future frame from frame 101 of the video scene 100 and includes a scaled image 308b of athlete 308a as athlete 308a in picture-in-picture area 304. In some embodiments, picture-in-picture area 304 can be placed in any area on the scene 100 and the size and scaling features can be adjusted by the user. Text information can be provided in area 304 in the scene 100 or outside area 304 in area 306. The text information can include statistics such as the time of the game. Frames can be provided around other objects of interest in the scene 100 other than the athletes. Although only one scaled image is shown in frames 101, 200, and 300, multiple objects of interest can be selected for enhanced video features in some embodiments.

[0049] refer to Figure 5 , video processing system 14( Figure 1 ) executes process 400 to provide enhanced video. Process 400 includes an initial object of interest selection operation 402, followed by operation 404 by the selector 20 ( Figure 1) and tracker 22 perform object detection and tracking processes. In operation 402, a user selects a video enhancement mode for an object of interest or a region of interest. In operation 404, the object detection and tracking process may use deep learning and convolutional neural networks, metadata flags, voice processing, multimodal signal processing, feature extractors, etc. to detect and track the object of interest or the region of interest.

[0050] At operation 404, a frame is provided for display with a frame of superimposed bounding boxes for each potential object of interest detected and tracked by operation 404. At operation 408, a selection of an object is received and video enhancement is provided by the video enhancer 24 for the selected object. In some embodiments, the video enhancement includes object size and position adjustments. At operation 410, tracking of the enhanced video with the selected object is initiated. At operation 412, a frame containing zoomed features of the selected object of interest or a picture-in-picture window containing the selected object (e.g., Figure 3 In some embodiments, the selected object of interest is provided in a region 304 in the video scene. In some embodiments, subsequent frames of the track include video enhancements of the selected object of interest until the user exits the object of interest or region mode or until the object of interest leaves the video scene. In some embodiments, if the object of interest reenters the scene, a new track for the object of interest with enhanced video features is initiated in operation 410.

[0051] refer to Figure 6 The video processing system 14 includes a video decoder 62 that receives a compressed data stream 72, an audio decoder 64 that receives a compressed audio bitstream 74, a post-processing engine 66 that receives a decompressed frame 80, a neural network engine 68 that receives sound and direction data 84, a scaled frame 78, and object filtering parameters 86, and a graphics engine 70 that receives a bounding box 88 and a frame 82. The selector 20, the tracker 22, and the video enhancer 24 cooperate to perform a reference Figure 6 The video processing operations described herein can be performed at a video player or set-top box unit. Figure 6 The described operation.

[0052] The compressed data stream 72 consists of video frames of the scene extracted at the start of the tracking process. Each video frame in the compressed data stream 72 is decoded by the video decoder 62 to provide a decompressed frame 80. The size and pixel format of each decoded video frame of the decompressed frames 80 are adjusted using the post-processing engine 66 to match the input size and pixel format of the object detector or selector 20. According to some embodiments, the post-processing engine 66 performs operations including, but not limited to, scaling, cropping, color space conversion, bit depth conversion, etc.

[0053] The neural network engine 68 runs object detection on each of the scaled frames 78 and outputs a list of detected objects with bounding boxes 88. The object list may be filtered by predefined object size, object type, etc., as well as sound identification and direction generated from the audio decoder 64 based on the compressed audio bitstream 74. The processing is a background process in parallel with normal video processing and display, or is performed when the video display is paused. The filtered bounding boxes 88 are superimposed on top of the decoded frames 82 to provide frames with detected bounding boxes 90 in the enhancer 24. The video associated with the frames with detected bounding boxes 90 is displayed on the monitor 16 ( Figure 1 ) for the user to select which object or area to track via the user interface 12.

[0054] In some embodiments, compressed data stream 72 (e.g., a video bitstream) is a high dynamic range (HDR) video bitstream, and video decoder 62 parses HDR parameters from compressed data stream 72 that are provided to graphics engine 70. Overlay graphics including bounding box 88 are adjusted according to the HDR parameters.

[0055] refer to Figure 7 The video processing system 14 includes a video decoder 702 that receives a compressed data stream 714, a post-processing engine 704 that receives a decompressed frame 716, a neural network engine 706 that receives object filtering parameters based on a user profile 712 and a zoomed frame 718, and a local storage device 708 that receives tracking information 720. The selector 20, the tracker 22, and the video enhancer 24 ( Figure 1 ) collaborate to perform as referenced Figure 7 In some embodiments, the tracking information is pre-generated and calculated or saved as a metadata file on a local device (e.g., local storage device 708) during the recording process. In some embodiments, the tracking information is downloaded or streamed (along with the video) from a cloud source as a metadata stream.

[0056] refer to Figure 8 The video processing system 14 includes a video decoder 802 that receives a compressed data stream 810, a post-processing engine 804 that receives a decompressed frame 812 and a frame scaling parameter 814, a graphics engine 806 that receives a bounding box 822 and a frame 820, and a processor 808 that receives tracking metadata 826 and a scaled frame 816. The selector 20, the tracker 22, and the video enhancer 24 ( Figure 1 ) collaborate to perform reference Figure 8The video processing operations described. The user can select which track to follow based on the metadata file. The tracking information in the metadata file (e.g., tracking metadata 826) contains information about all tracks of interest, such as some specific object types, which can be derived from the user profile explicitly or implicitly based on previous user selection history. The tracking information metadata file has the following fields for each track of interest, including but not limited to frame number, timestamp, track ID, object ID, and bounding box coordinates. Reference Figure 9 , tracking information may be mixed into the electronic program guide (EPG) display 900 to show whether tracking information is available for this program and, if so, what type of tracking information is available.

[0057] refer to Figure 10A -B, video processing system 14( Figure 1 ) Process 1000 is executed for each frame 1002 to provide enhanced video for tracking. In some embodiments, process 1000 performs a tracking process that includes the following three components: motion modeling, appearance modeling, and object detection. In some embodiments, the motion model is used to predict the object's motion trajectory.

[0058] At operation 1004, the video processing system 14 performs shot transition detection to detect a scene change or cross-fade. If the frame contains a scene change or a cross-fade or is part of one, the track is terminated at operation 1007. At operation 1006, if the frame 1002 does not contain a scene change or a cross-fade or is not part of one, the video processing system 14 proceeds to operation 1008. In some embodiments, at operation 1008, a motion model is used to predict the position of the object of interest in the next frame, the next region of interest, or a region associated with the object of interest.

[0059] At operation 1010, the video processing system 14 determines whether the frame 1002 is scheduled for updating with object detection. If the frame is scheduled for updating with object detection, the process 1000 proceeds to operation 1024. At operation 1024, the predicted object of interest or region of interest is used and a detection miss counter is incremented by one. If the frame is not scheduled for updating with object detection, the process 1000 proceeds to operation 1012 and the video processing system 14 detects an object proximate to the predicted object or region of interest. Since the selector 20 ( Figure 1 If the current frame is not scheduled to be updated by object detection due to the processing capacity limit of the detector of the set-top box unit, the predicted object of interest or region of interest is directly output as the object of interest or region of interest of the current frame in operation 1012. Otherwise, object detection is run on the current frame to find objects close to the predicted region of interest.

[0060] At operation 1014, the video processing system 14 determines whether the object detection process has returned an object list on time. If the object detection process has returned an object list on time, then process 1000 proceeds to operation 1016. If the object detection process has not returned an object list on time, then process 1000 proceeds to operation 1024. To speed up detection, object detection is performed only on a portion of the current frame surrounding the predicted object of interest or region of interest. If the object is not found on time in operation 1014, then the predicted object of interest or region of interest is used in operation 1024 and a detection miss counter is incremented by one.

[0061] At operation 1016, if the overlap is greater than T0, then the video processing system 14 merges the detections, where T0 is a threshold. In some embodiments, the list of detected objects is examined in operation 1016 and detections with significant overlap are merged.

[0062] At operation 1018, the video processing system 14 obtains an embedding for the detection. At operation 1022, the video processing system 14 uses the embedding to determine whether the detection best matches the predicted region of interest. If the detection best matches the predicted region of interest, process 1000 proceeds to operation 1028. In some embodiments, the embedding vector from operation 1018 is used to calculate a similarity score between the detection and the target. The detection that best matches the predicted object of interest or region of interest using bounding box overlap and similarity score is selected as a match. If a match is found, the matched detection is used to update the motion model and output the updated object of interest or region of interest.

[0063] In operation 1022, if the detection does not best match the predicted region of interest, flow 1000 proceeds to operation 1024. At operation 1024, the predicted object of interest or region of interest is used and a detection miss counter is incremented by one.

[0064] After operation 1024, the video processing system 14 determines whether the miss counter is greater than T1, where T1 is a threshold value. If the miss counter is not greater than T1, then the process 1000 proceeds to operation 1030. If the miss counter is greater than T1, then the process 1000 proceeds to operation 1007 and the tracking is terminated. Therefore, if it is detected in operation 1024 that the miss counter is greater than the given threshold value T1, then the tracking process is terminated.

[0065] At operation 1028, the video processing system 14 updates the motion model using the matched detection region of interest. At operation 1030, the video processing system 14 calculates a moving average of the center position of the region of interest 1034. In some embodiments, a moving average of the center position of the object of interest or the region of interest is calculated to smooth the tracked object trajectory.

[0066] refer to Figure 11 The video processing system 14 includes a video decoder 1102 that receives a compressed data stream 1112, a post-processing engine 1104 that receives a decompressed frame 1114, a detection region of interest 1116, and a display object of interest or a display region of interest 1118, a host processor 1107 that receives a bounding box and embedding 1124, a graphics engine 1108 that receives a frame 1128, and a neural network engine 1106 that receives object filtering parameters 1122 and a scaled frame 1126. The selector 20, the tracker 22, and the video enhancer 24 ( Figure 1 ) Collaborative Execution Reference Figure 11 Video processing operations described.In some embodiments, a video processing system 14 is provided at the player or set-top box unit.

[0067] The host processor 1107 uses the motion model to generate a predicted object of interest or region of interest and derives a detection region of interest 1116 based on the predicted object of interest or region of interest. The host processor 1107 sends the result (e.g., detection region of interest 1116) to the post-processing engine 1104. The post-processing engine 1104 uses the detection region of interest 1116 to generate a zoomed frame (e.g., zoomed frame 1126) surrounding the predicted object of interest or region of interest for the neural network engine 1106. The neural network engine 1106 performs an object detection process and sends the resulting bounding box and embedding 1124 to the host processor 1107 for target matching. The host processor 1107 uses the bounding box and embedding 1124 to find the best match with the target. Based on the matched results and the magnification ratio, a display object of interest or display region of interest 1118 is derived. The object of interest or region of interest 118 is sent to the post-processing engine 1104 to extract the pixels to be displayed. In some embodiments, upon termination of a track, video processing system 14 may pause or gracefully restore the original full-size window at the last updated frame containing the target.

[0068] refer to Figure 12 The video processing system 14 includes a video decoder 1202 that receives a compressed data stream 1212, a post-processing engine 1204 that receives a decompressed frame 1214, a detected object or region of interest 1224, and a display region of interest 1226, a host processor 1206 that receives a bounding box and embedding 1218, a graphics engine 1208 that receives a key frame 1220 and a picture-in-picture 1222, and a neural network engine 1210 that receives object filtering parameters 1222 and a scaled frame 1126. The selector 20, the tracker 22, and the video enhancer 24 ( Figure 1 ) collaborate to perform reference Figure 12 Video processing operations described.In some embodiments, a video processing system 14 is provided at the player or set-top box unit.

[0069] In some embodiments, the video processing system 14 provides enhanced video in a picture-in-picture mode. After the host processor 1206 determines the object or area of ​​interest 1226, the host processor 1206 sends the determined object or area of ​​interest 1226 to the post-processing engine 1204 to extract the image block of the tracked object. By default, the target image block is displayed as a picture-in-picture window (for example, in some embodiments, using picture-in-picture 1222 and main frame 1220). The user can also swap the main window and picture-in-picture window (for example, displaying the target image block as the main window and the original image as the picture-in-picture window). When the track ends, the video processing system 14 can pause at the last updated frame containing the target or picture-in-picture window or fade out gracefully when the main window continues to play.

[0070] refer to Figure 13 , system 1300 includes a set-top box device 1304 containing a local storage device 1308 and is configured to collect and track user data and share tracking information. In some embodiments, when user 1302 starts the tracking process, the set-top box device 1304 collects a snapshot image with a time stamp of the tracked object. The collected data is stored in a local storage device 1308 (e.g., a flash drive) or a storage server in the cloud 1306 or other network together with the identification of the user 1302. When sent to the cloud 1306, the user data can be encrypted, for example, using a homomorphic encryption algorithm. In some embodiments, the encrypted user data can be analyzed and classified without decryption.

[0071] refer to Figure 14 , system 1400 includes a set-top box device 1402, a cloud database 1408 in a cloud 1404, and a set-top box device 1406. The set-top box device 1406 provides a content identification 1442 to the cloud 1404 and receives tracking information 1446 from the cloud 1404. The set-top box device 1402 provides the content identification, the user identification, and the tracking information to the cloud 1404. The system 1400 is configured to collect and track user data and metadata file information, and upload the metadata file information to the cloud for sharing with other users.

[0072] Tracking information metadata files can be uploaded to the cloud 1404 along with the user ID and content ID. The operator maintains a tracking information metadata database 1410. Other clients can request this metadata from the cloud using the content ID and play back the area or object of interest based on the downloaded metadata. Tracking-related information can also be generated or collected in the cloud 1404. For example, tracking information for a movie can be generated or collected in the cloud 1404. This information can include scene changes, character labels in the scene, object-related information, etc. In some embodiments, this information is embedded in the video service stream or sent as metadata via a side channel to the players or set-top box devices 1402 and 1406.

[0073] refer to Figure 15 , system 1500 includes a set-top box device 1502 and a monitor 1504. In some embodiments, monitor 1504 is coupled to set-top box device 1502 via a high-definition media interface cable. In some embodiments, monitor 1504 is a television. Tracking information is sent to monitor 1504 via the cable as part of frame metadata. In some embodiments, monitor 1504 uses the information to enhance the video, for example, to highlight the tracking target area.

[0074] In some embodiments, the video processing system 14 provides digital video recorder trick play operations on the OOI and ROI. During trick play operations, a direction flag is added to the motion model that indicates whether the current motion model is in the forward or backward direction. During trick play operations, if the direction of the motion model is different from the trick play direction (for example, if the direction of the motion model is forward and the user wants to play backward), the motion model is first inverted by multiplying all motion components by -1 and the inverted motion model is used to predict the next object of interest or region of interest.

[0075] refer to Figure 16 In some embodiments, the index format 1600 includes an index file 1610, a stream file 1620, and a track information metadata file 1630. The video processing system 14 ( Figure 1 ) uses an index format 1600. The index format 1600 provides a configuration that can be used to quickly locate corresponding track information in the metadata file 1630 and frame data in the stream file 1620 associated with the video stream. The index format 1600 can be used with the video processing system 14 to facilitate trick play in OOI / ROI mode (e.g., see below). Figure 17 ).

[0076] Stream file 1620 includes frame n data 1622, frame n+1 data 1624, and frame n+2 data 1626. Frame data 1622, 1624, and 1626 are derived from corresponding frame n index data 1612, frame n+1 index data 1614, and frame n+2 index data 1616. Each of frame n index data 1612, frame n+1 index data 1614, and frame n+2 index data 1614 includes frame data, frame offset data, and track information offset data. Track information metadata file 1630 includes metadata 1632, 1634, and 1636. Each of metadata 1632, 1634, and 1636 includes frame data, track data, and bounding box data for each corresponding frame n, n+1, and n+2.

[0077] refer to Figure 17 , the video processing system 14 is configured to use the index file 1610 ( Figure 16 ) to quickly locate the corresponding track information in the metadata file 1630 and the frame data 1622 in the stream file 1620. The video processing system 14 includes a video decoder 1702 that receives a compressed data stream 1712, a post-processing engine 1704 that receives a decompressed frame 1714 and a frame scaling parameter 1716, and a post-processing engine 1704 that receives a scaled frame 1718 and a frame based on the index file 1610 ( Figure 16 ), a host processor 1708 receiving the extracted trajectory information 1724 of the image processing unit 1701, a graphics engine 1706 receiving the frame 1720 and the bounding box 1726, and a local storage device receiving the extracted frame data based on the index file 1610. The local storage device stores the stream data (e.g., the stream file 1620), the index file 1610, and the metadata file 1630. In some embodiments, Figure 17 The video processing system 14 is configured to operate in DVR trick mode on a selected video object having a bounding box. In some embodiments, the operation is not only a frame indexing operation, but also an object positioning operation in each frame. The local storage device provides the extracted trajectory information based on the index file 1610 to the processor 1708. The video decoder 1702 provides decompressed frames 1714 and the post-processing engine 1704 provides frames 1720 and scaled frames 1718. The processor 1708 uses the extracted trajectory information based on the index file 1610 ( Figure 16 )'s extracted trajectory information 1724 to provide a bounding box 1726.

[0078] It should be noted that certain paragraphs of this disclosure may refer to terms such as "first" and "second" in relation to devices, operating modes, frames, streams, objects of interest, etc., to identify or distinguish one from another or others. These terms are not intended to relate entities solely in time or according to sequence (e.g., a first device and a second device), although in some cases, these entities may include such a relationship. These terms also do not limit the number of possible entities (e.g., devices) that can operate within a system or environment.

[0079] It should be understood that the systems described above may provide multiples of any or each of those components and that these components may be provided on a standalone machine or, in some embodiments, on multiple machines in a distributed system. Furthermore, the systems and methods described above may be provided as one or more computer-readable programs or executable instructions embodied on or in one or more articles of manufacture. The articles of manufacture may be a floppy disk, hard disk, CD-ROM, flash memory card, PROM, RAM, ROM, or magnetic tape. Generally, the computer-readable program may be implemented in any programming language, such as LISP, PERL, C, C++, C#, PROLOG, or in any bytecode language, such as JAVA. The software program or executable instructions may be stored as object code on or in one or more articles of manufacture.

[0080] While the foregoing written description of the methods and systems enables one of ordinary skill in the art to make and use what is presently believed to be the best mode thereof, one of ordinary skill in the art will understand and appreciate that there are variations, combinations, and equivalents of the specific embodiments, methods, and examples herein. Accordingly, the present methods and systems should not be limited to the embodiments, methods, and examples described above, but rather to all embodiments and methods within the scope and spirit of the present disclosure.

Claims

1. A method comprising: providing a first video stream for display; receiving a user selection of an object of interest; and A second video stream is provided that points to the same video content as the first video stream, wherein the second video stream includes enhanced video content of the object of interest indicated by the user selection, wherein the second video stream is provided by a graphics engine coupled to a neural network engine, wherein a host processor is configured to use a motion model to predict a position of the object of interest in a next frame and generate a detection area of ​​interest based on the predicted position of the object of interest, wherein the neural network engine is configured to perform object detection in a zoomed frame surrounding the detection area of ​​interest to detect the object of interest, provide a bounding box for the object of interest, and send the provided bounding box and embedding vector for the object of interest to the host processor for target matching, wherein the host processor is further configured to use the bounding box and the embedding vector to find a matched result for the object of interest and derive a display object of interest based on the matched result, and wherein the graphics engine is configured to provide the enhanced video content based on the display object of interest.

2. The method according to claim 1, further comprising: performing object detection to provide a visual indication of potential objects of interest in the first video stream; and Wherein the object of interest is selected using far-field voice, push-to-talk or remote control selection and the visual indication. The method of claim 2 , wherein the object detection uses target locations in previous frames to derive the detection region of interest. The method of claim 3 , wherein the detection region of interest is used to crop a frame.

5. The method of claim 2, wherein the neural network engine receives and uses sound and direction data, zoomed frames, and object filtering parameters to detect potential objects of interest. The method of claim 5 , wherein trajectory information metadata is provided, the trajectory information metadata comprising a frame number, a trajectory identification number, and bounding box coordinates. 7 . The method according to claim 1 , wherein an index file is provided, the index file containing track information offset data of each frame to quickly locate the track information in the metadata file.

8. The method of claim 1, wherein the enhanced video content includes a zoom feature, wherein a level of the zoom feature is selected by a user.

9. The method of claim 1, further comprising providing the enhanced video content in a picture-in-picture area using a set-top box unit, and wherein the picture-in-picture area is faded out if the object of interest leaves a scene defined by the first video stream.

10. The method of claim 1, wherein the augmented video content is an athlete, an actor, a scenic feature, or a shopping item, wherein the augmented video content includes a three-dimensional image of the object of interest.

11. The method according to claim 1 , further comprising: A trick play operation is used when viewing the second video stream, wherein a motion model inversion process is used to predict the position of an object of interest when the direction of the motion model is different from the trick play direction.

12. The method of claim 1, further comprising: performing object detection to provide a visual indication of potential objects of interest in the first video stream; and Bit conversion is provided to match the pixel format of the first video stream to an input format of an object detector used when performing object detection.

13. The method of claim 1, further comprising: Preselecting video streams of interest based on user profiles; and The video stream of interest is provided in an electronic program guide.

14. The method of claim 1, further comprising: The transition detection or object tracking score is used to terminate the augmented video content in response to a scene change or a cross-fade.

15. The method of claim 1, further comprising: Homomorphic encryption is used when uploading trace information to support data analysis in the encrypted domain.

16. The method of claim 1, further comprising: Tracking information is generated at the edge device and a tracking information metadata database indexed by unique content identification and user identification is used to share the results through a portal.

17. The method of claim 1, further comprising: Tracking information metadata is provided as part of a frame to a television over a high-definition multimedia interface to allow the television to use the tracking information metadata to provide the enhanced video content.

18. The method of claim 1, further comprising: receiving an index file for a trick play mode; and Receives a separate track information metadata file.

19. A video processing system comprising: Circuitry configured to perform object detection to detect visual indications of potential objects of interest in a video scene, the circuitry comprising a processor configured to receive a selection of an object of interest from the potential objects of interest, and the processor configured to provide enhanced video content of the object of interest indicated by the selection within the video scene, the circuitry comprising a host processor configured to use a motion model to predict a position of the object of interest in a next frame and to generate a detection region of interest based on the predicted position of the object of interest, wherein the circuitry comprises a graphics engine coupled to a neural network engine, wherein the neural network engine is configured to perform the object detection in a zoomed frame surrounding the detection region of interest, provide a bounding box for the object of interest, and send the provided bounding box and embedding vector for the object of interest to the host processor for target matching, wherein the host processor is further configured to use the bounding box and the embedding vector to find a matched result for the object of interest and derive a display object of interest based on the matched result, and wherein the graphics engine is configured to provide the enhanced video content based on the display object of interest.

20. The video processing system of claim 19, further comprising an interface configured to receive the selection from a user, wherein the selection is provided using far-field voice, push-to-talk, or a remote control interface.

21. An entertainment system for providing videos for viewing by a user, the entertainment system comprising: an interface configured to receive a selection; and One or more processors, one or more circuits, or any combination thereof, configured to: Providing visual indications of potential objects of interest in a video scene; receiving a selection of a subject of interest from among the potential subjects of interest; and providing enhanced video content of the object of interest indicated by the selection within the video scene, wherein the enhanced video content includes a zoomed, higher definition, three-dimensional, or higher contrast image of the object of interest and textual information related to the object of interest, the textual information including parameters related to a current state of the object of interest in the video scene, wherein the one or more processors include a host processor configured to use a motion model to predict a position of the object of interest in a next frame and generate a detection region of interest based on the predicted position of the object of interest, wherein the one or more processors include a graphics engine coupled to a neural network engine, wherein the neural network engine is configured to perform object detection in zoomed frames surrounding the detection region of interest to detect the object of interest, provide a bounding box for the object of interest, and send the provided bounding box and embedding vector for the object of interest to the host processor for target matching, wherein the host processor is further configured to use the bounding box and the embedding vector to find a matched result for the object of interest and derive a display object of interest based on the matched result, and wherein the graphics engine is configured to provide the enhanced video content based on the display object of interest.

Citation Information

Patent Citations

  • System and Method for Improved General Object Detection Using Neural Networks

    US20170154425A1

  • Systems and methods for selective object-of-interest zooming in streaming video

    WO2018152437A1