Addition of Augmented Reality to the Sub-View of a High-Resolution Central Video Feed

By capturing high-resolution central video feeds and splitting them into sub-views with adjustable augmented reality features, the system addresses the limitations of conventional camera tracking systems, enhancing the viewing experience and reducing resource needs.

JP7699712B2Active Publication Date: 2025-06-27エクソス·アイピー·リミテッド ライアビリティ カンパニー
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024503815
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-21
Filing Date
2022-05-17
Publication Date
2025-06-27
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

Conventional camera tracking systems require significant resources and are prone to losing sight of objects if they move out of the camera's line of sight or if multiple similar objects are present, lacking individual viewer tailoring.

Method used

The system captures a high-resolution central video feed and splits it into sub-views, using augmented reality to enhance the viewing experience by adding features like scrimmage lines and statistics, which can be adjusted based on user interests and content.

Benefits of technology

This approach improves the efficiency of capturing and displaying sports events by providing clear, focused views of objects of interest with enhanced AR features, reducing resource requirements and improving viewer engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699712000001
    Figure 0007699712000001
  • Figure 0007699712000002
    Figure 0007699712000002
  • Figure 0007699712000003
    Figure 0007699712000003
Patent Text Reader

Abstract

Techniques are disclosed for adding augmented reality to a sub-view of a high-resolution central video feed. In various embodiments, a central video feed is received from a first camera on a first iteration basis and time-stamped position information is received from a tracking system on a second iteration basis. The central video feed is calibrated to a spatial region covered by the central video feed. A first sub-view of the central video feed is defined using the received time-stamped position information and a determined number of tiles associated with at least one frame of the central video feed. The first sub-view and a homography defining the placement of an augmented reality element on the at least one frame of the central video feed are provided as output to a device configured to display the first sub-view using the first sub-view and the homography.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Conventional camera tracking systems typically track an object either by manually operating the camera focused on the object of interest or by analyzing the subject captured by each camera. In the first typical method, the object is manually tracked by a camera operator, and the operator ensures that the object is always within the view frame. In the second typical method, a series of images are captured by the camera, and these images are analyzed to determine the optical characteristics of the object to be tracked, such as identifying the color associated with the object or the silhouette of the object. These optical characteristics are recognized in further images, enabling the object to be tracked through the progression of the series of images.

[0002] In the first example method, many resources (such as devices and camera operators) are required to effectively track different types of objects. In the second example method, conventional systems are prone to losing sight of the object if the object suddenly jumps out of the camera's line of sight or if multiple objects optically similar to the object of interest are present within the camera's line of sight. Both example methods are typically not tailored to the interests of individual viewers and are more commonly seen in broadcast media for the general audience. Therefore, there is a need for an improved object tracking and display system.

Brief Description of the Drawings

[0003] The various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.

[0004]

Figure 1

[0005]

Figure 2A

[0006]

Figure 2B

[0007]

Figure 2C

[0008]

Figure 2D

[0009]

Figure 3

[0010]

Figure 4A

[0011]

Figure 4B

[0012]

Figure 5

[0013]

Figure 6

[0014]

Figure 7

Best Mode for Carrying Out the Invention

[0015] The present invention can be implemented in various forms, including a process, an apparatus, a system, a composition of matter, a computer program product embodied on a computer-readable storage medium, and / or a processor (a processor configured to execute instructions stored in and / or provided by a memory connected to the processor). In this specification, these embodiments or any other form that the present invention can take may be referred to as techniques. Generally, the order of the steps of the disclosed processes may be changed within the scope of the present invention. Unless otherwise specified, components such as processors or memories described as being configured to perform tasks may be implemented as general components temporarily configured to perform the tasks at a certain time or as specific components manufactured to perform the tasks. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data such as computer program instructions.

[0016] Hereinafter, with reference to the drawings showing the principles of the present invention, a detailed description of one or more embodiments of the present invention will be given. The present invention is described in relation to such embodiments, but is not limited to any of them. The scope of the present invention is limited only by the claims, and the present invention includes many alternatives, modifications, and equivalents. In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. These details are for illustrative purposes only, and the present invention can be practiced according to the claims without some or all of these specific details. For the sake of simplicity, technical matters well known in the technical field related to the present invention are not described in detail so as not to make the present invention unnecessarily difficult to understand.

[0017] Techniques are disclosed for adding augmented reality to a sub - view of a high - resolution central video feed. In various embodiments, the high - resolution central video feed is captured by a high - resolution camera having a full view of the playing space. The high - resolution central video feed can be split into one or more sub - views. In various embodiments, the sub - views are isolated shots of a portion of the playing space and can be focused on one or more objects of interest (such as a ball, a player, or a group of players). Using the example of a stadium or an American football game, the high - resolution central video feed captures at least the entire playing field. In various embodiments, the high - resolution central video feed also captures areas (such as benches or other areas) along the sides of the field of interest where players or other objects of interest may be present when not on the field. The isolated shots can follow a ball, a particular player, a group of players, and / or other objects of interest. To enhance the viewing experience of spectators (sometimes also referred to as users of the disclosed system) of a football game, augmented reality can be added to the isolated shots. The augmented reality can more clearly show the scrimmage line, statistics related to what is happening during the game (such as the speed of the ball, shooting rate, etc.), advertising content, etc. Different from conventional augmented reality components added to television - broadcast sports events, the augmented reality components added to the isolated shots according to the disclosed technology are adjustable according to the content of the isolated shots and / or the user's interests. The augmented reality can be distributed to the client device after being added by a central server or can be added locally by the client device using metadata sent by the central server.

[0018] The examples in this specification mainly use a stadium or an American football game, but the technology of benefiting from the capture of video of an area and providing augmented reality to a sub - view of that area can be applied to various sports (and non - sports) events, so this is merely an example and is not intended to be limiting.

[0019] The disclosed technology is applicable in various situations, but not limited to, quickly and easily changing shots, defining a shot until an event occurs to an object of interest (e.g., a ball is caught), and then tracking the object of interest (e.g., a player with the ball). This improves the efficiency of capturing a sports event by complementing or replacing at least a part of the conventional method of capturing a video of a sports event.

[0020] FIG. 1 is a flowchart showing an embodiment of a process for adding augmented reality to a sub-view of a high-resolution central video feed. This process may be executed in (such as the system shown in FIG. 5). In various embodiments, the process is executed by the cooperation of a separate shot engine 514 and an augmented reality engine 512.

[0021] The process begins by receiving the transmission of the central video feed from the first camera on a first iteration basis (step 100). The central video feed is a series of one or more frames of video that captures a full view of the game space. In other words, the central video feed covers the entire playing field and captures the entire scene. In various embodiments, the central video feed also captures areas (such as benches or other areas) adjacent to or related to the game space where players or other objects of interest may be present when not on the field.

[0022] Referring to FIG. 7, the central video feed is received, for example, from the first camera 740. In some embodiments, the camera 740 is a fixed camera (e.g., the camera is restricted in movement along at least one axis). For example, in some embodiments, the camera 740 is fixed such that while the camera can have variable tilt, pan, and / or zoom, it cannot be physically moved to another location. In some embodiments, the camera 750 is fixed such that the camera cannot move along any axis and / or cannot zoom. In some embodiments, the camera is not fixed and the field registration data can be utilized to obtain the orientation for the purpose of applying the split camera effect.

[0023] The camera can be placed in various positions and orientations, such as, in particular, horizontally at the first end of the playing field (e.g., the half court line, the 50-yard line) or vertically at the second end of the field (e.g., the end zone, the goal). As further described herein, the camera may be configured to capture a high-resolution image (high-resolution central video feed) such that when the view is split into subviews, the subviews have sufficient resolution to be displayed on the user's device (such as a smartphone). The camera 740 communicates with a network to communicate with one or more devices and systems of the present disclosure.

[0024] In some embodiments, the central video feed includes and / or is included in a plurality of central video feeds or other video feeds, and each video feed is generated by one or more cameras arranged and oriented to generate video of at least a portion of the playing field. In some embodiments, the central video feed and / or another video feed may be generated, at least in part, by combining video data generated by a plurality of cameras (such as a composite video or video otherwise combined or synthesized). In some embodiments, a plurality of cameras may be provided around the environment, and each camera covers the entire playing field. Some central video feeds may be generated to enable play scenes from different angles and viewpoints. The central video feed may include time stamps from each data stream to calibrate / align the composition of the scene.

[0025] The process receives respective time-stamped position information from the tracking system at a second iteration criterion (step 102). In various embodiments, the time-stamped position information corresponds to a player or other object of interest. For example, a player can be a point in space, and the amount of padding around that point defines the level of zoom around the player as further described herein.

[0026] Any tracking system may be used, and the system described for this example is merely an example and is not intended to be limiting. In various embodiments, the tracking system includes tracking devices worn by corresponding subjects (e.g., players) participating in the game in a spatial region or associated with other objects of interest (e.g., balls). Respective time-stamped position information is received from each tracking device. For example, each tracking device transmits position information describing the time-stamped position of the corresponding subject in the spatial region.

[0027] An example of a tracking system and the transmission of timestamped location information is further described with respect to FIG. 7. Referring to FIG. 7, in an example of a tracking system, an array of anchor devices (e.g., anchor device 720-1, anchor device 720-2, ···, anchor device 720-Q) receives telemetry data from one or more tracking devices associated with respective subjects or objects of interest in a game. The subjects or objects of interest (represented by squares and circles) may have one or more tracking devices attached to their bodies or otherwise monitoring their movement / behavior.

[0028] The process calibrates the central video feed with respect to the spatial region covered by the central video feed (step 104). The process performs the calibration, such as by dividing at least one frame of the central video feed into a plurality of tiles at a first resolution and determining a homography that defines the placement of augmented reality elements on at least one frame of the central video feed.

[0029] The central video feed is calibrated with respect to a spatial region represented in at least two dimensions covered by the central video feed. In some embodiments, the spatial region is the region captured by an array of anchor devices 720. The spatial region can be the playing field of a live sports event. In some embodiments, calibration of the central video feed includes determining an equivalent portion of the central video feed with respect to the coordinate system used by the position information (e.g., telemetry data). Since standard playing fields for competitive sports include boundary lines that are in accordance with the rules, i.e., of uniform length and thickness / width (e.g., out-of-bounds lines, half-court lines, yard lines, etc.), these lengths and thicknesses can be utilized to determine coordinate positions within the video feed. For example, if it is known that a line on the playing field has a uniform thickness (e.g., 6 centimeters thick) and it is determined that the thickness of the line in the central video feed linearly decreases from a first thickness to a second thickness, the exact position of the subject with respect to the line can be determined within the central video feed.

[0030] In various embodiments, the process divides at least one frame of the central video feed into a plurality of tiles of a first resolution. Referring briefly to FIG. 5, the central video feed captures the entire field of view and is divided into nine tiles (separate shots labeled in the figure). Each of the tiles captures a sub-view of the entire field of view. The number of tiles is merely an example and is not intended to be limiting. Other examples include dividing an 8K frame of a high-resolution camera feed into 16 (or more generally n) segments depending on the desired resolution. In various embodiments, the tiles are stored on a server, a client requests a tile or a subject / object of interest, and the corresponding tile is supplied to the client. An example of dividing at least one frame of the central video feed into a plurality of tiles of a first resolution is further described with respect to FIG. 5.

[0031] In various embodiments, camera calibration information may be utilized to determine a homography that defines the placement of augmented reality elements on at least one frame of a central video feed. Referring briefly to FIG. 5, an AR engine 512 determines a homography, and a split shot engine 514 outputs information to a client device 550 using the determined homography. The client renders an image that includes augmented reality elements incorporated within the image using the homography data or metadata. An example of the process of utilizing a homography to place augmented reality elements on a frame of a central video feed is further described with respect to FIG. 3.

[0032] The process uses the received timestamped position information and a plurality of tiles associated with at least one frame of the central video feed to define a first subview of the central video feed (step 106). In various embodiments, the first subview includes a portion of a plurality of tiles associated with at least one frame of the central video feed. The first subview is associated with a first set of one or more subjects. For example, the first subview may be associated with (may include / may display) a first subject among a plurality of subjects within a spatial region. The first subview includes a corresponding subframe associated with the first set of subjects for each of the plurality of frames that make up the central video feed. In a non-limiting example, the subview defines frames around a single player or multiple players, can track an object of interest (such as a ball), and can always track one or more players even when the player is not within the playing field (e.g., the player is on the bench), or can track other subjects / objects of interest (such as game officials). By using metadata, a player or object of interest can be tracked. The metadata is used to continuously generate split shots that create a visual effect of the camera tracking the player or object of interest.

[0033] For example, in some embodiments, the process is based on timestamp data including received location information and location information associated with each timestamp (e.g., the XYZ coordinates of subject A), and at least partially based on corresponding camera / video calibration data, a mathematical transformation is applied to each of a plurality of sequential frames of video data to determine a portion or part of each sequential frame associated with the corresponding location information of subject A. The determined portion / part of the sequential frame is used to provide a sub-view of the central video feed associated with subject A.

[0034] The sub-view, in various embodiments, has a different resolution than the central video feed. Despite having a different resolution, the difference in quality is not necessarily noticeable to the average viewer so that the viewing experience remains enjoyable. For example, the central video feed is provided at a first resolution (e.g., the native resolution of camera 140) such as between 2K and 12K. Up to this point, in some embodiments, the central video feed includes a plurality of full two-dimensional frames (e.g., a first frame associated with a first time point, a second frame associated with a second time point, ···, an nth frame associated with an nth time point). Each of the plurality of full two-dimensional frames has a first dimension size and a second dimension size (e.g., a horizontal size and a vertical size such as the number of horizontal pixels and the number of vertical pixels). The first sub-view includes a corresponding sub-frame for each of the plurality of full two-dimensional frames. Each corresponding sub-frame is a part of the corresponding full frame (e.g., sub-view / separated shot 1 and sub-view / separated shot 2 of FIG. 5 show immediate sub-frames of the central video feed full frame (entire field of view) of FIG. 5).

[0035] As described herein, the first sub-view of the central video feed can be defined at a second resolution that is smaller than the first resolution (the resolution of the tile). For example, the first resolution can be at least 4 times, 6 times, or 8 times the pixel resolution of the second resolution of the video divided from the central video feed.

[0036] The sub-views can have various zoom levels. The zoom can be defined around the player or the object of interest by setting padding around the player, as further described with respect to FIGS. 2A-2D.

[0037] In some embodiments, each sub-frame has a third dimension size and a fourth dimension size. Further, the third dimension size can be a fixed ratio of the first dimension size, and the fourth dimension size can be a fixed ratio of the second dimension size. For example, the fixed ratio of the first dimension size and the fixed ratio of the second dimension size of the same ratio (e.g., 10%, 20%, 30%, ···, 90%). Similarly, the fixed ratio of the first dimension size can be a first ratio, and the fixed ratio of the second dimension size can be a second ratio different from the first ratio (e.g., the central video feed is captured horizontally and each sub-view is divided vertically). In non-limiting examples, (i) the first dimension size is 7680 pixels, the third dimension size is 3840 pixels, the second dimension size is 4320 pixels, the fourth dimension size is 2160 pixels, or (ii) the first dimension size is 8192 pixels, the third dimension size is 3840 pixels, the second dimension size is 4320 pixels, the fourth dimension size is 2160 pixels. In some embodiments, each of the plurality of full two-dimensional frames includes from at least 10 megapixels to 40 megapixels. In some embodiments, a sub-view (e.g., the first sub-view) includes a corresponding sub-frame that includes from less than 5 megapixels to 15 megapixels for each of the plurality of full two-dimensional frames.

[0038] The coordinates of the center of the first sub-view within the central video feed change over time without human intervention according to the change in the position of the first subject determined from the repetition of reception performed according to the second repetition criterion by duplication. In some embodiments, the center of the first sub-view is associated with position coordinates (e.g., XYZ) generated by a tracking device worn on or otherwise associated with the subject. In some embodiments, the subject may wear multiple tracking devices, and the first sub-view is centered based on a set of coordinates generated based on tracking data from the multiple devices. For example, device data from multiple tracking devices worn by the subject may be correlated, for example, based on timestamp data, and a geometric or other set of center coordinates may be calculated based on the coordinates generated by each tracking device.

[0039] In some embodiments, the first sub-view of the central video feed is communicated to a remote device (e.g., the client device 550 of FIG. 5) independently of the central video feed. In response, the communication causes the remote device to display the first sub-view of the central video feed. In non-limiting examples, the remote device is a handheld device (such as a smartphone, tablet, gaming console, etc.), a fixed computer system (such as a personal home computer, etc.), and the like. Further, the communication can be performed wirelessly (e.g., via a network).

[0040] In various embodiments, at least a first subject within a portion of the subject is selected. The selection of at least the first subject can be performed, for example, by an operator of a computer system (e.g., a video production professional, a producer, a director, etc.), an end user of each remote device (e.g., via each user device 550), or automatically by the computer system. For example, the first subject can be automatically selected based at least in part on proximity (being within a threshold distance) to a ball or other subject (e.g., a previously selected subject to which the subject is related, such as in a one-on-one game or in relation to complementary positions (such as offensive and defensive linemen against each other)). Further, the sub-view can be selected from a broader group of sub-views (e.g., a list of available sub-views, a preview of available sub-views, etc.). The broader group of sub-views includes sub-views for each player participating in the game (e.g., 22 sub-views for an American football game). This end-user selection enables each user to select one or more subjects as desired. For example, if an end user has a list of favorite subjects across multiple teams, the end user can view sub-views of each of these favorite subjects on a single remote device and / or display.

[0041] In some embodiments, the identity of the first subject is received at the remote device. For example, the first sub-view includes information related to the identity of the first subject (e.g., the name of the first subject). This identity of each subject enables the end user to quickly identify different sub-views when browsing two or more sub-views. In some embodiments, the tracking device is attached to (e.g., embedded within) a ball used in a competitive sport in a spatial area. Thus, the identity of the first subject is determined without human intervention based on determining which of the plurality of subjects is currently closest to the ball using each transmission of timestamped position information from each tracking device.

[0042] The process outputs the first sub-view and the determined homography to a device configured to display the first sub-view using the first sub-view and the homography (step 108). In various embodiments, the client device generates a composite image showing a combination of players on the field together with the AR component using a blank scene (empty field), player information (e.g., tracking information), and AR component information (the determined homography). An example of the process for displaying the first sub-view including the AR component is further described with respect to FIGS. 3 and 4.

[0043] In various embodiments, one or more steps of the process of FIG. 1 are performed during a live game in which a plurality of subjects participate. However, the present disclosure is not limited thereto. For example, the communication may be performed after the live game (e.g., viewing highlights of the live game or replays of the live game, etc.).

[0044] Figure 2A shows an example of a subview obtained in various embodiments of the present disclosure. As background, the entire field is shown by the dashed line in Figure 2A. The subview (labeled as a split shot) is specified by a box around a portion of the player on the field. In this example, the object of interest is the player represented by a circle at the center of the split shot.

[0045] Figure 2B shows an example of a subview at a first zoom level obtained in various embodiments of the present disclosure. The zoom is centered on the object of interest labeled in Figure 2A. The object of interest is a point (defined by coordinates x, y, or x, y, z), and the subview is centered on that point in various embodiments. The level of zoom is defined by the amount of padding (here x'') surrounding the object of interest.

[0046] Figure 2C shows an example of a subview at a second zoom level obtained in various embodiments of the present disclosure. Compared to Figure 2B, the subview in this figure is further zoomed out so that the player appears smaller / less detailed. Here, the padding around the object of interest is a different value (x) than the padding in Figure 2B, which makes the zoom level appear different.

[0047] Figure 2D shows an example of a subview at a third zoom level obtained in various embodiments of the present disclosure. Compared to Figure 2C, the subview in this figure is further zoomed out so that the player appears smaller / less detailed. Here, the padding around the object of interest is a different value (x') than the padding in Figure 2C, which makes the zoom level appear different.

[0048] Figure 3 is a flowchart showing one embodiment of a process for synthesizing a subview including augmented reality. The process can be executed by the system of Figure 5. The process of Figure 3 is described with reference to Figures 4A and 4B.

[0049] Figure 4A shows an example of components for constructing a synthetic image according to various embodiments of the present disclosure. The components include a blank scene 400, a player 420, and an augmented reality component 430. The blank scene 400 shows a playing field with no players on the field. For example, the player 420 can be determined by subtracting or otherwise removing the blank scene 400 from a frame that captures the scene of the player on the field (tracking data further described herein may also be used). The AR component 430 includes any component that enhances the frame of the video. In this example, the AR component is the scrimmage line 432. The scrimmage line can be displayed to be visually distinguishable, such as with an accent color, to help the user see more clearly.

[0050] Figure 4B shows an example of a synthetic image obtained by combining the components according to various embodiments of the present disclosure. This synthetic image is obtained by combining the blank scene 400, the player 420, and the AR component 430.

[0051] Returning to Figure 3, the process begins by capturing a blank scene (step 300). The blank scene 400 can be captured by a camera (such as camera 740) before any player enters the field. The blank scene can be a baseline or reference frame for other frames (such as frames of a video showing various states of the play of the game).

[0052] The process determines the augmented reality components to add to the scene (step 302). The augmented reality components can be determined based on predetermined settings or user interests. For example, for the sake of enhancing the user experience of all users, this AR component can be determined for all frames. The component can be determined based on the player 420's position by determining the position where the ball was placed after the most recent play ended based on the tracking data and taking into account any penalty yards. The user can be interested in other information (such as the statistics of a particular player). The statistics are determined and can be added to the scene as an AR component, for example, at a corner of the display. Referring briefly to FIG. 5, the AR engine 512 is configured to determine AR components in various embodiments.

[0053] The process reconstructs the image by synthesizing a blank scene, the determined augmented reality components, and the player (step 304). The process creates a composite image by combining (e.g., overlaying) the player 420 onto the blank scene 400 and any AR component 430 on top.

[0054] Referring briefly to FIG. 5, the client device 550 is configured to create a composite image in various embodiments. Alternatively, the server 510 is configured to create the composite image and send the composite image and / or related data to the client device 550.

[0055] FIG. 5 is a block diagram showing an embodiment of a system for adding augmented reality to a sub-view of a high-resolution central video feed. The system includes a server 510 and a client device 550. The client device 550 can be a smartphone, a computer, or other device on which one or more video frames are rendered.

[0056] Server 510 includes an AR engine 512 and a split shot engine 514. A user selection store 516 is configured to store user selections and / or profiles and may be provided locally on the server or remotely as shown in the figure. The AR engine 512 is configured to determine one or more AR components to be displayed on a frame of video data. The AR components can be based on a user selection of the viewer or a known selection. For example, a scrimmage line can be rendered to visualize the current state of play and determined as an AR component and added to the frame of the video. Other AR components can be more user specific depending on the interests of the user (such as a user who is a fan of a particular player or group of players). The split shot engine 514 is configured to execute a process (such as the process of FIG. 1) to determine a sub-view centered on a player / object of interest.

[0057] FIG. 6 is a diagram showing an example of a graphical user interface for adding augmented reality to a sub-view of a high-resolution central video feed according to an embodiment of the present disclosure. The graphical user interface includes a play panel 610, a video panel 650, and a player panel 680.

[0058] The play panel 610 displays various plays related to the sports event currently being shown on the video panel 650. The sports event can be viewed live or after the event has ended. By selecting the corresponding play within the play panel, a specific play can be viewed. In this example, the user is viewing a specific play identified by play ID 195. Related information such as the start time of the play and the status of the play, such as which player (if any) is holding the ball, is displayed. In some embodiments, the play start time and / or play status information is automatically determined, for example, by processing the video content using artificial intelligence, machine learning, and / or related technologies. In some embodiments, the play start time and / or play status information may be entered, in whole or in part, as input by a human operator.

[0059] The video panel 650 displays a video of the sports event. The video can be a sub-view generated using the disclosed technology. The sub-view can be a synthetic image that includes players and AR components. The video can be maximized to be displayed across the entire screen by selecting an icon in the lower right corner of the video. Also, various options are displayed along the top of the video panel. In this example, the user can select the angle of the video image. Here, the video is from an 8K camera located in the upper left corner of the field. Other video feeds with different resolutions and / or at other locations around the field may be available. Another option is the type of filter to apply to the video. In this example, the default is broadcast shading. Other filters include black and white or other color schemes. Filtering can be performed locally on the client device. Through the AR drop-down menu, the user can select one or more AR components to be rendered on the video panel 650.

[0060] Although not shown in this figure, AR components such as a scrimmage line can be displayed on the video panel. The AR components can be customized according to the user. For example, the color of the scrimmage line can follow the user's preference. Different from the conventional scrimmage line displayed in television broadcasts, the scrimmage line determined according to the technology of the present disclosure is more accurate because the viewport (which constitutes tiles and is considered a virtual camera) moves.

[0061] The position and number of the menus are merely examples and are not intended to be limiting. For example, the menus may alternatively be displayed on both sides or the bottom of the video panel.

[0062] The player panel 680 shows at least a part of the team roster. Here, player 15 (the quarterback of KC) is emphasized because the user is interested in this player. The video panel 650 displays a sub - view centered on player 15. The user can view sub - views related to other players by selecting one or more other players. The user can reset the customization / personalization by selecting "Clear All".

[0063] FIG. 7 shows an example of an environment including a playing field with tracking components according to an embodiment of the present disclosure. The system is an example of a system that can capture the central video feed used in step 100 and collect the time - stamped position information used in step 102.

[0064] The environment 700 includes a playing field 702 where a game (e.g., a football game) is being played. The environment 700 includes an area 704 that includes the playing field 702 and an area directly surrounding the playing field (e.g., an area that includes subjects not participating in the game such as subject 730-1). The environment 700 includes an array of anchor devices 720 (e.g., anchor devices 720-1, anchor devices 720-2, ···, anchor devices 720-Q) that receive telemetry data from one or more tracking devices associated with each subject of the game. As shown in FIG. 7, in some embodiments, the array of anchor devices communicates with a telemetry parsing system. Further, in some embodiments, one or more cameras 740 capture images and / or videos of a sports event that are used to form a virtual replay.

[0065] The camera 740 is a high-resolution camera that can capture video at a high resolution (such as 12K or other resolutions available on the market). In various embodiments, the camera may have various lenses, such as those that are suitable for minimizing distortion in subviews by dividing a central video feed into subviews. In various embodiments, the camera captures an image of the field at a relatively high camera angle.

[0066] As described herein, the central video feed can be divided into subviews, where padding defines the zoom level of the subviews. Since the central video is high resolution, the subviews can be displayed on a user device at various levels of zoom that are not too coarse. In FIG. 7, the square markers represent subjects of the first team of the game, and the circular markers represent subjects of the second team of the game.

[0067] Each transmission of the timestamped location information (e.g., telemetry data 230) is received from each tracking device 300 among the plurality of tracking devices. The repetition criterion when receiving the transmission of the timestamped location information can be the ping value of each tracking device 300 (e.g., the instant ping value 310 in FIG. 3). In some embodiments, the transmission of the timestamped location information from each tracking device among the plurality of tracking devices is performed with a bandwidth greater than 500 MHz or a specific bandwidth ratio of 0.20 or more. In a non-limiting example, the transmission of the timestamped location information from each tracking device among the plurality of tracking devices is in the range of 3.4 GHz to 10.6 GHz, each tracking device 300 among the plurality of tracking devices has a signal refresh rate between 1 Hz and 60 Hz, and / or the repetition criterion is between 1 Hz and 60 Hz. Each tracking device 300 among the plurality of tracking devices transmits a unique signal that is received by the receiving side and identifies each tracking device. Each tracking device can transmit biometric data (e.g., biometric telemetry 236) unique to each subject associated with each tracking device when biometric data is collected.

[0068] Each tracking device 300 is worn by the corresponding subject among the plurality of subjects participating in the competition in the spatial area. Further, each tracking device 300 transmits location information (e.g., telemetry data 230) describing the timestamped location of the corresponding subject in the spatial area. In some embodiments, there are at least two tracking devices 300 worn by each subject among the plurality of subjects. Each additional tracking device 300 associated with the corresponding subject reduces the amount of error in predicting the actual location of the subject.

[0069] In some embodiments, the plurality of subjects includes a first team (e.g., a home team) and a second team (e.g., an away team). In some embodiments, the first team and / or the second team is included in a league of teams (e.g., a football league, a basketball association, etc.). The first team includes a first plurality of players (e.g., the players on the first roster), and the second team includes a second plurality of players (e.g., the players on the second roster). Throughout various embodiments of the present disclosure, the first team and the second team are participating in a competitive game (e.g., a live sports event), such as a football game or a basketball game. Accordingly, the spatial region is a playing field of a competitive game, such as a football field or a basketball court. In some embodiments, the subjects of the present disclosure are players, coaches, referees, or combinations thereof related to the game.

[0070] In some embodiments, each timestamped position of the plurality of independent timestamped positions for each player of the first or second plurality of players includes the xyz coordinates of the respective player with respect to the spatial region. For example, in some embodiments, the spatial region is mapped such that a central portion of the spatial region (e.g., a half court, a 50-yard line, etc.) is the origin of the axis and a boundary region of the spatial region (e.g., an out-of-bounds line) is the maximum or minimum coordinate of the axis. In some embodiments, the xyz coordinates have an accuracy of ±5 centimeters, ±7.5 centimeters, ±10 centimeters, etc.

[0071] The above embodiments have been described in some detail for ease of understanding, but the present invention is not limited to the provided details. There are many alternative ways to implement the present invention. The disclosed embodiments are illustrative and not intended to be limiting.

Claims

**Claim 1** A system comprising: A communication interface configured to: Receive transmission of a central video feed from a first camera based on a first iteration criterion, and Receive respective timestamped position information from a tracking system based on a second iteration criterion; A processor connected to the communication interface, the processor configured to: Calibrate the central video feed with respect to a spatial region covered by the central video feed by splitting at least one frame of the central video feed into a plurality of tiles at a first resolution and determining a homography that defines placement of augmented reality elements on the at least one frame of the central video feed; Define a first sub-view of the central video feed using the received timestamped position information and the plurality of tiles associated with at least one frame of the central video feed, the first sub-view including: A portion of the plurality of tiles associated with at least one frame of the central video feed, Being associated with a first subject among a plurality of subjects within the spatial region, and Output the first sub-view and the determined homography to a device configured to display the first sub-view using the first sub-view and the homography; The system. **Claim 2** The system of claim 1, wherein the first subject among the plurality of subjects within the spatial region includes at least one player or object of interest. **Claim 3** The system of claim 1, wherein the first subject among the plurality of subjects within the spatial region includes a plurality of subjects. **Claim 4** The system of claim 1, wherein the portion of the plurality of tiles is selected based at least in part on the first subject. **Claim 5** The system of claim 1, wherein the central video feed includes a full view of the spatial region, and the spatial region includes a game space. **Claim 6** The system of claim 1, wherein the first sub-view of the central video feed includes a split shot of a portion of the spatial region. **Claim 7** The system according to claim 1, wherein the first sub-view includes a plurality of video frames depicting the first subject.

8. The system according to claim 1, wherein the first sub-view includes a plurality of video frames centered on the first subject.

9. The system according to claim 1, wherein the first sub-view is zoomable, and the zoom level is at least partially based on the padding around the first subject.

10. The system according to claim 1, wherein dividing at least one frame of the central video feed into a plurality of tiles of a first resolution includes dividing the at least one frame into a predetermined number of tiles at least partially based on a desired resolution.

11. The system according to claim 1, wherein the augmented reality element includes a skrim line.

12. The system according to claim 1, wherein the augmented reality element includes statistical values related to a game corresponding to the central video feed.

13. The system according to claim 1, wherein the augmented reality element includes customized advertisement content for the user.

14. The system according to claim 1, wherein the processor is configured to capture a blank scene and determine augmented reality elements to add to the blank scene.

15. The system according to claim 1, wherein the device is configured to display the first sub-view using the first sub-view and the homography, such as by reconstructing an image by synthesizing a blank scene, the augmented reality element, and the first subject.

16. The system according to claim 1, wherein the first camera is fixed.

17. The system according to claim 1, wherein the first camera captures the central video feed at a resolution of at least 12K.

18. The system according to claim 1, wherein the device is configured to display a graphical user interface including a play panel, a video panel including the first sub-view, and a player panel, and to display the first sub-view using the first sub-view and the homography.

19. A method comprising: Receiving transmission of a central video feed from a first camera at a first iteration criterion; Receiving respective time-stamped position information from a tracking system at a second iteration criterion; Dividing at least one frame of the central video feed into a plurality of tiles of a first resolution; Calibrating the central video feed with respect to a spatial region covered by the central video feed by determining a homography that defines an arrangement of augmented reality elements on the at least one frame of the central video feed; Using the received time-stamped position information and the plurality of tiles associated with at least one frame of the central video feed to define a first sub-view of the central video feed, the first sub-view comprising: Including a portion of the plurality of tiles associated with at least one frame of the central video feed; Being associated with a first subject among a plurality of subjects within the spatial region; Providing the first sub-view and the determined homography for output to a device configured to display the first sub-view using the first sub-view and the homography; A method comprising.

20. A computer program product embodied in a non-transitory computer-readable medium, Computer instructions for receiving transmission of a central video feed from a first camera at a first iteration criterion; Computer instructions for receiving respective time-stamped position information from a tracking system at a second iteration criterion; Computer instructions for dividing at least one frame of the central video feed into a plurality of tiles of a first resolution and calibrating the central video feed with respect to a spatial region covered by the central video feed by determining a homography that defines an arrangement of augmented reality elements on the at least one frame of the central video feed; Computer instructions for defining a first sub-view of the central video feed using the received timestamped position information and the plurality of tiles associated with at least one frame of the central video feed, and the first sub-view is comprising a portion of the plurality of tiles associated with at least one frame of the central video feed, associated with a first subject among the plurality of subjects within the spatial region, computer instructions for providing the first sub-view and the determined homography to a device configured to display the first sub-view using the first sub-view and the homography, A computer program product comprising.

Citation Information

Patent Citations

  • Video processing device, video processing method, and program

    JP2010183302A

  • Camera calibration device, method, and program

    JP2018028864A

  • Imaging apparatus, imaging system, mobile, and chip

    JP2018186455A

  • Image processing device, image processing method, and program

    JP2020086983A

  • Encapsulation of image data

    JP2020127244A