Automatic Alignment of Video Streams

The method aligns video streams from diverse sources using visually distinctive activities, ensuring synchronized and accurate processing of live events, particularly sports events, through techniques like optical character recognition and computer vision.

JP2025521947AActive Publication Date: 2025-07-10GENIUS SPORTS SS LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025500371
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-01
Filing Date
2023-07-03
Publication Date
2025-07-10
Estimated Expiration
2043-07-03

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for aligning video streams from multiple sources, particularly those captured by portable devices like smartphones, which are not addressed by conventional broadcast camera solutions.

Method used

A method and system for determining the time offset between video streams by identifying visually distinctive activities such as game clock changes, camera flashes, electronic displays, ball movement, and participant positions, using optical character recognition and computer vision techniques to synchronize and align the streams.

Benefits of technology

Provides robust and accurate alignment of video streams, enabling spatio-temporal pattern recognition and generation of synchronized extended video content from multiple viewpoints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521947000001_ABST
    Figure 2025521947000001_ABST
Patent Text Reader

Abstract

A method and system for determining a time offset between a first video stream and a second video stream depicting a sports event. Identify depictions of a first type of visually distinctive activity in the first and second video streams. Determine the time offset between the two video streams, at least in part, by comparing the depiction of the first type of visually distinctive activity in the first video stream with the depiction of the first type of visually distinctive activity in the second video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods and systems for processing video, and in particular, to methods and systems for time-aligning video streams, i.e., methods and systems for determining the time offset between video streams. In the embodiments disclosed herein, the video streams depict live events, particularly sporting events such as sporting events at the university and professional levels.

Background Art

[0002] Live events, particularly sporting events such as those at the university and professional levels, continue to grow in popularity and revenue as individual universities and franchises earn billions of dollars in revenue each year. Understanding the time offset between video streams depicting live events is important (or even essential) for performing various types of processing on such video streams, for example, to process the video stream to generate an analysis of the event (e.g., in the case of a sporting event, an analysis of the game, team, and / or athlete), or to process the video stream to generate extended video content of the event, such as a single video showing highlights of the game from multiple angles.

Summary of the Invention

[0003] According to a first aspect of the present disclosure, identifying one or more depictions of a first type of visually distinctive activity in a first video stream depicting a sports event, identifying one or more depictions of the first type of visually distinctive activity in a second video stream depicting the sports event, and determining a time offset between the first video stream and the second video stream, wherein the determination of the time offset includes comparing one or more depictions of the first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream, a video processing method is provided.

[0004] In embodiments, the method includes processing at least one of the first and second video streams based on the time offset. In an example, the processing includes performing spatio-temporal pattern recognition based on the first video stream, the second video stream, and the time offset between the first video stream and the second video stream. Additionally or alternatively, the processing includes generating video content (e.g., extended video content) based on the first video stream, the second video stream, and the time offset between the first video stream and the second video stream.

[0005] In embodiments, the method includes receiving the first video stream from a first portable device including a camera. The first portable device can be, for example, a smartphone. The method may include receiving the second video stream from a second portable device including a camera. The second portable device can be, for example, a smartphone.

[0006] According to a second aspect of the present disclosure, a video processing system comprising a memory storing a plurality of computer-executable instructions and one or more processors executing the computer-executable instructions, wherein the computer-executable instructions cause the one or more processors to identify one or more depictions of a first type of visually distinctive activity in a first video stream depicting a sports event, identify one or more depictions of the first type of visually distinctive activity in a second video stream depicting the sports event, and determine a time offset between the first video stream and the second video stream, and the system determines the time offset by comparing at least one or more depictions of the first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream. A video processing system is provided.

[0007] In an embodiment, the system is configured to receive a first video stream from a first portable device comprising a camera. The first portable device can be, for example, a smartphone. The system may also be configured to receive a second video stream from a second portable device comprising a camera. The second portable device can be, for example, a smartphone.

[0008] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium including a set of computer-readable instructions, wherein the set of computer-readable instructions, when executed by one or more processors, causes the one or more processors to identify one or more depictions of a first type of visually distinctive activity in a first video stream depicting a sports event, identify one or more depictions of the first type of visually distinctive activity in a second video stream depicting the sports event, and determine a time offset between the first video stream and the second video stream, and the system determines the time offset by comparing at least one or more depictions of the first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream. A non-transitory computer-readable storage medium is provided.

[0009] In embodiments of the second and third aspects of the present disclosure, the computer-executable instructions further cause the one or more processors to process at least one of the first and second video streams based on the time offset. In an example, this processing includes performing spatio-temporal pattern recognition based on the first video stream, the second video stream, and the time offset between the first video stream and the second video stream. Additionally, or alternatively, this processing includes generating video content (e.g., extended video content) based on the first video stream, the second video stream, and the time offset between the first video stream and the second video stream.

[0010] In embodiments of the second and third aspects of the present disclosure, computer-executable instructions cause one or more processors to identify one or more depictions of a second type of visually distinctive activity in a first video stream and one or more depictions of a second type of visually distinctive activity in a second video stream, and the system determines a time offset by comparing, at least, one or more depictions of a second type of visually distinctive activity in the first video stream with one or more depictions of a second type of visually distinctive activity in the second video stream.

[0011] In the above aspects and embodiments, the video stream depicts a sports event, but in the methods and systems according to additional aspects, the event may be a non-sports live event such as a concert, comedy show, or play.

[0012] Additional features and advantages will become apparent from the following description written with reference to the accompanying drawings, which are given by way of example only.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7A

Figure 7B

Figure 8A

Figure 8B

Figure 9A

Figure 9B

Figure 10A

Figure 10B

Figure 11

[0014] Embodiments of the present application relate to the automatic alignment of video streams, particularly video streams depicting live events such as sports events.

[0015] First, refer to FIG. 1. FIG. 1 is a flowchart showing a video processing method 100 according to an exemplary embodiment. As shown, method 100 includes a step 101 of identifying one or more depictions of a first type of visually distinctive activity in a first video stream. In many of the examples presented herein, the first video stream depicts a sports event such as a soccer game, a basketball game, a tennis game, etc., but it is expected that the techniques described herein may be equally applicable to video streams depicting non-sports events.

[0016] As further shown in FIG. 1, method 100 further includes step 102 of identifying one or more depictions of the same type of visually distinctive activity in a second video stream depicting the same sports event.

[0017] As will be described later with respect to the examples shown in FIGS. 2-10B, method 100 of FIG. 1 may use various types of visually distinctive activities to determine the time offset between two video streams. Visually distinctive activities of the type described later with respect to FIGS. 2A-10B may be characterized, for example, as being visible from multiple positions and orientations of a sports event and having visually apparent changes that may be temporally localized, e.g., having visually apparent changes that are temporally localized because the appearance changes abruptly from one frame of the video stream to the next (e.g., due to changes in position, motion, and / or patterning), and / or being characterized as having an appearance that changes frequently throughout the progression of the sports event. Thus (or for another reason), depictions of such visually distinctive activities can be easily and / or robustly identified within two (or more than two) video streams of the same sports event.

[0018] Video streams of a sports event can be received from various sources. In particular, it is expected that method 100 of FIG. 1 can be used when the first and / or second video streams are generated by a portable device equipped with a camera, such as a smartphone or tablet device. Compared to video streams generated by conventional broadcast cameras, there does not appear to be a solution for aligning video streams generated by such devices. Nevertheless, method 100 of FIG. 1 can also be used when the first and / or second video streams are generated by a conventional broadcast camera instead.

[0019] Return to FIG. 1. It can be seen that method 100 further includes step 106 of determining a time offset between the first video stream and the second video stream. As shown in block 108, this step 106 includes comparing a depiction of a first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream.

[0020] In some examples, such as the examples described with respect to FIGS. 3A-6, this comparison may include comparing the intensity (e.g., average or median intensity) of some or all of the pixels corresponding to the depiction of the first type of visually distinctive activity in the first video stream with the intensity (e.g., average or median intensity) of some or all of the pixels corresponding to the depiction of the first type of visually distinctive activity in the second video stream. Alternatively (or in addition), as in the case of the examples described with respect to FIGS. 7A-10B, the motion and / or position of the depiction of the first type of visually distinctive activity in the first video stream can also be compared with the motion and / or position of the depiction of the first type of visually distinctive activity in the second video stream.

[0021] As further shown in FIG. 1, method 100 can include an additional video processing step 110 after step 106 of determining a time offset between a first video stream and a second video stream. In this additional video processing step 110, the thus determined time offset is utilized in the processing of one or both of the first and second video streams. The additional processing step 110 can include, for example, performing spatio-temporal pattern recognition based on the first video stream, the second video stream, and the time offset between the first video stream and the second video stream. In addition or alternatively, step 110 can include generating video content (e.g., extended video content) based on the first video stream, the second video stream, and the time offset between the first video stream and the second video stream. For example, video content that simultaneously shows two views of the same scene of a sports event (two views corresponding to the first and second video streams), and the videos of the two views of the scene are synchronized can be generated.

[0022] Next, attention is paid to FIGS. 2A to 10B. FIGS. 2A to 10B show various types of visually distinctive activities that can be identified in steps 101 and 102 of the video processing method 100 of FIG. 1 for determining the time offset between two video streams.

[0023] Alignment using a game clock Referring to FIGS. 2A and 2B, a first example illustrating a suitable type of visually distinctive activity is the change in time shown on the clock of a sports event. In a particular example, this clock can be a game clock, i.e., a clock that shows the elapsed time or remaining time of a game or a playing period of a game.

[0024] Figures 2A and 2B show video frames from two different video streams of a sports event, which in the particular example shown is a basketball game. As can be seen, even though the video streams correspond to significantly different viewpoints of the game, the game clock 200 is visible in both video streams. The game clock is, by design, intended to be visible from multiple positions and orientations of the sports event. The function of the game clock is to enable the audience to know the current time of the game. In addition, the game clock can be characterized in that the way it looks changes abruptly from one frame of the video stream to the next, such that the change in time on the game clock represents a very narrow time window. Further, the game clock originally changes its appearance frequently throughout the course of the sports event. For these reasons (and / or other reasons), using the change in time shown on the clock of a sports event as the first type of visually distinctive activity of method 100 of FIG. 1 may provide a robust and / or accurate alignment of the video streams.

[0025] In steps 101 and 102 of method 100 of FIG. 1, the depiction of the changing time on the clock can be identified, for example, by using an optical character recognition algorithm for each frame of the video stream. In a particular example, the optical character recognition algorithm can be applied to a particular portion of each frame that is expected to contain the clock. The portion of the frame that contains the clock can be determined, for example, using knowledge of the real-world shape, position, and orientation of the clock in the venue, as well as knowledge of the position and orientation of the camera providing the video stream (e.g., by calibrating the camera). Alternatively (or in addition), the portion of the frame that contains the clock can also be determined using a segmentation algorithm that utilizes a suitably trained neural network such as Mask R-CNN.

[0026] In step 106, comparing the depiction of the changing time on the clock in the first video stream with the depiction of the changing time on the clock in the second video stream can include, for example, comparing the frame numbers (or the times within each video stream) of each video stream at which one or more time transitions of the game clock occur (e.g., the transition from 12:00 to 12:01 or the transition from 1:41 to 1:42). Generally, performing such comparisons for multiple transitions of the game clock can give a more accurate estimate of the time offset between the video streams.

[0027] Alignment Using Camera Flashes A further example of a suitable kind of visually distinctive activity is the occurrence of camera flashes during a sports event. Considering that a very high-intensity light is emitted by the camera flash unit, the camera flash can be characterized as being visible from multiple positions and orientations of the sports event. In addition, considering the short duration of the camera flash (usually shorter than the time between two frames of the video stream) and the rapid increase in luminance generated by the camera flash unit, the appearance of the camera flash changes suddenly from one frame of the video stream to the next. Furthermore, the camera flash can be characterized as changing in appearance frequently throughout the progression of the sports event. For these reasons (and / or other reasons), using the camera flash as the first kind of visually distinctive activity in method 100 of FIG. 1 can provide a robust and / or accurate alignment of the video streams.

[0028] Next, refer to FIGS. 3A - 3D. FIGS. 3A - 3D show two video frames from each of two video streams of a sports event. As is apparent, in the particular example shown, this sports event is a basketball game. In indoor sports events such as basketball or ice hockey games, one may notice that camera flashes are visible. Thus, camera flashes can be a particularly suitable option for visually distinctive activities to identify when time-aligning video streams of indoor sports events using the method 100 of FIG. 1.

[0029] FIGS. 3A and 3B respectively show a first frame in which no camera flash is emitted and a second frame in which one or more camera flashes are emitted, and these frames are both from the same video stream and thus are taken from the same perspective. Similarly, FIGS. 3C and 3D also show the first and second frames from a second video stream, with no camera flash emitted in the first frame (FIG. 3C) and one or more camera flashes emitted in the second frame (FIG. 3D). As is apparent, the second video stream (shown in FIGS. 3C and 3D) is taken from a different perspective than the first video stream (shown in FIGS. 3A and 3B).

[0030] In steps 101 and 102 of method 100 of FIG. 1, the depiction of the camera flash can be identified in a given video stream by, for example, analyzing the pixel intensity (e.g., total or average pixel intensity) for some or all of the frames in the video stream. The inventors of the present invention have confirmed that the frames depicting the camera flash are significantly overexposed. Thus, frames having particularly high pixel intensities are likely to depict the camera flash. This is shown in FIGS. 4A and 4B, which are graphs showing the pixel intensities of the video streams shown in FIGS. 3A and 3B and the video streams shown in FIGS. 3C and 3D, respectively, over time for the respective frames. As is clear, each graph includes a peak of extremely short duration, which has a pixel intensity significantly higher than that of the frames immediately before or after. This peak in pixel intensity corresponds to the camera flash depicted in FIGS. 3B and 3D.

[0031] In step 106, comparing the depiction of the camera flash in the first video stream with the depiction of the camera flash in the second video stream can include, for example, comparing the frame numbers of the frames in each video stream in which a particularly short duration and a particularly large absolute value peak of pixel intensity occur. A simple technique that may be utilized in some examples is to assume that the flash is very short compared to the frame rate, such that each flash impinges on only one frame (or, if the flash is emitted during the "dead time" of the camera sensor, e.g., when the shutter is closed, does not impinge on any frame), and each flash very significantly increases the average image intensity for that frame only.

[0032] Other examples can utilize more complex techniques that have a known short duration with a flash, such that there is a potential for the flash to span two or more consecutive frames. Such techniques can further consider whether the camera has a global shutter or a rolling shutter, and use the camera specifications / information to determine when the sensor opens or closes. From this and the average frame intensity (in the case of a global shutter) or per-line intensity (in the case of a rolling shutter), this model can estimate the most likely flash start / end times.

[0033] Regardless of whether a simple or more complex technique is employed to determine the timing of each flash, performing a comparison of multiple camera flashes can sometimes provide a more accurate estimate of the time offset between video streams.

[0034] Alignment using an electronic display Further examples of visually distinctive activities suitable for aligning video streams are changes in images and / or patterns shown on one or more electronic displays of a sports event. In a specific example, this electronic display is a billboard, but it can also be a "large screen" that shows highlights or other video content to the spectators of a sports event.

[0035] An electronic display for a sports event is designed to be visible from multiple positions and orientations of the sports event. The electronic display is intended to be visible to most, if not all, of the spectators. In addition, the electronic display can be characterized in that the appearance changes abruptly from one frame of the video stream to the next. Further, the electronic display inherently changes its appearance frequently throughout the course of the sports event. For these reasons (and / or other reasons), using the changes in the images and / or patterns shown on one or more electronic displays of a sports event as the first type of visually distinctive activity of method 100 of FIG. 1 may provide a robust and / or accurate alignment of the video stream.

[0036] In steps 101 and 102 of method 100 of FIG. 1, knowledge of the real-world shape, position, and orientation of the electronic display, as well as knowledge of the real-world position and orientation of the camera providing the video stream, may be used to identify the depiction of the electronic display within the video stream. The electronic display for a sports event rarely moves given its large size and typically has a simple, regular shape (e.g., the electronic display is rectangular). Thus (or for another reason), knowledge of the shape, position, and orientation of the electronic display can be obtained and maintained relatively easily. Knowledge of the position and orientation of the camera providing the video stream can be obtained by calibrating the internal and external parameters of each camera. A wide range of camera calibration techniques can be used, including the technique disclosed in U.S. Patent No. 10,600,210B1, assigned to the assignee of the present invention, the disclosure of which is incorporated herein by reference.

[0037] In other examples, alternative techniques can be used to identify the depiction of the electronic display within the video stream, for example, using a segmentation algorithm that utilizes a suitably trained neural network such as Mask R-CNN.

[0038] In step 106 of method 100, comparing the depiction of the electronic display in the first video stream with the depiction of the electronic display in the second video stream can include, for example, comparing the change over time of the pixel intensity of some or all of the pixels identified as depicting the electronic display in the first video stream at step 101 with the change over time of the pixel intensity of some / all of the pixels identified as depicting the electronic display in the second video stream at step 102.

[0039] An example of such an approach is shown in FIGS. 5A-5D and FIG. 6. FIGS. 5A-5D show two video frames from each of two video streams of a sports event. In the particular example shown, the sports event is a soccer game. There are several billboards 150 at the end of the stadium. As is apparent, even if the video streams correspond to significantly different viewpoints of the game, the billboards 150 are visible both within the first video stream (shown in FIGS. 5A and 5B) and within the second video stream (shown in FIGS. 5C and 5D). FIG. 6 is a line graph showing the pixel intensity of a subset 510 of the pixels depicting the billboard over time.

[0040] In the example shown in FIGS. 5A and 5B, a subset 510 of pixels of the first video stream (shown in FIGS. 5A and 5B) corresponds to the same real-world location as a subset 510 of pixels of the second video stream (shown in FIGS. 5C and 5D). This can be achieved using knowledge of the real-world shape, position, and orientation of the billboard 150 discussed above and / or using image feature matching. Since the same part of the display is sampled from each video stream, such an approach may provide a more robust and / or accurate determination of the time offset. However, it is not essential that the two subsets 510 of pixels correspond to the same real-world location. This is because, for example, even if the sampled subset 510 of pixels corresponds to different parts of the display, large-scale transitions (e.g., transitions where the display goes black or changes between one advertisement and the next) are still distinguishable within the pixel intensity time series of the two video streams. Such large-scale transitions are shown in FIGS. 5A-5D by the hatching of the display panel 150 in FIGS. 5B and 5D. The hatching is not visible in FIGS. 5A and 5C.

[0041] Returning to FIG. 6, it can be seen that the first line 610 shows the pixel intensity of the video stream whose frames are shown in FIGS. 5A and 5B, and the second line 620 shows the pixel intensity of the video stream whose frames are shown in FIGS. 5C and 5D over time. To find a promising time offset between these two time series, the two time series can be compared using, for example, a cross-correlation function.

[0042] Alignment Using the Movement of a Ball A further example of a type of visually distinctive activity suitable for aligning video streams is the movement of a ball, pack, or similar sporting object during a sports event.

[0043] Similar to the other types of visually distinctive activities discussed above, the ball (or other sports object) can be seen from multiple positions and orientations during a sports event and (particularly when kicked, thrown, caught, etc. by a player) appears to change suddenly from one frame of a video stream to the next and is characterized by frequent changes in appearance throughout the progression of the sports event. For these reasons (and / or other reasons), using the movement of the ball (or other sports object) as the first type of visually distinctive activity of method 100 of FIG. 1 may provide a robust and / or accurate alignment of the video stream.

[0044] In steps 101 and 102 of method 100, the depiction of the ball within the video stream may be identified using various techniques. For example, various object detection algorithms can be utilized, including neural network algorithms such as Faster R-CNN or YOLO, and non-neural network algorithms such as SIFT.

[0045] In step 106, comparing the depiction of the movement of the ball in the first video stream with the depiction of the movement of the ball in the second video stream can include, for example, comparing the change in the movement of the ball depicted in the first video stream with the change in the movement of the ball depicted in the second video stream. As can be seen from FIGS. 7A and 7B showing two video frames from a video stream of a sports event, the movement of the ball can change suddenly between frames within the video stream. As is evident, this sports event is a soccer game. In FIG. 7A, the soccer ball is moving towards the player, and then the player kicks the ball, as a result of which the ball suddenly moves in the opposite direction as shown in FIG. 7B. Such a sudden change in movement is associated with a very narrow time window and tends to be visually differential, thus assisting in time alignment.

[0046] In addition to, or instead of, comparing the depiction of the movement of the ball in the first video stream with the depiction of the movement of the ball in the second video stream can include, for example, comparing the horizontal or vertical movement of the ball in the first video stream with the corresponding movement of the ball in the second video stream. In this context, the horizontal movement of the ball in the video stream means the left-right movement within the frame of the video stream, and the vertical movement means the up-down movement within the frame of the video stream.

[0047] If the cameras generating the first and second video streams are calibrated (such that the internal and external parameters of the cameras are known), it is possible to determine how the perspectives of the two video streams differ. (A wide range of camera calibration techniques can be used, including the techniques disclosed in U.S. Patent No. 10,600,210B1, assigned to the assignee of the present invention.) If the perspectives are relatively similar, appropriate results may be obtained simply by comparing the 2D horizontal movement in the first video stream (i.e., the left-right movement of the ball within the frame of the first video stream) with the 2D horizontal movement in the second video stream (i.e., the left-right movement of the ball within the frame of the second video stream).

[0048] A more general approach where the video streams do not need to have similar viewpoints is to transform the movement of the ball depicted in the second video stream into 3D movement and then determine the horizontal (or vertical) component of such 3D movement in the first video stream. This way, a like-for-like comparison of the movement can be performed. Transforming the 2D movement depicted in the video stream into 3D movement can be achieved in various ways. In one example, this 3D movement can be determined by triangulation of the ball using the second video stream in combination with a third video stream. The third video stream is time-aligned with the second video stream (i.e., the time offset between the second video stream and the third video stream is known). In addition, to assist in performing the triangulation, the cameras that generate the second and third video streams may be calibrated (such that the external and internal parameters of the cameras are known).

[0049] Figures 8A and 8B show an example where such an approach is applied to the video streams whose frames are shown in Figures 7A and 7B. Figure 8A shows the change in horizontal movement between one frame and the next frame of the video streams whose frames are shown in Figures 7A and 7B. As described above, the 2D movement in the second video stream is transformed into 3D movement using triangulation with a synchronized third video stream. Then, the horizontal component of such 3D movement in the first video stream (shown in Figures 7A and 7B) is determined. The inter-frame change of the components of the 3D movement thus determined is shown in Figure 8B. As can be seen, both Figures 8A and 8B show peaks with a very short duration. To find the promising time offset between these two time series, the two time series can be compared using, for example, the cross-correlation function.

[0050] The above approach converts 2D motion in the second video stream to 3D motion, but it should be noted that instead, the 2D motion in the first video stream can be converted to 3D motion. In fact, considering that method 100 of FIG. 1 treats the first video stream and the second video stream symmetrically, it is essentially arbitrary whether a given one of the two video streams is called the "first" video stream or the "second" video stream.

[0051] Although horizontal and vertical motion was described above, this was simply for simplicity, and it should also be noted that in step 106, 2D motion in other directions can also be compared.

[0052] Alignment using head position A further example of a type of visually distinctive activity suitable for aligning video streams is the change in the position of the head of a participant in a sports event (e.g., the change in the position of the head of an athlete and / or referee in a sports event). The head of the participant can be seen from multiple positions and orientations in the sports event and can be characterized as frequently changing in appearance throughout the progression of the sports event. For these reasons (and / or other reasons), using the change in the position of the head of the participant as the first type of visually distinctive activity in method 100 of FIG. 1 may provide robust and / or accurate alignment of the video streams.

[0053] In steps 101 and 102 of method 100, the depiction of the head of the participant within the video stream may be identified using various techniques. For example, various computer vision algorithms, such as various computer vision algorithms that utilize neural networks such as Faster R-CNN, YOLO, or OpenPose, or various computer vision algorithms that utilize feature detection algorithms such as SIFT or SURF, can be utilized.

[0054] In step 106 of method 100, comparing the change in the depiction of the position of the participant's head in the first video stream with the change in the depiction of the position of the participant's head in the second video stream can include, for example, converting the (2D) head positions depicted in some of the frames of the second video stream to 3D head positions and then reprojection of these 3D positions onto one of the frames of the first video stream. Converting the 2D head positions to 3D head positions can be performed in a similar manner as the method of converting the 2D movement of the ball to 3D movement of the ball described in the section "Alignment Using the Movement of the Ball" above.

[0055] Reprojecting the head position from the second video stream enables comparison of the head position under the same conditions. The frame having the closest match between the identified head position 910 and the reprojected head position 920 indicates the possible time offset between the two video streams. This is shown in FIGS. 9A and 9B, which show the same frame from the first video stream of a soccer game and thus the same identified head position 910. However, each of FIGS. 9A and 9B shows a reprojected head position 920 corresponding to a different frame from the second video stream. As is clear, the reprojected head position 920 shown in FIG. 9A matches much better than the reprojected head position 920 shown in FIG. 9B.

[0056] To provide further estimation of the time offset and thus a more accurate final determination of the time offset between the video streams, this process of converting the head position to 3D, reprojecting, and finding the closest matching frame from the first video stream can be repeated for additional frames from the first video stream.

[0057] Alignment Using Posture Still further examples of types of visually distinctive activities suitable for aligning video streams are changes in the postures of participants in a sports event (e.g., changes in the positions and orientations of the heads, limbs, and torso of athletes and / or referees in a sports event). The postures of the participants can be seen from multiple positions and orientations of the sports event and can be characterized as frequently changing in appearance over the course of the sports event. For these reasons (and / or other reasons), using changes in the postures of the participants as the first type of visually distinctive activity in method 100 of FIG. 1 may provide a robust and / or accurate alignment of the video stream.

[0058] In steps 101 and 102 of method 100, the depictions of the postures of the participants within the video stream may be identified using various techniques. For example, various computer vision algorithms may be utilized, such as various computer vision algorithms that utilize neural networks such as Mask R-CNN, OpenPose, AlphaPose, etc. For example, a computer vision algorithm can be used to identify the key points 1010 of each participant's body, such as the head, shoulders, elbows, wrists, hips, knees, and ankles of the participant 1020 in question, as shown in FIGS. 10A and 10B. The set of positions of these key points characterizes the posture of the participant.

[0059] Step 106 of method 100 may include sub-steps similar to the sub-steps described in the section "Alignment Using Head Position" above. Specifically, the 2D positions of the body key points in several frames of the second video stream can be converted to 3D positions and then re-projected onto the frames from the first video stream to find the most closely matching frame.

[0060] As is apparent by comparing FIG. 10B with FIG. 10A, the position of the pose tends to change more rapidly compared to the position of the participant's head. The head position moves only a small amount between the two frames shown in FIGS. 10A and 10B, respectively, while the positions of other key points such as the elbows and wrists move significantly. Therefore, aligning the video stream using the pose may provide relatively accurate results. Moreover, even when only one participant is visible in the video frame, the participant can obtain relatively accurate results because multiple key points that can be used for comparison are provided.

[0061] Combination To provide even more robust temporal alignment of the video streams, it may be expected that two or more of the above-described techniques may be combined. Referring to FIG. 1 again in this regard. As shown, method 100 includes an optional step 103 of identifying one or more depictions of a second different type of visually distinctive activity in a first video stream, and an optional step 104 of identifying one or more depictions of a second type of visually distinctive activity in a second video stream. Correspondingly, as shown in FIG. 1, determining the time offset 106 between the first video stream and the second video stream may include comparing 109 a depiction of a second type of visually distinctive activity in the first video stream with a depiction of a second type of visually distinctive activity in the second video stream. The second type of visually distinctive activity may be, for example, any one of the examples illustrated above with respect to FIGS. 2A-10B.

[0062] In an example, step 109 of comparing a depiction of a second type of visually distinctive activity in a first video stream with a depiction of the second type of visually distinctive activity in a second video stream may provide an estimate of one time offset, and step 108 of comparing a depiction of a first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream may provide another estimate of the time offset. In such an example, at step 106, the time offset can be determined based on both of these estimates.

[0063] FIG. 1 shows steps 101 and 102 as being performed before optional steps 103 and 104, but it should be noted that in other embodiments, steps 101 and 102 can be performed after optional steps 103 and 104, or simultaneously with steps 103 and 104. In fact, steps 101-104 may be performed in any suitable order.

[0064] Next, refer to FIG. 11. FIG. 11 is a schematic diagram of a video processing system according to an exemplary additional embodiment. As shown, video processing system 1100 includes a memory 1110 and a processor 1120. Memory 1110 stores a plurality of computer-executable instructions 1112-1118, and when the plurality of computer-executable instructions 1112-1118 are executed by processor 1120, cause the processor to perform the following:

[0065] Identifying one or more depictions of a first type of visually distinctive activity in a first video stream depicting a sports event 1112,

[0066] Identifying one or more depictions of the first type of visually distinctive activity in a second video stream depicting a sports event 1114, and

[0067] Determining a time offset between a first video stream and a second video stream 1116.

[0068] Determining the time offset includes comparing one or more depictions of a first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream 1118.

[0069] Any of the types of visually distinctive activities described above with respect to FIGS. 2A - 10B can be utilized by video processing system 1100. Moreover, the techniques for identifying and comparing depictions of visually distinctive activities described above with respect to FIGS. 2A - 10B can be implemented within video processing system 1100.

[0070] Memory 1110 may be of any suitable type, such as including one or more of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), and flash memory.

[0071] Although only one processor 1120 is shown in FIG. 11, it is understood that system 1100 can include multiple processors. In particular, when system 1100 includes multiple processors, (without limitation), system 1100 may include a GPU and / or a neural network accelerator.

[0072] The video stream of a sports event can be received from various sources. In particular, when the first and / or second video streams are generated by a portable device 1140, 1150 equipped with a camera such as a smartphone or a tablet device, it is expected that the system 1100 of FIG. 11 can be used. Compared with the video stream generated by a conventional broadcast camera, the inventor of the present invention believes that there is no solution for aligning the video stream generated by such a device. However, when the first and / or second video streams are generated by a conventional broadcast camera, the method 100 of FIG. 1 can also be equally used.

[0073] The video stream can be received via any suitable communication system. For example, the video stream can be transmitted through the Internet, through an intranet (including, for example, the video processing system 1100), through a cellular network (especially when the video stream is generated by the portable devices 1140, 1150), or through the bus of the video processing system 1100.

[0074] It should be noted that the video streams of the embodiments described above with respect to FIGS. 1-11 depict sports events, but many embodiments, in particular, the method 100 described above with respect to FIG. 1 and the system 1100 described above with respect to FIG. 11, may be adapted to process video streams depicting other types of live events such as concerts, comedy shows, or plays.

[0075] Any feature described in relation to any one embodiment may be used alone or in combination with any other feature described, and any feature described in relation to any one embodiment may be used in combination with one or more features of other embodiments, or in any combination of other embodiments whatsoever. Furthermore, equivalents and modifications not described above may be used, without departing from the scope of the invention as defined by the appended claims.

Claims

1. identifying one or more depictions of a first type of visually distinctive activity in a first video stream depicting a sports event; identifying one or more depictions of the first type of visually distinctive activity in a second video stream depicting the sports event; determining a time offset between the first video stream and the second video stream The method comprising the steps of: determining the time offset includes comparing one or more depictions of the first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream, a video processing method.

2. The method of claim 1, wherein the first type of visually distinctive activity corresponds to a change in time indicated on a clock of the sports event.

3. The method of claim 1, wherein the first type of visually distinctive activity corresponds to a camera flash emitted during the sports event.

4. The method of claim 1, wherein the first type of visually distinctive activity corresponds to a change in an image shown on one or more electronic displays of the sports event.

5. The method of claim 4, wherein the one or more electronic displays include billboards.

6. The method of claim 1, wherein the first type of visually distinctive activity corresponds to the movement of a ball or pack.

7. The method of claim 1, wherein the first type of visually distinctive activity corresponds to the movement of one or more participants in the sports event.

8. The method of claim 7, wherein the first type of visually distinctive activity corresponds to the movement of the head of the one or more participants in the sports event.

9. The method of claim 7, wherein the first type of visually distinctive activity corresponds to a change in the respective posture of the one or more participants in the sports event.

10. The method according to any one of claims 1 to 9, comprising receiving the first video stream from a first portable device equipped with a camera.

11. The method of claim 10, wherein the first portable device is a smartphone.

12. identifying one or more depictions of a second type of visually distinctive activity in the first video stream; identifying one or more depictions of the second type of visually distinctive activity in the second video stream including, wherein determining the time offset includes comparing the one or more depictions of the second type of visually distinctive activity in the first video stream with the one or more depictions of the second type of visually distinctive activity in the second video stream, the method according to any one of claims 1 to 11. **Claim 13** a memory storing a plurality of computer-executable instructions; one or more processors for executing the computer-executable instructions comprising, wherein the computer-executable instructions identifying one or more depictions of a first type of visually distinctive activity in a first video stream depicting a sports event; identifying one or more depictions of the first type of visually distinctive activity in a second video stream depicting the sports event; determining a time offset between the first video stream and the second video stream causing the one or more processors to execute, the system determining the time offset by at least comparing the one or more depictions of the first type of visually distinctive activity in the first video stream with the one or more depictions of the first type of visually distinctive activity in the second video stream, a video processing system. **Claim 14** The system according to claim 13, wherein the first type of visually distinctive activity corresponds to a change in time indicated by a clock of the sports event. **Claim 15** The system according to claim 13, wherein the first type of visually distinctive activity corresponds to a camera flash emitted during the sports event. **Claim 16** The system according to claim 13, wherein the first type of visually distinctive activity corresponds to a change in an image shown on one or more electronic displays of the sports event. **Claim 17** The system according to claim 16, wherein the one or more electronic displays include billboards. **Claim 18** The system according to claim 13, wherein the first type of visually distinctive activity corresponds to the movement of a ball or pack during play of the sports event. **Claim 19** The system according to claim 13, wherein the visually distinctive activity of the first type corresponds to the movement of one or more participants in the sports event.

20. The system according to claim 19, wherein the visually distinctive activity of the first type corresponds to the movement of the head of the one or more participants in the sports event.

21. The system according to claim 19, wherein the visually distinctive activity of the first type corresponds to a change in the respective posture of the one or more participants in the sports event.

22. The system according to any one of claims 13 to 21, wherein the system is configured to receive the first video stream from a first portable device comprising a camera.

23. The system according to claim 22, wherein the first portable device is a smartphone.

24. The computer-executable instructions identifying one or more depictions of a second type of visually distinctive activity in the first video stream identifying one or more depictions of the second type of visually distinctive activity in the second video stream and causing the one or more processors to execute the system to determine the time offset by comparing at least the one or more depictions of the second type of visually distinctive activity in the first video stream with the one or more depictions of the second type of visually distinctive activity in the second video stream. The system according to any one of claims 13 to 23.

25. A non-transitory computer-readable storage medium comprising a set of computer-readable instructions, the set of computer-readable instructions when executed by one or more processors identifying one or more depictions of a first type of visually distinctive activity in a first video stream depicting a sports event identifying one or more depictions of the first type of visually distinctive activity in a second video stream depicting the sports event determining a time offset between the first video stream and the second video stream Causing the one or more processors to execute, wherein the system determines the time offset by comparing at least one or more depictions of the first type of visually distinctive activity in the first video stream with one or more depictions of the first type of visually distinctive activity in the second video stream. A non-transitory computer-readable storage medium. Claim 26 The non-transitory computer-readable storage medium according to claim 25, wherein the first type of visually distinctive activity corresponds to a change in time indicated by a clock of the sports event. Claim 27 The non-transitory computer-readable storage medium according to claim 25, wherein the first type of visually distinctive activity corresponds to a camera flash emitted during the sports event. Claim 28 The non-transitory computer-readable storage medium according to claim 25, wherein the first type of visually distinctive activity corresponds to a change in an image shown on one or more electronic displays of the sports event. Claim 29 The non-transitory computer-readable storage medium according to claim 28, wherein the one or more electronic displays include billboards. Claim 30 The non-transitory computer-readable storage medium according to claim 25, wherein the first type of visually distinctive activity corresponds to the movement of a ball or pack during play of the sports event. Claim 31 The non-transitory computer-readable storage medium according to claim 25, wherein the first type of visually distinctive activity corresponds to the movement of one or more participants of the sports event. Claim 32 The non-transitory computer-readable storage medium according to claim 31, wherein the first type of visually distinctive activity corresponds to the movement of the head of the one or more participants of the sports event. Claim 33 The non-transitory computer-readable storage medium according to claim 31, wherein the first type of visually distinctive activity corresponds to a change in the respective posture of the one or more participants of the sports event. Claim 34 The non-transitory computer-readable storage medium according to any one of claims 25 to 33, wherein the system is configured to receive the first video stream from a first portable device including a camera. Claim 35 The non-transitory computer-readable storage medium according to claim 34, wherein the first mobile device is a smartphone. **Claim 36** The computer-executable instructions identifying one or more depictions of a second type of visually distinctive activity in the first video stream; and identifying one or more depictions of the second type of visually distinctive activity in the second video stream are executed by the one or more processors, the one or more processors determining the time offset by comparing at least the one or more depictions of the second type of visually distinctive activity in the first video stream with the one or more depictions of the second type of visually distinctive activity in the second video stream. The non-transitory computer-readable storage medium according to any one of claims 25 to 35.

Citation Information

Patent Citations

  • Server-side support for seamless rewind and playback of video streaming.

    JP2012517160A