Sports image processing and object tracking
By extracting and transmitting relevant sports event metadata from a local device to a remote processor, the system addresses bandwidth limitations and delays in cloud-based sports assistive technology, ensuring timely decision-making.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2026-03-19
AI Technical Summary
The use of cloud-based processing for sports assistive technology is hindered by network bandwidth limitations and potential delays, especially in fast-moving sports with high-resolution image capture, which is critical for real-time decision-making.
A local device extracts relevant image information, such as positions of players, officials, and the ball, and transmits this metadata over a network to a remote device for processing, reducing the data volume and minimizing delays.
This approach reduces network bandwidth requirements and minimizes delays, making cloud-based processing more viable for real-time sports assistance by focusing on essential image data transmission.
Smart Images

Figure GB2025051893_19032026_PF_FP_ABST
Abstract
Description
[0001] SPORTS IMAGE PROCESSING AND OBJECT TRACKING
[0002] BACKGROUND
[0003] Field of the Disclosure
[0004] The present disclosure relates to sports image processing and object tracking.
[0005] Description of the Related Art
[0006] The “background” description provided is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in the background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly or impliedly admitted as prior art against the present disclosure.
[0007] Various types of sports assistive technology are available for spectating and / or officiating sports events. Many of these use captured images of a sports event which are then processed to determine particular outputs. For example, images of a soccer match may be captured and processed to generate an output indicating whether a player is offside, whether an illegal handball has occurred or whether the ball has crossed the goal line (indicating a goal has been scored). In another example, images of a tennis match may be captured and processed to generate an output indicating whether the ball is outside the boundary lines of the tennis court.
[0008] Processing these images involves the accurate tracking of objects (e.g. player(s), official(s), a ball, etc.) and is a processor-intensive task. This is particularly the case for fast-moving sports for which many high resolution images are captured each second, often from multiple cameras. Traditionally, this has required specialist, high performance image processing hardware to be located onsite. However, there now is a desire to take advantage of the power and efficiency of cloud-based processing.
[0009] Cloud-based processing allows a client party (e.g. the party responsible for providing the sports assistive technology to the sports event) to remotely borrow the computing power of a serving party (e.g. a specialist in providing high performance hardware). It involves the client party sending input data (e.g. captured images) to be processed to computing device(s) of the serving party and receiving output data (e.g. data indicating whether a goal has been scored) from those computing device(s) over a network. This alleviates the need for the client party to provide, maintain and regularly upgrade their own hardware.
[0010] However, due to the often limited network bandwidth available at sports event locations and the large amount of data associated with captured images that must be processed, there is a risk of unacceptable delays when using cloud-based processing for sports assistive technology. This can be a significant problem, especially when using the technology for sports officiating, where both sport participants and spectators expect decisions made using sports assistive technology to be made with minimal delay.
[0011] There is therefore a desire to address this.
[0012] SUMMARY
[0013] The present disclosure is defined by the claims.
[0014] BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Non-limiting embodiments and advantages of the present disclosure are explained with reference to the following detailed description taken in conjunction with the accompanying drawings, wherein:
[0016] Fig. 1 schematically shows an example system;
[0017] Figs. 2A and 2B schematically show example data processing apparatuses;
[0018] Fig. 3 shows an example captured image;
[0019] Fig. 4 shows an example processed version of the captured image;
[0020] Fig. 5 show example steps implemented by the system; and
[0021] Figs. 6A and 6B show example methods.
[0022] Like reference numerals designate identical or corresponding parts throughout the drawings.
[0023] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Fig. 1 shows an example system 100. The system 100 comprises one or more cameras (in this example, four cameras 101A-D), a first, local, data processing apparatus / device 102, a network 103 and a second, remote, data processing apparatus / device 104.
[0025] The cameras 101A-D are each configured to capture images (e.g. successive video frames) of a sports event.
[0026] The captured images are input to the local device 102 (sports image processing apparatus). The local device 102 is configured to extract information from each captured image for generating sports assistance information. For example, if the sports assistive technology comprises tracking of the position of objects (e.g. objects corresponding to one or more predetermined and / or detectable object types, such as sports player(s), sport officials (s) and / or a ball), the extracted information comprises portion(s) of the image comprising I corresponding to the object(s) to be tracked. The extracted information may also include metadata. The metadata is defined for each captured image frame and comprises, for example, a timestamp of the image frame (or another type of frame identifier allowing the order of successively received frames to be ascertained). The metadata may also comprise other information such as information indicating pixel position(s) of each of extracted image portion(s) of each image frame. The extracted information is transmitted to the remote device 104 (sports object tracking apparatus) over the network 103. The network 103 is a data communications network such as the internet.
[0027] The remote device 104 is configured to process the extracted information to generate, as an output, the sports assistance information. For example, if the sports assistive technology comprises tracking, the generated sports assistance information comprises a position with respect to a predefined sports playing area (e.g. a soccer pitch or tennis court) of the player(s), official(s) and / or ball being tracked. In an example, the position of a tracked object is defined for each image capture time (e.g. with multiple cameras being temporally synchronised to capture corresponding images at substantially the same time) in a three-dimensional virtual space calibrated to the dimensions of the sports playing area. In an example, the generated sports assistance information is returned, over the network 103, to the local device 102. It may also be transmitted to one or more other devices (not shown).
[0028] In an example, the local device 102 is shown as a separate entity which processes images from each of the cameras 101A-D. In this case, the local device 102 is connected to each camera 101A-D via dedicated high-speed wired connection, for example. However, the local device 102 may instead be comprised within one or more of the cameras 101A-D. There may also be more than one local device 102. In particular, each camera may comprise a respective local device 102 for processing images captured by that camera.
[0029] The extracted information obtained by the local device 102 is information necessary for generating the desired sports assistance information rather than data representing full captured images. For instance, for tracking, the extracted information comprises only the image portions of the player(s), official(s) and / or ball rather than the entire image frame. The required network bandwidth for transmitting the extracted information is thus less than that required for transmitting the full captured images. This enables the remote device 104 to be hosted remotely with respect to the location of the cameras 101A-D and local device 102 (e.g. as a cloud-based service) while alleviating the risk of unacceptable network delays in transmission of the extracted information. Subsequent delays to the generation and transmission of the sports assistance information by the remote device 104 using the extracted information are thus also alleviated. This makes the use of cloud-based processing more technically suited for the implementation of sports assistance technology.
[0030] Figs. 2A and 2B show example components of the local device 102 and remote device 104, respectively.
[0031] Local device 102 comprises a processor 201 for executing electronic instructions, a memory 202 (e.g. volatile memory) for storing the electronic instructions to be executed and electronic input and output information associated with the electronic instructions, a storage medium 203 (e.g. non-volatile memory) for long term (persistent) storage of information, a camera interface 204 for receiving image data captured by an image sensor (e.g. Complementary Metal-Oxide- Semiconductor, CMOS, sensor) of one or more of the cameras 101A-D, a network interface 205 for sending information to and / or receiving information from one or more other apparatuses (e.g. remote device 104) over the network 103 and a user interface 206 (e.g. a touch screen, a nontouch screen, buttons, a keyboard and / or a mouse) for receiving commands from and / or outputting information to a user. Each of the processor 201 , memory 202, storage medium 203, camera interface 204, network interface 205 and user interface 206 are implemented using appropriate circuitry, for example. The processor 201 controls the operation of each of the memory 202, storage medium 203, camera interface 204, network interface 205 and user interface 206.
[0032] Optionally, the local device 102 may comprise one or more further interfaces (not shown) for receiving additional captured data, such as data captured from stereoscopic sensors (e.g. two cameras separated by a predetermined distance to determine depth), laser dot projection data (to help increase the accuracy of camera pose perception), time-of-flight (TOF) sensors (which, for example, use perception of phase shifted light pulses to determine depth and motion), inertial measurement unit (IMU) data and / or global navigation satellite system (GNSS) data. The IMU and / or GNSS data may be captured from sensor(s) worn by participant(s) of the sports event and / or comprised within a ball, for example.
[0033] Remote device 104 comprises a processor 207 for executing electronic instructions, a memory 208 (e.g. volatile memory) for storing the electronic instructions to be executed and electronic input and output information associated with the electronic instructions, a storage medium 209 (e.g. non-volatile memory) for long term (persistent) storage of information, a network interface 210 for sending information to and / or receiving information from one or more other apparatuses (e.g. local device 102) over the network 103 and a user interface 211 (e.g. a touch screen, a nontouch screen, buttons, a keyboard and / or a mouse) for receiving commands from and / or outputting information (e.g. the generated sports assistance information) to a user. Each of the processor 207, memory 208, storage medium 209, network interface 210 and user interface 211 are implemented using appropriate circuitry, for example, memory 208, storage medium 209, network interface 210 and user interface 211.
[0034] Fig. 3 shows an example original image 301 captured by one of the cameras 101A-D and received by the camera interface 204 of the local device 102. The image 301 is of a soccer match and comprises both information relevant for generating sports assistance information and irrelevant for generating sports assistance information.
[0035] In this example, the desired sports assistance information comprises a position of each soccer player 303 and the referee 307, a position of the ball 306, a position of the pitch lines 304 defining the soccer pitch and a position of the goal 305. For ease of explanation, in this example, the entire structure of the goal 305 is considered (including the frame of the goal and the net). However, alternatively, only the goal line (marked as part of the pitch boundary) and / or the vertical posts and / or horizontal crossbar of the goal 305 may be considered. The relevant information in this case therefore comprises each portion of the image 301 corresponding to one of the soccer players 303 and referee 307, the portion of the image corresponding to the ball 305, each portion of the image corresponding to one of the pitch lines 304 and the portion of the image corresponding to the goal 305. The irrelevant information comprises all other features of the image, such as the grass 302 of the soccer pitch and any crowd or stadium detail (not shown in image 301).
[0036] Fig. 4 shows a processed version 301 ’ of image 301 . The processed image 301 ’ is then generated by the processor 201 of the local device 102. Here, all relevant information has been retained (by retaining the pixel values in the captured image 301 of the image portion(s) containing relevant information) while all remaining, irrelevant information has been removed. In this case, portion(s) of the image comprising irrelevant information have been coloured black (e.g. by replacing the pixel value of each pixel in these portion(s) with zero). The effect is that the relevant information has been extracted from the original image 301 and is therefore extracted information. The extracted information can thus be transmitted to the remote device 104 over the network.
[0037] Any suitable image processing technique(s) may be implemented by the processor 201 to detect and extract the relevant information. In one example, edge detection (using any suitable known edge detection technique) is used to detect the pitch lines 304 in the image 301 . Instance and / or semantic segmentation (using any suitable known instance and / or semantic segmentation technique) is then used to detect the players 303 and referee 307, ball 306 and goal 305 in the image 301. The image segments corresponding to the detected pitch lines 304, players 303 and referee 307, ball 306 and goal 305 are then retained as extracted information while all remaining parts of the image are discarded. For example, the parts of the image to be discarded are overwritten with a predetermined pixel value (such as a zero pixel value so they are coloured black).
[0038] There may be more than one predetermined pixel value used to override parts of the image which are to be discarded. The predetermined pixel value used for each discarded part may then depend on a characteristic of the discarded part. For instance, a discarded part corresponding to grass may be overridden with a first predetermined pixel value (e.g. zero) whereas a discarded part corresponding to the crowd may be overridden with a second, different, predetermined pixel value (e.g. one). This allows, for instance, basic information about different discarded image parts to be ascertained by the remote device 104.
[0039] The type of object(s) detectable by the processor 201 (e.g. the image segmentation model used) may depend on the sport. For instance, for soccer, a “soccer” image segmentation model may be used which has been trained on past soccer image data (to improve the accuracy of player, referee and ball detection, for example). On the other hand, for motor racing, a “motor racing” image segmentation model may be used which has been trained on past motor racing image data (to improve the accuracy of vehicle and vehicle part detection, for example). It may also depend on whether or not a particular incident is detected as occurring during the sports event. For instance, if an altercation between players is detected as an incident, a facial expression detection model may be used to determine which player(s) have an “angry” or “sad” facial expression (and thus player face(s) are detected as additional object types). Such player(s) may then be determined as higher priority objects for tracking (and image portion(s) corresponding to those player(s) thus transmitted at shorter periodic time intervals, for example).
[0040] In a first example, the extracted information is transmitted to the remote device 104 by compressing and transmitting the entire processed image 301 ’. For instance, any suitable known compression technique (such as conversion of the processed image 301 ’ to a Joint Photographic Experts Group, JPEG, format) which reduces the data size of an image by exploiting redundancies in the image may be used. Use of such a compression technique on the processed image 301 ’ will result in a reduced data size compared to use of the same compression technique on original image 301 due to the grass 302 in the original image being replaced with zero value pixels.
[0041] In a second example, the extracted information is transmitted to the remote device 104 by transmitting only the portions of the processed image 301 ’ corresponding to the detected pitch lines 304, players 303 and referee 307, ball 306 and goal 305 (that is, only the portions of the processed image 301 ’ corresponding to the retained relevant information). For example, only the pixel values of the pixels of the image segments corresponding to the detected pitch lines 304, players 303 and referee 307, ball 306 and goal 305 may be transmitted (and, optionally, may be compressed using a suitable compression technique such as JPEG or the like). The amount of data which must be transmitted is thus reduced compared to transmitting the full original image 301.
[0042] In the second example, metadata indicating a pixel position (e.g. lower-most, left-most pixel position) of each segment of the original image 301 (or processed image 301 ’, which has the same number of pixels as the original image 301 in this example, even though some of them are converted to zero value pixels) is transmitted with the pixel values to enable the remote device 104 to determine the location of each segment (and therefore the location of the relevant information) in the original image 301. This allows the remote device 104 to then determine the position of each of the pitch lines 304, players 303 and referee 307, ball 306 and goal 305 on the pitch. The metadata also has only a small data size, thereby retaining the reduction in the amount of data transmitted compared to transmitting the original image 301. In an example, a predetermined calibration of the cameras 101A-D mapping the pixel positions in the images captured by each camera with a corresponding location on the soccer pitch allows tracking by the remote device 104. The calibration and tracking may be carried out using any known suitable technique(s), such as those available from Hawk-Eye Innovations ®. Such known calibration and tracking techniques are not discussed in detail here.
[0043] In the second example, to transmit position information even more efficiently, extracted image segments may be grouped together. For example, a first segment and second segment may be grouped together when at least one pixel of the first segment is directly adjacent a pixel of the second segment. A pixel position of the entire group (e.g. lower-most, left-most pixel position) in the original image 301 may then be transmitted. In an example, a box is positioned around the group of segments (e.g. the smallest box in which all segments in the group will fit). The box (containing the grouped image segments) is then transmitted as an image portion togetherwith a pixel position of the box (e.g. the lower-most, left-most pixel position of the box) in the original image 301 . The group of segments is thus treated as a single object and the box containing the segments is treated as the corresponding image portion which is transmitted.
[0044] In both the first and second examples, data representative of one or more extracted portions of the captured image 301 and position information of each extracted portion in the captured image is thus transmitted. In particular, in the first example, the extracted image portion(s) remain in their original pixel positions in the transmitted processed image 301 ’ (and the processed image 301 ’ is itself therefore indicative of the position of each extracted portion). In the second example, the pixel position of each extracted image portion(s) is indicated by the transmitted metadata. In an example, the local device 102 may transmit data according to the technique of the first example in a first mode and according to the technique of the second example in a second mode.
[0045] The way the extracted information is extracted in the example of Figs. 3 and 4 is only an example and it will be appreciated any suitable technique which enables only the relevant portions of an image to be extracted by the local device 102 and transmitted to the remote device may be used 104. For example, instead of, or in addition to, using image segmentation and edge detection, green-coloured pixels (e.g. those a particular range in YCbCr or RGB space corresponding to green shades) may be detected and replaced with zero value pixels to remote information associated with the grass 302.
[0046] The local device 102 may also be configured to use different extraction techniques (e.g. different machine learning (ML) models such as different instance I semantic segmentation models) for different sports and / or different lighting and / or environmental conditions. For instance, for tennis, a first semantic segmentation ML model trained using images of previous tennis matches may be used whereas, for soccer, a second semantic segmentation ML model trained using images of previous soccer matches may be used. Different ML models (different ML models being defined by different ML model architecture and / or model parameters, for example) may also be used depending on whether the current sports event is occurring during the day or at night, whether it is sunny or raining, etc. Data representative of each trained model may be stored in the storage medium 203 of the local device 102 in advance and selected (e.g. by a human operator or automatically depending on current weather data obtained via the network 103, for example) prior to the start of the sports event.
[0047] The extracted information, the way the extracted information is extracted and / or the way the extracted information is transmitted (which may be collectively referred to as the extracted information configuration) may also be adjusted depending on a number of factors. This adjustment may be configured in advance of the sports event (so the extracted information configuration remains the same throughout the transmission of extracted information from all images captured during the sports event) or during the sports event (so the transmitted extracted information of a first portion of images captured during the sports event has a different extracted information configuration to that of a second, later, portion of images captured during the sports event). For adjustment during the sports event, a signal (e.g. flag or code) may be transmitted with the extracted information (e.g. as part of metadata) from the local device 102 to the remote device 104 to indicate the extracted information configuration. This allows the remote device 104 to use the extracted information correctly in performing tracking or the like.
[0048] For instance, a first signal (e.g. setting of a flag bit to “0”) may indicate the compression and transmission of the entire processed image 301 ’. The remote device 104 thus knows it is able to apply tracking straight away based on the received image data. This is because the received image data represents a full image (e.g. compressed processed image 301 ’) in which the position of each pixel of each extracted portion of the original image 301 has been retained.
[0049] On the other hand, a second signal (e.g. setting the flag bit to “1”) may indicate only the extracted portions of the image are being transmitted together with respective pixel position information of each extracted portion. The remote device 104 thus knows it must first reconstruct the processed image 301 ’ from this information (e.g. by starting with an image only containing zero value pixels and then replacing those pixels with the received extracted image portions at their respective pixel positions) before applying tracking.
[0050] One example factor for adjusting the extracted information configuration is the type of sports event itself. For instance, for a sports event in which several participants are likely to be within each captured image (e.g. soccer), it may be more computationally efficient to compress and transmit the entire processed image 301 ’ rather than to separately transmit each extracted portion of the image (together with the respective pixel position information of each extracted portion). On the other hand, for a sports event with only one or two participants in each captured image (e.g. singles tennis), it may be more computationally efficient to transmit only the extracted portions of the image (together with the respective pixel position information of each extracted portion).
[0051] This is because, for example, the extracted portions of an image of a sport with only a small number of participants are expected to be smaller and / or fewer in number. For instance, in singles tennis, only extracted portion(s) corresponding to one or both players are required and these portions will together typically correspond to only smaller proportion of the pixels of the original captured image. On the other hand, the extracted portions of an image of a sport with a larger number of participants are expected to be larger and / or greater in number. For instance, in soccer, extracted portion(s) corresponding to up to all 22 players on the pitch (plus, potentially, the referee and / or line official(s)) are required and these portions will together thus correspond to a greater proportion of the pixels of the original captured image.
[0052] In an example, the extracted information configuration may change from the transmission of compressed full processed images 301 ’ to the transmission of only extracted portion(s) of each processed image 301’ depending on the proportion of pixels of the captured I processed image corresponding to those extracted portions. This may occur on an image-by-image basis. For instance, a threshold may be set (e.g. 10%, 15% or 20%) such that, if the proportion of pixels of the image corresponding to extracted portion(s) of the image exceeds this threshold, the entire processed image is compressed and transmitted. On the other hand, if the proportion of pixels does not exceed this threshold, only the extracted portion(s) (together with respective pixel position information of each portion) are transmitted. The extracted information configuration is thus dynamically adjusted depending on the content of each image to improve computational efficiency.
[0053] Another example factor is the speed of movement of object(s) during the sports event. For example, in a soccer match, for successively captured image frames of a given camera, there may be one player in the image moving more quickly (e.g. running with the ball) whereas another player in the image (e.g. a player behind the player running with the ball) moves less quickly. In this case, an extracted image portion associated with the quickly moving player is deemed higher priority than an extracted image portion associated with the slowly moving player.
[0054] In an example, when extracted portions of the processed image 301 ’ (rather than the entire processed image 301 ’) are transmitted, the extracted portion corresponding to the faster moving player may be transmitted more frequently than the extracted portion corresponding to the slower moving player. This is applicable to different players at the same time (e.g. one faster moving and one slower moving in the same set of successively captured image frames) or to the same player at different times (e.g. a player moving faster in one set of successively captured image frames but slower in another set of successively captured image frames captured at a different time).
[0055] For instance, the extracted portion corresponding to the faster moving player may be transmitted for every image successively captured by the camera. On the other hand, the extracted portion corresponding to the slower moving player may be transmitted for only a subset of the images successively captured by the camera (e.g. transmitted only every 2nd, 3rdor 4thimage or, more generally, every nthimage where n < 1). Thus, extracted portions corresponding to different objects may be transmitted at different periodic intervals (e.g. every frame for a shorter periodic interval or every nthframe where n > 1 for successively longer periodic intervals). This further helps reduce the transmission data rate while ensuring important information (e.g. the extracted portion corresponding to the fast moving player who, because of their faster movement, is more likely to be playing an important part in the match at the current time) is provided to the remote device 104.
[0056] In an example, the local device 102 may distinguish a faster moving extracted image portion from a slower moving extracted image portion using any suitable known technique. For example, motion vectors of each extracted image portion may be determined between successive image frames. An extracted image portion is considered to correspond to a faster moving object if a magnitude of a motion vector associated with that extracted image portion exceeds a predetermined threshold. On the other hand, if the threshold is not exceeded, the extracted image portion is considered to correspond to a slower moving object.
[0057] In an example, a plurality of thresholds are defined to enable each extracted image portion in a set of captured to be transmitted to the remote device 104 at an appropriate timing. For instance, two thresholds (e.g. a first, lower, threshold and a second, higher, threshold) may be used to classify the inter-frame speed of each extracted image portion as slow (e.g. not exceeding the first threshold), moderate (e.g. exceeding the first threshold but not exceeding the second threshold) or fast (exceeding the second threshold). Image portions classified as fast may then be transmitted for every successively captured frame while image portions classified as moderate are transmitted every nthframe (where n > 1) and image portions classified as slow are transmitted every mthframe (where m > n). For tracking, the remote device 104 then uses the most recently received version of each image portion. This allows an improved balance between tracking accuracy and data transmission reduction.
[0058] Another example factor is whether or not any incident of interest (e.g. from a predetermined list I set of detectable incident(s)) is currently occurring during the sports event. If an incident of interest is occurring, then the amount of extracted information which is transmitted by the local device 102 may be increased to provide improved tracking accuracy at the remote device 104. On the other hand, if no event of interest is occurring, then the amount of extracted information which is transmitted may be reduced to alleviate the bandwidth requirement.
[0059] An example of an incident of interest may be an approach of the ball 306 to certain other object(s) (e.g. one of the pitch lines 304 or the goal 305) detected in the captured image 301. For instance, the local device 102 may be configured to determine a start of an incident of interest if the portion of the image corresponding to the ball 306 (ball image portion) is within a predetermined distance (e.g. predetermined number of pixels) of the portion of the image corresponding to one of the pitch lines 304 (pitch line image portion) or the goal 305 (goal image portion). An end of the incident of interest is then determined once the ball image portion is no longer within the predetermined distance.
[0060] During the incident of interest, the extracted information (in particular, the ball image portion and the pitch line image portion or goal image portion) may be transmitted for every successive frame. This improves the accuracy of the tracking of the remote device 104. Before and after the incident of interest, however, the extracted information may be transmitted less frequently (e.g. every nthframe where n > 1) to reduce the required data transmission rate. The data transmission rate is thus increased temporarily only when necessary (e.g. when tracking may be required to determine a critical decision in the game, such as whether or not a goal was scored or whether or not the ball travelled outside the pitch boundary). Again, this helps provide an improved balance between tracking accuracy and data transmission reduction.
[0061] This is only an example of an incident of interest and it will be appreciated the local device 102 may be configured to detect other characteristic(s) of the captured image 301 to determine the start and / or end of an incident of interest. For instance, if a predetermined number of detected players (e.g. 5, 7 or 10) begin moving quickly towards a particular common location (e.g. the location of an intersection of motion vectors of the respective players), this may be indicative of an incident and the amount of extracted information which is transmitted may again be temporarily increased (e.g. by transmitting the extracted information for every successive frame instead of every nthframe where n > 1). The amount of extracted information which is transmitted may then be reduced again after a predetermined period of time (e.g. 30, 45 or 60 seconds) or when the incident is deemed to be over. For instance, the incident may be deemed to be over once a predetermined number (e.g. 3, 5 or 8) of the players whose movement triggered the start of the incident are detected as moving away from the common location.
[0062] To further assist with managing available bandwidth when images are captured by a plurality of cameras, the local device 102 may be configured to dynamically allocate the available bandwidth between the cameras depending on a characteristic of the set of images captured by each camera. For example, for a first camera currently capturing images containing a smaller amount of relevant information, a lower proportion of the available bandwidth may be allocated to that camera. On the other hand, for a second camera currently capturing images containing a larger amount of relevant information, a higher proportion of the available bandwidth may be allocated to that camera. In other words, a portion of the bandwidth allocated to the first camera is reallocated to the second camera.
[0063] In an example, a camera may be determined as capturing images containing a smaller amount of relevant information (and thus have bandwidth reallocated from it) if, for each of a predetermined number of images consecutively captured by the camera, the proportion of pixels belonging to an extracted portion of that image (with respect to the total number of pixels of the image) is less than a predetermined threshold (e.g. less than 5%, 10% or 15%). On the other hand, a camera may be determined as capturing images containing a larger amount of relevant information (and thus have bandwidth reallocated to it) if, for each of a predetermined number of images consecutively captured by the camera, the proportion of pixels belonging to an extracted portion of that image (with respect to the total number of pixels of the image) is greater than a predetermined threshold (e.g. greater than 25%, 30% or 35%).
[0064] The local device 102 then adjusts the amount of extracted information which is transmitted from each camera depending on the available bandwidth. For instance, if the bandwidth available to a particular camera is reduced, extracted information may only be sent for a smaller number of frames per unit time (e.g. every nthframe, where n > 1) and / or only a portion of extracted information in each frame (e.g. only the extracted image portion corresponding to the ball 306 and no other extracted image portions) may be transmitted to ensure the transmission rate of extracted information does not exceed the available bandwidth. On the other hand, if the bandwidth available to a particular camera is increased, extracted information may be sent for a larger number of frames per unit time (e.g. every frame) and / or more extracted information in each frame (e.g. all extracted image portions or, at least, the extracted image portion corresponding to the ball 306 and the extracted image portions corresponding to the lines 304 and / or goal 305) may be transmitted (within the constraints of the newly available additional bandwidth).
[0065] It is noted that, if the overall bandwidth available to the system 100 fluctuates (e.g. due to general network demand), adjustment of the amount of extracted information transmitted by each camera may be carried out in the same way. This is applicable even if the local device 102 an equal portion of the available bandwidth to each respective camera. In an example, acceptable latency may also be taken into account. For example, if low latency is to be prioritised over bandwidth efficiency (e.g. temporarily during an important part of a live sports event), one or more portions of captured images that would ordinarily not be deemed relevant may nonetheless be transmitted with the extracted image portion(s). This may be effective to reduce latency if the processing time required to distinguish such portion(s) from the portion(s) to be extracted and transmitted (e.g. to distinguish people in the crowd from players in the pitch) is sufficiently high to result in an undesirably high latency. In this case, such portion(s) may be transmitted while other, more easily distinguishable non-relevant portions (e.g. the grass of the pitch) are not transmitted. This helps achieve an appropriate balance between improved bandwidth efficiency and latency reduction.
[0066] If each of a plurality of cameras comprises a respective local device 102, one of the local devices 102 may be designated a master local device configured to control each of the other local devices (e.g. via control signals transmitted and received via the network interface 205 of each local device). This allows the master local device to control the amount of extracted information transmitted from each camera in accordance with the bandwidth allocation for that camera.
[0067] In an example, the local device 102 comprises override functionality to enable a human operator to manually select one or more of a plurality of cameras whose images are to be prioritised for the transmission of extracted information. This allows, for example, camera(s) with the best view of an unexpected incident (e.g. an incident the local device 102 is not configured to detect) to be quickly allocated more bandwidth and for the amount of extracted information from images captured by those camera(s) to be increased. This facilitates, for example, improved tracking of the players 303, referee 307 and ball 306 during the incident by the remote device 104. The override functionality is provided via the user interface 206 (e.g. via a quick-access interactive menu which allows the desired camera(s) to be selected), for example.
[0068] Fig. 5 shows example steps implemented by the system 100. These include an encoding step 501 (e.g. carried out by the cameras 101A-D and local device(s) 102), streaming optimisation step 502 (e.g. carried out by the local device(s) 102) and decoding and tracking step 503 (e.g. carried out by remote device 104).
[0069] The encoding step 501 comprises a plurality of sub-steps. Sub-step 501 A comprises controlling each of the cameras 101A-D to capture respective sets of images (image frames) 301. Sub-step 501 B comprises detecting, in each image frame, relevant information for implementing tracking. For example, the image portions corresponding to each of the players 303 and referee 307, the pitch lines 304, the ball 306 and the goal 305 are detected. The detected relevant information may also be prioritised (e.g. as exemplified above to prioritise image portion(s) corresponding to a faster moving player above image portion(s) corresponding to a slower moving player). For example, the processed image frame 301 ’ is generated at sub-step 501 B. Sub-step 501 C comprises encoding the relevant information (e.g. by compressing the entire processed image frame 301’ or extracting (and, optionally, compressing) only the image portion(s) corresponding to the detected relevant information). The streaming optimisation step 502 comprises a plurality of sub-steps. Sub-step 502A comprises dynamically adjusting the encoding strategy (e.g. by controlling the encoding technique used in sub-step 501 C). For example, depending on the proportion of pixels in the current image corresponding to extracted portion(s) of the image, the encoding strategy may be selected as either compression of the entire processed image frame 301 ’ or extraction (and, optionally, compression) of only the image portion(s) corresponding to the detected relevant information. At sub-step 502C, the encoded information is transmitted from the local device 102 to the remote device 104 over the network 103.
[0070] The decoding and tracking step 503 comprises a plurality of sub-steps. Sub-step 503A comprises receiving and decoding the received encoded information.
[0071] For example, if a compressed version of the entire processed image frame 301 ’ is transmitted (e.g. in a JPEG format), the received image is decoded to determine the pixel value of each pixel in the processed image frame 301 ’. This allows tracking to be implemented based only on the non-zero pixels at sub-step 503B.
[0072] On the other hand, if only extracted portion(s) of the processed image frame 301 ’ are transmitted (togetherwith metadata indicating the position of the pixels of each extracted portion in the original image 301 I processed image 301 ’), the decoding comprises determining the position of each extracted portion based on the metadata to reconstruct the processed image frame 301’.
[0073] In an example, the reconstructed processed image frame is initialised as an image of a size corresponding to that of the original image 301 I processed image 301 ’ with all pixels being zero value pixels. Then, if, say, the lower-most, left-most pixel position in the original image 301 I processed image 301 ’ of each extracted portion is indicated, the pixel value of the pixel at this position in the reconstructed image is updated to match that of the extracted portion. The values of the remaining pixels in the reconstructed image can then also be updated accordingly (since the position of each of the other pixels in the extracted portion can be determined with respect to the lower-most, left-most pixel position). This allows tracking to be implemented based only on the non-zero pixels of the reconstructed image at sub-step 503B.
[0074] The received extracted portion(s) may also be compressed at sub-step 501 C (e.g. in a JPEG format) and thus first decoded at sub-step 503A to determine the pixel value at each pixel position. In this case, each extracted portion(s) is represented as a standalone compressed image transmitted over the network 103. The standalone compressed image may be rectangular whereas the extracted portion(s) are unlikely to be correspondingly rectangular (especially if, for instance, they are image segments of irregular shape). In this case, the pixels of the rectangular image not corresponding to a pixel of the extracted portion(s) in the original image 301 1 processed image 301 ’ are set as zero value pixels, for example.
[0075] Figs. 6A and 6B show example methods.
[0076] Fig. 6A is an example method executed by the processor 201 of local device 102. At step 601 , captured images of a sports event are received (e.g. via camera interface 204).
[0077] At step 602, one or more objects in each image to be tracked are detected.
[0078] At step 603, one or more portions of each image corresponding to the one or more detected objects are extracted.
[0079] At step 604, data representative of the one or more extracted portions and a position of each extracted portion in each captured image is transmitted (e.g. via network interface 205) to a sports object tracking apparatus (e.g. remote device 104) over a network (e.g. network 103).
[0080] Fig. 6B is an example method executed by the processor 201 of local device 102.
[0081] At step 605, data representative of one or more extracted portions of a captured image of a sports event and a position of each extracted portion in the captured image is received (e.g. via network interface 210) over a network (e.g. network 103). The one or more extracted portions corresponding to one or more objects in the image to be tracked.
[0082] At step 606, object tracking is performed using the received data.
[0083] Embodiment(s) of the present disclosure are defined by the following numbered clauses:
[0084] 1 . A sports image processing apparatus comprising circuitry configured to: receive captured images of a sports event; detect one or more objects in each image to be tracked; extract one or more portions of each image corresponding to the one or more detected objects; and transmit, over a network, data representative of the one or more extracted portions and position information of each extracted portion in each captured image to a sports object tracking apparatus.
[0085] 2. A sports image processing apparatus according to clause 1 , wherein the circuitry is configured to transmit the data according to a first mode or a second mode, wherein: the first mode comprises transmitting, as the data, a processed version of each image in which pixel values of pixels corresponding to the one or more extracted portions are retained and pixel values of pixels not corresponding to the one or more extracted portions are set to a predetermined value; and the second mode comprises transmitting, as the data, one or more image portions corresponding to the one or more extracted portions and metadata indicating position information of each extracted portion in each image. 3. A sports image processing apparatus according to clause 2, wherein the circuitry is configured to transmit the data according to the first mode or second mode depending on a type of the sports event.
[0086] 4. A sports image processing apparatus according to clause 2, wherein the circuitry is configured to transmit the data according to the first mode or second mode depending on a proportion of pixels in each image corresponding to the one or more extracted portions.
[0087] 5. A sports image processing apparatus according to any one of clauses 2 to 4, wherein, when the data is transmitted according to the second mode, the circuitry is configured to transmit extracted portions corresponding to different objects at different periodic intervals.
[0088] 6. A sports image processing apparatus according to clause 5, wherein, for: a first set of extracted portions corresponding to a first object travelling at a speed exceeding a predetermined threshold, the circuitry is configured to transmit the extracted portions at a first periodic interval; and a second set of extracted portions corresponding to a second object travelling at a speed not exceeding a predetermined threshold, the circuitry is configured to transmit the extracted portions at a second periodic interval longer than the first periodic interval.
[0089] 7. A sports image processing apparatus according to any preceding clause, wherein the circuitry is configured to: determine, using the images, whether an incident is occurring in the sports event; during a time in which the incident is occurring, transmit the data at a first periodic interval; and during a time in which the incident is not occurring, transmit the data at a second periodic interval longer than the first periodic interval.
[0090] 8. A sports image processing apparatus according to any preceding clause, wherein the captured images comprise sets of captured images captured by respective cameras, and the circuitry is configured to: allocate a portion of available bandwidth of the network for transmitting the data associated with each set of captured images depending on a characteristic of the captured images in each set; and adjust an amount of the data associated with each set of captured images which is transmitted according to the portion of bandwidth allocated to the set of captured images.
[0091] 9. A sports image processing apparatus according to any preceding clause, wherein the circuitry is configured to detect one or more types of object depending on the sports event and / or whether an incident is occurring in the sports event.
[0092] 10. A camera comprising a sports image processing device according to any preceding clause. 11. A sports object tracking apparatus comprising circuitry configured to: receive, over a network, data representative of one or more extracted portions of a captured image of a sports event and position information of each extracted portion in the captured image, the one or more extracted portions corresponding to one or more objects in the image to be tracked; and perform object tracking using the received data.
[0093] 12. A sports object tracking apparatus according to clause 11 , wherein the received data is transmitted according to a first mode or a second mode, wherein the first mode comprises transmitting, as the data, a processed version of each image in which pixel values of pixels corresponding to the one or more extracted portions are retained and pixel values of pixels not corresponding to the one or more extracted portions are set to a predetermined value, and the second mode comprises transmitting, as the data, one or more image portions corresponding to the one or more extracted portions and metadata indicating position information of each extracted portion in each image, and wherein the circuitry is configured to: if the received data is transmitted according to the first mode, perform object tracking using the retained pixel values of the one or more extracted portions of the received processed version of each image; and if the received data is transmitted according to the second mode, perform object tracking using the received one or more extracted portions and metadata of each image.
[0094] 13. A sports object tracking apparatus according to clause 12, wherein, when the received data is transmitted according to the second mode, extracted portions corresponding to different objects are transmitted at different periodic intervals, and the circuitry is configured to use the most recently received extracted portion and metadata of each object to be tracked to track the object.
[0095] 14. A system comprising a sports object image processing apparatus according to clause 1 and a sports object tracking apparatus according to clause 11 .
[0096] 15. A computer-implemented sports image processing method comprising: receiving captured images of a sports event; detecting one or more objects in each image to be tracked; extracting one or more portions of each image corresponding to the one or more detected objects; and transmitting, over a network, data representative of the one or more extracted portions and position information of each extracted portion in each captured image to a sports object tracking apparatus.
[0097] 16. A computer-implemented sports object tracking method comprising: receiving, over a network, data representative of one or more extracted portions of a captured image of a sports event and position information of each extracted portion in the captured image, the one or more extracted portions corresponding to one or more objects in the image to be tracked; and performing object tracking using the received data.
[0098] 17. A program for controlling a computer to perform a method according to clause 15 or 16.
[0099] 18. A computer-readable storage medium storing a program according to clause 17.
[0100] Numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that, within the scope of the claims, the disclosure may be practiced otherwise than as specifically described herein.
[0101] In so far as embodiments of the disclosure have been described as being implemented, at least in part, by one or more software-controlled information processing apparatuses, it will be appreciated that a machine-readable medium (in particular, a non-transitory machine-readable medium) carrying such software, such as an optical disk, a magnetic disk, semiconductor memory or the like, is also considered to represent an embodiment of the present disclosure. In particular, the present disclosure should be understood to include a non-transitory storage medium comprising code components which cause a computer to perform any of the disclosed method(s).
[0102] It will be appreciated that the above description for clarity has described embodiments with reference to different functional units, circuitry and / or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, circuitry and / or processors may be used without detracting from the embodiments.
[0103] Described embodiments may be implemented in any suitable form including hardware, software, firmware or any combination of these. Described embodiments may optionally be implemented at least partly as computer software running on one or more computer processors (e.g. data processors and / or digital signal processors). The elements and components of any embodiment may be physically, functionally and logically implemented in any suitable way. Indeed, the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the disclosed embodiments may be implemented in a single unit or may be physically and functionally distributed between different units, circuitry and / or processors.
[0104] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to these embodiments. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in any manner suitable to implement the present disclosure.
Claims
CLAIMS1 . A sports image processing apparatus comprising circuitry configured to: receive captured images of a sports event; detect one or more objects in each image to be tracked; extract one or more portions of each image corresponding to the one or more detected objects; and transmit, over a network, data representative of the one or more extracted portions and position information of each extracted portion in each captured image to a sports object tracking apparatus.
2. A sports image processing apparatus according to claim 1 , wherein the circuitry is configured to transmit the data according to a first mode or a second mode, wherein: the first mode comprises transmitting, as the data, a processed version of each image in which pixel values of pixels corresponding to the one or more extracted portions are retained and pixel values of pixels not corresponding to the one or more extracted portions are set to a predetermined value; and the second mode comprises transmitting, as the data, one or more image portions corresponding to the one or more extracted portions and metadata indicating position information of each extracted portion in each image.
3. A sports image processing apparatus according to claim 2, wherein the circuitry is configured to transmit the data according to the first mode or second mode depending on a type of the sports event.
4. A sports image processing apparatus according to claim 2, wherein the circuitry is configured to transmit the data according to the first mode or second mode depending on a proportion of pixels in each image corresponding to the one or more extracted portions.
5. A sports image processing apparatus according to claim 2, wherein, when the data is transmitted according to the second mode, the circuitry is configured to transmit extracted portions corresponding to different objects at different periodic intervals.
6. A sports image processing apparatus according to claim 5, wherein, for: a first set of extracted portions corresponding to a first object travelling at a speed exceeding a predetermined threshold, the circuitry is configured to transmit the extracted portions at a first periodic interval; and a second set of extracted portions corresponding to a second object travelling at a speed not exceeding a predetermined threshold, the circuitry is configured to transmit the extracted portions at a second periodic interval longer than the first periodic interval.
7. A sports image processing apparatus according to claim 1 , wherein the circuitry is configured to: determine, using the images, whether an incident is occurring in the sports event;during a time in which the incident is occurring, transmit the data at a first periodic interval; and during a time in which the incident is not occurring, transmit the data at a second periodic interval longer than the first periodic interval.
8. A sports image processing apparatus according to claim 1 , wherein the captured images comprise sets of captured images captured by respective cameras, and the circuitry is configured to: allocate a portion of available bandwidth of the network for transmitting the data associated with each set of captured images depending on a characteristic of the captured images in each set; and adjust an amount of the data associated with each set of captured images which is transmitted according to the portion of bandwidth allocated to the set of captured images.
9. A sports image processing apparatus according to claim 1 , wherein the circuitry is configured to detect one or more types of object depending on the sports event and / or whether an incident is occurring in the sports event.
10. A camera comprising a sports image processing device according to claim 1 .
11. A sports object tracking apparatus comprising circuitry configured to: receive, over a network, data representative of one or more extracted portions of a captured image of a sports event and position information of each extracted portion in the captured image, the one or more extracted portions corresponding to one or more objects in the image to be tracked; and perform object tracking using the received data.
12. A sports object tracking apparatus according to claim 11 , wherein the received data is transmitted according to a first mode or a second mode, wherein the first mode comprises transmitting, as the data, a processed version of each image in which pixel values of pixels corresponding to the one or more extracted portions are retained and pixel values of pixels not corresponding to the one or more extracted portions are set to a predetermined value, and the second mode comprises transmitting, as the data, one or more image portions corresponding to the one or more extracted portions and metadata indicating position information of each extracted portion in each image, and wherein the circuitry is configured to: if the received data is transmitted according to the first mode, perform object tracking using the retained pixel values of the one or more extracted portions of the received processed version of each image; and if the received data is transmitted according to the second mode, perform object tracking using the received one or more extracted portions and metadata of each image.
13. A sports object tracking apparatus according to claim 12, wherein, when the received data is transmitted according to the second mode, extracted portions corresponding to different objects are transmitted at different periodic intervals, and the circuitry is configured to use themost recently received extracted portion and metadata of each object to be tracked to track the object.
14. A system comprising a sports object image processing apparatus according to claim 1 and a sports object tracking apparatus according to claim 11 .
15. A computer-implemented sports image processing method comprising: receiving captured images of a sports event; detecting one or more objects in each image to be tracked; extracting one or more portions of each image corresponding to the one or more detected objects; and transmitting, over a network, data representative of the one or more extracted portions and position information of each extracted portion in each captured image to a sports object tracking apparatus.
16. A computer-implemented sports object tracking method comprising: receiving, over a network, data representative of one or more extracted portions of a captured image of a sports event and position information of each extracted portion in the captured image, the one or more extracted portions corresponding to one or more objects in the image to be tracked; and performing object tracking using the received data.
17. A program for controlling a computer to perform a method according to claim 15 or 16.
18. A computer-readable storage medium storing a program according to claim 17.
Citation Information
Patent Citations
Smart-court system and method for providing real-time debriefing and training services of sport games
US20150018990A1
System and method for real-time processing of ultra-high resolution digital video
US20160205341A1
Computer-implemented method for automated detection of a moving area of interest in a video stream of field sports with a common object of interest
US20200404174A1