Onboard device for targeted video recording of sporting activities
Patent Information
- Application Number
- FR2023009595
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-12
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-09-12
Abstract
Description
Title of the invention: On-board device for targeted video recording of sporting activities Technical field
[0001] The present invention relates to the technical field of video recording systems for sporting events and in particular to devices and methods for analyzing and processing images to monitor and highlight key actions and moments of sporting activities.
[0002] In the above field, it is known to use fixed cameras or cameras controlled by an operator to capture sporting events. However, these methods have several major drawbacks.
[0003] First, fixed cameras cannot effectively track moving actions on the field, which can lead to poor video recording quality and make it difficult to analyze sports performances. In addition, fixed cameras can miss key moments of the match due to their limited field of vision.
[0004] On the other hand, operator-driven cameras require constant human intervention to track the actions on the field. This method can lead to tracking errors and be costly in terms of human resources. In addition, the operator may also miss important moments of the match due to the need to focus on specific areas of the field.
[0005] Furthermore, existing video recording systems are not always able to automatically and accurately determine regions of interest in captured images, which can hamper the effectiveness of analyzing and broadcasting key moments of sporting events.
[0006] Finally, current solutions may not be suitable for effectively dealing with different lighting conditions, terrain specificities or the particularities of amateur sports, particularly team sports, which may affect the overall quality of video recording and analysis of sports performance. Furthermore, existing solutions often require a plurality of cameras, which considerably increases the final cost of the product as well as greater complexity in using the device.
[0007] In the field of video recording systems for sporting events, existing shortcomings in the state of the art pose significant problems. For example, currently established methods for tracking sporting actions do not always manage to provide complete and accurate coverage of events. on the pitch. This can lead to gaps in performance analysis and difficulties for spectators who want to focus on key moments or specific actions.
[0008] In addition, existing video recording systems are often not flexible enough to adapt to different types of sports, fields and lighting conditions. This limitation can reduce the quality of the captured images and the relevance of the analyses provided, making it difficult for coaches, analysts and spectators to identify the most important elements.
[0009] Another problem concerns the automatic determination of regions of interest in captured images. Current systems may lack accuracy and efficiency in this task, which can impact the quality of performance analysis and the broadcasting of key moments in sporting events. In addition, this can lead to a waste of time and resources when reviewing video recordings.
[0010] Furthermore, current video recording devices can be expensive and complex to set up and maintain, especially when multiple cameras or operators are required to cover the sporting event. These costs and complexities can be a barrier for sports organizations and teams that want to record and analyze their performances, especially for organizations with limited budgets.
[0011] Finally, existing video recording systems may not provide an optimal user experience for spectators who wish to follow and replay key moments of sporting events. The image streams generated by these systems can be difficult to navigate, analyze, and share, which can harm spectator engagement and the popularity of the sports concerned.
[0012] Currently, there is no accessible and inexpensive solution for automatically and accurately capturing sports actions and key moments of sporting events without constant human intervention, while adapting to different lighting conditions, specificities of fields and the particularities of amateur sport. In addition, there is also no inexpensive embedded system capable of extracting key moments of a sports match, including selectively recording video sequences from a past moment. There is therefore a need for an innovative solution that addresses these problems and fills the gaps in the state of the art.
[0013] Furthermore, major shortcomings of the state of the art include the need to use multiple cameras to capture the stream of an amateur sporting event to be broadcast in other images. These systems are often not well suited to extracting key moments from a match or competition. In addition, another significant drawback of these devices is their high cost, making their acquisition difficult for amateur sports clubs. Despite their technological potential, these solutions do not always manage to meet the specific requirements of amateur sport State of the art
[0014] The invention therefore falls within this context and seeks to resolve all of the aforementioned drawbacks. Thus, the invention seeks to propose an on-board device and a method for video recording of sporting activities, making it possible to extract regions of interest from the continuous flow of images captured by a single camera and to generate an image flow, without aberration and optimized for detailed performance analysis. It also facilitates more efficient and targeted broadcasting of significant moments of sporting events. Presentation of the invention.
[0015] The subject of the invention is a device for video recording a sporting activity on a sports field, comprising a camera equipped with a short focal length lens and a computing unit embedded in the device and capable of receiving at least one stream of images recorded by said camera equipped with a short focal length lens; said computing unit is arranged to: a. for each image, called processed image, of at least part of the received image stream, predicting the location of the current action of the sporting activity on the field; b. for each image of each group of successive images of the received image stream containing a processed image, determining a region of interest in this image whose position depends on the location predicted from said processed image; c. for each image of each group of images, extracting a sub-image of said image corresponding to said determined region of interest; d. generate a second image stream comprising the set of extracted sub-images.
[0016] Advantageously, the camera equipped with a short focal length lens makes it possible to capture a large part of the sports field, or even the entirety of said sports field.
[0017] If desired, said camera equipped with a short focal length lens may in particular be a wide-angle camera, preferably a “fish-eye” type camera.
[0018] The device for video recording a sporting activity on a sports field according to the invention allows better analysis and precise monitoring of the current action on the field thanks to the prediction of the location of a region of interest. This allows you to optimize event coverage by focusing only on relevant areas.
[0019] Advantageously, determining a region of interest in each image makes it possible to target key moments of the action, improving the quality of the recorded content and facilitating subsequent analysis.
[0020] Furthermore, the extraction of a sub-image corresponding to the region of interest makes it possible to reduce the weight of processed and stored data, making the device more efficient and economical in terms of resources and storage space.
[0021] Therefore, the generation of a second image stream comprising the set of extracted sub-images provides a clearer and more concentrated visualization of the sporting action, improving the experience for users of the device.
[0022] By doing so, this invention allows the use of an inexpensive on-board device to perform video recording of sporting activities. This makes this technology particularly attractive to sports enthusiasts who have limited means but still wish to capture and analyze their performances on the field. Consequently, this invention contributes to democratizing access to high-quality video analysis tools for the general public.
[0023] In the context of the present invention, the term "video recording device" means any device capable of capturing and processing digital visual sequences. This may include, but is not limited to, smartphones and digital tablets equipped with at least one camera, or action cameras or drones equipped with at least one camera.
[0024] If desired, the prediction of the location of the current action on the field may be achieved using different image processing and machine learning algorithms, thus providing alternative solutions to adapt to various sporting activities and lighting conditions. These algorithms may also take into account the movements of players and objects, such as balls, to improve the accuracy of the prediction.
[0025] Advantageously, the determination of a region of interest in each image can be carried out by taking into account additional factors, such as the speed of movement of the players, areas of high player density or key moments in the game.
[0026] In addition, the size and shape of the region of interest can be dynamically adjusted based on these factors, providing a flexible solution adapted to each situation.
[0027] Advantageously, the extraction of a sub-image corresponding to the region of interest can be carried out using different compression and image processing techniques, thus making it possible to find an optimal balance between the quality of the image and reducing the amount of data processed and stored. For example, lossless compression algorithms can be used to maintain image quality while reducing data size.
[0028] Advantageously, the generation of a second image stream comprising the set of extracted sub-images may be carried out using different editing and presentation methods, thus offering a variety of solutions for improving the experience of spectators and sports analysts. For example, the sub-images may be assembled into a single panoramic view, or presented in mosaic form, so as to facilitate the comparison and analysis of actions on the field.
[0029] In one embodiment of the invention, the calculation unit is capable of receiving information on the type of sporting activity and the calculation unit is arranged to, upon receipt of said information on the type of sporting activity, implement one or more image processing algorithms on all or part of at least one initial image recorded by the camera equipped with a short focal length lens to detect an outline of a terrain, the calculation unit selecting said algorithm(s) from a set of given algorithms as a function of said information on the type of sport, and to generate a binary mask, called a “terrain” mask, the dimensions of which correspond to the dimensions of the images recorded by the camera equipped with a short focal length lens and defining a selection zone corresponding to the outline of the detected terrain.
[0030] According to these features, the information on the type of sporting activity allows automatic tracking of the players, where the camera can detect the characteristic movements of each sport and adjust camera parameters, such as the zoom level and / or colorimetric parameters, accordingly. In addition, these features ensure optimal coverage of the players, ensuring that key moments of the game are captured clearly.
[0031] By doing so, the camera can automatically adjust its settings to provide shots more suited to the type of sport being played. For example, during a handball or basketball match, it can choose a wider wide-angle portion of the image, thus making it possible to capture a more significant part of the field. For football or rugby, the camera can take a smaller wide-angle portion of the image, focusing on specific areas of the field. The information on the type of sport thus makes it possible to retrieve areas adapted to the specific needs of each type of sport, without changing the camera's viewing angle.
[0032] Advantageously, the device allows optimal adaptation to different types of sporting activities by receiving information on the type of sport and by implementing specific image processing algorithms. This flexibility improves the detection of field contours, thus ensuring better coverage and analysis of the sporting event. In addition, the selection of algorithms based on the type of sport contributes to more accurate and reliable contour detection, which enhances the quality of the processed images. Finally, the generation of a binary field mask, whose dimensions correspond to the dimensions of the images recorded by the camera equipped with a short focal length lens, makes it possible to define a selection area corresponding to the contour of the detected field, thus facilitating the visualization and analysis of actions on the field by focusing on the relevant areas.
[0033] According to an exemplary embodiment of the invention, the reception of information on the type of sporting activity may be obtained automatically. The calculation unit may, for example, be arranged to analyze the images recorded by the camera equipped with a short focal length lens and / or to analyze information from additional sensors, in particular a geolocation device, to detect characteristics specific to the sport practiced. This will allow the device to adapt quickly and autonomously to the different sporting activities, without requiring manual intervention.
[0034] According to another exemplary embodiment of the invention, the device is capable of receiving information on a geographical location of the device, in particular provided by the user, so that upon receipt of said geographical location the calculation unit executes an instruction to deploy a menu for selecting the type of sporting activity allowing the user to choose the type of sport recorded.
[0035] Advantageously, the computing unit can be arranged to infer information on the type of sporting activity from the geographical location. This functionality makes it possible to improve the accuracy and efficiency of the image processing algorithms and of the entire device, by optimizing the monitoring and analysis of actions on the field for different sports.
[0036] Alternatively, the identification of the geographic location of the device may be carried out using other location technologies, such as positioning by wireless local area network (Wi-Fi), triangulation of mobile telephone signals or low energy Bluetooth beacons.
[0037] As a further variant, the inference of information on the type of sporting activity from the geographical location could be improved by using enriched geographical databases containing specific information on the sports grounds and facilities allowing the device to better identify the type of sport practiced and to adapt more precisely and quickly to different situations.
[0038] As a further variant, the inference of information on the type of sporting activity from the geographical location could be combined with other sources information, such as sensors built into the device or data from other connected devices.
[0039] Alternatively, the computing unit may be arranged to use algorithms based on machine learning techniques to infer information on the type of sporting activity from the geographic location and other contextual data such as calendars.
[0040] In one embodiment of the invention, the calculation unit is arranged to implement at least one colorimetric segmentation algorithm of said initial image to determine a set of regions in said initial image, each region having substantially the same color, and to select from said set of regions a region corresponding to the terrain, said selection zone being defined from the selected region.
[0041] Advantageously, by implementing at least one colorimetric segmentation algorithm of the initial image, the calculation unit makes it possible to determine a set of regions having substantially the same color, thus improving the accuracy of the detection of the terrain, in particular when the latter consists mainly of a single color such as a rugby or football field. This approach facilitates the selection of a region corresponding to the terrain from among all the identified regions, which makes it possible to define the selection zone more precisely and to improve the quality of the extracted images for a more detailed and relevant analysis.
[0042] Alternatively, the implementation of a colorimetric segmentation algorithm may be supplemented or replaced by the implementation, by the computing unit, of other image segmentation techniques, such as segmentation based on textures, contours or shapes. These alternatives may offer equivalent or superior performance for certain situations or certain types of terrain, by adapting better to variations in color, lighting or specific characteristics of the terrain.
[0043] Alternatively, the computing unit may implement several colorimetric segmentation algorithms in parallel or in series, combining or comparing their results to improve the detection and selection of the region corresponding to the terrain. This approach makes it possible to offer a more robust and reliable solution in the face of changing lighting conditions or terrains with complex colors or patterns.
[0044] Alternatively, the computing unit may implement one or more algorithms based on deep learning techniques, such as convolutional neural networks, to perform colorimetric segmentation and selection of the region corresponding to the terrain. These techniques are particularly robust in the face of variations in color, lighting or specific characteristics of the terrain.
[0045] If desired, the determination of the selection area from the selected region may be carried out by the calculation unit taking into account other factors or contextual information, such as the position of the players, the direction of the game or key moments of faction. This approach makes it possible to dynamically adjust the selection area and improve the relevance of the extracted images.
[0046] In a cumulative embodiment of the invention, the calculation unit is arranged to implement at least one panoptic segmentation algorithm of all or part of said initial image to determine a set of regions in said initial image, each region being associated with an object class, and to select from said set of regions a region whose class corresponds to a “terrain” class, said selection zone being defined from the selected region.
[0047] Advantageously, by implementing at least one panoptic segmentation algorithm of the initial image, the calculation unit makes it possible to determine a set of regions associated with different classes of objects. This approach improves the accuracy of terrain detection by identifying objects relevant for sports analysis in the initial image. By selecting a region whose class corresponds to a "terrain" class from among all the regions, the calculation unit precisely defines the selection zone in the initial image, thus improving the quality and relevance of the extracted sub-images.
[0048] Advantageously, the terrain selection step uses panoptic segmentation to analyze and identify specific areas based on vertical position criteria. More specifically, this technical step comprises the calculation of two relevant values, which correspond to specific percentiles of the vertical position in the area recovered by the panoptic segmentation. The first value, associated with a higher percentile, provides an indicator of the maximum height of the terrain, while the second value corresponds to a percentile located approximately at the vertical midpoint of the terrain. In practice, the area selected for subsequent operations is the one that happens to be closest, in terms of absolute value, to this second value.Thus, this method enables optimized terrain selection based on robust and accurate analysis of the vertical attributes of the image, providing a significant improvement in the accuracy and efficiency of the terrain selection process.
[0049] Alternatively, the computing unit may be arranged to implement several panoptic segmentation algorithms in parallel or in series, by combining or comparing their results to improve the detection and selection of the region corresponding to the terrain. This approach makes it possible to offer a more robust and reliable solution in the face of changing lighting conditions or terrains presenting complex objects or patterns.
[0050] Alternatively, the determination of the selection area, by the calculation unit, from the selected region may be carried out taking into account other factors or contextual information, thus making it possible to dynamically adjust the selection area.
[0051] Alternatively, the selection of the region corresponding to the terrain may be based on additional or alternative criteria, such as the size, shape or texture of the detected regions. These criteria make it possible to more precisely identify the region of the terrain among all the regions.
[0052] Advantageously, the computing unit may be arranged to select algorithms based on information on the type of sport by combining several algorithms or by using meta-algorithms to select and adjust the algorithms best suited to each situation. This will make it possible to further improve contour detection and adapt to more complex or unexpected scenarios.
[0053] Advantageously still, the generation of a binary field mask by the calculation unit may be supplemented by the creation of additional masks, for example to distinguish the different areas of the field, the teams or the key players. These masks could be used to facilitate the analysis of actions on the field and to personalize the presentation of the images according to the specific needs of spectators or sports analysts.
[0054] In an alternative or cumulative embodiment of the invention, the calculation unit is arranged to predict the location of the current action of the sporting activity, from each processed image and said binary terrain mask. The prediction may be carried out directly from the processed image, combined with the binary terrain mask or from a mask obtained from the binary terrain mask, or alternatively be carried out from data obtained from the processed image, these data being able to be combined with the binary terrain mask or from a mask obtained from the binary terrain mask.
[0055] Advantageously, by arranging the computing unit to detect and estimate the position of the people in each processed image, the device improves the quality of the sports analysis by taking into account the movements and position of the players. Furthermore, by predicting the location of the current action of the sporting activity from each processed image, the positions of the detected people and the binary terrain mask, the device offers a more precise and dynamic representation of the sporting event. This approach makes it possible to improve the relevance of the extracted images and to optimize the analysis of the sporting actions for a better understanding and a more precise evaluation of the performances of the players.
[0056] Alternatively, the detection and estimation of the position of people could be performed using different object tracking techniques, such as correlation tracking, appearance tracking, or particle filter tracking. These alternatives could offer equivalent or superior performance for certain situations or types of sports, by better adapting to the movements and interactions of players on the field.
[0057] Alternatively, the computing unit could detect and estimate not only the position of people, including players, but also other objects or elements relevant to sports analysis, such as the ball, goals or field lines.
[0058] In a cumulative embodiment of the invention, the calculation unit is arranged to detect foreground objects in each processed image and to generate a binary mask, called "foreground" mask, defining a selection zone for each detected foreground object; to generate a final binary mask by multiplying said terrain binary mask and said foreground binary mask; and to predict the location of the current action of the sporting activity from each processed image and said final binary mask.
[0059] Advantageously, by arranging the computing unit to detect foreground objects in each processed image and generate a foreground binary mask defining a selection area for each detected foreground object, the device allows for more accurate identification of important elements of the sporting action. In addition, by generating a final binary mask by multiplying the terrain binary mask and the foreground binary mask, the device improves the accuracy of localizing sporting events by focusing on relevant areas of the image. Thus, by predicting the localization of the current action of the sporting activity from the final binary mask, the device provides a better representation of the key moments and important actions of the event, which improves the analysis and understanding of the performance of players and teams.
[0060] Advantageously, the calculation unit may be arranged to implement a machine learning algorithm to detect foreground objects in each processed image and to generate said foreground binary mask, said machine learning algorithm being trained to detect and estimate the position of a foreground object in an image.
[0061] Alternatively, the computing unit may be arranged to implement a two-step approach for image processing. The first step may be based on the use of a first algorithm, which may be used to segment each processed image. This algorithm will be able to accurately distinguish the foreground from the background of the image, thus making it possible to delimit the areas of interest for subsequent operations.
[0062] After this initial segmentation, the computing unit can call upon a second algorithm. The latter can be used to detect specific objects in the foreground of each processed image. It should be noted that this foreground will have been previously segmented by the first algorithm. This combination of two algorithms will be able to provide a fine-grained image.
[0063] The last phase of image processing may consist of extracting the features. This extraction may be done by superimposing the foreground mask and the original image. By superimposing these two elements, a richer set of information may be obtained from the image, which will allow for a more detailed analysis and a better understanding of its content.
[0064] In one embodiment of the invention, said machine learning algorithm trained to detect and estimate the position of an object in the foreground in an image may use a neural network, in particular a convolutional neural network.
[0065] Alternatively, the detection of foreground objects may be performed by the computing unit using alternative image segmentation methods, such as texture-based segmentation, region-growing segmentation, or supervised classification segmentation.
[0066] As a variant of the operation of multiplying the binary terrain and foreground masks to generate a final binary mask, other operations may be used by the calculation unit such as mask combination operations, in particular resulting from the composition of one or more operations such as intersection, union or difference of images.
[0067] In an exemplary embodiment of the invention, the calculation unit may be arranged to implement at least one machine learning algorithm, from each processed image and said final binary mask, the machine learning algorithm being trained to predict the location of a current action of a sporting activity from an image. For example, each processed image may be superimposed with said final binary mask, the active pixels resulting from this superposition being highlighted relative to the inactive pixels, in particular by being colored red. The image resulting from this superposition may thus be provided as input to the machine learning algorithm trained to predict the location of the current action of the sporting activity.
[0068] For example, said machine learning algorithm may use a ResNet50 type neural network. In this case, the ResNet50 type neural network used by the algorithm implemented by the computing unit may, for example, comprise a convolutional input layer, provided in particular with filters of size 7x7 and a step size of 2. This first layer may be followed by a batch normalization function, a ReLU type activation function and finally a maximum grouping layer of size 3x3 and step size 2. This combination of operations allows low-level features to be extracted from the input image while reducing the spatial size for further processing.
[0069] Preferably, the ResNet50 type neural network employed by the algorithm implemented by the computing unit comprises, following the convolutional input layer, a plurality of main stages, such as four main stages, each composed of several residual blocks. The fundamental principle behind these residual blocks is the concept of residual connection or jump, which makes it possible to combat the problem of gradient disappearance in deep networks.
[0070] The first stage may, for example, comprise a convolutional block and two identity blocks. The second stage may, for example, comprise a convolutional block and three identity blocks. The third stage may, for example, comprise a convolutional block and five identity blocks. The fourth stage may comprise a convolutional block and two identity blocks. It may be envisaged that the dimensions of the filters used in these blocks increase progressively, making it possible to extract increasingly complex characteristics.
[0071] Still preferably, the ResNet50 type neural network used by the algorithm implemented by the computing unit comprises, after the last main stage, a global average pooling layer followed by a dense layer, for example having a softmax activation function. The last global pooling layer is used to reduce the spatial dimension to 1x1 while the output dense layer is used to produce the output class probabilities. This last step makes it possible to make the final prediction based on the characteristics extracted by the network.
[0072] In an alternative or cumulative embodiment of the invention, the calculation unit is arranged to detect and estimate the position of people in each processed image; to exclude the positions of the detected people which are located outside the selection zone of said binary terrain mask; and to predict the location of the current action of the sporting activity from the remaining positions of the detected people.
[0073] Advantageously, excluding the positions of detected persons located outside the selection area of the binary field mask allows to focus only on the persons actually involved in the sporting activity on the field. Thus, this approach improves the accuracy and relevance of the prediction of the location of the current action of the sporting activity, avoiding potential interference or distractions caused by irrelevant persons or objects located outside the field. In addition, this also allows to optimize the computing resources and reduce the processing time, because the computing unit does not have to process unnecessary or irrelevant information for the analysis of the current sporting activity.
[0074] Advantageously, the calculation unit may be arranged to implement a machine learning algorithm to detect and estimate the position of people in each processed image, said machine learning algorithm being trained to detect and estimate the position of a person in an image.
[0075] In one embodiment of the invention, said machine learning algorithm trained to detect and estimate the position of a person in an image may use a “Faster R-CNN” type neural network.
[0076] In an exemplary embodiment of the invention, the Faster R-CNN type neural network used by the algorithm implemented by the computing unit comprises a first network of the Region Proposal Network type followed by a second detection network. It may be provided that the neural network comprises a deep convolutional neural network provided upstream of the first network, making it possible to extract the feature maps from the input images. This deep convolutional neural network may be a pre-trained model, such as ResNet50, to produce a feature map from an image provided as input.
[0077] Preferably, the first region proposal network is arranged to receive the feature map of said deep convolutional neural network and to generate a series of proposals for regions of interest of said image, each likely to contain an object, each associated with an objectivity score or a probability that said region of interest actually contains an object.
[0078] This first network could, for example, move a grid of variable size on the feature map provided as input and, for each position of the grid, generate, from the grid and the feature map, one or more bounding boxes likely to contain a person and, for each box, a probability that the box contains a person rather than a background element. It may be possible to consider using a set of anchor boxes, with various scales and aspect ratios, to adapt to people of various sizes and shapes.
[0079] Preferably, it may be provided that the first region proposal network comprises a stage for grouping region proposals into regions of interest. This stage makes it possible to guarantee that the proposed regions are of a fixed size before moving on to the following layers. For example, the grouping of regions of interest may be carried out by maximum grouping on inputs of variable sizes to produce feature maps of fixed size, which are then sent to a dense neural network.
[0080] Advantageously, the second detection network receives as input the fixed-size feature maps of the clustering of the regions of interest and comprises a series of fully connected layers, arranged to classify the objects and determine whether a region of interest proposed by the first network comprises a person. The classification layer can be expected to apply a softmax function to calculate the probability distribution over different object classes, including the background. These layers can also be arranged to optimize the coordinates of the object's bounding box, in order to obtain more precise localization.
[0081] Thus, due to its architecture, the Faster R-CNN network offers a robust mechanism for object detection, in this case, for detecting players on the field in the context of this invention.
[0082] Alternatively, instead of excluding the positions of detected persons located outside the selection area of the terrain binary mask, the computing unit may implement a person tracking algorithm that assigns a unique identifier to each detected person and tracks their position over time. This makes it possible to filter out persons not involved in the sporting activity on the field based on their trajectory and behavior, thus providing an alternative for predicting the location of the current action of the sporting activity accurately.
[0083] Alternatively, instead of excluding the positions of detected persons located outside the selection zone of said binary terrain mask, the calculation unit may implement a weighting algorithm which assigns weights to the positions of the detected persons according to their proximity to the terrain and their apparent involvement in the sporting activity. This approach makes it possible to give more importance to the most relevant persons for the analysis of the sporting activity, while considering the information coming from persons located outside the terrain, thus offering an alternative for predicting the location of the current action of the sporting activity.
[0084] Alternatively, instead of excluding the positions of detected persons located outside the selection zone of said binary terrain mask, the computing unit may implement a supervised learning algorithm using a trained neural network to predict the position of persons positioned only on the terrain.
[0085] In one embodiment of the invention, the computing unit is arranged to implement at least one machine learning algorithm to predict the location of the current action of the sporting activity, from the remaining positions of the detected persons, the machine learning algorithm being trained to predict the location of a current action of a sporting activity from a set of coordinates.
[0086] Advantageously, the use of at least one machine learning algorithm to predict the location of the current action of the sporting activity from the positions of the people detected and present inside the binary mask of field, improves the accuracy and efficiency of prediction. Indeed, the machine learning algorithm can adapt to the different situations encountered in a sporting activity and learn to recognize the patterns and behaviors characteristic of each sport, making the device more versatile and efficient in tracking the action taking place on the sports field. In addition, this approach makes it easier to take into account various situations, including those that might be difficult to handle with traditional methods, such as overlapping player positions or the rapid and complex movements of athletes.
[0087] In one embodiment of the invention, said machine learning algorithm trained to predict the location of the current action of the sporting activity may employ a neural network architecture designed specifically to predict the central position of the action in a sports game from the locations of the players. This architecture may be adapted to different sports, each having a specific number of input coordinates. For example, for handball, the architecture takes into account 192 input coordinates, while for basketball, this number is reduced to 144.
[0088] Preferably, the neural network may have a perceptron-type architecture comprising at least two dense layers. This type of architecture makes it possible to process the input coordinates to produce the position of the center of the action. Thus, the output of this network is a pair of coordinates indicating the center of the action. This prediction is notably carried out by evaluating the position of the players with different temporalities.
[0089] Alternatively, the computing unit may be arranged to implement several machine learning algorithms, each specifically designed to manage particular aspects of the sporting activity or adapted to different types of sports. This approach makes it possible to offer a more personalized solution adapted to the specific needs of each sport, by exploiting the strengths of each algorithm for better prediction of the location of the current action.
[0090] Alternatively, the computing unit may implement one or more algorithms based on online learning techniques, allowing the machine learning algorithm to adapt and improve in real time, based on data received during a sporting event. This approach contributes to more accurate prediction and better adaptation to changing conditions during the event, such as changes in lighting, changes in weather or changes in team game strategies.
[0091] If desired, the computing unit may also be arranged to implement a combination of machine learning algorithms and techniques traditional image processing techniques to predict the location of the current action. This hybrid approach leverages the strengths of each method, for example, using image processing techniques to quickly detect significant changes in the scene and machine learning algorithms to refine the prediction based on contextual information and player behaviors.
[0092] In a cumulative embodiment of the invention, the calculation unit is arranged to predict the location of the current action of the sporting activity, from at least two successive processed images and said binary terrain mask. For example, the calculation unit may be arranged to implement at least one machine learning algorithm trained to predict the location of the current action of the sporting activity, from at least two successive processed images and said binary terrain mask.
[0093] Advantageously, the use of at least one machine learning algorithm to predict the location of the current action of the sporting activity from at least two successive processed images, and in particular from the positions of the people detected from these successive processed images, makes it possible to obtain a more precise and reliable prediction of the evolution of the action. Indeed, by analyzing the movements and interactions between the players from several successive images, the device can better understand the dynamics of the game and anticipate future actions. Consequently, this contributes to a smoother and more natural monitoring of sporting events, thus improving the viewing experience for spectators and facilitating the analysis of the athletes' performances for coaches and analysts.
[0094] Alternatively, the computing unit may be arranged to implement several machine learning algorithms in parallel, each being specialized in predicting a specific aspect of the current action, such as player movements, ball trajectories or interactions between players. The predictions of these algorithms may then be combined to obtain a more robust overall prediction of the location of the current action.
[0095] Alternatively, the computing unit may be arranged to implement one or more algorithms based on unsupervised or semi-supervised learning techniques to predict the location of the current action, thereby making it possible to adapt the algorithms to specific sports or game situations without requiring labeled training data. This approach could facilitate the implementation of the device in various sporting contexts, in particular for less common sports or amateur levels of competition.
[0096] As a further variant, the algorithms implemented by the computing unit can integrate temporal data, such as the speed and acceleration of the players or moving objects, to predict the location of the current action. Using this additional information allows for more precise predictions and improved tracking accuracy of sports action, especially in situations where movements are fast and unpredictable.
[0097] Alternatively, the computing unit may be configured to implement deep neural network algorithms, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), to predict the location of the current action. These types of algorithms are particularly suited for processing complex visual and temporal data and could provide improved prediction performance compared to traditional machine learning algorithms.
[0098] In a cumulative embodiment of the invention, the calculation unit is arranged to determine the location of the current action of sporting activity on the ground in the form of a pair of numbers each representing a horizontal and / or vertical coordinate in the selection zone of the binary ground mask and each defining a proportion between a normalized minimum value and a normalized maximum value.
[0099] Advantageously, the use of a pair of numbers to represent the location of the current action of sporting activity on the field allows for a simplified and accurate representation of the positions on the field, thus facilitating the processing and analysis of the data. By expressing the horizontal and vertical coordinates as proportions between a standardized minimum value and a standardized maximum value, the device is able to easily adapt to the different dimensions of sports fields, thus providing greater flexibility and better compatibility with various sporting activities. In addition, this standardized approach facilitates the comparison of action locations between different instances of the device or between different sports, thus allowing for a more consistent and robust analysis of the collected data.
[0100] Alternatively, the computing unit may be arranged to determine the location of the current sporting activity action on the field using a polar rather than Cartesian coordinate system. In this case, the position of the action will be represented by an angle and a distance relative to a fixed reference point on the field, thus providing an alternative way of expressing the position of actions in a precise manner that is adaptable to different sports and fields.
[0101] Alternatively, the location of the current sporting activity action by the computing unit may be determined using a three-dimensional coordinate system, taking into account not only horizontal and vertical coordinates, but also height. This approach allows for better analysis of actions involving overhead movements, such as jumps or throws, thus providing a more complete representation of the location of the action for certain sports.
[0102] Alternatively, the computing unit may be configured to determine the location of the current action based on predefined areas of the field rather than continuous coordinates. In this case, the field is divided into sections or grids, and the position of the action would be represented by the index of the corresponding section. This method simplifies data processing and analysis, while allowing for a flexible representation adaptable to different sports and fields.
[0103] In a cumulative embodiment of the invention, the calculation unit is arranged to segment so-called “limit” zones in the initial image; to compare each predicted location with said limit zones, and to update, as a function of said comparison, the prediction of said location of the action.
[0104] Advantageously, the segmentation of the so-called "boundary" zones in the initial image allows the device to identify the edges of the field or other areas important for tracking the action. By comparing each predicted location with the boundary zones, the device can adjust and refine the prediction of the location of the action based on this information. This leads to an improvement in the accuracy and reliability of the predictions, allowing a better understanding of the dynamics of the game and the movements of the players on the field. Furthermore, updating the predictions based on the comparison with the boundary zones ensures that the data generated by the device is more relevant and reliable.
[0105] In a cumulative embodiment of the invention, for each image of each group of successive images of the received image stream containing a processed image, the calculation unit is arranged to determine a region of interest in this image whose position depends on the location predicted from said processed image and whose dimensions depend on the dimensions of the selection zone of the binary terrain mask.
[0106] Advantageously, the use of a region of interest in each image of each group of successive images makes it possible to optimize the computing resources by focusing only on the relevant areas. By determining the position of the region of interest based on the location predicted from the processed image, the device can track the action more accurately and efficiently. In addition, by adapting the dimensions of the region of interest based on the dimensions of the selection area of the terrain binary mask, the device ensures that the region of interest covers an area large enough to capture the relevant action while avoiding processing unnecessary data. This improves the overall performance of the device and provides more accurate and informative sports analyses.
[0107] Alternatively, the determination of the region of interest by the computing unit may be based on other methods, such as motion detection or trajectory analysis, to more accurately track dynamic actions in the image stream. This allows for better targeting of relevant areas without negatively affecting the overall performance of the device.
[0108] Alternatively, rather than using only the dimensions of the selection area of the terrain binary mask to determine the dimensions of the region of interest, the calculation unit may be arranged to also take into account other factors, such as the density of players or objects in the area. This makes it possible to create a region of interest that adapts more precisely to the specific situation on the field and to provide more relevant analysis results.
[0109] Alternatively, the computing unit may be arranged to use multiple regions of interest instead of just one to better cover the entire field and capture multiple simultaneous actions. The regions of interest may be adapted based on specific events on the field, such as player duels, passes, or shots. This allows for a more comprehensive and detailed analysis of the ongoing sporting activity.
[0110] In a cumulative embodiment of the invention, the stream of images recorded by said camera equipped with a short focal length lens comprises a succession of groups of successive images. Where appropriate, the calculation unit is arranged to predict the location of the current action of the sporting activity on the field from the first image only of each group of successive images, this image forming said processed image; and the calculation unit is arranged to determine said region of interest in each of the images of the same group from the location predicted from said first image and to extract a sub-image of said image corresponding to said determined region of interest.
[0111] Advantageously, by using only the first image of each group of successive images to predict the location of the current action, the device can reduce the processing load and the resources required while maintaining acceptable accuracy. In addition, by determining the region of interest in each of the images of the same group from the location predicted from the first image, the device can follow the evolution of the action without having to process each image individually, which allows for gains in efficiency and speed. Extracting a sub-image corresponding to the determined region of interest also allows the analysis to be focused on the relevant areas of the image, thus improving the quality of the information provided and reducing noise and interference.
[0112] Alternatively, instead of predicting the location of the current action only from the first image of each group of successive images, the calculation unit can be arranged to predict localization using a representative subset of images from the group, thereby maintaining a reduction in processing load while potentially improving localization accuracy.
[0113] Alternatively, the region of interest may be determined by the computing unit using an approach for tracking objects in the group of successive images, thereby enabling dynamic adaptation of the region of interest based on movements of objects of interest, such as players or the ball, within the group of images.
[0114] Alternatively, the extraction of the sub-image can be carried out by the computing unit by applying a compression or resolution reduction algorithm to the initial images, thus making it possible to reduce the quantity of data processed without sacrificing the quality of the information extracted from the region of interest.
[0115] In a cumulative embodiment of the invention, the calculation unit is arranged to carry out image processing operations on each extracted sub-image so that the images of said second image stream are free from distortions generated by the camera equipped with a short focal length lens.
[0116] Advantageously, performing image processing operations on each extracted sub-image makes it possible to eliminate the distortions generated by the camera equipped with a short focal length lens, thus improving the quality and precision of the images of the second stream. This facilitates the analysis and understanding of the sporting actions in progress by spectators or coaches, by providing them with images free from distortions and more representative of the reality on the ground. In addition, correcting these distortions can also improve the performance of machine learning algorithms by providing them with more precise data for the analysis and prediction of sporting actions.
[0117] Advantageously, the calculation unit may comprise a memory in which a camera distortion model is stored. This model may be a non-linear model, such as a polynomial radial distortion model. This model uses a series of coefficients to describe the manner in which the radial distance of each point from the center of the image changes as a function of the angle of incidence of the light ray. The calculation unit may be arranged to apply, to each pixel of each image of said second image stream, a geometric transformation depending on the coefficients of the distortion model and making it possible to remap the coordinates of each pixel in said image to obtain new coordinates in a rectified image. It may be provided that this geometric transformation is followed by an interpolation, in the case where this geometric transformation does not provide exact pixel coordinates in the rectified image.We can provide that the interpolation is bilinear or bicubic to determine the values of the pixels. We can also provide that this geometric transformation is . followed by a calibration to adjust the parameters of the distortion model and the geometric transformation according to the specificities of the camera and lens used. This calibration makes it possible to transform an image distorted by a camera equipped with a short focal length lens into a rectified image offering a more accurate and realistic representation of the original scene.
[0118] Alternatively, the computing unit may be arranged to apply different types of image processing to correct distortions caused by the camera equipped with a short focal length lens, such as geometric transformations, radial distortion corrections or tangential distortion corrections. These different methods may be chosen according to the specific characteristics of the camera used or the particular needs of the application.
[0119] Alternatively, the computing unit may be arranged to implement an adaptive approach for image processing, adjusting the correction parameters according to the intensity of the distortions observed in the extracted sub-images. This approach allows for more precise correction of the distortions according to the specific conditions of each image.
[0120] Alternatively, the computing unit may be arranged to implement deep learning algorithms to correct the distortions generated by the camera equipped with a short focal length lens. Convolutional neural networks may be trained to learn to correct these specific distortions, offering an alternative to conventional image processing methods.
[0121] In a cumulative embodiment of the invention, the calculation unit is arranged to apply a distortion operation to each extracted sub-image, said distortion operation being determined at least from the optical characteristics of the camera equipped with a short focal length lens.
[0122] Advantageously, by applying a distortion operation to each extracted sub-image based on the optical characteristics of the camera equipped with a short focal length lens, the device is able to accurately and appropriately correct the specific distortions caused by this camera. This customized approach makes it possible to obtain images of the second image stream that are free of distortions, thus improving the quality and fidelity of the images for better analysis and understanding of the ongoing sporting activity. Furthermore, by taking into account the optical characteristics of the camera, the device can be adapted to different types of wide-angle cameras, providing a flexible solution for various situations and needs.
[0123] Alternatively, the computing unit may be arranged to implement one or more machine learning-based distortion correction algorithms, where a pre-trained model is used to estimate and correct specific distortions. to the camera equipped with a short focal length lens. This approach allows dynamic adaptation to different wide-angle cameras, providing a more flexible and potentially more accurate solution for distortion correction.
[0124] Alternatively, the computing unit may be arranged to implement one or more distortion correction algorithms based on geometric methods, which take into account the geometry of the camera equipped with a short focal length lens and the placement of objects in the image to correct distortions. This approach offers a robust and efficient alternative for correcting distortions in extracted images.
[0125] As a further variant, the computing unit may be arranged to implement a combination of several distortion correction methods, thus making it possible to exploit the advantages of different approaches for optimal distortion correction. This solution offers better overall performance in terms of distortion correction for the images extracted from the second image stream.
[0126] In one embodiment of the invention, the distortion operation is determined from one or more pieces of information among: distance from the camera to the terrain, tilt angle of the camera, brightness of the image, current zoom of the camera, the size of the terrain or even the position in the initial image of the determined region of interest.
[0127] Advantageously, by determining the distortion operation from one or more pieces of information such as the distance of the camera to the terrain, the tilt angle of the camera, the brightness of the image, the current zoom of the camera and the size of the terrain, the device can more precisely and adaptively adjust the distortion correction for each extracted sub-image. This makes it possible to obtain images from the second image stream with improved visual quality and a more faithful representation of the real scene. In addition, by taking into account this specific information, the device can dynamically adapt to changes in recording conditions, thus ensuring optimal correction performance even when conditions vary.
[0128] Alternatively, the distortion operation could be determined from other information such as the camera's altitude relative to the terrain. This alternative approach would also allow distortions to be corrected adaptively based on the specific recording conditions.
[0129] Alternatively, the computing unit may be arranged to implement a machine learning-based distortion correction technique, which would be trained on a dataset representative of typical recording conditions. This approach makes it possible to predict and apply an optimal distortion operation for each sub-image, taking into account the optical characteristics of the camera. equipped with a short focal length lens and specific recording conditions.
[0130] The invention also relates to a system for video recording a sporting activity on a sports field, the system comprising a video recording device, and at least one mobile terminal remote from the device and comprising a control unit, a storage memory and a user interface. Where appropriate, the device and the mobile terminal each comprise wireless communication means so that the device transmits the second video stream to the mobile terminal; and the control unit of the mobile terminal is arranged to, in response to a command entered via the user interface, record in the storage memory of the mobile terminal, at least a portion of the second video stream starting at a time prior to the time of entry of the command.
[0131] Advantageously, the integration of a video recording device into a system for video recording a sporting activity makes it possible to obtain a high-quality recording with adaptive correction of the distortions generated by the camera equipped with a short focal length lens. The use of a remote mobile terminal, equipped with a control unit, a storage memory and a user interface, facilitates management and interaction with the system by users. Thanks to the wireless communication means integrated into the device and the mobile terminal, the second video stream is transmitted in real time to the terminal, allowing users to view and analyze the images live.The control unit of the mobile terminal makes it possible to record in the storage memory of the terminal, at least a part of the second video stream in response to a command entered via the user interface, thus facilitating the saving of important images or sequences at a time prior to the time of entering the command. This configuration considerably improves the flexibility and user-friendliness of the video recording system for users, while ensuring high-quality recordings for subsequent analysis and broadcasting.
[0132] Alternatively, instead of using a remote mobile terminal for management and interaction with the system, it may be possible to use a fixed terminal or a workstation, providing an alternative solution for situations where mobility is not required or for use in a centralized control center.
[0133] Alternatively, the system may integrate several video recording devices, thus making it possible to cover larger areas of the sports field or to record different viewing angles for more complete monitoring of the sporting activity.
[0134] Alternatively, the system may be equipped with wired communication means or a combination of wired and wireless communications to transmit the second video stream to the mobile terminal, thus providing an alternative depending on infrastructure needs and connectivity preferences.
[0135] Alternatively, the control unit of the mobile terminal may allow not only to record parts of the second video stream, but also to edit or annotate them, providing a more complete solution for the analysis and review of the captured sequences.
[0136] Alternatively, the video recording system may include sports analysis software integrated into the mobile terminal or video recording device, making it possible to automatically generate statistics, reports and visualizations based on the recorded sequences, thus providing a more complete solution for coaches and performance analysts.
[0137] In one embodiment of the invention, the mobile terminal may include a pre-recording trigger button forming said user interface. The button may be a mechanical button or a graphic element of an interface.
[0138] In one embodiment of the invention, the mobile terminal may integrate a pre-recording trigger button which will be part of the user interface. This button may be a mechanical button or a graphic element in a digital interface.
[0139] It may be provided that all or part of the second video stream received by the mobile terminal will be temporarily stored in a pre-recording buffer memory in the memory of the mobile terminal. This buffer memory may be designed to retain a limited quantity of content, for example the last 15 seconds of the video stream.
[0140] The second video stream will be recorded and broadcast in its entirety, but in parallel, a 15-second circular buffer will also allow selective recording of highlights. When a user activates the trigger button on the mobile terminal, the images of the second video stream stored in this pre-recording buffer will be retrieved by the control unit of the mobile terminal and added to the beginning of the current recording.
[0141] The stream will be continuously transmitted to the phone and will be stored only when the button is pressed, allowing the 15 seconds preceding the button press to be specifically recorded. This will give the user the opportunity to capture not only real-time moments, but also to revisit important moments immediately preceding them.
[0142] Of course, the various features, variants and embodiments of the invention may be combined with each other in various combinations to the extent that they are not incompatible, or exclusive, with respect to each other. Brief description of the figures
[0143] Other advantages and characteristics of the present invention are now described using examples which are purely illustrative and in no way limitative of the scope of the invention, and from the appended drawings, drawings in which the various figures represent:
[0144] [Fig-1] represents, schematically and partially, a method of generating a video stream executed by a computing unit comprising regions of interest of a sporting activity according to a first embodiment of the invention.
[0145] [Fig.2] represents, schematically and partially, a method for generating a video stream executed by a computing unit comprising regions of interest of a sporting activity according to a second embodiment of the invention.
[0146] [Fig.3] represents, schematically and partially, a first image of a sports field, a second image of a colorimetric segmentation of the same sports field and a third image of a binary mask of the same sports field according to an embodiment of the invention.
[0147] [Fig.4] represents, schematically and partially, the different stages of generation of a video stream from a succession of extracted sub-images corresponding to regions of interest according to an embodiment of the invention.
[0148] [Fig.5] represents, schematically and partially, a region of interest of an image of a sporting activity determined from the detection of the players and the ball, according to an embodiment of the invention.
[0149] [Fig.6] represents, schematically and partially, a device for video recording a sporting activity according to one embodiment of the invention.
[0150] [Fig.7] represents, schematically and partially, a method of pre-recording and broadcasting a video stream of a region of interest of a sporting activity according to an embodiment of the invention.
[0151] [Fig.8] represents, schematically and partially, the architecture of the convolutional neural network ResNet50.
[0152] [Fig.9] represents, schematically and partially, the architecture of the Faster R-CNN neural network.
[0153] In the following description, elements that are identical, by structure or by function, appearing in different figures retain, unless otherwise specified, the same references. Description of the embodiments
[0154] Furthermore, various other characteristics of the invention emerge from the appended description given with reference to the drawings which illustrate non-limiting forms of embodiment of the invention and where:
[0155] [Fig.l] illustrates a method for generating a video stream 106, according to a first embodiment, executed by a computing unit 112 of a video recording device 110 of a sporting activity on a sports field 200, as represented in [Fig.6].
[0156] The method begins with a step 1A of selecting the type of sport to be recorded from a menu and the geolocation of the device.
[0157] Step 1A comprises a sub-step 1.1A of providing said geolocation data by a user of said video recording device 110 through a user interface (not shown).
[0158] Step 1A includes a sub-step 1.2A of selecting the type of sport from a selection menu accessible from said user interface of the video recording device 110. The options presented by said selection menu being inferred from the geolocation data provided in step 1.1A.
[0159] Step 1A of the method is followed by a step 2A of capturing an initial video stream 100 by a camera equipped with a short focal length lens 111 of the video recording device 110 and sampling said initial video stream 100.
[0160] Step 2A includes a sub-step 2.1A of capturing the initial video stream by a camera equipped with a short focal length lens 111 of the “fish-eye” type and recording 60 images per second. The initial video stream 100 recorded is transmitted to the calculation unit 112, the latter receives said initial video stream and records it in a buffer memory (not shown) of the video recording device 110; said received video stream is a received image stream 100, the images of said received image stream 100 are ordered chronologically.
[0161] Step 2A comprises a sub-step 2.2A of sampling by the calculation unit 112 of said initial video stream received 100 so as to produce two image sub-streams F1 and F2. The first image sub-stream F1 corresponds to the extraction of one image out of five from the received image stream 100 and the second image sub-stream F2 corresponds to the complementary set of the image sub-stream F1 with respect to the initial image stream received 100. The creation of the video streams F1 and F2 is carried out so that the images of each video stream are chronologically indexed as a whole, so that the joining of the two video streams F1 and F2 preserves the temporal order of the events.
[0162] Sub-step 2.2A also comprises the transmission of the second image sub-stream F2 to a processing stack within the calculation unit 112.
[0163] Step 2A is followed by a step 3A of detecting the sports field 200 and generating a binary field mask 203 from the initial video stream 100 or from one of the image sub-streams F1 or F2. The algorithms which are implemented in this step 3A depend in particular on the information from step IA and are thus selected by the calculation unit 112, from a range of algorithms, based on this information and in particular the type of sport obtained at the end of step IA.
[0164] In the example described, for one or more initial images 201 of this image stream 100, F1 or F2, step 3A comprises a first sub-step 3.1A of determining a colorimetric segmentation 202 of this initial image 201 by the calculation unit 112. The colorimetric segmentation 202 produces a set of pixel zones, not necessarily connected, each of said zones being associated with the same color.
[0165] Step 3A comprises a second sub-step 3.2A of determining a panoptic segmentation (not shown) of the initial image 201 by the calculation unit 112 making it possible to identify the field 200, the players and the spectators present in the initial image 201.
[0166] Step 3A comprises a final sub-step 3.3A of generating a binary terrain mask 203 from the geolocation data provided in step 1A and the colorimetric 202 and panoptic segmentations obtained during steps 3.1A and 3.2A. The binary terrain mask 203 is obtained by first determining the areas of the colorimetrically segmented image whose surface area is larger than a threshold of 25% of the total surface area of the image.
[0167] In this particular case, two key values are calculated from specific percentiles of the vertical position in the area targeted by the panoptic segmentation. The first value, corresponding to a high percentile, gives an indication of the maximum height of the area. The second value, in turn, corresponds to a percentile approximately at the vertical midpoint of the area, thus providing an index of the central position of the area. In this case, the area finally selected for the following operations is the one that is closest, in terms of absolute value, to this second value.
[0168] Finally, the terrain binary mask 203 is created from the selected area. The binary mask 203 represents the precise area of the sports field 200 which will be retained for subsequent processing, in particular in order to exclude elements outside the sports field 200.
[0169] The method also comprises a step 4A of segmenting the foreground and the background of each image from the first image sub-stream FL and of generating a foreground binary mask (not shown). The steps which will thus be described are therefore repeated for each image of the first image sub-stream FL
[0170] Step 4A comprises a first sub-step 4.1A of segmentation by the calculation unit 112 of the foreground and the background of each image from the first image stream FL. The calculation unit 112 detects the objects in the foreground of each image using a machine learning algorithm in the form of a convolutional neural network trained to detect and estimate the position of a foreground object in an image.
[0171] Step 4A includes a second sub-step 4.2A of generation by the calculation unit 112 of a foreground binary mask 204 (not shown) defining a foreground selection zone.
[0172] The foreground binary mask is created based on the foreground segmentation performed in substep 4.1A. It represents the position area of the main objects in the image, in particular the players on the field.
[0173] The method comprises a step 5A of generating a final binary mask and superimposing said final binary mask with the sub-stream of images FL
[0174] Step 5A includes a first sub-step 5.1A of generating the final binary mask 205 obtained by multiplying the terrain binary mask 203 and the foreground mask 204.
[0175] Step 5A comprises a second sub-step 5.2A of superposition by the calculation unit 112 of the final binary mask 205 with each image of the first sub-stream of images F1 to obtain, for each of the images, a set of active pixels resulting from this superposition and highlighted with respect to the inactive pixels by a coloring in red (not shown). The resulting image 206 of the superposition of an image of the sub-stream of images F1 and the final binary mask 205 is capable of being provided as input to a machine learning algorithm trained to predict the location of the current action of the sporting activity. All of the images 206 thus generated form a third stream of images F3.
[0176] Step 5A is followed by a step 6A of providing the third stream of frames F3 as input to a ResNet50 neural network for locating the center of the sporting action. The ResNet50 neural network comprises a convolutional input layer, provided in particular with filters of size 7x7 and a step size of 2. This first layer can be followed by a batch normalization function, a ReLU type activation function and finally a maximum grouping layer of size 3x3 and step size 2.
[0177] The ResNet50 neural network used by the algorithm implemented by the calculation unit 112 comprises, following the convolutional input layer, four main stages, each composed of several residual blocks.
[0178] The first stage includes one convolutional block and two identity blocks. The second stage includes one convolutional block and three identity blocks. The third stage includes one convolutional block and five identity blocks. The fourth stage includes one convolutional block and two identity blocks.
[0179] Following the last main stage, the neural network has a global average pooling layer followed by a dense layer, having a softmax activation function.
[0180] This step also comprises the determination and extraction, by the calculation unit 112, of a region of interest 103; 104 from the image stream F3. The determination of the region of interest is carried out by means of the ResNet50 neural network trained to return as output the location of the center of the action and to determine from said location of the center of the action a region of interest, in the form of a set of coordinates on the image defining a rectangle.
[0181] Step 6A is followed by a step 7A of generating an image stream 105 corresponding to the extracted regions of interest 103; 104. The extraction of the region of interest is carried out on each image of the image streams F1 and F2 to recompose a stream corresponding to the initial stream 100. All of the regions of interest thus determined form an image stream 105 composed solely of regions of interest.
[0182] Step 7A is followed by a step 8A of generating a stream of rectified images 106 from the stream 105. The rectification of the distortion generated by the camera equipped with a short focal length lens 111 of the video recording device 110 on all of the images forming the stream of images 105. The rectification of the image is carried out by applying a geometric transformation to each of the sub-images to correct the distortion introduced by the optics of the “fish-eye” camera 111.
[0183] First, a mathematical model of the distortion is determined from the specific characteristics of the optics of the camera equipped with a short focal length lens 111, such as the focal length, the position of the principal point and the radial and tangential distortion coefficients. This information is obtained by a procedure of calibrating the camera by capturing images of a known standard pattern and estimating the camera parameters which minimize the difference between the projected positions of the points of the pattern and their positions observed in the images.
[0184] Next, for each sub-image, a set of corresponding points is identified between the distorted image and an idealized version of the undistorted image. These points are used to calculate the geometric transformation that corrects the distortion. This transformation is generally a non-linear function that moves each pixel of the distorted image to its corrected position in the rectified image. Finally, the transformation is applied to the entire image to produce a rectified sub-image.
[0185] Finally, step 8A thus makes it possible to generate a stream of images 106 corresponding to the rectified sub-images of the regions of interest 103; 104.
[0186] [Fig.2] illustrates a method for generating a video stream 106, according to a second embodiment, executed by a computing unit 112 of a device video recording 110 of a sporting activity on a sports field 200. The steps identical to the first embodiment retain the same references.
[0187] The method begins with a step 1A of selecting the type of sport to be recorded from a menu and the geolocation of the device 110.
[0188] Step 1A comprises a sub-step 1.1A of providing said geolocation data by a user of said video recording device 110 through a user interface (not shown).
[0189] Step 1A includes a sub-step 1.2A of selecting the type of sport from a selection menu accessible from said user interface of the video recording device 110. The options presented by said selection menu being inferred from the geolocation data provided in step 1.1A.
[0190] Step 1A of the method is followed by a step 2A of capturing an initial video stream 100 by a camera equipped with a short focal length lens 111 of the video recording device 110 and sampling said initial video stream 100.
[0191] Step 2A includes a sub-step 2.1A of capturing the initial video stream by a camera equipped with a short focal length lens 111 of the “fish-eye” type and recording 60 images per second. The initial video stream 100 recorded is transmitted to the calculation unit 112, the latter receives said video stream and records it in a buffer memory (not shown) of the video recording device 110; said received video stream is a received image stream 100, the images of said received image stream 100 are ordered chronologically.
[0192] Step 2A comprises a sub-step 2.2A of sampling by the calculation unit 112 of said initial video stream received 100 so as to produce two image sub-streams F1 and F2. The first image sub-stream F1 corresponds to the extraction of one image out of five from the received image stream 100 and the second image sub-stream F2 corresponds to the complementary set of the image sub-stream F1 with respect to the received image stream 100. The creation of the video streams F1 and F2 is carried out so that the images of each video stream are chronologically indexed as a whole, so that the joining of the two video streams F1 and F2 preserves the temporal order of the events.
[0193] Sub-step 2.2A also comprises the transmission of the second image sub-stream F2 to a processing stack within the calculation unit 112.
[0194] Step 2A is followed by a step 3A of detecting the sports field 200 and generating a binary field mask 203 from the initial video stream 100 or from one of the image sub-streams F1 or F2. The algorithms that are implemented in this step 3A depend in particular on the information from step IA and are thus selected by the calculation unit 112, from a range of algorithms, from this information and in particular from the type of sport obtained at the end of step IA.
[0195] In the example described, for one or more images 201 of this image stream 100, F1 or F2, step 3A comprises a first sub-step 3.1A of determining a colorimetric segmentation 202 of this initial image 201 by the calculation unit 112. The colorimetric segmentation 202 produces a set of pixel zones, not necessarily connected, each of said zones being associated with the same color.
[0196] Step 3A comprises a second sub-step 3.2A of determining a panoptic segmentation (not shown) of the initial image 201 by the calculation unit 112 making it possible to identify the sports field 200, the players and the spectators present in the initial image 201.
[0197] Step 3A comprises a final sub-step 3.3A of generating a binary terrain mask 203 from the geolocation data provided in step 1A and the colorimetric 202 and panoptic segmentations obtained during steps 3.1A and 3.2A. The binary terrain mask 203 is obtained by first determining the areas of the colorimetrically segmented image whose surface area is larger than a threshold of 25% of the total surface area of the image.
[0198] In this particular case, two key values are calculated from specific percentiles of the vertical position in the area targeted by the panoptic segmentation. The first value, corresponding to a high percentile, gives an indication of the maximum height of the area. The second value, in turn, corresponds to a percentile approximately at the vertical midpoint of the area, thus providing an index of the central position of the area. In this case, the area finally selected for the following operations is the one that is closest, in terms of absolute value, to this second value.
[0199] Finally, the terrain binary mask 203 is created from the selected area. The binary mask 203 represents the precise area of the sports field 200 which will be retained for subsequent processing, in particular in order to exclude elements outside the sports field 200.
[0200] In this second embodiment of the invention, step 3A is followed by a step 4B of determining the positions of the players on the sports field 200 from each image of the first image sub-stream FL. The steps which will thus be described are therefore repeated for each image of the first image sub-stream FL.
[0201] In a first sub-step 4.1B, the image is provided as input to a neural network whose architecture is based on that of the Faster R-CNN network to detect and determine the coordinates of the players on the sports field. The Faster R-CNN network was previously trained on the COCO dataset, which consists of more than 330,000 images divided into 80 object categories and 91 background element categories, with a total of 1.5 million object instances. Each image also has 5 captions, and there are key points for 250,000 people in the dataset.
[0202] Then, in a second sub-step 4.2B, the calculation unit 112 removes possible false positives of the players from the binary field mask 203, by excluding the positions located outside the limits of the field defined by this mask. The calculation unit can also remove other positions, depending on the information of the type of sport provided in step 1A by counting the number of players and referees provided in this type of sport. Thus, a set of positions of the players on the field is generated.
[0203] Step 4B is followed by a step 5B of determining a region of interest 103; 104 from the image stream FL, from all the positions of the players determined in step 4B as well as from the binary terrain mask 203. The determination of the region of interest is carried out by means of a perceptron-type neural network (not shown) trained to receive as input the coordinates of the positions of the players, determined in the previous step, and to return as output the center of the action, in the form of a pair of coordinates relative to the dimensions of the image and defining the center of the sporting action. From the center of the sporting action determined, the calculation unit 112 determines regions of interest 103; 104 on all the images of the stream FL
[0204] Step 5B is followed by a step 7A of generating an image stream 105 corresponding to the regions of interest 103; 104 extracted in step 5B. The extraction of the region of interest is carried out on each image of the image streams F1 and F2 to recompose an image stream corresponding to the initial stream 100. All of the regions of interest thus determined form an image stream 105 formed solely of the regions of interest.
[0205] Step 7A is followed by a step 8A of generating a stream of rectified images 106 from the stream 105. The rectification of the distortion generated by the camera equipped with a short focal length lens 111 of the video recording device 110 on all of the images forming the stream of images 105. The rectification of the image is carried out by applying a geometric transformation to each of the sub-images to correct the distortion introduced by the optics of the camera equipped with a short focal length lens 111 of the “fish-eye” type.
[0206] Firstly, a mathematical model of the distortion is determined from the specific characteristics of the optics of the camera equipped with a short focal length lens 111, such as the focal length, the position of the principal point and the radial and tangential distortion coefficients. This information is obtained by a procedure of calibrating the camera by capturing images of a standard pattern known and the estimation of camera parameters that minimize the difference between the projected positions of the pattern points and their observed positions in the images.
[0207] Then, for each sub-image, a set of corresponding points is identified between the distorted image and an idealized version of the undistorted image. These points are used to calculate the geometric transformation that corrects the distortion. This transformation is generally a non-linear function that moves each pixel of the distorted image to its corrected position in the rectified image. Finally, the transformation is applied to the entire image to produce a rectified sub-image.
[0208] Finally, step 8A thus makes it possible to generate a second stream of images 106 corresponding to the rectified sub-images of the regions of interest 103; 104.
[0209] It should be noted that although the various steps of this method have been presented in a specific order, they may be performed in any order that remains consistent with the objectives of the method. In addition, certain steps may be omitted or combined with others, depending on the specific needs of the user and the capabilities of the video recording apparatus 110.
[0210] [Fig.3] presents a series of three images illustrating different images used in determining the binary mask of a sports field. The first image 201 depicts an initial panorama of the sports field captured in its entirety and the elements around it such as the stadium, the stands and the spectators.
[0211] The second image 202 represents the same image after having been the subject of a colorimetric segmentation allowing the different zones to be displayed distinctly, in particular of the field, thanks to their specific color range. Finally, a third image 203 corresponding to a binary mask of a sports field obtained by joint application of the detection and localization algorithms allowing an image to be displayed by panoptic segmentation of the field on the one hand, and by the information obtained by the colorimetric segmentation on the other hand.
[0212] [Fig.4], we observe a sequence of drawings illustrating the different sub-steps of a video processing process according to the embodiment of [Fig.2].
[0213] The first part of the sequence shows the image stream 100 of a sports field obtained by the camera equipped with a short focal length lens 111.
[0214] The second part illustrates the process of sampling said first image stream 100 to form the sub-streams F1 and F2.
[0215] The third part illustrates the step where the terrain mask 203 is generated, by the calculation unit 112, from an initial image 201 of the first sub-stream F1, the step where the calculation unit 112 determines the positions of the people present on each image of the first sub-stream F1 and the step where certain positions are excluded, in particular by taking into account the mask 203 to arrive at a set of coordinates (x;, y;).
[0216] The fourth part illustrates the step where the calculation unit 112 predicts the location of the current action on the ground, from each image of the first sub-stream Fl and the set of coordinates (xi5 yO, and thus determines regions of interest 103; 104.
[0217] The fifth part of the sequence illustrates the formation of an image stream 105 where each image is a sub-image extracted from an image of the sub-stream F1 or F2, corresponding to a region of interest 103; 104, the stream 105 thus corresponding to the initial stream 100.
[0218] Finally, the last part illustrates the step where the calculation unit corrects, in each image of the stream 105, the distortions introduced by the short focal length lens 111 of the “fish-eye” type in each image of the initial stream 100.
[0219] In [Fig.5], we observe a specific graphical representation of a region of interest 103 of a sporting activity in progress on a sports field 200 from an image taken by a camera equipped with a short focal length lens.
[0220] This image highlights several players in the middle of the game, captured at the most crucial moment of the action. The human forms are distinctly highlighted, resulting from the identification of the players on the field.
[0221] Each player is framed by a rectangular box 300 highlighting their specific position on the field. In addition, each box is accompanied by a label, providing additional information about the type of object identified (not shown).
[0222] Also shown are the velocity vectors 302 of the players indicating the direction and sense of movement, as well as the speed of movement of the players. This information forms part of the input data of the machine learning algorithm responsible for predicting the region of interest in the images of the sampled video stream.
[0223] Within said region of interest 103, the ball, the central object of the sporting activity, can also be located. The ball is also highlighted by a circular box 301 indicating its precise position on the field.
[0224] [Fig.6] is a schematic representation of a video recording system of a sporting activity on a sports field according to an exemplary embodiment of the invention. The system comprises a video recording device 110 of a sporting activity. The device 110 integrates a computing unit 112 and a camera comprising a short focal length lens 111, both embedded in the device.
[0225] The camera equipped with a short focal length lens 111 is a fish-eye camera and captures a stream of images 100 of the sports field 200 and making it possible to record all the sporting activity on the field 200.
[0226] The computing unit 112, in close interaction with the camera 111, processes the captured image stream 100. The computing unit 112 implements the method of [Fig.l] and segments the images, identifies the regions of interest 103; 104, predicts the actions and locations, and performs all the operations necessary to create a second image stream 106 free of distortions and focused on the regions of interest 103; 104.
[0227] Furthermore, the device integrates a wireless communication element 113, of the Wi-Fi type. This element allows the device 110 to transmit the second image stream 106, generated by the calculation unit 112 at the end of step 8A of the method of [Fig.l] or [Fig.2], to a mobile terminal 116.
[0228] The mobile terminal 116 is provided with a control interface comprising a trigger button 115 for recording in memory the second image stream 105 transmitted by the device 110 via the wireless communication element and this from an instant preceding the instant of pressing the trigger button 115, in this case the 15 seconds preceding said instant of pressing the trigger button 115.
[0229] In [Fig.7], we have illustrated a procedure that continuously records a sporting activity from the second video stream 105 on the mobile terminal 116, highlighting a specific region of interest. Step 0 refers to the actuation of the trigger button 115. This does not initiate the recording, but rather marks a point of interest in the continuous video stream. More specifically, pressing the trigger button 115 signals an instruction to retain the last fifteen seconds of the second video stream 105 preceding that moment, until the next actuation of the same button.
[0230] After step 0, the process continues with a series of steps, indicated by a block, corresponding to steps 1A to 8A described either in the first embodiment of the invention described in [Fig.l], or in the second embodiment of the invention described in [Fig.2]. It should be noted that these steps are constantly being executed, capturing the entire sports match and transmitting it to the mobile terminal 116, independently of the actuation of the trigger button 115. The trigger button 115 is mainly used to mark moments of interest to be reviewed later, although the entire sequence of the match can be retrieved thanks to a specific functionality.
[0231] The process is completed by a step 9 of post-processing the second video stream 105, which makes it possible to add graphic elements, such as for example score tracking and player identification, to the images of the video stream.
[0232] In [Fig.8] we have represented the architecture of the ResNet50 neural network.
[0233] In [Fig.9], we have represented the architecture of the Faster R-CNN neural network, used to determine the region of interest.
[0234] The network is composed of a feature extraction layer 400, performed by a pre-trained ResNet50 network, arranged to extract a feature map 401 from an image 101 provided as input to the network.
[0235] Next, a region proposal network 402 is applied to the feature map 401. The region proposal network, or RPN, 402 uses a convolutional network with a filter of size 3x3 to produce two outputs: the objectness scores (not shown) and the regression bounding boxes 403.
[0236] The RPN object proposals then pass through a region of interest grouping mechanism 404.
[0237] Finally, a detection network 405 is applied, which generally comprises two dense layers of size 4096. Two outputs are produced by this detection network 405: an object classification layer and a bounding box regression layer.
[0238] The dimensions may vary depending on the size of the input image 101, the network used for feature extraction, and the number of object classes.
[0239] Training the Faster R-CNN follows a process comprising a four-step alternative training regime. Training begins by training the region proposal network as an independent network for generating object proposals. Subsequently, the proposals generated by the region proposal network are used to train the detection network. Next, the region proposal network is trained again, this time being initialized with the parameters from the detection network. Finally, the detection network is retrained once more using the common convolution layers with the region proposal network.
[0240] In terms of training data, for handball, a database comprising 279,656 images was used. For basketball, the training database has 121,000 images. These databases are then augmented to enhance the robustness of the network training. For validation, 25% of the training database is reserved. These images are annotated with the position of the center of the action.
[0241] Regarding optimization, the ADAM optimizer is used for all models. The cost function used to evaluate the performance of the models is the mean square cost which measures the average of the squares of the errors or deviations between the predictions and the actual values.
[0242] The foregoing description clearly explains how the invention achieves its stated objectives, namely to improve the quality and efficiency of recording and monitoring sporting activities, by proposing an intelligent system which, through the use of a camera equipped with a short focal length lens and of a sophisticated computing unit, is capable of predicting and tracking the location of the current action of the sporting activity. In addition, the invention provides for the use of a mobile terminal allowing convenient user interaction and management of video recordings.
[0243] In any event, the invention cannot be limited to the embodiments specifically described in this document, and extends in particular to any equivalent means and to any technically operative combination of these means. In particular, it will be possible to envisage the use of different types of sensors or cameras, the application of other machine learning algorithms, the prediction of location based on additional data such as historical activity data, or even the extension of the use of the system to applications other than sport, such as security monitoring or the analysis of movements in the context of artistic performances. In addition, it would be possible to envisage the integration of additional functionalities in the mobile terminal, such as the possibility of annotating videos in real time, or of instantly sharing video extracts on social networks.
Claims
Claims
1. Device (110) for video recording of a sporting activity on a sports field (200) comprising a camera equipped with a short focal length lens (111) and a computing unit (112) embedded in the device (110) and capable of receiving at least one image stream (100) recorded by said camera equipped with a short focal length lens (111), characterized in that the computing unit (112) is arranged to: a. for each image, called processed image, of at least part of the received image stream (100, Fl), predicting the location of the current action of the sporting activity on the field; b. for each image of each group of successive images of the received image stream (100) containing a processed image, determining a region of interest (103; 104) in this image whose position depends on the location predicted from said processed image; c. for each image of each group of images, extracting from said image a sub-image corresponding to said determined region of interest (103; 104); d. generating a second image stream (105, 106) comprising the set of extracted sub-images; and in that the calculation unit (112) is arranged to carry out image processing operations on each extracted sub-image so that the images of said second image stream (106) are free from distortions generated by the camera equipped with a short focal length lens (111).
2. Device (110) for video recording a sporting activity on a sports field (200) according to claim 1, characterized in that the calculation unit (112) is capable of receiving information on the type of sporting activity and in that the calculation unit (112) is arranged to, upon receipt of said information on the type of sport, implement one or more image processing algorithms on all or part of at least one initial image (201) recorded by the camera equipped with a short focal length lens (111) to detect an outline of a field (200), the calculation unit (112) selecting said algorithm(s) from a set of given algorithms based on said information on the type of sport, and to generate a binary mask (203), called “field mask”, the dimensions of which correspond to the dimensions of the images recorded by the camera equipped with a short focal length lens (111) and defining a selection zone corresponding to the contour of the detected field (200).
3. Device (110) for video recording a sporting activity on a sports field (200) according to claim 2, characterized in that the calculation unit (112) is arranged to implement at least one colorimetric segmentation algorithm of said initial image to determine a set of regions in said initial image, each region having substantially the same color, and to select from said set of regions a region corresponding to the field (200), said selection zone being defined from the selected region.
4. Device (110) for video recording a sporting activity on a sports field (200) according to one of claims 2 to 3, characterized in that the calculation unit (112) is arranged to implement at least one panoptic segmentation algorithm of all or part of said initial image to determine a set of regions in said initial image, each region being associated with an object class, and to select from said set of regions a region whose class corresponds to a “ground” class, said selection zone being defined from the selected region.
5. Device (110) for video recording a sporting activity on a sports field (200) according to one of claims 2 to 4, characterized in that the calculation unit (112) is arranged to predict the location of the current action of the sporting activity, from each processed image and said binary field mask (203).
6. Device (110) for video recording a sporting activity on a sports field (200) according to the preceding claim, characterized in that the calculation unit (112) is arranged to detect foreground objects in each processed image and to generate a binary mask (204), called "foreground" mask, defining a selection zone for each detected foreground object; to generate a final binary mask (205) by multiplying said terrain binary mask (203) and said first binary mask plan (204); and to predict the location of the current action of the sporting activity from each processed image and said final binary mask.
7. Device (110) for video recording a sporting activity on a sports field (200) according to one of claims 5 to 6, characterized in that the calculation unit (112) is arranged to detect and estimate the position of people in each processed image; to exclude the positions of the detected people which are located outside the selection zone of said binary field mask (203); and to predict the location of the current action of the sporting activity from the remaining positions of the detected people.
8. Device (110) for video recording a sporting activity on a sports field (200) according to one of claims 5 to 7, characterized in that the calculation unit (112) is arranged to determine the location of the current action of sporting activity on the field in the form of a pair of numbers each representing a horizontal and / or vertical coordinate in the selection zone of the binary field mask (203) and each defining a proportion between a standardized minimum value and a standardized maximum value.
9. System for video recording a sporting activity on a sports field, the system being characterized in that it comprises a device (110) for video recording a sporting activity on a sports field (200) according to one of the preceding claims, and at least one mobile terminal (116) remote from the device and comprising a control unit, a storage memory and a user interface, in that the device and the mobile terminal each comprise wireless communication means so that the device transmits the second video stream (106) to the remote mobile terminal (116), and in that the control unit of the mobile terminal is arranged to, in response to a command entered via the user interface, record in the storage memory of the mobile terminal, at least a part of the second video stream (106) starting at a time prior to the time of entry of the command.