System for identifying sports events in real time and automatically generating playback
By designing a system that recognizes and automatically generates back-watching during sports events, using AI intelligent recognition and statistics module and video processing module, the problem of inefficient generation of traditional back-watching videos is solved, and fast, accurate and comprehensive back-watching video generation is achieved, improving the fairness and viewing of the game.
Patent Information
- Application Number
- CN202411149553.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The generation method of videos for traditional sports events is inefficient and error-prone, and cannot meet the needs of modern sports events for fast, accurate and comprehensive video generation.
Design a system that recognizes and automatically generates back-watching during sports events, including a collection module, AI intelligent recognition and statistics module, video storage and processing module, video synthesis unit and video output and display module. By analyzing video streams in real time, automatically identifying gestures, calculating scores and generating playback videos, fast and accurate return-watching video generation is achieved.
Automatic real-time recognition and scoring statistics are realized, the efficiency of returning videos is improved, manual intervention and error is reduced, the fairness and viewing of the game are enhanced, and controversy caused by misjudgment or omissions are avoided.
Smart Images

Figure CN120126207A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sports event video recognition, and specifically relates to a system for real-time recognition of sports events and automatic generation of replays. Background Art
[0002] With the globalization of sports events and the continuous improvement of the audience's requirements for the quality of the games, replay videos have also become an indispensable part of sports competitions. It can not only help referees make more accurate judgments, but also enhance the viewing experience of the audience. In replay videos, accurately capturing and judging the gestures of players and referees is an important trigger condition for replay videos.
[0003] Currently, traditional gesture recognition mainly relies on manual observation and judgment. This method is not only inefficient, but also easily affected by subjective factors, resulting in misjudgments. Moreover, the traditional method of generating replay videos also mainly relies on manual operations, which has problems such as low efficiency and easy errors. Especially when there are disputes during the game, referees cannot quickly view the videos related to scoring actions, which undoubtedly increases the difficulty of dispute resolution. Therefore, traditional devices cannot meet the requirements of modern sports events for fast, accurate, and comprehensive generation of replay videos. Thus, a system for real-time recognition of sports events and automatic generation of replays is proposed to solve the above problems. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] To solve the problems raised in the above background art, the present invention provides a system for real-time recognition of sports events and automatic generation of replays, which has the advantages of real-time analysis of video images, score statistics update, and real-time overlay of replay videos, and solves the problem that traditional devices cannot meet the requirements of modern sports events for fast, accurate, and comprehensive generation of replay videos.
[0006] (2) Technical Solutions
[0007] To achieve the above object, the present invention provides the following technical solutions: A system for real-time recognition of sports events and automatic generation of replays, comprising:
[0008] An acquisition module for real-time acquisition of video stream data of referees and players;
[0009] An AI intelligent recognition and statistics module, including a processing unit and a statistics unit;
[0010] The processing unit is coupled to the acquisition module, receives the video stream data, performs image processing, gesture separation, feature extraction, and pattern matching, and automatically identifies the gesture types in the video stream;
[0011] The statistical unit is coupled to the processing unit, calculates scores according to the gesture recognition results, and records and updates the scores, achievements, or penalty information of the scoring party and the violating party;
[0012] It further includes:
[0013] A video storage and processing module, coupled to the AI intelligent recognition and statistics module, is used to automatically save video stream segments and generate a playback video of X duration when a scoring or violating gesture type is recognized;
[0014] A video synthesis unit that synthesizes and overlays the live picture and the generated playback video into a single video picture;
[0015] A video output and display module, used to output the video picture to a display device for display.
[0016] In the above technical solution, preferably, the processing unit includes:
[0017] A preprocessing module that receives the video stream data and obtains an image set through image denoising, image conversion, and image enhancement processing respectively;
[0018] A gesture detection and separation module that identifies the area of the gesture that meets the threshold range in the image set according to the preset skin color threshold range, and separates the gesture from the background to obtain a gesture image set;
[0019] A feature extraction and recognition module that receives the gesture image set, recognizes and extracts a gesture feature set, and detects key points of the gesture according to the extracted gesture feature set to obtain a geometric feature image set;
[0020] A gesture analysis module that matches the extracted geometric feature image set with predefined actions and gesture patterns to identify and analyze the gesture type.
[0021] In the above technical solution, preferably, the specific steps for the preprocessing module to perform image denoising, image conversion, and image enhancement processing to obtain an image set include:
[0022] Image denoising: Use a filtering algorithm to process image pixels to remove image noise;
[0023] Image conversion: Convert the image after removing image noise from the BGR color space to the HSV color space.
[0024] Image enhancement: Adjust the gray distribution of the converted image or enhance the contrast in the image.
[0025] In the above technical solution, preferably, the filtering algorithm includes Gaussian filtering and median filtering, and the image enhancement includes histogram equalization and contrast enhancement.
[0026] In the above technical solution, preferably, the gesture detection and separation module identifies and separates the gesture image set to obtain the gesture image set, which specifically includes:
[0027] Skin detection: According to the preset skin color threshold range, the pixels in the image that meet the range are identified, and morphological operations are performed on the identified skin pixels to optimize the boundaries of the skin area;
[0028] Gesture segmentation: Separate gestures from the background through image segmentation algorithms;
[0029] The image segmentation algorithm includes threshold segmentation, edge detection and morphological operation.
[0030] In the above technical solution, preferably, the step of the feature extraction and recognition module identifying and extracting the gesture feature set specifically includes:
[0031] Gesture feature set extraction: Use a deep learning model to extract features from gesture image sets and generate a feature map containing human body structure information;
[0032] Specifically, the deep learning model includes a convolutional neural network (CNN);
[0033] The convolutional neural network (CNN) feature extraction process is as follows:
[0034] The preprocessed gesture image set is input into the convolutional neural network (CNN) model;
[0035] The image data is converted into a feature map containing key information of gestures through convolutional layers, activation functions, and pooling layers.
[0036] The feature map after repeated convolution and pooling is input into the fully connected layer, which maps it to the sample label space for classification or regression tasks to extract the feature map containing human body structure information.
[0037] In the above technical solution, preferably, the convolutional neural network (CNN) model includes:
[0038] at least one convolutional layer for extracting local features from the input image;
[0039] Activation function, connected after each convolutional layer;
[0040] At least one pooling layer to reduce the dimensionality of the feature map and retain important features;
[0041] At least one fully connected layer, located after the convolutional layer and the pooling layer, is used to map the learned feature representation to the sample label space;
[0042] Among them, the convolutional layer includes multiple convolutional kernels, and each convolutional kernel performs a convolution operation on the input image through a sliding window;
[0043] The activation function includes the ReLU function;
[0044] The pooling layer includes a max-pooling operation;
[0045] The fully connected layer flattens the feature map into a one-dimensional vector and outputs the image classification result.
[0046] In the above technical solution, preferably, the steps of the feature extraction and recognition module for detecting the key points of the gesture to obtain the geometric feature image set specifically include:
[0047] Geometric feature image set extraction: Using a deep learning model based on the extracted feature map, detecting the key points of the gesture, and generating a geometric feature image set containing the key point information;
[0048] Specifically, the deep learning model includes the OpenPose model;
[0049] The extraction process of the OpenPose model is as follows:
[0050] The preprocessed gesture feature set is input into the OpenPose model.
[0051] For the input gesture feature set, predict the positions of the key points of the human body in the image through the feature map and the confidence map, and analyze the relationship between the key points through the part affinity field to detect the connection and arrangement of each part of the human body and its limbs;
[0052] Generate a geometric feature image set containing the key point information.
[0053] In the above technical solution, preferably, the gesture analysis module for identifying and analyzing the gesture type specifically includes:
[0054] Match the key points to the corresponding gesture structure to form a complete gesture posture, and identify the gesture as a scoring or violation gesture;
[0055] Through the facial features in the gesture image, match whether the referee / player is in the front / back position;
[0056] Through the relative position features of the hand and the body in the gesture image, match and distinguish the left and right hand positions of the referee's gesture, and output the left and right hand judgment result;
[0057] Through the relative position relationship between the finger joint points in the gesture image, match and determine the number of the referee's fingers.
[0058] (III) Beneficial effects
[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0060] Through the AI intelligent recognition and statistics module, the present invention can accurately identify key information, such as referee actions, competitor actions, etc., and then record in real time information such as the scoring or violating party, the number of scores or violations, etc., achieving the effect of automatic real-time recognition and scoring statistics;
[0061] Moreover, by analyzing the video stream images in real time, it can automatically identify referee actions, competitor actions, etc., accurately grasp the timing of triggering conditions and generate a review video in real time, meeting the requirements for quick response during the game, greatly improving the generation efficiency, reducing manual intervention and errors, and also helping the referee to make accurate judgments in a timely manner during the game, improving the fairness and viewing pleasure of the game, and avoiding disputes caused by misjudgments or omissions;
[0062] At the same time, when automatically generating the review video, it can also superimpose the highlight video on the live or broadcast screen in real time, providing a more intuitive and vivid viewing experience for users and increasing the viewing pleasure of the game. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a system diagram of the present invention;
[0064] Figure 2 is a flowchart of the present invention;
[0065] Figure 3 is a system flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0067] As Figures 1 to 3 shown, the present invention provides a system for real-time recognition and automatic generation of reviews in sports competitions, which is applicable to various scenarios such as jiu-jitsu competitions, ball games, combat sports, and competitive sports arenas, and specifically includes:
[0068] A collection module that obtains referee and competitor video stream data in real time;
[0069] Specifically, the collection module includes multiple cameras, one of which is dedicated to following the referee and capturing his gesture actions in real time; the remaining cameras are used to follow the competitors and capture the real-time images of the game site. Each camera is connected to the AI intelligent recognition and statistics module through a network or data cable to achieve synchronous transmission of video data;
[0070] The AI intelligent recognition and statistics module includes a processing unit and a statistics unit;
[0071] The processing unit is coupled to the acquisition module, receives the video stream data, performs image processing, gesture separation, feature extraction and pattern matching, and automatically identifies the gesture types in the video stream;
[0072] The processing unit includes:
[0073] The preprocessing module receives the video stream data and obtains an image set through image denoising, image conversion and image enhancement processing respectively;
[0074] Preferably, the specific steps for the preprocessing module to perform image denoising, image conversion and image enhancement processing to obtain an image set include:
[0075] Image denoising: Use a filtering algorithm to process the image pixels to remove image noise;
[0076] The filtering algorithm includes Gaussian filtering and median filtering, and the image quality can be improved through various transition operations;
[0077] Image conversion: Convert the image after removing image noise from the BGR color space to the HSV color space, which is convenient for subsequent skin detection and gesture localization.
[0078] Image enhancement: Adjust the gray distribution of the converted image or enhance the contrast in the image.
[0079] Image enhancement includes histogram equalization and contrast enhancement, which can improve the image contrast and make the gesture clearer.
[0080] The gesture detection and separation module identifies the area where the gesture in the image set meets the threshold range according to the preset skin color threshold range, and separates the gesture from the background to obtain a gesture image set;
[0081] Preferably, the gesture detection and separation module to identify and separate to obtain a gesture image set specifically includes:
[0082] Skin detection: In the HSV color space, according to the preset skin color threshold range, identify the pixels in the image that meet the range, and perform morphological operations on the identified skin pixels to optimize the boundary of the skin area;
[0083] Gesture segmentation: Separate the gesture from the background through an image segmentation algorithm;
[0084] The image segmentation algorithm includes threshold segmentation, edge detection and morphological operations.
[0085] The feature extraction and recognition module receives a set of gesture images, recognizes and extracts a set of gesture features, and detects key points of the gesture according to the extracted set of gesture features to obtain a set of geometric feature images;
[0086] Preferably, the steps of the feature extraction and recognition module for recognizing and extracting a set of gesture features specifically include:
[0087] Extraction of gesture feature set: Use a deep learning model to extract features from the set of gesture images to generate a feature map containing human body structure information;
[0088] Specifically, the deep learning model includes a convolutional neural network (CNN);
[0089] The feature extraction process of the convolutional neural network (CNN) is as follows:
[0090] The preprocessed set of gesture images is input into the convolutional neural network (CNN) model;
[0091] The image data is converted into a feature map containing key gesture information through successive convolution, activation, and pooling operations through convolutional layers, activation functions, and pooling layers;
[0092] The feature map after repeated convolution and pooling is input into the fully connected layer, and the fully connected layer maps it to the sample label space for classification or regression tasks to extract a feature map containing human body structure information, such as edges, corners, textures, etc.
[0093] Preferably, the convolutional neural network (CNN) model includes:
[0094] At least one convolutional layer for extracting local features from the input image;
[0095] An activation function connected after each convolutional layer;
[0096] At least one pooling layer for reducing the dimension of the feature map and retaining important features;
[0097] At least one fully connected layer located after the convolutional layer and the pooling layer for mapping the learned feature representation to the sample label space;
[0098] Among them, the convolutional layer includes multiple convolutional kernels, and each convolutional kernel performs a convolution operation on the input image through a sliding window;
[0099] The activation function includes the ReLU function;
[0100] The pooling layer includes a max pooling operation;
[0101] The fully connected layer flattens the feature map into a one-dimensional vector and outputs the image classification result.
[0102] Before the above - constructed Convolutional Neural Network (CNN) model is actually applied, it also includes:
[0103] Collect a large number of image or video data containing various gestures, and these data cover a wide range of scenarios, lighting conditions, and gesture types;
[0104] Clean the collected data, remove noisy, blurred, or duplicate images to ensure data quality;
[0105] Label the gesture images to clarify the key - point positions of each gesture;
[0106] Increase the data volume through methods such as rotation, scaling, and cropping to improve the generalization ability of the model;
[0107] Train the Convolutional Neural Network (CNN) model with the large amount of data collected above. Through the diversity of training data, improve the robustness of the model;
[0108] Preferably, the steps of the feature extraction and recognition module for detecting the key points of the gesture to obtain the geometric feature image set specifically include:
[0109] Geometric feature image set extraction: Use a deep - learning model based on the extracted feature maps to detect the key points of the gesture and generate a geometric feature image set containing key - point information;
[0110] Specifically, the deep - learning model includes the OpenPose model;
[0111] The extraction process of the OpenPose model is as follows:
[0112] The pre - processed gesture feature set is input into the OpenPose model.
[0113] Predict the positions of each key point of the human body in the input gesture feature set through the feature map and confidence map, and analyze the relationship between key points through the part affinity field to detect the connection and arrangement of each part of the human body and its limbs;
[0114] Generate a geometric feature image set containing key - point information
[0115] The gesture analysis module matches the extracted geometric feature image set with predefined action and gesture patterns to identify and analyze the gesture type;
[0116] Preferably, the gesture analysis module for identifying and analyzing the gesture type specifically includes:
[0117] Match the key points to the corresponding gesture structure to form a complete gesture posture, and identify the gesture as a scoring or violating gesture;
[0118] Match the positions of the referee / player (front or back) based on the facial features in the gesture image;
[0119] Match and distinguish the left and right hand positions of the referee's gesture based on the relative position features between the hand and the body in the gesture image, and output the judgment result of the left and right hands;
[0120] Match and determine the number of the referee's fingers based on the relative position relationship between the finger joint points in the gesture image.
[0121] Specifically, based on the predicted key point positions and PAFS, match the key points to the corresponding gesture structures to form a complete gesture posture. Calculate the position of the referee (front or back) through the facial position, so as to distinguish the left and right hands, and further distinguish the competing teams; determine the number of fingers by calculating the relative position relationship between the finger joint points, and then calculate the score value;
[0122] The statistical unit is coupled to the processing unit, calculates the score according to the gesture recognition result, and records and updates the scores and penalty information of the scoring party and the violating party;
[0123] Among them, there is also a database for storing and managing the game data to ensure the accuracy, integrity and traceability of the data;
[0124] It also includes:
[0125] A video storage and processing module, coupled to the AI intelligent recognition and statistics module, is used to automatically save video stream segments and generate a playback video of X duration when a scoring or violating gesture type is recognized;
[0126] Among them, the generated playback video uses an efficient video compression algorithm and storage technology to ensure the clarity and storage efficiency of the video data, and the playback video is saved to the local storage or the cloud server;
[0127] Specifically, automatically save the video stream segment, and save the wonderful moment video by backing up X duration from the recognized ruling action of the referee.
[0128] And this X duration can be set by the system. For different events, the required viewing duration is different. Some scoring actions are longer, and some scoring actions are shorter. The X duration is also the cache record duration of this AI system. The video of this cache record duration can be used for real-time viewing and subsequent viewing;
[0129] A video synthesis unit that synthesizes and overlays the live picture and the generated playback video into a single video picture;
[0130] A video output and display module for outputting the video picture to a display device for display;
[0131] Output multiple video signals to a display device through a video output and display module;
[0132] Specifically, it includes live on-site pictures, score or violation action pictures, and superimposed video pictures, supports multiple output formats and resolutions, and meets the video display requirements in different scenarios and needs;
[0133] The display device includes large screens on the competition site, TVs, mobile phones, screens and other electronic devices that can receive live broadcast pictures.
[0134] This device has completely changed the way of generating video replays in traditional sports events. Previously, manual participation was required, but now it is fully automated. This not only reduces the tediousness and errors of manual operations, but also makes it possible to record scoring actions in real time. When there are disputes during the game, the referee can quickly review the video with disputes and make fair and accurate judgments.
[0135] The working principle and usage process of the present invention:
[0136] Receive the video stream obtained by the real-time acquisition module in real time, perform real-time image processing on the obtained video stream to obtain an image set, obtain a gesture image set through operations such as action recognition, skin detection, and gesture segmentation, use a deep learning model to recognize the referee's gesture actions on the gesture image set, when a gesture is recognized, match it with a predefined action and gesture pattern, recognize the gesture type representing scoring or violation, calculate the specific score value or violation type by recognizing the number of fingers, judge the left or right hand attribute of the gesture to determine the scoring or violating party, and then trigger the video storage and processing module to generate and store a playback video of X duration. After combining it with the live picture and synthesizing it into a video picture, it is output to the display device for display;
[0137] Specifically, the score, score or violation action picture, real-time picture, score / foul action superimposed picture and the original on-site picture are all transmitted to the production control desk and then pushed to the live streaming system;
[0138] The score, score or violation action picture, real-time picture, and score / foul action superimposed picture are displayed on the on-site large screen.
[0139] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0140] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A system for real-time identification and automatic generation of replays of sports events, characterized in that: include: The acquisition module obtains the referee and player video stream data in real time; AI intelligent recognition and statistics module, including processing unit and statistics unit; The processing unit is coupled to the acquisition module, receives the video stream data, performs image processing, gesture separation, feature extraction and pattern matching, and automatically identifies the gesture type in the video stream; The statistical unit is coupled to the processing unit, calculates the score according to the gesture recognition result, and records and updates the score or penalty information of the scoring party and the violating party; Also includes: The video storage and processing module is coupled to the AI intelligent recognition and statistics module, and is used to automatically save the video stream segment and generate a playback video of X length when a scoring or illegal gesture type is recognized; A video synthesis unit synthesizes and superimposes the live picture and the generated playback video into one video picture; The video output and display module is used to output the video image to the display device for display.
2. A system for real-time identification and automatic generation of replays of sports events according to claim 1, characterized in that: The processing unit comprises: A preprocessing module receives the video stream data and obtains an image set by performing image denoising, image conversion and image enhancement respectively; The gesture detection and separation module identifies the gesture area in the image set that meets the threshold range according to the preset skin color threshold range, and separates the gesture from the background to obtain a gesture image set; The feature extraction and recognition module receives the gesture image set and recognizes and extracts the gesture feature set, and extracts the key points of the gesture according to the extracted gesture feature set to obtain a geometric feature image set; The gesture analysis module identifies and analyzes the gesture type by matching the extracted geometric feature image set with the predefined action and gesture patterns.
3. A system for real-time identification and automatic generation of replays of sports events according to claim 2, characterized in that: The specific steps of the preprocessing module performing image denoising, image conversion and image enhancement processing to obtain an image set include: Image denoising: Use filtering algorithms to process image pixels to remove image noise; Image conversion: Convert the image after removing image noise from BGR color space to HSV color space. Image enhancement: Adjust the grayscale distribution of the converted image or enhance the contrast in the image.
4. A system for real-time identification and automatic generation of replays of sports events according to claim 3, characterized in that: The filtering algorithm includes Gaussian filtering and median filtering, and the image enhancement includes histogram equalization and contrast enhancement.
5. A system for real-time identification and automatic generation of replays of sports events according to claim 2, characterized in that: The gesture detection and separation module identifies and separates the gesture image set, which specifically includes: Skin detection: According to the preset skin color threshold range, the pixels in the image that meet the range are identified, and morphological operations are performed on the identified skin pixels to optimize the boundaries of the skin area; Gesture segmentation: Separate gestures from the background through image segmentation algorithms; The image segmentation algorithm includes threshold segmentation, edge detection and morphological operation.
6. A system for real-time identification and automatic generation of replays of sports events according to claim 2, characterized in that: The step of extracting a gesture feature set by the feature extraction and recognition module specifically includes: Gesture feature set extraction: Use a deep learning model to extract features from gesture image sets and generate a feature map containing human body structure information; Specifically, the deep learning model includes a convolutional neural network (CNN); The convolutional neural network (CNN) feature extraction process is as follows: The preprocessed gesture image set is input into the convolutional neural network (CNN) model; The image data is converted into a feature map containing key information of gestures through convolutional layers, activation functions, and pooling layers. The feature map after repeated convolution and pooling is input into the fully connected layer, which maps it to the sample label space for classification or regression tasks to extract the feature map containing human body structure information.
7. A system for real-time identification and automatic generation of replays of sports events according to claim 6, characterized in that: The convolutional neural network (CNN) model includes: at least one convolutional layer for extracting local features from the input image; Activation function, connected after each convolutional layer; At least one pooling layer to reduce the dimensionality of the feature map and retain important features; At least one fully connected layer, located after the convolutional layer and the pooling layer, is used to map the learned feature representation to the sample label space; Among them, the convolution layer includes multiple convolution kernels, each of which performs a convolution operation on the input image through a sliding window; Activation functions include ReLU functions; The pooling layer includes the maximum pooling operation; The fully connected layer flattens the feature map into a one-dimensional vector and outputs the image classification result.
8. A system for real-time identification and automatic generation of replays of sports events according to claim 2, characterized in that: The step of extracting the key points of the gesture detected by the feature extraction and recognition module to obtain a geometric feature image set specifically includes: Extraction of geometric feature image sets: Use a deep learning model to detect the key points of gestures based on the extracted feature maps and generate a geometric feature image set containing key point information; Specifically, the deep learning model includes an OpenPose model; The OpenPose model extraction process is as follows: The preprocessed gesture feature set is input into the OpenPose model. The input gesture feature set is used to predict the positions of key points of the human body in the image through feature maps and confidence maps, and the relationship between key points is analyzed through component affinity fields to detect the connection and arrangement of various parts of the human body and its limbs; Generate a set of geometric feature images containing key point information.
9. The method for real-time scoring and automatic replay generation according to claim 2, characterized in that: The gesture analysis module recognizes and analyzes the gesture type specifically including: Match the key points to the corresponding gesture structure to form a complete gesture posture, and identify the gesture as a scoring or illegal gesture; Match the referee / player's front / back position through facial features in the gesture image; Through the relative position features of the hand and body in the gesture image, the left and right hand positions of the referee's gesture are matched and distinguished, and the left and right hand judgment results are output; The number of referee fingers is determined by matching the relative position relationship between the finger joints in the gesture image.
Citation Information
Patent Citations
Method and device for realizing picture-in-picture playing function
CN103607657A
Gesture recognition method based on fused skin color region segmentation and machine learning algorithm and application thereof
CN108846359A
Competition live broadcast display method and device, equipment and storage medium
CN114339368A
Competition video key clip extraction and description method based on referee gesture behavior
CN118015510A
Play segment extraction method and play segment extraction device
US20180039825A1