Automatic scoring method and device for sports competition based on lightweight AI model
By using lightweight AI models to automatically score in sports where two-person matches, the problems of high hardware costs and high resource consumption in the existing technology are solved, and efficient and accurate automatic scoring is achieved.
Patent Information
- Application Number
- CN202510264713.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-20
AI Technical Summary
The existing automatic scoring technology has high requirements for video image quality, high hardware costs and high resource consumption, making it difficult to achieve efficient automatic scoring in sports where two-person matches.
An automatic scoring method based on a lightweight AI model is adopted to obtain video frames in preset competition areas, and a pre-trained behavior recognition model is used to identify character behaviors, and the score is determined based on the behavior recognition results.
Automatic scoring of designated ball games is realized, which reduces hardware costs and resource consumption, improves identification efficiency and accuracy, and can score without relying on manual labor.
Smart Images

Figure CN120182890A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine vision, and in particular, to an automatic scoring method and device for sports competitions based on a lightweight AI model. Background Art
[0002] Currently, in sports with two-player duels, manual scoring is usually required. For example, if Party A and Party B play a table tennis match, usually a third person who understands the game rules needs to score the match between the two. Then, it requires labor costs and depends on manual experience.
[0003] Therefore, the sports industry has also developed image analysis technology to automatically count the scores of competitions. The existing image analysis technology for automatic scoring often has high requirements for the quality of video images and requires a professional camera to capture high-definition high-frame-rate video images. Not only is the hardware cost high, but the analysis of video images also consumes a lot of resources. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide an automatic scoring method and device for sports competitions based on a lightweight AI model, so as to achieve automatic scoring of a specified ball game, without relying on manual labor, and reduce hardware costs and resource consumption. The specific technical solutions are as follows:
[0005] In a first aspect, the embodiments of the present application provide an automatic scoring method for sports competitions based on a lightweight AI model, and the method includes:
[0006] Obtain each to-be-detected video frame for a preset competition area; wherein, the preset competition area is set for a specified ball game between two participating parties;
[0007] According to the time sequence of each to-be-detected video frame, for each obtained to-be-detected video frame, use a pre-trained behavior recognition model to respectively perform behavior recognition on each person included in the to-be-detected video frame, and obtain a behavior recognition result corresponding to each person; wherein, the behavior recognition model is trained based on a first sample video frame including a person and a first label indicating the behavior of the person in the first sample video frame;
[0008] After recognizing the first behavior recognition result representing a serving behavior, if any behavior recognition result represents a serving behavior, determine the identity authentication result of the person corresponding to the behavior recognition result;
[0009] Add points to the participating party indicated by the obtained identity authentication result.
[0010] Optionally, the method further includes:
[0011] If, after any behavior recognition result indicating a serving behavior is recognized, and no behavior recognition result indicating the next serving behavior is recognized in more than a preset number of video frames to be detected, the participating party indicated by the identity authentication result of the person corresponding to the obtained behavior recognition result is scored again to obtain the final score.
[0012] Optionally, before using the pre-trained behavior recognition model to separately perform behavior recognition on each person included in the video frame to be detected and obtain the behavior recognition result corresponding to each person, the method further includes:
[0013] Input the video frame to be detected into a pre-trained person detection model to obtain the image regions occupied by each person in the video frame to be detected; wherein, the person detection model is trained based on a second sample video frame including a person and a second label indicating the image regions occupied by the person in the second sample video frame.
[0014] The step of using the pre-trained behavior recognition model to separately perform behavior recognition on each person included in the video frame to be detected and obtain the behavior recognition result corresponding to each person includes:
[0015] Input the image region occupied by each person in the video frame to be detected into the pre-trained behavior recognition model to obtain the behavior recognition result corresponding to the person in each image region.
[0016] Optionally, the behavior recognition model is trained in the following manner:
[0017] Obtain a first sample video frame including a person;
[0018] Input the first sample video frame into an initial behavior recognition model to obtain the behavior recognition result of the person in the first sample video frame; wherein, the behavior recognition result is used to represent a serving behavior, a receiving behavior, or a non-playing behavior.
[0019] Calculate a first model loss value based on the difference between the obtained behavior recognition result and the first label;
[0020] Adjust the model parameters of the behavior recognition model based on the first model loss value.
[0021] Optionally, the step of inputting the video frame to be detected into a pre-trained person detection model to obtain the image regions occupied by each person in the video frame to be detected includes:
[0022] Input the video frame to be detected into a pre-trained person detection model to obtain the detection boxes of each person in the video frame to be detected as detection boxes to be utilized.
[0023] For each detection box to be utilized, determine the image region occupied by the expanded detection box in the video frame to be detected, and obtain the image regions occupied by the respective individuals in the video frame to be detected.
[0024] Optionally, if any behavior recognition result indicates a serving behavior, determining the identity authentication result of the person corresponding to the behavior recognition result includes:
[0025] If any behavior recognition result indicates a serving behavior, input the image region occupied by the person corresponding to the behavior recognition result into the feature extraction network in the pre-trained identity authentication model to obtain the feature vector of the person corresponding to the behavior recognition result as the feature vector to be authenticated; wherein, the identity authentication model is trained based on the third sample video frame including a person and the third label indicating the identity of the person in the third sample video frame.
[0026] From the candidate vector library, determine the candidate feature vector with the highest similarity to the feature vector to be authenticated, and determine the identity of the person to which the determined candidate feature vector belongs as the identity authentication result of the person corresponding to the behavior recognition result; wherein, each candidate feature vector included in the candidate vector library is obtained by extracting features from the participants in the specified ball game using the feature extraction network.
[0027] Optionally, the identity authentication model is trained in the following manner:
[0028] Obtain a third sample video frame including a person;
[0029] Input the third sample video frame into the initial identity authentication model to obtain the identity authentication result of the person in the third sample video frame;
[0030] Based on the difference between the obtained identity authentication result and the third label, calculate the second model loss value;
[0031] Based on the second model loss value, adjust the model parameters of the identity authentication model.
[0032] Optionally, the method further includes:
[0033] For each behavior recognition result indicating a serving behavior, determine the first behavior recognition result indicating a non-playing behavior recognized after the behavior recognition result, and use the video frame to which the behavior recognition result indicating a non-playing behavior belongs as the end video frame corresponding to the behavior recognition result indicating a serving behavior.
[0034] If there are multiple consecutive behavior recognition results with the same identity authentication result for the corresponding person among all the behavior recognition results that detect serving behaviors, then select multiple video frames between the video frame to which the first behavior recognition result among the multiple consecutive behavior recognition results belongs and the end video frame corresponding to the last behavior recognition result among the multiple consecutive behavior recognition results;
[0035] Based on the selected multiple video frames, generate an exciting game clip of the person corresponding to the multiple consecutive behavior recognition results.
[0036] In a second aspect, an embodiment of the present application provides an automatic scoring device for a sports game based on a lightweight AI model, and the device includes:
[0037] An acquisition module, configured to acquire each video frame to be detected for a preset game area; wherein, the preset game area is set for a specified ball game between two participating parties;
[0038] A behavior recognition module, configured to, according to the time sequence of each video frame to be detected, for each acquired video frame to be detected, use a pre-trained behavior recognition model to perform behavior recognition on each person included in the video frame to be detected, and obtain a behavior recognition result corresponding to each person; wherein, the behavior recognition model is trained based on a first sample video frame including a person and a first label indicating the behavior of the person in the first sample video frame;
[0039] An identity authentication module, configured to, after recognizing the first behavior recognition result that represents a serving behavior, if any behavior recognition result represents a serving behavior, determine the identity authentication result of the person corresponding to the behavior recognition result;
[0040] A score addition module, configured to add points to the participating party represented by the obtained identity authentication result.
[0041] Optionally, the device further includes:
[0042] A score determination module, configured to, if after recognizing any behavior recognition result that represents a serving behavior, no next behavior recognition result that represents a serving behavior is recognized in more than a preset number of video frames to be detected, add points to the participating party represented by the identity authentication result of the person corresponding to the behavior recognition result again to obtain the final score.
[0043] Optionally, the device further includes:
[0044] A person detection module, configured to input the video frame to be detected into a pre-trained person detection model before the behavior recognition module uses a pre-trained behavior recognition model to respectively perform behavior recognition on each person included in the video frame to be detected, so as to obtain the image regions occupied by the respective persons in the video frame to be detected; wherein, the person detection model is trained based on a second sample video frame including persons and a second label indicating the image regions occupied by the persons in the second sample video frame.
[0045] The behavior recognition module is specifically configured to:
[0046] Input the image region occupied by each person in the video frame to be detected into a pre-trained behavior recognition model to obtain the behavior recognition result corresponding to the person in each image region.
[0047] Optionally, the behavior recognition model is trained in the following manner:
[0048] Obtain a first sample video frame including persons;
[0049] Input the first sample video frame into an initial behavior recognition model to obtain the behavior recognition result of the persons in the first sample video frame; wherein, the behavior recognition result is used to represent a serving behavior, a receiving behavior or a non-playing behavior.
[0050] Calculate a first model loss value based on the difference between the obtained behavior recognition result and the first label;
[0051] Adjust the model parameters of the behavior recognition model based on the first model loss value.
[0052] Optionally, the person detection module includes:
[0053] A detection sub-module, configured to input the video frame to be detected into a pre-trained person detection model to obtain the detection boxes of the respective persons in the video frame to be detected as the detection boxes to be utilized;
[0054] A determination sub-module, configured to determine, for each detection box to be utilized, the image region occupied by the detection box after expansion in the video frame to be detected, so as to obtain the image regions occupied by the respective persons in the video frame to be detected.
[0055] Optionally, the identity authentication module includes:
[0056] A feature extraction sub-module, which is configured to, if it is recognized that any behavior recognition result represents a serving behavior, input the image region occupied by the person corresponding to the behavior recognition result into a feature extraction network in a pre-trained identity authentication model, and obtain a feature vector of the person corresponding to the behavior recognition result as a to-be-authenticated feature vector; wherein, the identity authentication model is trained based on a third sample video frame including a person and a third label indicating the identity of the person in the third sample video frame.
[0057] A matching sub-module, which is configured to determine, from a candidate vector library, a candidate feature vector with the highest similarity to the to-be-authenticated feature vector, and determine the identity of the person to which the determined candidate feature vector belongs as the identity authentication result of the person corresponding to the behavior recognition result; wherein, each candidate feature vector included in the candidate vector library is obtained by performing feature extraction on the participants in the specified ball game using the feature extraction network.
[0058] Optionally, the identity authentication model is trained in the following manner:
[0059] Obtain a third sample video frame including a person.
[0060] Input the third sample video frame into an initial identity authentication model to obtain an identity authentication result of the person in the third sample video frame.
[0061] Based on the difference between the obtained identity authentication result and the third label, calculate a second model loss value.
[0062] Based on the second model loss value, adjust the model parameters of the identity authentication model.
[0063] Optionally, the device further includes:
[0064] An end video frame determination module, which is configured to, for each behavior recognition result representing a serving behavior, determine the first behavior recognition result representing a non-playing behavior recognized after the behavior recognition result, and use the video frame to which the behavior recognition result representing the non-playing behavior belongs as the end video frame corresponding to the behavior recognition result representing the serving behavior.
[0065] A selection module, which is configured to, if there are multiple consecutive behavior recognition results with the same identity authentication result of the corresponding person among all the behavior recognition results representing serving behaviors detected, select multiple video frames between the video frame to which the first behavior recognition result in the multiple consecutive behavior recognition results belongs and the end video frame corresponding to the last behavior recognition result in the multiple consecutive behavior recognition results.
[0066] A generation module, configured to generate an exciting game clip of a person corresponding to the obtained continuous multiple behavior recognition results based on the selected multiple video frames.
[0067] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0068] A memory, configured to store a computer program;
[0069] A processor, configured to implement the automatic scoring method for a sports game based on a lightweight AI model according to any one of the above when executing the program stored in the memory.
[0070] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program implements the automatic scoring method for a sports game based on a lightweight AI model according to any one of the above when being executed by a processor.
[0071] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes executable instructions, and when the executable instructions are executed on a computer, the computer is caused to execute the automatic scoring method for a sports game based on a lightweight AI model according to any one of the above.
[0072] Beneficial effects of the embodiments of the present application:
[0073] The solution provided by the embodiments of the present application performs behavior recognition on each person included in each video frame to be detected according to the time sequence of the video frames to be detected, and behavior recognition results of each person in each video frame to be detected can be obtained. After the first behavior recognition result representing a serving behavior is recognized, if any behavior recognition result represents a serving behavior, it can be determined that the participating party to which the serving behavior belongs is the participating party that won the previous small round. Therefore, after the identity authentication result corresponding to the behavior recognition result representing the serving behavior is determined, points can be added to the participating party represented by the obtained identity authentication result, thereby realizing automatic scoring. It can be seen that this solution can realize automatic scoring for a specified ball game, without relying on manual labor, reducing labor costs; and, compared with the method of predicting the landing point based on the sphere trajectory to realize automatic scoring, it does not require a high-frame-rate and high-resolution video acquisition device to capture the fast-moving sphere, reducing hardware costs; only the behavior and identity of the person need to be recognized to determine whether to score, and the behavior of the person is not instantaneous and usually lasts for a period of time, making it relatively simple to recognize the behavior of the person, with high recognition efficiency and accuracy; in addition, since the area occupied by the person is larger than that of the sphere, and the behavior of the person is slower than that of the moving sphere, a lightweight AI model can be used to complete the behavior recognition of the person, with low resource consumption.
[0074] Of course, it is not necessary for any product or method implementing this application to achieve all of the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0076] Figure 1 Flowchart of an automatic scoring method for a sports game based on a lightweight AI model provided by an embodiment of this application;
[0077] Figure 2A Schematic diagram of a video frame to be detected provided by an embodiment of this application;
[0078] Figure 2B Schematic diagram of a detection box to be utilized provided by an embodiment of this application;
[0079] Figure 3 Flowchart for implementing step S103 in the automatic scoring method for a sports game based on a lightweight AI model provided by an embodiment of this application;
[0080] Figure 4 Flowchart of another automatic scoring method for a sports game based on a lightweight AI model provided by an embodiment of this application;
[0081] Figure 5 Flowchart of a specific example for implementing the automatic scoring method for a sports game based on a lightweight AI model provided by an embodiment of this application;
[0082] Figure 6 Another flowchart of a specific example for implementing the automatic scoring method for a sports game based on a lightweight AI model provided by an embodiment of this application;
[0083] Figure 7A Schematic diagram of a detection box after expansion provided by an embodiment of this application;
[0084] Figure 7B Schematic diagram of another detection box after expansion provided by an embodiment of this application;
[0085] Figure 8A Schematic diagram of a serving behavior provided by an embodiment of this application;
[0086] Figure 8B Schematic diagram of a receiving behavior provided by an embodiment of this application;
[0087] Figure 8C Another schematic diagram of the catching behavior provided by the embodiment of the present application;
[0088] Figure 8D A schematic diagram of a non-playing behavior provided by the embodiment of the present application;
[0089] Figure 8E Another schematic diagram of a non-playing behavior provided by the embodiment of the present application;
[0090] Figure 9 A schematic diagram of the structure of an automatic scoring device for a sports competition based on a lightweight AI model provided by the embodiment of the present application;
[0091] Figure 10 A block diagram of an electronic device for implementing an automatic scoring method for a sports competition based on a lightweight AI model provided by the embodiment of the present application. Detailed implementation manners
[0092] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0093] An automatic scoring method for a sports competition based on a lightweight AI model provided by the embodiment of the present application can be applied to various electronic devices, such as AI (Artificial Intelligence) cameras, personal computers, servers, and other devices with data processing capabilities. In addition, it can be understood that the automatic scoring method for a sports competition based on a lightweight AI model provided by the embodiment of the present application can be implemented in a software, hardware, or a combination of software and hardware manner.
[0094] As Figure 1 shown, the automatic scoring method for a sports competition based on a lightweight AI model provided by the embodiment of the present application includes steps S101 - S104:
[0095] S101, obtaining each to-be-detected video frame for a preset competition area; wherein, the preset competition area is set for a specified ball game between two participating parties;
[0096] In this embodiment, the specified ball game is a ball game in which two participating parties play against each other, such as a table tennis game, a badminton game, or a tennis game, etc. Exemplarily, if the specified ball game is a table tennis game, the preset competition area can be the venue range set for the table tennis game in a sports stadium, an indoor sports hall, or other sports venues.
[0097] Exemplarily, in practical applications, video capture devices for capturing videos of scenes within a preset competition area can be installed in sports venues such as stadiums and indoor gymnasiums. The electronic device that executes the automatic scoring method for sports competitions based on the lightweight AI model provided in this embodiment can obtain video frames from the video capture device at a preset time interval as the video frames to be detected. Exemplarily, if the video capture device captures videos at a frame rate of 24 frames / s, the preset time interval can be 1 / 24 second, 1 / 12 second, 1 / 6 second, etc. The present application embodiment does not limit the preset time interval. In practical applications, the preset time interval can be set by relevant technicians according to requirements.
[0098] It can be understood that in practical applications, before a specified ball game, relevant staff can turn on the video capture device for shooting the preset competition area. After the video capture device is turned on, it captures videos of the scenes within the preset competition area. The electronic device that executes the automatic scoring method for sports competitions based on the lightweight AI model provided in this embodiment obtains each video frame to be detected from the video stream captured by the video capture device.
[0099] The electronic device that executes the automatic scoring method for sports competitions based on the lightweight AI model provided in this embodiment can be the video capture device that captures videos of the scenes within the preset competition area. That is, this electronic device has a video capture function and can not only capture videos but also execute the automatic scoring method for sports competitions based on the lightweight AI model provided in this embodiment. Of course, this electronic device can also be other devices with data processing capabilities that communicate with the video capture device, which are all reasonable.
[0100] In addition, the method for obtaining each video frame to be detected can either obtain the video frames to be detected from the video stream captured by the video capture device at a preset time interval or, after the video capture device captures the complete competition video, obtain each video frame to be detected from each video frame included in the complete competition video at a certain interval. For example, it is reasonable to select a video frame to be detected every two video frames.
[0101] S102. According to the time sequence of each video frame to be detected, for each obtained video frame to be detected, use the pre-trained behavior recognition model to perform behavior recognition on each person included in the video frame to be detected, and obtain the behavior recognition result corresponding to each person. Among them, the behavior recognition model is trained based on the first sample video frames including people and the first labels indicating the behaviors of the people in the first sample video frames.
[0102] Exemplarily, the first tag representing the behavior of the person in the first sample video frame can be obtained by manual annotation, and the behaviors of the person can include "serving", "receiving the ball", and "not playing the ball".
[0103] In this embodiment, the time sequence of each video frame to be detected is the order of acquisition time of each video frame to be detected. After obtaining each video frame to be detected, each video frame to be detected is detected in the order of the video frames to be detected, so as to perform behavior recognition on each person included in each video frame to be detected.
[0104] It can be understood that after the start of the game, the participating personnel of the two participating parties will appear in the preset game area. At this time, performing behavior recognition on each person included in the video frame to be detected can recognize the behaviors of each participating personnel, and thus determine whether each participating personnel is serving, receiving the ball, or in other non-playing behaviors.
[0105] Exemplarily, the method of performing behavior recognition on each person included in each video frame to be detected can be to first identify the image area occupied by each person in the video frame to be detected, and then use the behavior recognition model to perform behavior recognition on the person included in each image area.
[0106] It can be understood that since the area occupied by the person in the video frame to be detected is larger than that of the ball, and the behavior of the person is slower than that of the moving ball, the behavior recognition of the person can be completed by using a lightweight AI model. Therefore, in this embodiment, the behavior recognition model can adopt a lightweight AI model. Exemplarily, the behavior recognition model can adopt the network structure of resnet50 (residual neural network) or lightweight mobilenet (mobile neural network), etc., and the output channels of the network can be modified to adapt to the prediction of three categories of classification, that is, the output channels are 3. For the clarity of the scheme layout, the specific implementation method of performing behavior recognition on each person included in the video frame to be detected in the following embodiments will be introduced, and will not be elaborated here.
[0107] S103, after recognizing the first behavior recognition result representing the serving behavior, if any behavior recognition result represents the serving behavior, determine the identity authentication result of the person corresponding to the behavior recognition result;
[0108] The game rule of the specified ball game in this embodiment is that the winner serves, that is, the participating party that wins a small round gets the serving opportunity for the next small round. It can be understood that the specified ball game is usually divided into multiple small rounds, and one of the participating parties needs to get the specified score to determine the final game result. For example, in a table tennis game, usually one of the participating parties wins the game when reaching 11 points and leading the other participating party by at least 2 points.
[0109] It can be understood that since the serving side of the first small round at the start of the game is randomly determined, for example, by means of a lottery, etc., and the serving side of any small round after the first small round is the participating party that won the previous small round of that small round. Therefore, in this embodiment, the winning or losing situation of the two participating parties in the previous small round can be determined based on any behavior recognition result representing a serving behavior after the first behavior recognition result representing a serving behavior. Thus, points can be added to each participating party according to the winning or losing situation.
[0110] Exemplarily, if after performing behavior recognition on each person included in a video frame to be detected, the behavior recognition result of person A is a serving behavior, and this behavior recognition result is not the first behavior recognition result representing a serving behavior, then the identity authentication result of person A is determined. It can be understood that in practical applications, the identity authentication result of any person can be obtained by performing identity authentication on the person corresponding to the behavior recognition result after recognizing the behavior recognition result representing a serving behavior, or the identity authentication of each person in the video frame to be detected can be performed in advance, and when the serving behavior of person A is recognized, the identity authentication result of person A can be obtained from the previously obtained identity authentication results.
[0111] Exemplarily, before performing behavior recognition on the persons in the video frame to be detected, the identity authentication of each person in the video frame to be detected can be performed first and then stored in a preset storage address. Thus, when determining the identity authentication result of the person corresponding to the behavior recognition result other than the first behavior recognition result representing a serving behavior, it can be obtained from this preset storage address. For example, if a video frame to be detected includes person A, person B, person C, and person D, the identity authentication of each person can be performed first. If the identity authentication result of person A indicates that person A is Xiaoming, the identity authentication result of person B indicates that person B is Xiaohong, the identity authentication result of person C indicates that person C is Xiaoqiang, and the identity authentication result of person D indicates that person D is Xiaofang, then the identity authentication results of each person can be stored in the preset storage address first. After recognizing that the behavior recognition result of person B represents a serving behavior, the identity authentication result of this person B can be obtained from the preset storage address, so as to determine that the person B corresponding to this behavior recognition result is Xiaohong.
[0112] It can be understood that in practical applications, before the game, each participating person of the two participating parties in a specified ball game can first input information that can represent the identity of each participating person, such as the overall image, team uniform information, or facial information of each participating person, etc. So that when performing identity authentication on each person, the identity authentication result of each person can be determined according to the pre-input information. The method of performing identity authentication on a person will be introduced in the following embodiments and will not be elaborated here.
[0113] S104, add points to the participating party indicated by the obtained identity authentication result.
[0114] It can be understood that since the game rule of the specified ball game is that the winner serves, and the serving party of the first small round at the beginning of the game is randomly determined, therefore, after recognizing the first behavior recognition result representing a serving behavior, if any behavior recognition result is recognized as representing a serving behavior, it can be determined that the participating party to which the serving behavior belongs is the participating party that won the previous small round. Therefore, after obtaining the identity authentication result of the person corresponding to the behavior recognition result representing a serving behavior, points can be added to the participating party indicated by the obtained identity authentication result.
[0115] Exemplarily, if the two participating parties in the specified ball game are Party A and Party B respectively, the participating personnel of Party A are Xiaoming and Xiaohong, and the participating personnel of Party B are Xiaoqiang and Xiaofang. When the person corresponding to the behavior recognition result representing a serving behavior is recognized as Xiaofang, and this behavior recognition result is not the first behavior recognition result representing a serving behavior recognized, then points can be added to Party B.
[0116] It can be understood that in practical applications, the participating party to which each participating personnel belongs can be entered when entering the information representing the identity of each participating personnel, so that when the identity authentication results of each person are determined, the participating parties indicated by each identity authentication result can be determined. In addition, when adding points to the participating party indicated by the obtained identity authentication result, 1 point or 2 points, etc. can be added. The specific score for adding points can be set according to the specific game rules of the specified ball game, and the embodiments of the present application do not limit this.
[0117] The solution provided by the embodiments of the present application performs behavior recognition on each person included in each to-be-detected video frame according to the time sequence of each to-be-detected video frame, and the behavior recognition results of each person in each to-be-detected video frame can be obtained. After recognizing the first behavior recognition result representing a serving behavior, if any behavior recognition result is recognized as representing a serving behavior, it can be determined that the participating party to which the serving behavior belongs is the participating party that won the previous small round. Therefore, after determining the identity authentication result of the person corresponding to the behavior recognition result representing a serving behavior, points can be added to the participating party indicated by the obtained identity authentication result, thereby realizing automatic scoring. It can be seen that this solution can realize automatic scoring for the specified ball game, does not rely on manual work, and reduces labor costs.
[0118] Moreover, compared with the method of predicting the landing point based on the spherical trajectory to achieve automatic scoring, this solution does not require high-frame-rate and high-resolution video acquisition equipment to capture the fast-moving sphere, reducing the hardware cost. It only needs to identify the behavior and identity of the person to determine whether a score is achieved. Since the behavior of a person is not instantaneous and usually lasts for a period of time, it is relatively simple to identify the behavior of a person, with high recognition efficiency and accuracy. In addition, since the area occupied by a person is larger than that of the sphere and the behavior of a person is slower than that of the moving sphere, a lightweight AI model can be used to complete the behavior recognition of a person, with low resource consumption.
[0119] In another embodiment of the present application, based on the embodiment shown in Figure 1 the above method further includes:
[0120] If, after recognizing the behavior recognition result of any serve behavior, no behavior recognition result of the next serve behavior is recognized in more than a preset number of video frames to be detected, then the participating party indicated by the identity authentication result of the person corresponding to the obtained behavior recognition result is scored again to obtain the final score.
[0121] It can be understood that during the game, after a small round ends, usually the next small round will quickly start. If, after recognizing the behavior recognition result of any serve behavior, no behavior recognition result of the next serve behavior is recognized in more than a preset number of video frames to be detected, it indicates that the specified ball game has ended. At this time, it can be determined that the participating party to which the person corresponding to the behavior recognition result of the serve behavior belongs has won the current small round again. For example, in a table tennis game, usually a participating party wins the game when it reaches 11 points and leads the other participating party by at least 2 points. If Party A and Party B are playing a game and it is recognized that Party A is serving at present, then the winner of the previous small round is Party A. At this time, if Party A does not win the current small round, then there must be a next small round. Therefore, when identifying each video frame to be detected in sequence, if no behavior recognition result of the next serve behavior is recognized in more than a preset number of video frames to be detected, it can be determined that Party A has won the game again.
[0122] Therefore, if, after recognizing the behavior recognition result of any serve behavior, no behavior recognition result of the next serve behavior is recognized in more than a preset number of video frames to be detected, then the participating party indicated by the identity authentication result of the person corresponding to the behavior recognition result is scored again to obtain the final score.
[0123] Exemplarily, in practical applications, the preset number of video frames to be detected can be the number of video frames whose acquisition time exceeds the halftime duration specified in a designated ball game. For example, if the halftime duration is 2 minutes, then the preset number can be the number of all video frames that can be acquired in 2 minutes. The embodiments of the present application do not limit the preset number.
[0124] It can be seen that through this solution, the final score can be obtained automatically and accurately.
[0125] In another embodiment of the present application, before using the pre-trained behavior recognition model in the above step S102 to perform behavior recognition on each person included in the video frame to be detected and obtain the behavior recognition result corresponding to each person, it further includes:
[0126] Input the video frame to be detected into a pre-trained person detection model to obtain the image regions occupied by each person in the video frame to be detected; wherein, the person detection model is trained based on a second sample video frame containing people and a second label indicating the image regions occupied by people in the second sample video frame.
[0127] In this embodiment, the person detection model can be an object detection model such as YOLOv8 (Real-time Object Detection Algorithm Version 8) or YOLOv11 (Real-time Object Detection Algorithm Version 11). After obtaining the video frame to be detected, input the video frame to be detected into the pre-trained person detection model, and the pre-trained person detection model can detect the human body regions in the video frame to be detected, so as to obtain the image regions occupied by each person.
[0128] Exemplarily, the training process of the person detection model can include: First, obtain a second sample video frame containing people. For example, obtain video data of a badminton game or a table tennis game from an open-source website, and then use the video frames containing people in the video data as the second sample video frames. After obtaining the second sample video frames, the image regions occupied by each person in the second sample video frames can be manually labeled to obtain a second label indicating the image regions occupied by each person in the second sample video frames. Then, input the second sample video frames into the initial person detection model to obtain the image regions occupied by each person detected by the person detection model. Then, according to the coordinate differences between each image region detected by the model and the second label corresponding to the image region, calculate the model loss value by means of an L1 (Mean Absolute Error) loss function or an L2 (Mean Squared Error) loss function, etc. Finally, adjust the model parameters by the backpropagation method of minimizing the model loss value until the model converges, and end the training. Otherwise, continue to train the person detection model using the second sample video frames.
[0129] In one implementation, inputting the video frame to be detected into a pre-trained human detection model to obtain the image regions occupied by each person in the video frame to be detected may include steps A1 - A2:
[0130] A1. Input the video frame to be detected into a pre-trained human detection model to obtain the detection boxes of each person in the video frame to be detected, serving as the detection boxes to be utilized.
[0131] A2. For each detection box to be utilized, determine the image region occupied in the video frame to be detected after expanding the detection box to be utilized, so as to obtain the image regions occupied by each person in the video frame to be detected.
[0132] In this implementation, after inputting the video frame to be detected into the human detection model, the human detection model detects the human body regions in the video frame to be detected and obtains the detection boxes of each person, serving as the detection boxes to be utilized. Exemplarily, as shown in the video frame to be detected Figure 2A and after inputting the video frame to be detected into the human detection model, the output is as shown in Figure 2B wherein the dotted rectangular boxes in Figure 2B are the detection boxes to be utilized.
[0133] It can be understood that since the detection boxes detected by the human detection model usually only contain the human body regions, and in order to accurately identify the behaviors of each person, the positions of objects such as rackets and balls and their relationships with each person should also be considered. Therefore, after obtaining each detection box to be utilized, the detection box to be utilized can also be expanded, and the image region occupied in the video frame to be detected after expanding the detection box to be utilized is determined as the image regions occupied by each person in the video frame to be detected.
[0134] Exemplarily, the method of expanding the detection box to be utilized can be taking the center point of the detection box to be utilized as the reference point and enlarging the detection box to be utilized by a certain multiple, such as 1.5 times or 2 times, etc. Additionally, it can be understood that in another implementation, the detection box obtained by the human detection model detecting the human body region in the video frame to be detected can be directly used as the image region occupied by the person, and this is all reasonable.
[0135] Correspondingly, in this embodiment, in step S102 above, using the pre-trained behavior recognition model to respectively perform behavior recognition on each person included in the video frame to be detected to obtain the behavior recognition result corresponding to each person includes:
[0136] Input the image region occupied by each person in the video frame to be detected into the pre-trained behavior recognition model to obtain the behavior recognition result corresponding to the person in each image region.
[0137] In this embodiment, after obtaining the image regions occupied by each person, the image regions are input into the trained behavior recognition model to use the behavior recognition model to recognize the behavior of the person included in each image region, and obtain the behavior recognition result corresponding to the person in each image region.
[0138] In one implementation manner, the behavior recognition model is trained according to steps B1 - B4:
[0139] B1. Obtain the first sample video frame containing a person;
[0140] B2. Input the first sample video frame into the initial behavior recognition model to obtain the behavior recognition result of the person in the first sample video frame; wherein, the behavior recognition result is used to represent a serving behavior, a receiving behavior, or a non - playing behavior;
[0141] B3. Calculate the first model loss value based on the difference between the obtained behavior recognition result and the first label;
[0142] B4. Adjust the model parameters of the behavior recognition model based on the first model loss value.
[0143] In this implementation manner, video data of a badminton game or a table tennis game can be obtained from an open - source website, and then the video frames containing people in the video data are used as the first sample video frames. After obtaining the first sample video frames, the behaviors of each person in the first sample video frames can be manually labeled to obtain the first labels representing the behaviors of the people in the first sample video frames. Then, the first sample video frames are input into the initial behavior recognition model to obtain the behavior recognition results of each person in the first sample video frames.
[0144] After obtaining the behavior recognition results of each person in the first sample video frames, according to the difference between the behavior recognition result of each person and the corresponding first label of this person, the first model loss value is calculated by using methods such as the L1 loss function or the L2 loss function. Then, the model parameters of the behavior recognition model are adjusted by the backpropagation method of minimizing the first model loss value until the model converges, and the training ends. Otherwise, step B1 can be returned to continue training this behavior recognition model.
[0145] It can be understood that by using the first sample video frames containing people and the first labels representing the behaviors of the people in the first sample video frames to train the behavior recognition model, the trained behavior recognition model can accurately recognize the behaviors of the people in each image region.
[0146] In another embodiment of the present application, as Figure 3As shown in the figure, if any behavior recognition result in the above step S103 represents a serving behavior, determining the identity authentication result of the person corresponding to the behavior recognition result may include steps S301 - S302:
[0147] S301, if any behavior recognition result represents a serving behavior, input the image region occupied by the person corresponding to the behavior recognition result into the feature extraction network in the pre-trained identity authentication model to obtain the feature vector of the person corresponding to the behavior recognition result as the feature vector to be authenticated; wherein, the identity authentication model is trained based on the third sample video frames containing people and the third labels representing the identities of the people in the third sample video frames;
[0148] In this embodiment, if a behavior recognition result representing a serving behavior is recognized, the image region occupied by the person corresponding to the behavior recognition result is input into the feature extraction network to extract the identity features of the person in the image region by using the feature extraction network, and the feature vector of the person is obtained as the feature vector to be authenticated.
[0149] Exemplarily, the identity authentication model can be a classification model based on network structures such as CNN (Convolutional Neural Network) or MLP (Multi-Layer Perceptron). The classification model includes a feature extraction network and a classification network. In practical applications, the first preset number of network layers in the identity authentication model can be used as the feature extraction network, and the network layers after the feature extraction network can be used as the classification network. The number of network layers of the feature extraction network is not limited in the embodiments of the present application. For example, in one implementation, the first 1 / 2 of the network layers in the identity authentication model can be used as the feature extraction network. In addition, the output channels of the classification network can be set according to actual needs. For example, if it is specified that there are 4 participants in total for the two participating parties in a ball game, the output channels of the classification network are 4, so as to achieve four-class classification; if it is specified that there are 2 participants in total for the two participating parties in a ball game, the output channels of the classification network are 2, so as to achieve two-class classification.
[0150] In one implementation, the identity authentication model is trained according to steps C1 - C4:
[0151] C1, obtain the third sample video frames containing people;
[0152] C2, input the third sample video frames into the initial identity authentication model to obtain the identity authentication result of the people in the third sample video frames;
[0153] C3, calculate the second model loss value based on the difference between the obtained identity authentication result and the third label;
[0154] C4. Adjust the model parameters of the identity authentication model based on the second model loss value.
[0155] In this implementation, video data of badminton games or table tennis games can be obtained from open source websites, and then the video frames containing people in the video data are used as the third sample video frames. After obtaining the third sample video frames, the identities of each person in the third sample video frames can be manually labeled to obtain the third labels representing the identities of the people in the third sample video frames. Then, the third sample video frames are input into the initial identity authentication model to obtain the identity authentication results of each person in the third sample video frames.
[0156] After obtaining the identity authentication results of each person in the third sample video frames, according to the difference between the identity authentication result of each person and the corresponding third label of that person, the second model loss value is calculated by using methods such as the L1 loss function or the L2 loss function. Then, the model parameters of the identity authentication model are adjusted through the backpropagation method of minimizing the second model loss value until the model converges and the training ends. Otherwise, step C1 can be returned to continue training the identity authentication model.
[0157] It can be understood that by using the third sample video frames containing people and the third labels representing the identities of the people in the third sample video frames to train the identity authentication model, the trained identity authentication model can accurately authenticate the identities of each person. Then, the feature extraction network in the trained identity authentication model can accurately extract the identity features of each person. So that subsequent matching from the candidate vector library using the extracted feature vector to be authenticated can accurately determine the identity of the person to whom the feature vector to be authenticated belongs.
[0158] S302. Determine the candidate feature vector with the highest similarity to the feature vector to be authenticated from the candidate vector library, and determine the identity of the person to whom the determined candidate feature vector belongs as the identity authentication result corresponding to the behavior recognition result; where each candidate feature vector included in the candidate vector library is obtained by using the feature extraction network to extract features from the participants in a specified ball game.
[0159] After obtaining the feature vector to be identified, the similarity between the feature vector to be identified and each candidate feature vector in the candidate vector library can be calculated, so as to determine the candidate feature vector with the highest similarity to the feature vector to be identified. Exemplarily, the cosine similarity or Euclidean distance between the feature vector to be identified and each candidate feature vector can be calculated respectively as the similarity between the feature vector to be identified and the candidate feature vector. Then, the identity of the person to whom the candidate feature vector with the highest corresponding similarity belongs is determined as the identity authentication result of the person to whom the feature vector to be identified belongs, that is, the identity authentication result of the person corresponding to the behavior recognition result is determined.
[0160] It can be understood that since each candidate feature vector in the candidate vector library is obtained by extracting features from each participant using a feature extraction network, therefore, the candidate feature vector with the highest similarity to the feature vector to be identified in the candidate vector library has the highest probability of being the same as the person to whom the feature vector to be identified belongs. Therefore, the identity of the person to whom the candidate feature vector with the highest similarity to the feature vector to be identified belongs can be determined as the identity authentication result of the person corresponding to the behavior recognition result.
[0161] In another embodiment of the present application, as Figure 4 shown, the above-mentioned automatic scoring method for sports competitions based on a lightweight AI model may further include steps S401-S403:
[0162] S401, for each behavior recognition result representing a serving behavior, determine the first behavior recognition result representing a non-playing behavior recognized after this behavior recognition result, and use the video frame to which the behavior recognition result representing the non-playing behavior belongs as the end video frame corresponding to the behavior recognition result representing the serving behavior;
[0163] It can be understood that the first behavior recognition result representing a non-playing behavior recognized after any behavior recognition result representing a serving behavior can be used as the end moment of the small round corresponding to this serving behavior. Therefore, the video frame to which the behavior recognition result representing the non-playing behavior belongs can be used as the end video frame corresponding to the behavior recognition result representing the serving behavior.
[0164] S402, if there are multiple consecutive behavior recognition results with the same identity authentication result of the corresponding person among all the behavior recognition results representing serving behaviors detected, then select multiple video frames between the video frame to which the first behavior recognition result in the multiple consecutive behavior recognition results belongs and the end video frame corresponding to the last behavior recognition result in the multiple consecutive behavior recognition results;
[0165] It can be understood that if there are multiple consecutive behavior recognition results with the same identity authentication result of the corresponding person among all the behavior recognition results representing serving behaviors, it means that the person has won multiple small rounds consecutively. At this time, multiple video frames within these consecutive small rounds can be selected. Since the video frame to which the first behavior recognition result among these multiple consecutive behavior recognition results belongs can represent the start moment of the first small round among these consecutive small rounds, and the end video frame corresponding to the last behavior recognition result among these consecutive behavior recognition results can represent the end moment of the last small round among these consecutive small rounds, therefore, multiple video frames between the video frame to which the first behavior recognition result among these multiple consecutive behavior recognition results belongs and the end video frame corresponding to the last behavior recognition result among these consecutive behavior recognition results are selected. Subsequently, the selected multiple video frames can be used to generate an exciting game segment of this person, that is, to generate an exciting moment of this person.
[0166] In practical applications, the selection method of these multiple video frames can be set by relevant technical personnel themselves, and the embodiments of this application do not limit this. For example, it is reasonable to select a preset number of consecutive video frames or all video frames from all the video frames between the video frame to which the first behavior recognition result among these multiple consecutive behavior recognition results belongs and the end video frame corresponding to the last behavior recognition result among these consecutive behavior recognition results.
[0167] S403, based on the selected multiple video frames, generate an exciting game segment of the person corresponding to these multiple consecutive behavior recognition results.
[0168] After obtaining the selected multiple video frames, these multiple video frames can be spliced in the order of acquisition time to generate an exciting game segment of the person corresponding to these multiple consecutive behavior recognition results.
[0169] To better understand the automatic scoring method for sports games based on a lightweight AI model provided by the embodiments of this application, a specific example of the automatic scoring method for sports games based on a lightweight AI model provided by the embodiments of this application will be introduced below.
[0170] This example is based on the recognition of the identity and behavior of a person for automatic scoring, and is applicable to ball games with two - player confrontation such as badminton, table tennis, tennis, etc., and the game rule is that the winner serves. For example Figure 5As shown, first, obtain the competition video and collect video frames from the obtained competition video at a preset interval. When collecting video frames, preprocessing can be performed on the video frames, such as identifying the position of the table or court in the video frame and filtering out irrelevant background information, so as to obtain video frames for the preset competition area (corresponding to the video frames to be detected in the above text). Then, use the pre-trained person detection model to perform object detection on the people in the video frame to obtain the detection boxes of each person in the video frame, and expand each obtained detection box, for example, expand it by 1.5 times, to obtain the image area occupied by each person in the video frame. Then, based on the image areas occupied by each person obtained, the bottom image library can be updated, and the behaviors of the people in each image area can be recognized. The behavior recognition results can include serving behavior, receiving behavior, and non-playing behavior. Among them, the receiving behavior and non-playing behavior can be called other behaviors. If it is recognized that a person in any image area is performing a serving behavior, then the identity of the person in that image area is authenticated, and the participating party represented by the authentication result is given points. For example, if there are two participating parties A and B, and the behavior recognition results representing serving behavior in the whole game are: "Initial serve, A serves, B serves, A serves, A serves, A serves, B serves, B serves, A serves", then according to the game rule of the winner serving, 5 points can be counted for A and 3 points for B. And, if no next serving behavior is recognized within a preset time after any serving behavior, then the serving party of that serving behavior can be given points again to obtain the final score. In addition, for identity authentication, that is, by comparing the image area with each image area stored in the bottom image library, the identity authentication result is determined.
[0171] Moreover, during the above scoring process, the exciting moments of each person can also be extracted. The specific process includes: after performing object detection and behavior recognition on each video frame, the behavior recognition results representing serving behavior, receiving behavior, and non-playing behavior can be obtained. Among them, the video frame to which the behavior recognition result representing serving behavior belongs is the video frame at the start moment of a small round, the video frame to which the behavior recognition result representing receiving behavior belongs is the video frame during a small round, and the video frame to which the first behavior recognition result representing non-playing behavior after the serving behavior belongs is the video frame at the end moment of that small round. Then, according to each behavior recognition result, the start and end moments of each small round can be determined. Thus, when it is recognized that A serves, the video frames in the previous round of A's serve can be determined as A's exciting moment (corresponding to the exciting video segment in the above text), and when it is recognized that B serves, the video frames in the previous round of B's serve can be determined as B's exciting moment. In addition, according to the above behavior recognition result, it can be known that A has a consecutive score of 3 times. At this time, these 3 small rounds can also be combined into A's exciting moment.
[0172] The following combines Figure 6A detailed introduction to the scoring process is as follows. Figure 6 As shown, first, obtain the video frames of the sports competition moment to get the video frames as Figure 2A shown. Then, use object detection models such as YOLOv11 (corresponding to the human detection model in the above text) to obtain the human body regions, and get two detection frames. As Figure 2B shown, the dotted rectangle frame located in the upper part of the video frame is detection frame A, and the other dotted rectangle frame located in the lower part is detection frame B. Then, expand each detection frame by a certain multiple, for example, 1.5 times, to include targets such as rackets and balls. Crop the image region in the expanded detection frame region in the video frame to obtain the image regions occupied by each person. Among them, after expanding detection frame A, the image region as Figure 7A shown is obtained, and after expanding detection frame B, the image region as Figure 7B shown is obtained. The obtained image regions can be used for human recognition (corresponding to identity authentication in the above text) and used as template data in the base library to establish the base library; or they can be used for subsequent behavior recognition.
[0173] Before performing behavior recognition, it is necessary to train the behavior recognition model. For example, first collect the competition videos of ball games such as badminton and table tennis, and then obtain the image regions occupied by each person according to the above method, and label the behaviors of the people in each image region. For example, label them as "serving behavior, receiving behavior, and non-playing behavior" to obtain sufficient training data to train the behavior recognition model. When training the behavior recognition model, the resnet50 network or the lightweight mobilenet network can be used to perform the training of the classification task, and modify the output channels of the network to adapt to the three-class prediction. The trained behavior recognition model can recognize the behavior of a person at any time. For example, as Figure 7B and Figure 8A shown, it is recognized as a serving behavior; as Figure 8B and Figure 8C shown, it is recognized as a receiving behavior; as Figure 8D and Figure 8E shown, it is recognized as a non-playing behavior.
[0174] Based on the behavior recognition results of each person, the start and end times of each small round in the whole game can be determined. For example, when a serving behavior is recognized, it can be determined as the start time of a small round. At this time, it can be judged whether it is the first recognized serving behavior to determine whether it is the start of the whole game. If so, it ends. If not, the identity of the person performing the serving behavior can be authenticated. According to the authentication result, it is determined whether it is Team A serving. If so, Team A scored in the previous small round. If not, Team B scored. In addition, in a small round of the game, the video frame to which the behavior recognition result representing the receiving behavior belongs represents the start time of a small round, and the video frame to which the first behavior recognition result representing a non-playing behavior recognized after the serving behavior belongs represents the end time of a small round. Therefore, based on the behavior recognition results, the start and end times of each small round can also be determined, so that the game process of the serving side in the previous small round can be extracted as the highlight moment of the serving side.
[0175] It can be understood that because the movement speeds of many balls are very fast and the balls themselves are relatively small, it is difficult for traditional vision models to recognize the balls, and the analysis and prediction of the ball trajectories are often not accurate enough. If a dedicated camera is used to shoot high-definition, high-frame-rate videos, the analysis will consume a lot of resources and the cost will be high. Since the game rules for fast-moving balls such as badminton, table tennis, and tennis are relatively clear, generally the winner serves, and there are singles for two people or doubles for four people. During the process after serving, playing, and ending, people's behaviors are not instantaneous and will last for a period of time. In this case, it becomes simple to recognize people's behavioral actions, and at the same time, the requirements for the video quality are not high, and an ordinary camera can also complete the recognition function. In addition, there are generally 2 or 4 people playing the ball, and it is also a simple matter to judge people's identities at this time. By establishing a base image library, the use of traditional CNN classification models can achieve this.
[0176] By establishing the information of both sides of the game before the start of the game, assuming Team A and Team B, except for the first small round of the game, the side that serves at other times scores, and all serving information can be easily obtained. Furthermore, the scoring situations of the two teams and the continuous scoring situations can be judged. And, based on the behavior recognition results, the start time and end time of each small round can be determined, so that the highlight moments can be extracted from a complete game video, and the highlight moments of each team's victory can also be extracted by team.
[0177] It can be seen that the strategy of this solution for judging scores based on human identity and behavior information has the advantages of not relying on high-definition videos, not occupying high resources, being fast, and having high accuracy. It can very accurately and efficiently judge the moments of serving, playing, and ending, so as to judge the start and end points of each small round, extract exciting games, and can judge scores, consecutive scores, etc. based on identity and behavior information. It is applicable to various competitions and can be used in sports fields, sports halls, and stadiums. It has low costs, does not rely on high-definition videos, does not require expensive equipment, and ordinary cameras can be used. The number of model parameters used is small, and the computational complexity is low, which can ensure a very fast recognition speed. At the same time, the recognition accuracy of identity and behavior will be very high, so the scoring is accurate, and the extraction of exciting moments is precise. It can use AI as a referee, and can also obtain the exciting scoring moments of each participant, and has a very broad application market and space in both formal sports competitions and amateur entertainment.
[0178] Corresponding to the above method embodiments, the embodiments of the present application further provide an automatic scoring device for a sports competition based on a lightweight AI model, as Figure 9 shown, the device includes:
[0179] An acquisition module 910, configured to acquire each to-be-detected video frame for a preset competition area; wherein, the preset competition area is set for a designated ball game between two participating parties;
[0180] A behavior recognition module 920, configured to, according to the time sequence of each to-be-detected video frame, for each acquired to-be-detected video frame, use a pre-trained behavior recognition model to respectively perform behavior recognition on each person included in the to-be-detected video frame, and obtain a behavior recognition result corresponding to each person; wherein, the behavior recognition model is trained based on a first sample video frame including a person and a first label representing the behavior of the person in the first sample video frame;
[0181] An identity authentication module 930, configured to, after recognizing a behavior recognition result representing a serving behavior for the first time, if any recognized behavior recognition result represents a serving behavior, determine the identity authentication result of the person corresponding to the behavior recognition result;
[0182] A score addition module 940, configured to add points to the participating party represented by the obtained identity authentication result.
[0183] Optionally, the device further includes:
[0184] A score determination module, configured to, if after recognizing any behavior recognition result representing a serving behavior, no behavior recognition result representing a serving behavior is recognized in more than a preset number of to-be-detected video frames, add points to the participating party represented by the identity authentication result of the person corresponding to the obtained behavior recognition result again to obtain the final score.
[0185] Optionally, the device further includes:
[0186] A person detection module, configured to input the video frame to be detected into a pre-trained person detection model before the behavior recognition module 920 uses a pre-trained behavior recognition model to respectively perform behavior recognition on each person included in the video frame to be detected, so as to obtain the image region occupied by each person in the video frame to be detected; wherein, the person detection model is trained based on a second sample video frame including a person and a second label indicating the image region occupied by the person in the second sample video frame;
[0187] The behavior recognition module 920 is specifically configured to:
[0188] Input the image region occupied by each person in the video frame to be detected into a pre-trained behavior recognition model, so as to obtain the behavior recognition result corresponding to the person in each image region.
[0189] Optionally, the behavior recognition model is trained in the following manner:
[0190] Obtain a first sample video frame including a person;
[0191] Input the first sample video frame into an initial behavior recognition model to obtain the behavior recognition result of the person in the first sample video frame; wherein, the behavior recognition result is used to represent a serving behavior, a receiving behavior or a non-playing behavior;
[0192] Calculate a first model loss value based on the difference between the obtained behavior recognition result and the first label;
[0193] Adjust the model parameters of the behavior recognition model based on the first model loss value.
[0194] Optionally, the person detection module includes:
[0195] A detection sub-module, configured to input the video frame to be detected into a pre-trained person detection model to obtain the detection frame of each person in the video frame to be detected as the detection frame to be used;
[0196] A determination sub-module, configured to determine, for each detection frame to be used, the image region occupied by the expanded detection frame in the video frame to be detected, so as to obtain the image region occupied by each person in the video frame to be detected.
[0197] Optionally, the identity authentication module 930 includes:
[0198] A feature extraction sub-module, which is configured to, if it is recognized that any behavior recognition result represents a serving behavior, input the image region occupied by the person corresponding to the behavior recognition result into a feature extraction network in a pre-trained identity authentication model to obtain a feature vector of the person corresponding to the behavior recognition result, and use it as a feature vector to be authenticated; wherein, the identity authentication model is trained based on a third sample video frame including a person and a third label indicating the identity of the person in the third sample video frame.
[0199] A matching sub-module, which is configured to determine, from a candidate vector library, a candidate feature vector with the highest similarity to the feature vector to be authenticated, and determine the identity of the person to which the determined candidate feature vector belongs as the identity authentication result of the person corresponding to the behavior recognition result; wherein, each candidate feature vector included in the candidate vector library is obtained by using the feature extraction network to extract features from the participants in the specified ball game.
[0200] Optionally, the identity authentication model is trained in the following manner:
[0201] Obtain a third sample video frame including a person.
[0202] Input the third sample video frame into an initial identity authentication model to obtain an identity authentication result of the person in the third sample video frame.
[0203] Based on the difference between the obtained identity authentication result and the third label, calculate a second model loss value.
[0204] Based on the second model loss value, adjust the model parameters of the identity authentication model.
[0205] Optionally, the device further includes:
[0206] An end video frame determination module, which is configured to, for each behavior recognition result representing a serving behavior, determine the first behavior recognition result representing a non-playing behavior recognized after the behavior recognition result, and use the video frame to which the behavior recognition result representing the non-playing behavior belongs as the end video frame corresponding to the behavior recognition result representing the serving behavior.
[0207] A selection module, which is configured to, if there are multiple consecutive behavior recognition results with the same identity authentication result of the corresponding person among all the behavior recognition results representing serving behaviors detected, select multiple video frames between the video frame to which the first behavior recognition result in the multiple consecutive behavior recognition results belongs and the end video frame corresponding to the last behavior recognition result in the multiple consecutive behavior recognition results.
[0208] A generation module for generating an exciting game clip of the person corresponding to the continuous multiple behavior recognition results based on the selected multiple video frames.
[0209] In the technical solution of this application, operations such as obtaining, storing, using, processing, transmitting, providing, and disclosing the game video and video frames containing people are all carried out under the condition of obtaining user authorization.
[0210] The embodiment of this application also provides an electronic device, such as Figure 10 shown, including:
[0211] A memory 1001 for storing a computer program;
[0212] A processor 1002 for implementing the steps of any of the above automatic scoring methods for sports games based on a lightweight AI model when executing the program stored on the memory 1001;
[0213] And the above electronic device may further include a communication bus and / or a communication interface, and the processor 1002, the communication interface, and the memory 1001 complete communication with each other through the communication bus.
[0214] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0215] The communication interface is used for communication between the above electronic device and other devices.
[0216] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0217] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0218] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned automatic scoring methods for sports competitions based on a lightweight AI model are implemented.
[0219] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the automatic scoring methods for sports competitions based on a lightweight AI model in the above embodiments.
[0220] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a Solid State Disk (SSD), etc.
[0221] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0222] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the device, electronic device, and computer-readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.
[0223] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.
Claims
1. An automatic scoring method for sports competitions based on a lightweight AI model, characterized in that: The method comprises: Obtaining each to-be-detected video frame for a preset game area; wherein the preset game area is set for two contestants to conduct a designated ball game; According to the time sequence of each video frame to be detected, for each acquired video frame to be detected, a pre-trained behavior recognition model is used to perform behavior recognition on each person contained in the video frame to be detected, and a behavior recognition result corresponding to each person is obtained; wherein the behavior recognition model is trained based on a first sample video frame containing a person and a first label representing the behavior of the person in the first sample video frame; After the first behavior recognition result representing the serving behavior is recognized, if any behavior recognition result is recognized to represent the serving behavior, then the identity authentication result of the person corresponding to the behavior recognition result is determined; Additional points are awarded to the contestant represented by the obtained identity authentication result.
2. The method according to claim 1, characterized in that The method further comprises: If after any behavior recognition result representing the serving behavior is recognized, if the next behavior recognition result representing the serving behavior is not recognized for more than a preset number of video frames to be detected, the contestant represented by the identity authentication result of the person corresponding to the behavior recognition result will be scored again to obtain the final score.
3. The method according to claim 1 or 2, characterized in that: Before performing behavior recognition on each person included in the to-be-detected video frame using the pre-trained behavior recognition model to obtain a behavior recognition result corresponding to each person, the method further includes: Input the video frame to be detected into a pre-trained person detection model to obtain the image area occupied by each person in the video frame to be detected; wherein the person detection model is trained based on a second sample video frame containing the person and a second label representing the image area occupied by the person in the second sample video frame; The pre-trained behavior recognition model is used to perform behavior recognition on each person contained in the video frame to be detected, and the behavior recognition result corresponding to each person is obtained, including: The image area occupied by each person in the video frame to be detected is input into a pre-trained behavior recognition model to obtain the behavior recognition result corresponding to the person in each image area.
4. The method according to claim 3, characterized in that The behavior recognition model is trained in the following manner: Obtain a first sample video frame containing a person; Inputting the first sample video frame into an initial behavior recognition model to obtain a behavior recognition result of the person in the first sample video frame; wherein the behavior recognition result is used to characterize a serving behavior, a receiving behavior, or a non-playing behavior; Calculating a first model loss value based on a difference between the obtained behavior recognition result and the first label; Based on the first model loss value, adjust the model parameters of the behavior recognition model.
5. The method according to claim 3, characterized in that: The step of inputting the video frame to be detected into a pre-trained person detection model to obtain image areas occupied by each person in the video frame to be detected includes: Input the video frame to be detected into a pre-trained person detection model to obtain the detection frame of each person in the video frame to be detected as the detection frame to be used; For each detection frame to be used, the image area occupied by the detection frame to be used after expansion in the video frame to be detected is determined, and the image area occupied by each person in the video frame to be detected is obtained.
6. The method according to claim 3, characterized in that If any behavior recognition result is identified as representing a serving behavior, then determining an identity authentication result of a person corresponding to the behavior recognition result includes: If any behavior recognition result is identified as representing a serving behavior, the image area occupied by the person corresponding to the behavior recognition result is input into a feature extraction network in a pre-trained identity identification model to obtain a feature vector of the person corresponding to the behavior recognition result as a feature vector to be identified; wherein the identity identification model is trained based on a third sample video frame containing the person and a third label representing the identity of the person in the third sample video frame; From the candidate vector library, determine the candidate feature vector with the highest similarity to the feature vector to be identified, and determine the identity of the person to which the determined candidate feature vector belongs as the identity identification result of the person corresponding to the behavior recognition result; wherein each candidate feature vector contained in the candidate vector library is obtained by extracting features of the participants of the designated ball game using the feature extraction network.
7. The method according to claim 6, characterized in that The identity authentication model is trained in the following manner: Obtain a third sample video frame containing a person; Inputting the third sample video frame into an initial identity authentication model to obtain an identity authentication result of the person in the third sample video frame; Calculating a second model loss value based on a difference between the obtained identity authentication result and the third label; Based on the second model loss value, a model parameter of the identity authentication model is adjusted.
8. The method according to claim 1 or 2, characterized in that: The method further comprises: For each behavior recognition result representing the serving behavior, determine the first behavior recognition result representing the non-playing behavior recognized after the behavior recognition result, and use the video frame to which the behavior recognition result representing the non-playing behavior belongs as the end video frame corresponding to the behavior recognition result representing the serving behavior; If there are multiple consecutive behavior recognition results with the same identity authentication result of the corresponding person among all the behavior recognition results representing the serving behavior detected, multiple video frames between the video frame to which the first behavior recognition result among the multiple consecutive behavior recognition results belongs and the ending video frame corresponding to the last behavior recognition result among the multiple consecutive behavior recognition results are selected; Based on the selected multiple video frames, a wonderful game clip of the person corresponding to the multiple consecutive behavior recognition results is generated.
9. An automatic scoring device for sports competitions based on a lightweight AI model, characterized in that: The device comprises: An acquisition module, used to acquire each to-be-detected video frame for a preset game area; wherein the preset game area is set for two contestants to conduct a designated ball game; A behavior recognition module is used to perform behavior recognition on each person contained in each video frame to be detected according to the time sequence of each video frame to be detected, using a pre-trained behavior recognition model, to obtain a behavior recognition result corresponding to each person; wherein the behavior recognition model is trained based on a first sample video frame containing a person and a first label representing the behavior of the person in the first sample video frame; An identity authentication module, for determining, after identifying the first behavior recognition result representing the serving behavior, an identity authentication result of the person corresponding to the behavior recognition result if any behavior recognition result is identified as representing the serving behavior; The scoring module is used to add points to the contestant represented by the obtained identity authentication result.
10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-8 when executing a program stored in a memory.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.