Data processing method, electronic device, storage medium and computer program product

By recognizing referee gestures in live basketball video streams and generating video streams containing referee information, the problem of viewers not being able to understand referee decisions in a timely manner has been solved, improving accuracy and timeliness, and enhancing user engagement and experience.

CN121334408APending Publication Date: 2026-01-13MIGU VIDEO TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511397748.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In live basketball broadcasts, viewers cannot obtain timely and accurate information about the referee's decisions. Current technology relies on commentators to capture the referee's gestures, which is subject to delays and errors, resulting in a poor user experience.

Method used

By receiving video streams, recognizing referee gestures, determining corresponding penalty information, and generating a video stream containing penalty information and an interactive interface, the system displays the video stream to the user in real time, allowing the user to interact with the interface.

Benefits of technology

It improves the accuracy and timeliness of judging gesture recognition, allowing users to obtain accurate judgment information in a timely manner and participate in interaction, thus enhancing user engagement and viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334408A_ABST
    Figure CN121334408A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: receiving a first video stream; identifying a penalty gesture of a referee in the first video stream; determining penalty information corresponding to the penalty gesture of the referee; a second video stream is sent or output, the second video stream comprises the first video stream, the penalty information and a first interactive interface, and the penalty information and the first interactive interface are used for being displayed on a playing interface of the first video stream. The first interaction interface is used for a user to interact with the penalty information and display an interaction result. According to the scheme, the accuracy and timeliness of identifying the penalty gesture can be improved, the user can obtain accurate penalty information in time, the user can interact with the penalty information through the first interaction interface and understand feedback of all users participating in interaction on the penalty information, and the participation sense and watching experience of the user are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a data processing method, electronic device, storage medium, and computer program product. Background Technology

[0002] In basketball live streams, neither commentary-driven nor non-commentary-driven systems can accurately convey the referee's hand signals to the audience in real time. Commentary-driven streams rely on commentators keenly observing the referee's hand signals and relaying the corresponding information to the audience. However, commentators cannot always capture the referee's signals, and there is often a delay in explaining the referee's hand signals, which is prone to errors. Non-commentary streams completely lack this information, making it difficult for viewers to understand the referee's decisions in a timely manner. Summary of the Invention

[0003] To address the related technical problems, embodiments of this application provide a data processing method, an electronic device, a storage medium, and a computer program product.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a data processing method, the method comprising:

[0006] Receive the first video stream;

[0007] Identify the referee's decision-making gestures in the first video stream;

[0008] Determine the penalty information corresponding to the referee's penalty gesture;

[0009] Send or output a second video stream, the second video stream including the first video stream, the penalty information and the first interactive interface, the penalty information and the first interactive interface being used to display on the playback interface of the first video stream, the first interactive interface being used to allow the user to interact with the penalty information and display the interaction result.

[0010] This application also provides an electronic device, including a processor and a memory for storing a computer program capable of running on the processor.

[0011] Wherein, when the processor is used to run the computer program, it executes the steps of any of the above-described data processing methods.

[0012] This application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of any of the above-described data processing methods.

[0013] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described data processing methods.

[0014] The data processing method, electronic device, storage medium, and computer program product provided in this application embodiment receive a first video stream; identify the referee's penalty gesture in the received first video stream; determine the penalty information corresponding to the referee's penalty gesture, which can identify the penalty gesture in the first video stream in real time and accurately; send or output a second video stream, the second video stream including the first video stream, the penalty information, and a first interactive interface, the penalty information and the first interactive interface being displayed on the playback interface of the first video stream, the first interactive interface being used for users to interact with the penalty information and display the interaction results, enabling users to promptly understand accurate penalty information, and to interact with the penalty information through the first interactive interface and obtain the interaction results displayed by the first interactive interface. Compared with related technologies that rely on commentators to capture the referee's penalty gestures, the above solution improves the accuracy and timeliness of identifying penalty gestures, allowing users to obtain accurate penalty information in a timely manner. Users can also participate in interaction through the first interactive interface, express their opinions on the penalty information in real time, and understand the feedback of all participating users on the penalty information, thereby enhancing user participation and viewing experience, and increasing user stickiness. Attached Figure Description

[0015] Figure 1 This is a schematic flowchart of a data processing method according to an embodiment of this application;

[0016] Figure 2 This is a schematic diagram of a first interactive interface according to an embodiment of this application;

[0017] Figure 3 This is a schematic diagram illustrating the gesture phase division in an embodiment of this application;

[0018] Figure 4 This is a schematic diagram illustrating the processing of a first spatiotemporal map sequence using a first model according to an embodiment of this application;

[0019] Figure 5 This is a schematic diagram of another first interactive interface according to an embodiment of this application;

[0020] Figure 6 This is a schematic diagram of yet another first interactive interface according to an embodiment of this application;

[0021] Figure 7 This is a schematic flowchart of a data processing method according to an application embodiment of this application;

[0022] Figure 8 This is a schematic diagram of a data processing device according to an embodiment of this application;

[0023] Figure 9 This is a schematic diagram of the electronic device structure according to an embodiment of this application. Detailed Implementation

[0024] During live basketball broadcasts, viewers cannot immediately access information about the referee's decisions or express their opinions on them. Furthermore, existing gesture recognition methods typically use isolated and static approaches, neglecting the complexity and continuity of gestures, resulting in low accuracy.

[0025] Based on this, in various embodiments of this application, a first video stream is received; the referee's penalty gesture is identified in the received first video stream; penalty information corresponding to the referee's penalty gesture is determined, enabling real-time and accurate identification of the penalty gesture in the first video stream; a second video stream is sent or output, the second video stream including the first video stream, the penalty information, and a first interactive interface, the penalty information and the first interactive interface being displayed on the playback interface of the first video stream, the first interactive interface being used for users to interact with the penalty information and display the interaction results, enabling users to promptly understand accurate penalty information, and to interact with the penalty information through the first interactive interface and obtain the interaction results displayed on the first interactive interface. Compared to related technologies that rely on commentators to capture the referee's penalty gestures, the above solution improves the accuracy and timeliness of identifying penalty gestures, allowing users to obtain accurate penalty information in a timely manner. Users can also participate in interaction through the first interactive interface, express their opinions on the penalty information in real time, and understand the feedback of all participating users on the penalty information, thereby enhancing user participation and viewing experience, and increasing user stickiness.

[0026] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.

[0027] This application provides a data processing method applied to an electronic device, including but not limited to a terminal and / or a server; the terminal may include a display screen or be connected to a display device to play video streams; the server includes but is not limited to one or more of the following: a server corresponding to a video playback application, a server corresponding to a video playback platform, a server corresponding to a live streaming application, and a server corresponding to a broadcasting platform. Figure 1 As shown, the method includes:

[0028] Step 101: Receive the first video stream.

[0029] Here, the first video stream can be a clean stream, which refers to a video stream that does not include narration. For example, an electronic device receives video streams simultaneously captured from multiple perspectives of a venue by multiple cameras positioned around the same location, and processes these video streams (such as video stream fusion, splicing, and / or transcoding) to obtain the first video stream. Here, "camera" can be replaced with "video camera." The venue refers to the physical space where a competition or event takes place, including competition venues and / or event venues; competition venues are specifically designed for hosting competitive activities, such as football fields, basketball courts, sports fields, athletic fields, swimming pools, boxing rings, etc.; these venues have clearly defined boundaries and standard dimensions, facilitating positioning and modeling.

[0030] The first video stream can also be a commentary stream. A commentary stream is a video stream that includes commentary content. For example, the received first video stream can be obtained by adding or inserting commentary content related to the video stream based on the video stream synchronously captured by multiple cameras set up around the same venue from multiple perspectives or shooting angles.

[0031] It should be noted that the first video stream can be received in real time. The first video stream can be understood as a live video stream, which includes, but is not limited to, live video streams of sports events, such as basketball live video streams, football live video streams, badminton live video streams, etc.

[0032] Step 102: Identify the referee's penalty gesture in the first video stream.

[0033] Here, upon receiving the first video stream, the electronic device processes the first video stream and identifies the referee's penalty gesture from it.

[0034] The received first video stream is processed to identify the referee's penalty gestures from the first video stream, including:

[0035] Based on image recognition technology, partial or all of the first video frames containing the referee's image can be identified in the received first video stream; partial or all of the second video frames related to the same gesture of the referee can be extracted from the partial or all of the first video frames containing the referee's image; and the referee's penalty gesture can be identified based on the partial or all of the second video frames related to the same gesture of the referee. The partial or all of the first video frames containing the referee's image can be arranged into a video frame sequence, in which all the first video frames are arranged in order of frame number or in order of the time of capture or reception of the video frames.

[0036] The identification of the referee's penalty gesture based on part or all of the second video frames related to the same gesture can include one or more of the following:

[0037] Image analysis or processing is performed on some or all of the second video frames related to the same gesture of the referee to determine the referee's gesture. The referee's gesture is compared with multiple standard penalty gestures to obtain the comparison results. Based on the comparison results, the standard penalty gesture to which the referee's gesture belongs is determined to obtain the referee's penalty gesture.

[0038] In the second video frame related to the same gesture of the referee, extract multiple images including the referee's gesture. These multiple images can be combined into an image sequence. Based on the multiple images including the referee's gesture and one or more images of multiple standard penalty gestures, calculate the similarity between the referee's gesture and the standard penalty gesture. Identify the standard penalty gesture corresponding to the highest similarity as the referee's penalty gesture.

[0039] The second video frames related to the same referee gesture are converted into multiple image frames played frame by frame to obtain an image frame sequence. Based on this image frame sequence, one or more spatiotemporal maps of the referee's same gesture are constructed. For example, based on the referee image in each image frame of the image frame sequence, the joints of the upper body of the referee are extracted to construct one or more spatiotemporal maps. The number of joints in the upper body includes, but is not limited to, 27. Based on the multiple spatiotemporal maps of the referee's same gesture, the referee's penalty gesture is determined. For example, based on the multiple spatiotemporal maps of the referee's same gesture, the dynamic gesture is determined. The referee's ruling gesture is obtained; or, based on one or more spatiotemporal maps of the referee's gesture and the spatiotemporal map of the standard ruling gesture, the referee's gesture is determined to belong to the standard ruling gesture, thus obtaining the referee's ruling gesture. This can be achieved by determining the referee's ruling gesture based on the similarity or matching degree between one or more spatiotemporal maps of the same referee's gesture and the spatiotemporal map of the standard ruling gesture. For example, the standard ruling gesture with the highest similarity or matching degree can be used as the referee's ruling gesture, or the standard ruling gesture with a similarity or matching degree greater than a set threshold can be used as the referee's ruling gesture. The spatiotemporal map of the gesture represents the positional changes of the hand joints at different times, dynamically indicating the trajectory of the gesture in time and space. In basketball, the standard ruling gesture, also called the reference ruling gesture or benchmark ruling gesture, is pre-defined. The standard ruling gesture can include the standard ruling gestures for the eight common violations and the seven common fouls described below.

[0040] It should be noted that the electronic device can perform penalty gesture recognition on the first video stream received in real time to identify the referee's penalty gestures included in the first video stream. The first video stream may include video frames of at least one penalty gesture by the referee, where "at least one" can be understood as one or more.

[0041] Step 103: Determine the penalty information corresponding to the referee's penalty gesture.

[0042] Here, when the electronic device recognizes the referee's penalty gesture, it can determine the penalty information corresponding to that gesture based on the correspondence or association between standard penalty gestures and penalty information. It should be noted that the correspondence or association between standard penalty gestures and penalty information can be stored in a local database or a cloud database. Different penalty gestures may correspond to different penalty information.

[0043] For example, in a basketball game, the information provided by the referee may include, but is not limited to, one or more of the eight common violations and seven common fouls. Specifically, the eight common violations may include traveling violation, double dribble violation, carrying the ball (palming) violation, 5-second inbound pass violation, 8-second half-court violation, 24-second offensive violation, backcourt violation, and kicking the ball; the seven common fouls may include technical foul, pulling foul, pushing foul, illegal screen by the offensive team, charging foul, blocking foul by the defensive team, and hand-checking foul.

[0044] Step 104: Send or output the second video stream.

[0045] The second video stream includes the first video stream, the penalty information, and the first interactive interface. The penalty information and the first interactive interface are used to display on the playback interface of the first video stream. The first interactive interface is used to allow users to interact with the penalty information and display the interaction results.

[0046] Here, after determining the penalty information corresponding to the referee's penalty gesture, a second video stream is generated based on the first video stream and the penalty information corresponding to the referee's penalty gesture, and then sent or output. For example, the penalty information and relevant information from the first interactive interface are inserted before, after, or between the video frames or image frames related to the referee's penalty gesture in the first video stream to obtain the second video stream, which is then sent or output.

[0047] Sending a second video stream includes sending the second video stream to a client, which may include a video playback client and / or a live streaming client. Outputting the second video stream includes playing the first video stream and displaying penalty information and a first interactive interface on the playback interface of the first video stream; and / or sending the second video stream to a display device or a terminal with a display screen to play the first video stream and display penalty information and a first interactive interface on the playback interface of the first video stream.

[0048] When the electronic device acts as a server, a second video stream is generated based on the first video stream and the ruling information corresponding to the referee's ruling gesture. This second video stream is then sent to the client, allowing the client to display the ruling information and a first interactive interface on the playback screen of the first video stream. This enables users to learn about the ruling information while watching the video and to rate and / or comment on the ruling through the first interactive interface, as well as to obtain the user's overall rating or evaluation of the ruling. The user's overall rating or evaluation of the ruling can be obtained by the server through statistics based on the ratings or evaluations sent by various clients for the same ruling.

[0049] When the electronic device is the terminal, outputting the second video stream can involve playing the first video stream and displaying, on the playback interface of the first video stream, the penalty information corresponding to the referee's penalty gestures identified in the first video stream, as well as displaying a first interactive interface. This allows users to obtain the penalty information while watching the video and to rate and / or comment on the penalty information through the first interactive interface, as well as to obtain the user's overall rating or evaluation of the penalty information. The user's overall rating or evaluation of the penalty information can be obtained by the server based on the ratings or evaluations of users on various terminals regarding the same penalty information.

[0050] It should be noted that users can rate the penalty information through the first interactive interface, such as 2-10 points, and can also view the real-time score of the penalty information displayed on the first interactive interface after rating.

[0051] The first interactive interface can be located in a corner of the video screen, such as the lower left corner, so that users can easily view and operate it.

[0052] For example, such as Figure 2 As shown, the lower left corner of the first video stream playback interface displays the penalty information and the first interactive interface. The first interactive interface can be understood as a rating option or rating button, which determines the user's rating of the penalty information based on the user's interaction with the rating button. For example, if the user selects... Figure 2 The five stars in the chart represent a penalty score of 10 or 5 points, meaning one star corresponds to one score.

[0053] The first interactive interface can also include a comment box and a real-time feedback area. The comment box can receive user-entered comments, and the real-time feedback area can display the overall score, average score, or fluctuating score of the ruling. Considering that in practical applications, a ruling may be made every few minutes during the game, and the duration of a ruling gesture is approximately one minute, in order not to interfere with the user's viewing of the video and to facilitate the user's timely access to ruling information, the display duration of one or more of the following can be set: ruling information is displayed for 10 seconds, the first interactive interface is displayed simultaneously with the ruling information, the first interactive interface is displayed for 15-20 seconds, and the content in the first interactive interface is displayed for 10 seconds. Of course, if there is no ruling information, the first interactive interface may not be displayed.

[0054] In this embodiment, the electronic device receives a first video stream; identifies the referee's penalty gesture in the received first video stream; determines the penalty information corresponding to the referee's penalty gesture, enabling real-time and accurate identification of the penalty gesture in the first video stream; and sends or outputs a second video stream, which includes the first video stream, the penalty information, and a first interactive interface. The penalty information and the first interactive interface are displayed on the playback interface of the first video stream. The first interactive interface allows users to interact with the penalty information and displays the interaction results, enabling users to promptly understand accurate penalty information and interact with the penalty information through the first interactive interface to obtain the interaction results displayed on the first interactive interface. Compared to related technologies that rely on commentators to capture the referee's penalty gestures, the above solution improves the accuracy and timeliness of penalty gesture identification. Users can obtain accurate penalty information in a timely manner, and users can also participate in interaction through the first interactive interface, expressing their opinions on the penalty information in real time and understanding the feedback from all participating users on the penalty information, thereby enhancing user participation and viewing experience and increasing user stickiness.

[0055] In one embodiment, identifying the referee's penalty gesture in the first video stream includes:

[0056] A first spatiotemporal graph sequence is constructed based on the first video stream. A first spatiotemporal graph sequence includes multiple spatiotemporal graphs of the same gesture of the referee.

[0057] Based on the first spatiotemporal graph sequence, the referee's penalty gesture is determined.

[0058] Here, the electronic device can use image recognition technology to identify all second video frames containing the same complete gesture in the received first video stream. Following the chronological order of reception, these second video frames containing the same complete gesture are converted into multiple image frames that are played sequentially, resulting in an image frame sequence, with one gesture corresponding to one image frame sequence. Alternatively, multiple image frames can be extracted from all second video frames containing the same complete gesture according to a set frame rate, a set frame interval (e.g., every other frame), or a set time interval (e.g., 0.1 seconds, 0.2 seconds), thus obtaining an image frame sequence. The set frame rate or frame interval can be set according to actual needs and is not limited here.

[0059] Based on the obtained image frame sequence, multiple spatiotemporal maps of the same gesture of the referee are constructed or extracted to obtain the first spatiotemporal map sequence. For example, based on the referee image in each image frame of the image frame sequence, the joint points of the upper body of the referee image are extracted to construct multiple spatiotemporal maps and obtain the first spatiotemporal map sequence. The spatiotemporal maps in the first spatiotemporal map sequence are arranged according to the arrangement order of the corresponding image frames. Based on the first spatiotemporal map sequence, the referee's judgment gesture is determined.

[0060] Among them, determining the referee's penalty gesture based on the first spatiotemporal graph sequence includes one or more of the following implementation methods:

[0061] Based on the first spatiotemporal graph sequence, a dynamic hand gesture is determined to obtain the referee's penalty gesture;

[0062] A dynamic hand gesture is determined based on the first spatiotemporal map sequence. The determined dynamic hand gesture is compared with multiple standard penalty gestures, and the referee's penalty gesture is determined based on the comparison results.

[0063] Based on the first spatiotemporal graph sequence and the spatiotemporal graph of the standard penalty gesture, the referee's gesture is determined to belong to the standard penalty gesture, thus obtaining the referee's penalty gesture. The referee's penalty gesture can be determined based on the similarity or matching degree between the first spatiotemporal graph sequence and the spatiotemporal graph sequence of the standard penalty gesture. For example, the standard penalty gesture with the highest similarity or matching degree can be used as the referee's penalty gesture, or the standard penalty gesture with a similarity or matching degree greater than a set threshold can be used as the referee's penalty gesture.

[0064] Feature information of key joints of hand gestures in each spatiotemporal map of the first spatiotemporal map sequence is extracted. Based on the feature information of key joints of hand gestures in each spatiotemporal map, the positions of key joints of hand gestures in each spatiotemporal map are predicted, resulting in multiple hand gesture joint position maps. Each hand gesture joint position map corresponds to one spatiotemporal map in the first spatiotemporal map sequence. The similarity between multiple hand gesture joint position maps and the hand gesture joint position maps of multiple standard penalty gestures is calculated, and the referee's penalty gesture is determined based on the similarity. For example, the standard penalty gesture with the highest similarity or matching degree is used as the referee's penalty gesture, or the standard penalty gesture with a similarity or matching degree greater than a set threshold is used as the referee's penalty gesture.

[0065] It should be noted that the first spatiotemporal atlas sequence can be understood as a set of atlas data constructed by jointly modeling the spatial position and temporal changes of multiple hand gesture joints in a gesture action. The first spatiotemporal atlas sequence can at least express the semantic content of the gesture; that is, the first spatiotemporal atlas sequence includes at least the spatiotemporal atlas corresponding to the gesturing stage mentioned below, and may also include the spatiotemporal atlases of other gesture stages besides the gesturing stage. The first spatiotemporal atlas sequence can not only reflect the spatial coordinate information of each hand gesture joint, but also reflect the trajectory of these joints over time.

[0066] A spatiotemporal graph can be represented as a vector matrix of dimension [C1, V], where C1 is the number of channels (which can be 3), representing the two-dimensional position coordinates of the gesture joints and the confidence level of the gesture joints, and V is the number of gesture joints. There are 27 upper body joints that can be selected as gesture joints. The order of the dimensions of C1 and V is not restricted here.

[0067] In the spatiotemporal graph, each node corresponds to a gesture joint. Nodes can be connected by edges to form a topological structure. The connection relationship between nodes can be represented by an adjacency matrix. For example, the connection relationship of nodes in the spatiotemporal graph can be defined according to the connection relationship of gesture joints in the human skeleton. If two joints in the human skeleton are connected, the adjacency matrix can be marked with 1 to indicate that there is an edge connecting the two corresponding nodes. If two joints in the human skeleton are not connected, the adjacency matrix can be marked with 0 to indicate that there is no edge connecting the two corresponding nodes. Of course, custom edges can also be added.

[0068] In this embodiment, a first spatiotemporal map sequence is constructed based on the first video stream. A first spatiotemporal map sequence includes multiple spatiotemporal maps of the same gesture of the referee. This can more accurately capture the overall dynamic features of a referee's gesture. At the same time, compared with directly recognizing the referee's penalty gesture based on the image frame sequence, it can reduce the influence of background noise in the image frame, reduce attention to irrelevant pixels, increase attention to the movement of the gesture joints over time, thereby reducing the amount of data extracted from the image frame sequence and improving the efficiency and accuracy of recognizing the referee's penalty gesture.

[0069] Considering that a referee's hand gestures are not sudden events but unfold over time, following a predictable phased pattern, the gestures can be segmented to further improve the accuracy of recognizing the penalty gestures. Based on this, in one embodiment, determining the referee's penalty gesture based on the first spatiotemporal graph sequence includes:

[0070] Based on the first spatiotemporal map sequence, the gesture action will be segmented to obtain multiple second spatiotemporal map sequences; wherein, the gesture action is at least divided into a preparation stage, a gesture stage and a finishing stage, and one second spatiotemporal map sequence corresponds to one stage of the gesture action;

[0071] Based on the multiple second spatiotemporal map sequences, the referee's penalty gesture is determined.

[0072] Here, the referee's hand gesture is not a sudden event, but rather unfolds over time. The gesture begins when the body part making the gesture leaves its stationary position and ends when that body part returns to its stationary position. This entire process can be considered a single gesture unit, with the stationary position also called the resting position. A gesture unit can be divided into a preparation phase, a gesture phase, and a finishing phase; the gesture unit itself is also called a hand gesture. For example... Figure 3 Gestures can be divided into at least a preparation stage, a gesture stage, and a finishing stage, and may also include a rest stage between different or the same gestures.

[0073] The set of gesture phases A can be represented as A = {Z, B, S, X}, where Z is the preparation phase, B is the gesturing phase, S is the closing phase, and X is the rest phase between gesture units. The preparation phase refers to the warm-up process of the referee's body movements before making a formal call. It typically involves the hand leaving its stationary position and entering the starting posture. The preparation phase can also be understood as the transition from the stationary position of the body part making the gesture to the most meaningful part of the gesture. The gesturing phase is the core part of the gesture, usually expressing the semantic content of the gesture. It typically includes actions by which the referee clearly expresses their intention to make a call, such as raising the arm to indicate a foul or spreading the arms to indicate a traveling violation. The closing phase is the process by which the referee returns the hand to a stationary state after completing the main gesture. It is used to confirm the completeness of the gesture; that is, the closing phase includes the movement of guiding the hand back to its stationary position.

[0074] Since a gesture can be divided into a preparation phase, a gesture phase, and a closing phase, an electronic device can obtain multiple second spatiotemporal map sequences based on multiple standard judgment gestures labeled with these phases, and by segmenting the gestures corresponding to the first spatiotemporal map sequence according to the first spatiotemporal map sequence. The standard judgment gestures labeled with these phases can be stored in the electronic device's local database or a non-local database. Each second spatiotemporal map sequence corresponds to one phase of the gesture. Multiple second spatiotemporal map sequences are arranged in the order of their corresponding phases; for example, multiple second spatiotemporal map sequences can be arranged in the order of preparation, gesture, and closing phases. The second spatiotemporal map sequences are subsets of the first spatiotemporal map sequences. The sum of multiple second spatiotemporal map sequences can be less than or equal to the first spatiotemporal map sequence, and different second spatiotemporal map sequences do not need to have overlapping spatiotemporal maps.

[0075] For example, if the overlap between at least two consecutive spatiotemporal maps in the first spatiotemporal map sequence and the image or spatiotemporal map of the preparation phase of any standard penalty gesture is less than 50% and greater than 20%, or if the overlap between at least two consecutive spatiotemporal maps in the first spatiotemporal map sequence and the starting portion of any standard penalty gesture is less than 50% and greater than 20%, then the corresponding at least two spatiotemporal maps are considered as a second spatiotemporal map sequence, i.e., the second spatiotemporal map sequence corresponding to the preparation phase. If the overlap between at least two spatiotemporal maps in the first spatiotemporal map sequence following the second spatiotemporal map sequence corresponding to the preparation phase and the image or spatiotemporal map of the gesture phase of any standard penalty gesture is greater than 50%, then the corresponding at least two spatiotemporal maps are considered as a second spatiotemporal map sequence, i.e., the second spatiotemporal map sequence corresponding to the preparation phase. In the case of %, at least two corresponding spatiotemporal maps are used as a second spatiotemporal map sequence to obtain the second spatiotemporal sequence of the gesture stage; if the overlap between at least two spatiotemporal maps located after the second spatiotemporal map sequence corresponding to the gesture stage in the first spatiotemporal map sequence and the image or spatiotemporal map of the closing stage of any standard judgment gesture is less than 50%, at least two corresponding spatiotemporal maps are used as a second spatiotemporal map sequence to obtain the second spatiotemporal sequence of the closing stage; of course, it is also possible to first determine the second spatiotemporal map sequences corresponding to the preparation stage and the closing stage in the first spatiotemporal map sequence, and determine the remaining spatiotemporal maps in the first spatiotemporal map sequence as the spatiotemporal map sequence of the gesture stage.

[0076] Considering that electronic devices can extract multiple image frames from all second video frames containing the same complete gesture action frame by frame, according to a set frame rate, set frame interval (e.g., every other frame), or set time interval (e.g., 0.1 seconds, 0.2 seconds), to obtain an image frame sequence, and construct a first spatiotemporal map sequence based on the image frame sequence, and that there is a one-to-one correspondence between the spatiotemporal map in the first spatiotemporal map sequence and the image frames in the image frame sequence, electronic devices can segment the gesture action corresponding to the first spatiotemporal map sequence according to a fixed time window, based on the first spatiotemporal map sequence, to obtain multiple second spatiotemporal map sequences. The duration of the fixed time window can be 0.2 seconds, or it can be set according to actual needs. The fixed time window can be converted to or from time interval, frame rate, or frame interval. One second spatiotemporal map sequence corresponds to one stage of the gesture action, and multiple second spatiotemporal map sequences are arranged in the order of the corresponding stages. For example, multiple second spatiotemporal map sequences can be arranged in the order of preparation stage, gesture stage, and finishing stage.

[0077] The process by which an electronic device segments gesture actions corresponding to a first spatiotemporal graph sequence according to a fixed time window can include:

[0078] Based on the conversion relationship between a fixed time window and time interval, frame rate, or frame interval, a first number of image frames or spatiotemporal maps corresponding to a fixed time window is determined, for example, the first number is N, where N is greater than or equal to 2; N consecutive spatiotemporal maps are determined from the first spatiotemporal map sequence, the N spatiotemporal maps may include the first spatiotemporal map in the first spatiotemporal map sequence; the overlapping portion between the N spatiotemporal maps and the image or spatiotemporal map of the preparation phase of any standard penalty gesture is determined, or the overlapping portion between the first number of spatiotemporal maps and the starting portion of any standard penalty gesture is determined; if the determined overlapping portion satisfies a first set condition, for example... If the overlap is less than 50% but greater than 20%, these N spatiotemporal maps are determined as a second spatiotemporal map sequence, resulting in the second spatiotemporal map sequence for the preparation stage. If the determined overlap does not meet the first set condition, N consecutive spatiotemporal maps are re-determined from the first spatiotemporal map sequence, and the overlap between the N spatiotemporal maps and the image or spatiotemporal map of the preparation stage of any standard penalty gesture is re-determined, or the overlap between the N spatiotemporal maps and the starting part of any standard penalty gesture is determined. Among these, at least one of the re-determined N spatiotemporal maps is different from the most recently determined N spatiotemporal maps.

[0079] After determining the second spatiotemporal map sequence corresponding to the preparation stage, N consecutive spatiotemporal maps are continuously determined in the spatiotemporal maps following the second spatiotemporal map sequence corresponding to the preparation stage in the first spatiotemporal map sequence. If the overlap between the N consecutive spatiotemporal maps and the image or spatiotemporal map of the gesture stage of any standard judgment gesture meets the second set condition, such as the overlap is greater than 50%, these N spatiotemporal maps are taken as a second spatiotemporal map sequence, thereby obtaining the second spatiotemporal sequence of the gesture stage.

[0080] In the first spatiotemporal atlas sequence, after the second spatiotemporal atlas sequence corresponding to the gesture stage, N spatiotemporal atlases are identified. If the overlap between the N spatiotemporal atlases and the image or spatiotemporal atlas of the closing stage of any standard judgment gesture meets the third set condition (e.g., the overlap is less than 50%), these N spatiotemporal atlases are used as a second spatiotemporal atlas sequence to obtain the second spatiotemporal atlas sequence of the closing stage. Alternatively, the second spatiotemporal atlas sequences corresponding to the preparation stage and the closing stage can be determined first in the first spatiotemporal atlas sequence, and the remaining spatiotemporal atlases in the first spatiotemporal atlas sequence can be determined as the spatiotemporal atlas sequence of the gesture stage.

[0081] In practical applications, N consecutive spatiotemporal images from the first spatiotemporal image sequence that overlap with the starting portion of the standard penalty gesture by less than 50% can be identified as the second spatiotemporal image sequence for the preparation phase; N consecutive spatiotemporal images from the first spatiotemporal image sequence that overlap with the standard penalty gesture by more than 50% can be identified as the second spatiotemporal image sequence for the gesture phase; N consecutive spatiotemporal images from the first spatiotemporal image sequence that overlap with the ending portion of the standard penalty gesture by less than 50% can be identified as the second spatiotemporal image sequence for the ending phase; and spatiotemporal images from the first spatiotemporal image sequence that do not overlap with the standard penalty gesture can be identified as the spatiotemporal image sequence for the rest phase. In practical applications, the first spatiotemporal image sequence can correspond to 1 second of the first video stream, and one second spatiotemporal image sequence can correspond to 0.2 seconds of the first video stream.

[0082] When multiple second spatiotemporal map sequences are identified, the electronic device determines the referee's penalty gesture based on these sequences. The determination of the referee's penalty gesture based on multiple second spatiotemporal map sequences includes:

[0083] Based on multiple second spatiotemporal map sequences and the corresponding gesture stages of these sequences, as well as the spatiotemporal map sequences of each gesture stage of multiple standard penalty gestures, a first similarity is calculated between the second spatiotemporal map sequence of each gesture stage and the spatiotemporal map sequence of each standard penalty gesture at the corresponding gesture stage. Based on multiple first similarities corresponding to each standard penalty gesture, the referee's penalty gesture is determined. For example, based on multiple first similarities corresponding to the same standard penalty gesture, a second similarity is calculated. The second similarity can be the mean or weighted average of multiple first similarities. When calculating the weighted average, the weight of the gesture stage can be maximized. Following this method, the second similarities corresponding to all standard penalty gestures are calculated. The standard penalty gesture corresponding to the largest second similarity is taken as the referee's penalty gesture, or any standard penalty gesture corresponding to a second similarity greater than a set threshold is determined as the referee's penalty gesture.

[0084] It should be noted that since the second spatiotemporal map sequence includes one or more spatiotemporal maps, when calculating the similarity between the second spatiotemporal map sequence of the same gesture stage and the spatiotemporal map sequence corresponding to the standard penalty gesture in the corresponding gesture stage, the similarity can be calculated for each spatiotemporal map, and then the average similarity or weighted average similarity can be calculated from the similarity of each spatiotemporal map to obtain the final similarity for this gesture stage. The similarity can be expressed as a score.

[0085] In this embodiment, the gesture is segmented based on the first spatiotemporal map sequence to obtain multiple second spatiotemporal map sequences. The gesture is divided into at least a preparation stage, a gesture stage, and a finishing stage, with one second spatiotemporal map sequence corresponding to one stage of the gesture. Based on the multiple second spatiotemporal map sequences, the referee's penalty gesture is determined. This allows for more refined recognition of penalty gestures at different gesture stages, improving the accuracy of recognizing complex and continuous penalty gestures. It also optimizes the shortcomings of traditional isolated and static gesture recognition, reducing the possibility of misidentifying sudden actions as complete gestures, thereby improving the accuracy and stability of penalty gesture recognition.

[0086] In one embodiment, determining the referee's penalty gesture based on the first spatiotemporal graph sequence includes:

[0087] The first spatiotemporal map sequence is input into the first model to obtain a first image sequence; wherein, each first image in the first image sequence includes the predicted position of each joint key point of the gesture, and one first image corresponds to one spatiotemporal map in the first spatiotemporal map sequence; the first model is used to predict the position of each joint key point of the gesture based on the input multiple spatiotemporal maps;

[0088] Based on the first image sequence and multiple second image sequences, the referee's penalty gesture is determined, and each second image in a second image sequence includes the labeled positions of various key joint points of the same standard penalty gesture.

[0089] Here, the electronic device can input the first spatiotemporal map sequence into the first model for processing to obtain the first image sequence output by the first model; that is, the electronic device can call the first model to process the first spatiotemporal map sequence to obtain the first image sequence, which can reflect the positional changes of each joint key point of the gesture; each first image in the first image sequence can also include the confidence level of the position of each joint key point of the gesture, one first image corresponds to one spatiotemporal map in the first spatiotemporal map sequence, and the order of the first images in the first image sequence is the same as the order of the corresponding spatiotemporal map in the first spatiotemporal map.

[0090] Given a first image sequence, and multiple second image sequences, the second image sequence with the highest matching degree can be determined, and the standard penalty gesture corresponding to the second image sequence with the highest matching degree can be determined as the referee's penalty gesture. The matching degree can be measured by similarity or score.

[0091] For example, based on a first image sequence and multiple second image sequences, a third similarity is calculated between the first image sequence and each of the second image sequences. The standard penalty gesture corresponding to the second image sequence with the highest third similarity is determined as the referee's penalty gesture. The third similarity can be calculated based on a fourth similarity between multiple first images in the first image sequence and multiple second images in the second image sequence.

[0092] For example, based on a first image sequence and multiple second image sequences, a first score is calculated between the first image sequence and each of the second image sequences. The standard penalty gesture corresponding to the second image sequence with the highest first score is determined as the referee's penalty gesture. The first score can be calculated based on the second score between every two first images in the first image sequence and every two second images in the second image sequence. For example, the first score can be the mean or variance calculated based on the second score. The second score can characterize the magnitude of the positional change of each joint key point in every two images; the greater the positional change, the larger the second score.

[0093] For example, using a Conditional Random Field (CRF) probability model, multiple conditional probabilities are calculated based on a first image sequence and multiple second image sequences, with each conditional probability corresponding to a second image sequence. The referee's final ruling gesture is then determined based on these multiple conditional probabilities. For instance, the standard ruling gesture corresponding to the second image sequence with the highest conditional probability is identified as the referee's ruling gesture. CRF is a powerful probability model used to leverage the dependencies between consecutive images for a first image sequence y. (1:n) Given a sequence *s*, we can obtain the set of all possible sequences *Ys*; *s* is a score sequence formed by the position transformation scores of joint keypoints in two adjacent images, and *Ys* is the set of all possible paths after matching the score sequence *s* with multiple second image sequences corresponding to all standard penalty gestures. The first image sequence *y* can be calculated using a CRF probability model. (1:n) Conditional probabilities CRFs(y) (1:n) The expression for the CRF probability model is as follows:

[0094]

[0095] Among them, CRFs(y (1:n) ) Represents the first image sequence y (1:n) Conditional probability, where n is the number of the first images included in the first image sequence; y (i-1) ,y (i) f(y) represents two adjacent images in the first image sequence; (i-1) ,y (i),s) represents the potential function score from the position of each joint key point in the (i-1)th image to the position of each joint key point in the i-th image in the first image sequence. The exponent is obtained by summing the transfer scores of all adjacent positions on the path; Two adjacent images in a second image sequence representing any standard penalty gesture; The potential function score characterizes the position of each joint key point in the (i-1)th image of the second image sequence to the position of each joint key point in the ith image. This can be calculated using the Forward dynamic programming algorithm, which represents the sum of the exponential scores of each path, used as the denominator of the formula to obtain the normalization constant. The conditional probabilities in the above equation can optimize the first image sequence corresponding to all possible paths. In the process of calculating multiple conditional probabilities using a conditional random field probability model and determining the referee's final judgment gesture based on these probabilities, the Viterbi algorithm can be used to solve the combinatorial search problem, determining the sequence with the highest cumulative probability.

[0096] It should be noted that the first model can be trained based on a first sample spatiotemporal map sequence of one or more sample judgment gestures and multiple second image sequences. It is used to predict the positions of key joints of the gesture action based on the input first sample spatiotemporal map sequence. The first model can be trained by an electronic device or other devices. One round of training for the first model is as follows:

[0097] The first sample spatiotemporal map sequence is input into the first model for processing to obtain the third image sequence output by the first model. One third image in the third image sequence corresponds to one sample spatiotemporal map in the first sample spatiotemporal map sequence. Each third image in the third image sequence includes the predicted positions of each joint key point of the same sample judgment gesture.

[0098] Based on the third image sequence and the corresponding second image sequence of standard judgment gestures, a loss value is calculated. The loss value represents the degree of difference between the predicted positions of each joint keypoint in the third image sequence and the labeled positions of each joint keypoint in the second image sequence. It should be noted that the conditional probability of the third image sequence can be calculated using the above formula, and the conditional probability CRFs(y) can be used to calculate the conditional probability. (1:n) The negative log-likelihood is taken as the loss value.

[0099] The model parameters of the first model are updated based on the calculated loss value.

[0100] Among them, if the set convergence conditions are met, the training of the first model is stopped. The set convergence conditions include, but are not limited to, the loss value being less than a set threshold, and / or, the training rounds reaching a set number.

[0101] The first model may include Spatial-Temporal Graph Convolutional Networks (ST-GCNs), a feature fusion module, and Fully Convolutional Networks (FCNNs). The feature fusion module includes, but is not limited to, a Transformer encoder. During the training of the first model, a spatiotemporal graph convolutional network is used to perform convolution processing on each sample spatiotemporal graph in the first sample spatiotemporal graph sequence to obtain a set of vector feature maps corresponding to the first sample spatiotemporal graph sequence. One vector feature map in the set of vector feature maps represents the features of the joint key points in a sample spatiotemporal graph. The feature fusion module is used to perform feature fusion processing on the set of vector feature maps corresponding to the first sample spatiotemporal graph sequence to obtain a set of feature vectors corresponding to the first sample spatiotemporal graph sequence. One feature vector in the set of feature vectors corresponds to a vector feature map. The set of feature vectors includes the position features of each joint key point of the sample penalty gesture. The set of feature vectors is used to indicate the position movement trajectory of each joint key point. A fully connected neural network is used to predict the position of the joint key points based on the set of feature vectors corresponding to the first sample spatiotemporal graph sequence to obtain a third image sequence. The third image sequence indicates a dynamic sample penalty gesture.

[0102] Understandably, ST-GCNs are a type of graph neural network that can effectively process data types such as gesture joint keypoints embedded in spatiotemporal graphs by extending the convolutional operations of traditional CNNs to graph structure data. ST-GCNs can capture the motion patterns of gesture joints in the first spatiotemporal graph sequence through the hierarchical representation of deep neural networks.

[0103] In this embodiment, the first spatiotemporal map sequence is input into the first model to obtain a first image sequence. Each first image in the first image sequence includes the predicted positions of the joint key points of the gesture, and one first image corresponds to one spatiotemporal map in the first spatiotemporal map sequence. The first model is used to predict the positions of the joint key points of the gesture based on the input multiple spatiotemporal maps. It can accurately extract the dynamic features of the gesture and accurately predict the predicted positions of the joint key points of each first image in the first image sequence. The first image sequence output by the first model displays the dynamic gesture corresponding to the first spatiotemporal map sequence. Based on the first image sequence and multiple second image sequences, the referee's penalty gesture is determined. Each second image in a second image sequence includes the labeled positions of the joint key points of the same standard penalty gesture. Higher accuracy recognition of the referee's penalty gesture can be achieved based on the predicted positions of the joint key points of each first image output by the first model.

[0104] In one embodiment, the first model includes a spatiotemporal graph convolutional network, a feature fusion module, and a fully connected neural network; the step of inputting the first spatiotemporal graph sequence into the first model to obtain a first image sequence includes:

[0105] The first spatiotemporal map sequence is input into the spatiotemporal map convolutional network for convolution processing to obtain a set of vector feature maps. One of the vector feature maps in the set of vector feature maps represents the features of the joint key points in a spatiotemporal map.

[0106] The set of vector feature maps are input into the feature fusion module for feature fusion processing to obtain a set of feature vectors, where one feature vector in the set of feature vectors corresponds to one vector feature map.

[0107] The set of feature vectors is input into the fully connected neural network to predict the position of joint key points, thereby obtaining the first image sequence.

[0108] Here, as Figure 4 As shown, the first model includes a spatiotemporal graph convolutional network, a feature fusion module, and a fully connected neural network.

[0109] The electronic device inputs a first spatiotemporal graph sequence into a spatiotemporal graph convolutional network (SPCRN). The SPCRN performs convolution processing on the first spatiotemporal graph sequence, resulting in a set of vector feature maps output by the SPCRN. Specifically, the SPCRN performs graph convolution operations on the first spatiotemporal graph sequence to update the features of key joints in the first spatiotemporal graph. It also fuses the features of key joints from adjacent spatiotemporal graphs through an attention mechanism, and generates a set of vector feature maps that fuse the spatiotemporal feature information of the key joints through multi-layer spatiotemporal convolution and pooling operations. The spatiotemporal feature information includes temporal and spatial features. The first spatiotemporal graph sequence x... (1:n) Inputting a spatiotemporal graph convolutional network (ST-GCNs) yields a set of vector feature maps. (1:n) , can be represented as e (1:n) =ST-GCNs(x (1:n) ), where n represents the number of spatiotemporal maps in the first spatiotemporal map sequence.

[0110] A set of vector feature maps corresponding to the first spatiotemporal atlas sequence is input into the feature fusion module for feature fusion processing, resulting in a set of feature vectors output by the feature fusion module. The feature fusion module converts a set of vector feature maps into a set of feature vectors to capture the spatiotemporal dependencies and semantic information of the joint keypoints in the vector feature maps; this set of feature vectors can be high-dimensional feature vectors, representing the positional features of each joint keypoint of the gesture at different times. The feature fusion module may include a Transformer encoder. The process of inputting a set of vector feature maps corresponding to the first spatiotemporal atlas sequence into the feature fusion module for feature fusion processing to obtain a set of feature vectors can be represented as u (1:n) =TransformerEncoder(e (1:n) ), where u (1:n) Represents a set of feature vectors, TransformerEncoder represents the feature fusion module, e (1:n) This is a set of vector feature maps corresponding to the first spatiotemporal map sequence input.

[0111] A set of feature vectors output by the feature fusion module is input into a fully connected neural network to predict the positions of joint keypoints, resulting in the first image sequence output by the fully connected neural network. This process can be represented as y (1:n) =FCNNs(u (1:n) ), where y (1:n) Let FCNNs represent the first image sequence. The fully connected neural network is used to predict the position of the joint key points and obtain the two-dimensional coordinate position of each joint key point in different first images in the first image sequence.

[0112] In this embodiment, the spatiotemporal graph sequence is convolved using a spatiotemporal graph convolutional network to extract discriminative features of key joints, making subsequent feature fusion and position prediction more accurate and thus improving the overall recognition performance of the model. The feature fusion module extracts the positional features of key joints, capturing spatiotemporal dependencies and semantic information in the sequence to determine the dynamic associations and long-term dependencies between key joints, enhancing the first model's ability to recognize key joints in complex hand gestures. Finally, a fully connected neural network comprehensively utilizes a set of input feature vectors to predict joint positions, resulting in a first image sequence dynamically displaying hand gestures.

[0113] Based on dividing the first spatiotemporal atlas sequence into multiple second spatiotemporal atlas sequences, in one embodiment, the first image sequence includes multiple first sub-sequences, and a second image sequence includes multiple second sub-sequences. Each second sub-sequence corresponds to the labeled positions of key joint points in a stage of a standard judgment gesture. The step of inputting the first spatiotemporal atlas sequence into a first model to obtain the first image sequence includes:

[0114] Input multiple second spatiotemporal map sequences corresponding to the first spatiotemporal map sequence into the first model to obtain multiple first subsequences, with each first subsequence corresponding to a second spatiotemporal map sequence.

[0115] Here, the electronic device can input multiple second spatiotemporal map sequences corresponding to the first spatiotemporal map sequence into the first model to obtain multiple first sub-sequences output by the first model. Among them, multiple first sub-sequences are merged to obtain a first image sequence. The first sub-sequence is an image sequence. One first sub-sequence corresponds to one second spatiotemporal map sequence. Each image in a first sub-sequence includes the predicted positions of each joint key point of a gesture action in a certain stage.

[0116] When an electronic device acquires multiple first sub-sequences (first image sequences) and multiple second image sequences, it determines the referee's penalty gesture based on these multiple first sub-sequences (first image sequences) and multiple second image sequences. A second image sequence comprises multiple second sub-sequences, each a sequence of images, and each second sub-sequence corresponds to the labeled positions of key joint points in a stage of a standard penalty gesture. Determining the referee's penalty gesture based on multiple first sub-sequences and multiple second image sequences may include:

[0117] For each second image sequence, based on the first images in multiple first sub-sequences and the second images in multiple second sub-sequences included in the second image sequence, a first matching degree is calculated between the first sub-sequence and the second sub-sequence corresponding to the same gesture stage. In this way, the first matching degree between the first sub-sequence and the second sub-sequence corresponding to each gesture stage can be calculated. Then, based on all the first matching degrees, a second matching degree is calculated between multiple first sub-sequences and multiple second sub-sequences included in a second image sequence. For example, the second matching degree can be the mean or weighted average of multiple first matching degrees. Following this method, a second matching degree between multiple first sub-sequences and multiple second sub-sequences included in each second image sequence can be calculated. Having calculated the second matching degrees for all second image sequences, the standard penalty gesture corresponding to the second image sequence with the highest second matching degree is determined as the referee's penalty gesture. The matching degree can be measured by similarity, score, or conditional probability. The specific implementation is similar to the implementation method described above for determining the referee's penalty gesture based on the first image sequence and multiple second image sequences, and will not be elaborated here.

[0118] It should be noted that, when the first model includes a spatiotemporal graph convolutional network, a feature fusion module, and a fully connected neural network, inputting multiple second spatiotemporal graph sequences corresponding to the first spatiotemporal graph sequence into the first model yields multiple first sub-sequences, including:

[0119] Multiple second spatiotemporal map sequences are input into a spatiotemporal map convolutional network for convolution processing to obtain multiple sets of vector feature maps. Each set of vector feature maps corresponds to a second spatiotemporal map sequence, and one vector feature map in a set of vector feature maps represents the features of the joint key points in a spatiotemporal map.

[0120] Multiple sets of vector feature maps are input into the feature fusion module for feature fusion processing to obtain multiple sets of feature vectors. One feature vector in a set of feature vectors corresponds to one vector feature map in a set of vector feature maps.

[0121] Multiple sets of feature vectors are input into a fully connected neural network to predict the location of joint key points, resulting in multiple first image sequences.

[0122] In the scenario where gesture actions are segmented, the data processing of the first model, including the spatiotemporal graph convolutional network, feature fusion module, and fully connected neural network, is similar to the data processing of inputting the first spatiotemporal graph sequence into the first model to obtain the first image sequence. Please refer to the relevant description above; it will not be repeated here. The spatiotemporal convolutional network captures joint motion patterns in multiple second spatiotemporal graph sequences, and the feature fusion module captures spatiotemporal dependencies and semantic information in multiple sets of vector feature maps, thereby improving the accuracy of recognizing complex and continuous penalty gestures.

[0123] It should be noted that the first model can be trained based on multiple second sample spatiotemporal map sequences and multiple second image sequences after segmenting the sample judgment gesture. One second sample spatiotemporal map sequence corresponds to one gesture stage of the sample judgment gesture. The training process of one round of the first model is as follows:

[0124] Multiple second sample spatiotemporal map sequences are input into the first model for processing to obtain multiple fourth image sequences output by the first model. A fourth image in a fourth image sequence corresponds to a sample spatiotemporal map in a second sample spatiotemporal map sequence. Each fourth image in the fourth image sequence includes the predicted positions of each joint key point of the same sample judgment gesture.

[0125] Based on multiple fourth image sequences and the corresponding standard judgment gestures, the second image sequence includes multiple second sub-sequences, and a loss value is calculated. The loss value represents the degree of difference between the predicted position of each joint key point in the multiple fourth image sequences and the labeled position of each joint key point in the multiple second sub-sequences included in the second image sequence.

[0126] The model parameters of the first model are updated based on the calculated loss value.

[0127] Among them, if the set convergence conditions are met, the training of the first model is stopped. The set convergence conditions include, but are not limited to, the loss value being less than a set threshold, and / or, the training rounds reaching a set number.

[0128] In the case where the first model includes a spatiotemporal graph convolutional network, a feature fusion module, and a fully connected neural network, the model training process is similar to the process of obtaining multiple first subsequences of gaze using the first model in the scenario of segmenting gesture actions, and will not be elaborated here.

[0129] In this embodiment, in the scenario of segmenting gesture actions, multiple second spatiotemporal map sequences corresponding to the first spatiotemporal map sequence are input into the first model to obtain multiple first sub-sequences. Each first sub-sequence corresponds to a second spatiotemporal map sequence. The first image sequence includes multiple first sub-sequences, and each second image sequence includes multiple second sub-sequences. Each second sub-sequence corresponds to the annotation position of each joint key point of a stage of a standard penalty gesture. This allows for more refined penalty gesture recognition at different gesture stages, improving the accuracy of recognizing complex and continuous penalty gestures, and optimizing the shortcomings of traditional gesture recognition's isolated and static recognition. Furthermore, the recognition of penalty gestures can only end when the preparation stage, gesture stage, and finishing stage of the gesture action are identified, ensuring the integrity of the recognized penalty gesture and reducing the possibility of misidentifying a sudden action as a complete gesture action.

[0130] In one embodiment, the method further includes:

[0131] Obtain one or more pieces of first information input through the first interactive interface, where each piece of first information represents a user's feedback information on the judgment information;

[0132] The second information is determined based on one or more pieces of first information, and the second information represents a comprehensive evaluation of the penalty information.

[0133] Send or output the second information, which is used to display in the first interactive interface.

[0134] Here, while watching the first video stream, users can express their opinions on the rulings through clicking, swiping, or typing on the first interactive interface, generating one or more "first messages." The electronic device acquires these one or more "first messages" from different users through the first interactive interface; each "first message" represents a user's feedback on the ruling. For example, during a basketball live stream, if the referee makes a traveling violation gesture, the electronic device will display the referee's ruling in a corner of the first video stream's playback interface and pop up a rating control, allowing users to select stars to rate the ruling.

[0135] When an electronic device receives one or more pieces of first information, it can determine second information based on all the received first information and send or output the second information. If the electronic device is a server, it can send the second information to a client so that the client can display the second information on the first interactive interface while playing the first video stream. If the electronic device is a terminal, it outputs the second information, that is, it outputs the second information in the playback interface of the first video stream, and the second information can be displayed on the first interactive interface. The methods of outputting the second information include, but are not limited to, numerical displays, chart displays, and text summaries. For example, below the penalty information, the current average score of 8.5 points can be displayed, along with a dynamic score trend chart, which is used to update the score situation over time in real time. In addition, the first interactive interface can also display comment summaries, such as most users believing the penalty was reasonable, and some users questioning the inconsistent standards for judging traveling.

[0136] The electronic device can determine a comprehensive feedback on the penalty information based on feedback from multiple users, such as average score and score distribution. In practical applications, the electronic device can statistically process the rating data of all users to generate visual information such as average score, highest score, lowest score, and score trend charts. Furthermore, the electronic device can combine the comments entered by users through the first interactive interface, use natural language processing technology to extract keywords, and form a concise text summary to help users quickly understand the viewpoints of all participants. Figure 5 As shown, in response to the user's rating input, the first interactive interface can display the current user's rating and the floating total score of all users.

[0137] When the second information is displayed on the first interactive interface, the commentator can adjust the commentary content for the first video stream and / or the penalty information based on the second information; the electronic device acquires the commentary content for the first video stream and / or the penalty information, generates or updates the second video stream based on the commentary content, the first video stream, and the penalty information, and sends or outputs the second video stream. Generating or updating the second video stream based on the commentary content, the first video stream, and the penalty information may include: inserting the acquired commentary content after the video frames in the first video stream that are related to the commentary content.

[0138] In this embodiment, one or more pieces of first information input through the first interactive interface are acquired. Each piece of first information represents a user's feedback on the ruling information. This allows for the timely collection of viewers' genuine opinions on the ruling information, increasing user participation and maintaining user interest and attention during the viewing of the first or second video stream. It also keeps users' attention focused on the playback interface and platform of the first or second video stream. Based on the one or more pieces of first information, second information is determined. This second information represents a comprehensive evaluation of the ruling information. The second information is then sent or output and displayed in the first interactive interface. This provides a more comprehensive reflection of users' attitudes towards the referee's rulings, improving user experience quality. Furthermore, it provides commentators with immediate user feedback for reference, thereby enhancing the richness and interactivity of the video content.

[0139] The following description, using a server as the executing entity as an example and in conjunction with application embodiments, will provide a more detailed description of this application.

[0140] like Figure 7 As shown, the data processing method includes the following steps:

[0141] Step 1: Receive the first video stream.

[0142] For the implementation process of step 1, please refer to the relevant description of step 101 above, which will not be repeated here.

[0143] Step 2: Construct a first spatiotemporal atlas sequence based on the first video stream. A first spatiotemporal atlas sequence includes multiple spatiotemporal atlases of the same gesture of the referee. Input the first spatiotemporal atlas sequence into the first model to obtain the first image sequence.

[0144] In one embodiment, the first image sequence includes multiple first sub-sequences, and a second image sequence includes multiple second sub-sequences. Each second sub-sequence corresponds to the labeled positions of key joint points in a stage of a standard penalty gesture. The first spatiotemporal atlas sequence is input into a first model to obtain the first image sequence, which includes:

[0145] Based on the first spatiotemporal map sequence, the gesture action will be segmented to obtain multiple second spatiotemporal map sequences; wherein, the gesture action is at least divided into a preparation stage, a gesture stage and a finishing stage, and one second spatiotemporal map sequence corresponds to one stage of the gesture action;

[0146] Multiple second spatiotemporal map sequences corresponding to the first spatiotemporal map sequence are input into the first model to obtain multiple first sub-sequences, with each first sub-sequence corresponding to one second spatiotemporal map sequence. The specific implementation process is as follows:

[0147] The server can use image recognition technology to identify all first video frames containing referee images in the received first video stream; extract all second video frames related to the same gesture of the referee from all the first video frames containing referee images; convert all the second video frames related to the same gesture of the referee into multiple image frames that are played frame by frame to obtain an image frame sequence; based on the referee images in each image frame in the image frame sequence, extract 27 joint points of the upper body of the referee image, construct n spatiotemporal maps, and form a first spatiotemporal map sequence, with each spatiotemporal map corresponding to one image frame.

[0148] The server can segment the gesture actions corresponding to the first spatiotemporal graph sequence according to a fixed time window, resulting in m second spatiotemporal graph sequences. Here, the fixed time window can be 0.2 seconds in length. All first spatiotemporal graphs corresponding to the time window in the first spatiotemporal graph sequence are arranged in their original order to form a second spatiotemporal graph sequence. One second spatiotemporal graph sequence corresponds to one stage of the gesture action. The m second spatiotemporal graph sequences are arranged in the order of preparation stage, gesture stage, and finishing stage. It can be understood that the number of spatiotemporal graphs included in the m second spatiotemporal graph sequences is n. The process of determining the gesture stage corresponding to each second spatiotemporal graph sequence is described above and will not be repeated here.

[0149] m second spatiotemporal map sequences are input into the first model. The spatiotemporal convolutional network of the first model performs convolution processing on the m second spatiotemporal map sequences to obtain m sets of vector feature maps. Each set of vector feature maps corresponds to one second spatiotemporal map sequence, and one vector feature map in each set represents the feature of a joint keypoint in a spatiotemporal map. The m sets of vector feature maps are then input into the feature fusion module to obtain m sets of feature vectors. One feature vector in each set of feature vectors corresponds to one vector feature map in each set of vector feature maps. The m sets of feature vectors are then input into a fully connected network to predict the position of the joint keypoints, resulting in m first subsequences. Here, the feature fusion module can be an encoder with a Transformer having four stacked layers.

[0150] The above process can be represented as e (1:m) =ST-GCNs(x (1:m) ), u (1:m) =TransformerEncoder(e (1:m) ), y (1:m) =FCNNs(u (1:m) ), where e (1:m) Let m be the feature maps of the vectors, and ST-GCNs represent the spatiotemporal convolutional network. (m) Let x represent the second spatiotemporal map sequence corresponding to the m-th time window. (1:m) Represents m second spatiotemporal map sequences, u (1:m) Let m be the m sets of vector feature maps, and let TransformerEncoder represent the feature fusion module. (1:m) Let y represent m first subsequences. (m) This represents the predicted location map of multiple gesture joint keypoints corresponding to the m-th time window, and FCNNs represents a fully connected network.

[0151] The first model is used to predict the positions of each joint key point of the gesture based on multiple spatiotemporal maps of the input; each first image in the first image sequence includes the predicted positions of each joint key point of the gesture, and one first image corresponds to one spatiotemporal map in the first spatiotemporal map sequence.

[0152] Step 3: Given multiple first sub-sequences, the server determines the referee's penalty gesture based on these first sub-sequences (first image sequences) and multiple second image sequences. Each second image sequence comprises multiple second sub-sequences, which are image sequences. Each second sub-sequence corresponds to the labeled positions of key joints in a stage of a standard penalty gesture. The specific implementation process for determining the referee's penalty gesture based on multiple first sub-sequences and multiple second image sequences is as follows:

[0153] For each second image sequence, based on the first images in multiple first sub-sequences and the second images in multiple second sub-sequences included in the second image sequence, a first matching degree is calculated between the first sub-sequences and the second sub-sequences corresponding to the same gesture phase. Based on all the first matching degrees, a second matching degree is calculated between multiple first sub-sequences and multiple second sub-sequences included in a second image sequence. Here, the second matching degree can be the average of multiple first matching degrees. Following this method, a second matching degree can be calculated between multiple first sub-sequences and multiple second sub-sequences included in each second image sequence. Having calculated the second matching degrees for all second image sequences, the standard penalty gesture corresponding to the second image sequence with the highest second matching degree is determined as the referee's penalty gesture. Here, the matching degree can be measured using conditional probability calculated by a conditional random field probability model, specifically implemented as follows:

[0154] The server calculates the first matching degree between the first and second subsequences corresponding to the same gesture phase using a conditional random field probability model. For the first subsequence y (1:k) Given a sequence *s*, we can obtain the set of all possible sequences *Ys*. *s* is a score sequence formed by the positional transformation scores of joint keypoints in two adjacent images within the same gesture stage. *Ys* is the set of all possible paths after matching the score sequence *s* with multiple second subsequences corresponding to all standard penalty gestures within the same gesture stage. The first subsequence *y* can be calculated using a CRF probability model. (1:k) Conditional probabilities CRFs(y) (1:k) The expression for the CRF probability model is as follows:

[0155]

[0156] Among them, CRFs(y (1:k) ) represents the first subsequence y (1:k) Conditional probability, k is the number of images in the first subsequence; y (i-1) ,y (i) f(y) represents two adjacent images in the first subsequence; (i-1) ,y (i) ,s) represents the potential function score from the position of each joint key point in the (i-1)th image of the first subsequence to the position of each joint key point in the i-th image. The exponent is obtained by summing the transfer scores of all adjacent positions on the path; The second subsequence represents two adjacent images in any standard penalty gesture during the same gesture phase; The potential function score characterizes the position of each joint key point in the (i-1)th image of the second subsequence to the position of each joint key point in the ith image. It can be calculated based on the Forward dynamic programming algorithm, which means summing up the exponential scores of each path and using them as the denominator of the formula to obtain the normalization constant.

[0157] The above formula can be used to calculate all the first matching degrees. Based on all the first matching degrees, the average of multiple first matching degrees can be calculated to obtain the second matching degree. After calculating the second matching degree for all the second image sequences, the standard penalty gesture corresponding to the second image sequence with the highest second matching degree is determined as the referee's penalty gesture.

[0158] Step 4: Determine the penalty information corresponding to the referee's penalty gesture.

[0159] Once the referee's penalty gesture is recognized, the server can determine the penalty information corresponding to the referee's penalty gesture based on the correspondence between standard penalty gestures and penalty information stored in the local database.

[0160] Step 5: Send a second video stream to the client. The second video stream includes the first video stream, the penalty information, and a first interactive interface. The penalty information and the first interactive interface are used to display on the playback interface of the first video stream. The first interactive interface is used to allow the user to interact with the penalty information and display the interaction results.

[0161] Once the referee's penalty gesture is identified and the corresponding penalty information is determined, the server can insert the penalty information and relevant information from the first interactive interface into the video frames related to the referee's penalty gesture in the first video stream to obtain a second video stream. The server then sends the second video stream to the client to play the first video stream and display the penalty information and the first interactive interface on the playback screen of the first video stream. The display duration of the first interactive interface can be set to 10 seconds.

[0162] Step 6: Obtain one or more pieces of first information input through the first interactive interface, where each piece of first information represents a user's feedback on the judgment information.

[0163] When the client receives the second video stream, it displays the referee's ruling information in a corner of the first video stream's playback interface and pops up a rating control, allowing users to select stars to rate the ruling information.

[0164] The server can obtain one or more initial information entries from different users through the first interactive interface to score the penalty information.

[0165] Step 7: Determine second information based on one or more pieces of first information, wherein the second information represents a comprehensive evaluation of the penalty information.

[0166] The server scores the first piece of information based on one or more pieces of user input regarding the penalty information. It can then calculate the total floating score of all users who participated in the scoring and use this total floating score as the second piece of information.

[0167] Step 8: Send the second message so that the terminal can display the second message on the playback interface of the first video stream.

[0168] The server can send a second message to the client so that the client can display the second message on the first interactive interface while playing the first video stream.

[0169] When the second information is displayed on the first interactive interface, the commentator can adjust the commentary content for the first video stream and / or the penalty information based on the first information and / or the second information; the server obtains the commentary content for the first video stream and / or the penalty information, generates or updates the second video stream based on the commentary content, the first video stream and the penalty information, inserts the obtained commentary content after the video frames related to the commentary content in the first video stream, and sends the updated second video stream to the client, thereby realizing the interaction between the commentator and the user in the studio.

[0170] Based on the first data processing method described above, embodiments of this application also provide a data processing apparatus, such as... Figure 8 As shown, the device includes:

[0171] The receiving unit 801 is used to receive the first video stream;

[0172] The recognition unit 802 is used to recognize the referee's penalty gesture in the first video stream;

[0173] The first determining unit 803 is used to determine the judgment information corresponding to the referee's judgment gesture;

[0174] The output unit 804 is used to send or output a second video stream, the second video stream including the first video stream, the penalty information and the first interactive interface, the penalty information and the first interactive interface are used to display on the playback interface of the first video stream, and the first interactive interface is used to allow the user to interact with the penalty information and display the interaction result.

[0175] In one embodiment, the identification unit 802 is further configured to:

[0176] A first spatiotemporal graph sequence is constructed based on the first video stream. A first spatiotemporal graph sequence includes multiple spatiotemporal graphs of the same gesture of the referee.

[0177] Based on the first spatiotemporal graph sequence, the referee's penalty gesture is determined.

[0178] In one embodiment, the identification unit 802 is further configured to:

[0179] Based on the first spatiotemporal map sequence, the gesture action will be segmented to obtain multiple second spatiotemporal map sequences; wherein, the gesture action is at least divided into a preparation stage, a gesture stage and a finishing stage, and one second spatiotemporal map sequence corresponds to one stage of the gesture action;

[0180] Based on the multiple second spatiotemporal map sequences, the referee's penalty gesture is determined.

[0181] In one embodiment, the identification unit 802 is further configured to:

[0182] The first spatiotemporal map sequence is input into the first model to obtain a first image sequence; wherein, each first image in the first image sequence includes the predicted position of each joint key point of the gesture, and one first image corresponds to one spatiotemporal map in the first spatiotemporal map sequence; the first model is used to predict the position of each joint key point of the gesture based on the input multiple spatiotemporal maps;

[0183] Based on the first image sequence and multiple second image sequences, the referee's penalty gesture is determined, and each second image in a second image sequence includes the labeled positions of various key joint points of the same standard penalty gesture.

[0184] In one embodiment, the first model includes a spatiotemporal graph convolutional network, a feature fusion module, and a fully connected neural network; the first spatiotemporal graph sequence is input into the first model to obtain a first image sequence, and the recognition unit 802 is further configured to:

[0185] The first spatiotemporal map sequence is input into the spatiotemporal map convolutional network for convolution processing to obtain a set of vector feature maps. One of the vector feature maps in the set of vector feature maps represents the features of the joint key points in a spatiotemporal map.

[0186] The set of vector feature maps are input into the feature fusion module for feature fusion processing to obtain a set of feature vectors, where one feature vector in the set of feature vectors corresponds to one vector feature map.

[0187] The set of feature vectors is input into the fully connected neural network to predict the position of joint key points, thereby obtaining the first image sequence.

[0188] In one embodiment, the first image sequence includes multiple first sub-sequences, and a second image sequence includes multiple second sub-sequences. Each second sub-sequence corresponds to the labeled positions of key joint points in a stage of a standard penalty gesture. The recognition unit 802 is further configured to:

[0189] The step of inputting the first spatiotemporal map sequence into the first model to obtain the first image sequence includes:

[0190] Input multiple second spatiotemporal map sequences corresponding to the first spatiotemporal map sequence into the first model to obtain multiple first subsequences, with each first subsequence corresponding to a second spatiotemporal map sequence.

[0191] In one embodiment, the device further includes:

[0192] The acquisition unit is used to acquire one or more pieces of first information input through the first interactive interface, wherein one piece of first information represents a user's feedback information on the judgment information;

[0193] The second determining unit is used to determine second information based on one or more pieces of first information, wherein the second information represents a comprehensive evaluation of the penalty information;

[0194] The output unit 804 is also used to send or output the second information, which is used to display in the first interactive interface.

[0195] In practical applications, the identification unit 802, the first determining unit 803, and the second confirming unit can be implemented by a processor in the data processing device, and the receiving unit 801, the output unit 804, and the acquisition unit can be implemented by a processor in the data processing device combined with a communication interface.

[0196] It should be noted that the data processing apparatus provided in the above embodiments is only illustrated by the division of the above program modules. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the apparatus can be divided into different program modules to complete all or part of the processing described above. In addition, the data processing apparatus and data processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0197] Based on the hardware implementation of the above program modules, and in order to implement the data processing method of the embodiments of this application, the embodiments of this application also provide an electronic device, such as... Figure 9 As shown, the electronic device 900 includes:

[0198] The communication interface 901 enables information exchange with other network nodes;

[0199] The processor 902 is connected to the communication interface 901 to enable information interaction with other network nodes and to execute the methods provided by one or more of the above-described technical solutions when running a computer program. The computer program is stored in the memory 903.

[0200] Specifically, the communication interface 901 is used to receive the first video stream;

[0201] The processor 902 is used to identify the referee's decision-making gestures in the first video stream;

[0202] Determine the penalty information corresponding to the referee's penalty gesture;

[0203] The communication interface 901 is also used to send or output a second video stream, the second video stream including the first video stream, the penalty information and the first interactive interface, the penalty information and the first interactive interface being used to display on the playback interface of the first video stream, and the first interactive interface being used to allow the user to interact with the penalty information and display the interaction result.

[0204] In one embodiment, the processor 902 is further configured to:

[0205] A first spatiotemporal graph sequence is constructed based on the first video stream. A first spatiotemporal graph sequence includes multiple spatiotemporal graphs of the same gesture of the referee.

[0206] Based on the first spatiotemporal graph sequence, the referee's penalty gesture is determined.

[0207] In one embodiment, the processor 902 is further configured to:

[0208] Based on the first spatiotemporal map sequence, the gesture action will be segmented to obtain multiple second spatiotemporal map sequences; wherein, the gesture action is at least divided into a preparation stage, a gesture stage and a finishing stage, and one second spatiotemporal map sequence corresponds to one stage of the gesture action;

[0209] Based on the multiple second spatiotemporal map sequences, the referee's penalty gesture is determined.

[0210] In one embodiment, the processor 902 is further configured to:

[0211] The first spatiotemporal map sequence is input into the first model to obtain a first image sequence; wherein, each first image in the first image sequence includes the predicted position of each joint key point of the gesture, and one first image corresponds to one spatiotemporal map in the first spatiotemporal map sequence; the first model is used to predict the position of each joint key point of the gesture based on the input multiple spatiotemporal maps;

[0212] Based on the first image sequence and multiple second image sequences, the referee's penalty gesture is determined, and each second image in a second image sequence includes the labeled positions of various key joint points of the same standard penalty gesture.

[0213] In one embodiment, the first model includes a spatiotemporal graph convolutional network, a feature fusion module, and a fully connected neural network, and the processor 902 is further configured to:

[0214] The first spatiotemporal map sequence is input into the spatiotemporal map convolutional network for convolution processing to obtain a set of vector feature maps. One of the vector feature maps in the set of vector feature maps represents the features of the joint key points in a spatiotemporal map.

[0215] The set of vector feature maps are input into the feature fusion module for feature fusion processing to obtain a set of feature vectors, where one feature vector in the set of feature vectors corresponds to one vector feature map.

[0216] The set of feature vectors is input into the fully connected neural network to predict the position of joint key points, thereby obtaining the first image sequence.

[0217] In one embodiment, the first image sequence includes multiple first sub-sequences, and a second image sequence includes multiple second sub-sequences. Each second sub-sequence corresponds to the labeled positions of key joint points in a stage of a standard penalty gesture. The processor 902 is further configured to:

[0218] Input multiple second spatiotemporal map sequences corresponding to the first spatiotemporal map sequence into the first model to obtain multiple first subsequences, with each first subsequence corresponding to a second spatiotemporal map sequence.

[0219] In one embodiment, the communication interface 901 is further configured to acquire one or more pieces of first information input through the first interactive interface, wherein one piece of first information represents a user's feedback information on the judgment information;

[0220] The processor 902 is further configured to determine second information based on one or more pieces of first information, wherein the second information represents a comprehensive evaluation of the penalty information;

[0221] The communication interface 901 is also used to send or output the second information, which is used to display in the first interactive interface.

[0222] It should be noted that the specific processing procedures of processor 902 and communication interface 901 can be understood by referring to the above method.

[0223] Of course, in practical applications, the various components in electronic device 900 are coupled together through bus system 904. It can be understood that bus system 904 is used to realize the connection and communication between these components. In addition to a data bus, bus system 904 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in... Figure 9 The general labeled all buses as Bus System 904.

[0224] The memory 903 in this embodiment is used to store various types of data to support the operation of the electronic device 900. Examples of such data include any computer program used to operate on the electronic device 900.

[0225] The methods disclosed in the embodiments of this application can be applied to, or implemented by, the processor 902. The processor 902 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry of the hardware or by instructions in software form within the processor 902. The processor 902 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 902 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, specifically a memory 903. The processor 902 reads information from the memory 903 and, in conjunction with its hardware, completes the steps of the aforementioned method.

[0226] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.

[0227] It is understood that the memory (memory 903) in this embodiment of the application can be volatile memory or non-volatile memory, or both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); the magnetic surface memory can be disk storage or magnetic tape storage. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.

[0228] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 903 storing a computer program, which can be executed by the processor 902 of the electronic device 900 to complete the steps described in the aforementioned data processing method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.

[0229] For example, this application also provides a computer program product, including a computer program that can be executed by a processor 902 of an electronic device 900 to complete the steps of the aforementioned data processing method.

[0230] It should be noted that terms such as "first" and "second" are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. "Multiple" can refer to two or more items, and "multiple" can refer to two or more items. The term "and / or" in this document merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, the term "one or more" in this document refers to any combination of at least two of the multiple elements. For example, including one or more of A, B, and C can represent including any one or at least two or more elements selected from the set consisting of A, B, and C.

[0231] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.

[0232] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.

Claims

1. A data processing method, characterized in that, The method includes: Receive the first video stream; Identify the referee's decision-making gestures in the first video stream; Determine the penalty information corresponding to the referee's penalty gesture; Send or output a second video stream, the second video stream including the first video stream, the penalty information and the first interactive interface, the penalty information and the first interactive interface being used to display on the playback interface of the first video stream, the first interactive interface being used to allow the user to interact with the penalty information and display the interaction result.

2. The method according to claim 1, characterized in that, The step of identifying the referee's decision-making gesture in the first video stream includes: A first spatiotemporal graph sequence is constructed based on the first video stream. A first spatiotemporal graph sequence includes multiple spatiotemporal graphs of the same gesture of the referee. Based on the first spatiotemporal graph sequence, the referee's penalty gesture is determined.

3. The method according to claim 2, characterized in that, The determination of the referee's penalty gesture based on the first spatiotemporal graph sequence includes: Based on the first spatiotemporal map sequence, the gesture action will be segmented to obtain multiple second spatiotemporal map sequences; wherein, the gesture action is at least divided into a preparation stage, a gesture stage and a finishing stage, and one second spatiotemporal map sequence corresponds to one stage of the gesture action; Based on the multiple second spatiotemporal map sequences, the referee's penalty gesture is determined.

4. The method according to claim 2 or 3, characterized in that, The determination of the referee's penalty gesture based on the first spatiotemporal graph sequence includes: The first spatiotemporal map sequence is input into the first model to obtain a first image sequence; wherein, each first image in the first image sequence includes the predicted position of each joint key point of the gesture, and one first image corresponds to one spatiotemporal map in the first spatiotemporal map sequence; the first model is used to predict the position of each joint key point of the gesture based on the input multiple spatiotemporal maps; Based on the first image sequence and multiple second image sequences, the referee's penalty gesture is determined, and each second image in a second image sequence includes the labeled positions of various key joint points of the same standard penalty gesture.

5. The method according to claim 4, characterized in that, The first model includes a spatiotemporal graph convolutional network, a feature fusion module, and a fully connected neural network; the step of inputting the first spatiotemporal graph sequence into the first model to obtain the first image sequence includes: The first spatiotemporal map sequence is input into the spatiotemporal map convolutional network for convolution processing to obtain a set of vector feature maps. One of the vector feature maps in the set of vector feature maps represents the features of the joint key points in a spatiotemporal map. The set of vector feature maps are input into the feature fusion module for feature fusion processing to obtain a set of feature vectors, where one feature vector in the set of feature vectors corresponds to one vector feature map. The set of feature vectors is input into the fully connected neural network to predict the position of joint key points, thereby obtaining the first image sequence.

6. The method according to claim 4, characterized in that, The first image sequence includes multiple first subsequences, and a second image sequence includes multiple second subsequences. Each second subsequence corresponds to the annotation position of each joint key point of a stage of a standard penalty gesture. The step of inputting the first spatiotemporal map sequence into the first model to obtain the first image sequence includes: Input multiple second spatiotemporal map sequences corresponding to the first spatiotemporal map sequence into the first model to obtain multiple first subsequences, with each first subsequence corresponding to a second spatiotemporal map sequence.

7. The method according to any one of claims 1 to 3, 5 to 6, characterized in that, The method further includes: Obtain one or more pieces of first information input through the first interactive interface, where each piece of first information represents a user's feedback information on the judgment information; The second information is determined based on one or more pieces of first information, and the second information represents a comprehensive evaluation of the penalty information. Send or output the second information, which is used to display in the first interactive interface.

8. An electronic device, characterized in that, This includes a processor and memory for storing computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 7.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.