System for reconstructing 3D player model and automatic recognition based on 2d video
Patent Information
- Application Number
- KR1020250114848
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-08-12
- Estimated Expiration
- 2045-08-19
Smart Images

Figure 112025094350855-PAT00005_ABST
Abstract
Description
Technology Field
[0001] This specification relates to a system that uses artificial intelligence to automatically determine whether a foul has occurred in a VAR situation by restoring a 3D player model from a 2D image in the field of sports match analysis technology. Background Technology
[0003] In modern sports, particularly in soccer, Video Assistant Referee (VAR) systems are being widely adopted to ensure the accuracy and fairness of officiating. While existing VAR pools are capable of providing highly precise analysis based on multiple high-performance broadcast cameras, dedicated fiber optic networks, high-performance servers, and control rooms, building and operating such systems requires massive initial installation and operating costs. Consequently, there is a limitation in that adoption is difficult unless the event is a large-scale league or international tournament with a sufficient budget.
[0004] Lightweight (VAR Lite) systems, developed to reduce costs, provide reading capabilities using relatively few pieces of equipment and personnel; however, because they rely on relay camera footage or a small number of simple recording devices, their image quality and field of view are limited. Consequently, in situations requiring high precision, such as determining offside, goals, or penalty kicks, reliance must be placed on visual inspection and auxiliary line tools, leading to the constant possibility of subjective judgment intervention and errors. Furthermore, due to limited automation features such as object recognition and 3D reconstruction, reading quality varies depending on the proficiency of the personnel, and analysis times become prolonged.
[0005] Under these circumstances, there is a growing need for technology that can bridge the gap between expensive VAR systems and low-cost VAR Lite systems while providing high reading precision and a level of automation. In particular, there is a demand for a system that can ensure the objectivity and reliability of match decisions while reducing the burden of infrastructure construction by reconstructing 2D images into 3D space and object dummies based on a single or small number of cameras and performing AI-based automatic readings. The problem to be solved
[0007] The purpose of this specification is to simultaneously address the high cost and complex infrastructure issues of existing VAR systems and the low read precision and limited automation level of VAR Lite systems.
[0008] In addition, the purpose of this specification is to provide consistent and objective results in various judgment situations, such as offside, whether a goal is scored, and whether a penalty kick is awarded, by utilizing a soccer domain-specific 3D pose database and inference algorithms to supplement and restore missing body information even in complex situations that occur during a match, such as overlap between players, non-visible areas, and partial joint recognition failures.
[0009] The technical problems that this specification aims to solve are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this specification belongs from the detailed description of the specification below. means of solving the problem
[0011] One aspect of the present specification is a method for restoring a 3D player model based on a 2D image and performing automatic reading, comprising: a step of collecting and preprocessing a 2D image of a stadium; a step of detecting a VAR (Video Assistant Referee) reading request; a step of tracking an object based on the VAR reading request, wherein the object includes a player, a referee, or a ball; a step of, if the object tracking is successful, extracting joint information based on the object and matching it based on training data, wherein the training data is a 3D pose database of a soccer domain, and a step of generating the 3D player model based on the matched joint information; and a step of performing a reading related to the VAR based on the 3D player model.
[0012] In addition, the step of collecting and preprocessing 2D images of the stadium may include the step of generating a field coordinate system of the stadium based on the 2D images.
[0013] Additionally, the step of tracking the object may include the step of maintaining the ID of the player, the referee, or the ball between frames.
[0014] Additionally, the matching step may include the step of comparing the joint information with the 3D pose database to select the most similar skeleton.
[0015] In addition, the 3D player model may include a joint structure projected onto the field coordinate system and a human body mesh.
[0016] In addition, if the object tracking fails, the method further includes the step of extracting partial joint information based on the object and matching based on the training data; and if the object tracking fails, the case may include the case where the ID is lost.
[0017] Additionally, if the object tracking fails, the step of extracting partial joint information based on the object and matching based on the training data may include: a step of recognizing only some joints of the object to extract partial joint information; a step of searching for the sample with the highest similarity by comparing the partial joint information with the training data; and a step of restoring the entire joint and mesh of the object based on the sample. Effects of the invention
[0019] According to the embodiments of this specification, the high cost and complex infrastructure issues of existing VAR systems and the low reading precision and limited level of automation of VAR Lite systems can be improved simultaneously.
[0020] In addition, according to the embodiments of the present specification, even in complex situations such as overlap between players, non-visible areas, and partial joint recognition failures that occur during a match, missing body information can be supplemented and restored by utilizing a soccer domain-specific 3D pose database and inference algorithms, thereby providing consistent and objective results in various judgment situations such as offside, whether a goal was scored, and whether a penalty kick was awarded.
[0021] The effects obtainable in this specification are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which this specification belongs from the description below. Brief explanation of the drawing
[0023] FIG. 1 is a block diagram for illustrating an electronic device related to the present specification. FIG. 2 is a block diagram of an AI device according to one embodiment of the present specification. FIG. 3 illustrates a system for 2D image-based 3D player model restoration and automatic reading to which the present specification may be applied. FIG. 4 illustrates a 2D image-based 3D player model restoration and automatic reading method to which the present specification may be applied. FIG. 5 illustrates in more detail the operation of collecting and preprocessing 2D images that can be applied to the present specification. FIG. 6 illustrates an object tracking method that can be applied to the present specification. FIG. 7 illustrates an object information extraction and training data-based matching method applicable to the present specification. FIG. 8 illustrates partial joint information to which the present specification can be applied, and training data-based alignment. FIG. 9 illustrates a method for generating and reading a 3D player model to which the present specification can be applied. Specific details for implementing the invention
[0024] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components regardless of drawing symbols are assigned the same reference number, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not have distinct meanings or roles in themselves. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the concept and technical scope of this specification.
[0025] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0026] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0027] A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0028] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0029] FIG. 1 is a block diagram for illustrating an electronic device related to the present specification.
[0030] The above electronic device (100) may include a wireless communication unit (110), an input unit (120), a sensing unit (140), an output unit (150), an interface unit (160), a memory (170), a control unit (180), and a power supply unit (190), etc. Since the components illustrated in FIG. 1 are not essential for implementing the electronic device, the electronic device described herein may have more or fewer components than those listed above.
[0031] More specifically, among the above components, the wireless communication unit (110) may include one or more modules that enable wireless communication between the electronic device (100) and a wireless communication system, between the electronic device (100) and another electronic device (100), or between the electronic device (100) and an external server. Additionally, the wireless communication unit (110) may include one or more modules that connect the electronic device (100) to one or more networks.
[0032] This wireless communication unit (110) may include at least one of a broadcast receiving module (111), a mobile communication module (112), a wireless internet module (113), a short-range communication module (114), and a location information module (115).
[0033] The input unit (120) may include a camera (121) or video input unit for inputting a video signal, a microphone (122) or audio input unit for inputting an audio signal, and a user input unit (123, e.g., a touch key, a mechanical key, etc.) for receiving information from a user. Voice data or image data collected from the input unit (120) may be analyzed and processed into a control command by the user.
[0034] The sensing unit (140) may include one or more sensors for sensing at least one of information within the electronic device, information about the surrounding environment surrounding the electronic device, and user information. For example, the sensing unit (140) may include at least one of a proximity sensor (141), an illumination sensor (142), a touch sensor, an acceleration sensor, a magnetic sensor, a gravity sensor (G-sensor), a gyroscope sensor, a motion sensor, an RGB sensor, an infrared sensor (IR sensor: infrared sensor), a fingerprint sensor (finger scan sensor), an ultrasonic sensor, an optical sensor (e.g., see camera (121)), a microphone (see 122), a battery gauge, an environmental sensor (e.g., a barometer, a hygrometer, a thermometer, a radiation detection sensor, a heat detection sensor, a gas detection sensor, etc.), and a chemical sensor (e.g., an electronic nose, a healthcare sensor, a biometric sensor, etc.). Meanwhile, the electronic device disclosed in this specification can utilize information sensed by at least two of these sensors in combination.
[0035] The output unit (150) is intended to generate output related to sight, hearing, or touch, and may include at least one of a display unit (151), an acoustic output unit (152), a haptic module (153), and an optical output unit (154). The display unit (151) may form a layered structure with a touch sensor or be formed integrally to implement a touch screen. Such a touch screen functions as a user input unit (123) that provides an input interface between the electronic device (100) and the user, and at the same time can provide an output interface between the electronic device (100) and the user.
[0036] The interface section (160) serves as a passage for various types of external devices connected to the electronic device (100). This interface section (160) may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port. In response to an external device being connected to the interface section (160), the electronic device (100) can perform appropriate control related to the connected external device.
[0037] Additionally, the memory (170) stores data that supports various functions of the electronic device (100). The memory (170) can store a number of application programs (or applications) running on the electronic device (100), data for the operation of the electronic device (100), and commands. At least some of these application programs may be downloaded from an external server via wireless communication. Also, at least some of these application programs may exist on the electronic device (100) from the time of shipment for the basic functions of the electronic device (100) (e.g., phone incoming and outgoing functions, message receiving and outgoing functions). Meanwhile, the application programs may be stored in the memory (170), installed on the electronic device (100), and driven by the control unit (180) to perform the operation (or function) of the electronic device.
[0038] In addition to operations related to the application program, the control unit (180) typically controls the overall operation of the electronic device (100). The control unit (180) can provide or process appropriate information or functions to the user by processing signals, data, information, etc. that are input or output through the components described above, or by running an application program stored in memory (170).
[0039] Additionally, the control unit (180) can control at least some of the components examined together with FIG. 1 in order to run an application program stored in memory (170). Furthermore, the control unit (180) can operate at least two or more of the components included in the electronic device (100) in combination with each other to run the application program.
[0040] The power supply unit (190) receives external power and internal power under the control of the control unit (180) and supplies power to each component included in the electronic device (100). This power supply unit (190) includes a battery, and the battery may be a built-in battery or a replaceable battery.
[0041] At least some of the above components may operate in cooperation with each other to implement the operation, control, or control method of an electronic device according to various embodiments described below. Additionally, the operation, control, or control method of the electronic device may be implemented on the electronic device by running at least one application program stored in the memory (170).
[0042] In this specification, the electronic device (100) may be composed of all or part of the system (300) described below and may be collectively referred to as a server, and the server may include a cloud server. Additionally, the terminal may include all or part of the electronic device (100) and may include a tablet PC.
[0043] FIG. 2 is a block diagram of an AI device according to one embodiment of the present specification.
[0044] The AI device (20) may include an electronic device including an AI module capable of performing AI processing, or a terminal including the AI module. Additionally, the AI device (20) may be configured to be included as at least a part of the configuration of the electronic device (100) shown in FIG. 1 to perform at least a part of the AI processing together.
[0045] The above AI device (20) may include an AI processor (21), memory (25) and / or a communication unit (27).
[0046] The above AI device (20) is a computing device capable of learning a neural network and can be implemented as various electronic devices such as a terminal, desktop PC, laptop PC, tablet PC, etc.
[0047] The AI processor (21) executes a program stored in memory (25) to learn an artificial intelligence model specialized for the sports domain, thereby automatically classifying and analyzing 2D video data and game situation metadata, and enabling the generation of optimized 3D player model restoration results and rule reading information. For example, a pre-training process may be performed on a 3D pose database (DB) specialized for the soccer domain. This database contains 3D joint coordinates and mesh shapes of various soccer movements (e.g., dribbling, shooting, tackling, jumping, etc.) collected through actual games or simulations and motion capture. This enables high-accuracy 3D restoration by matching with future joint extraction results. Additionally, the AI processor (21) can continuously improve 3D restoration accuracy and reading suitability by applying a reinforcement learning-based algorithm to learn referee feedback and game operation monitoring results.
[0048] For example, the AI processor (21) can analyze input game footage, player and ball position, speed, and pose information, game rule data, judgment history, etc., and automatically select 3D restoration parameters and judgment logic optimized for the situation. Then, by reflecting feedback from the referee or auxiliary judgment module (e.g., correction of judgment errors, correction of distance calculations, suggestions for overlapping correction, etc.) on the generated judgment results, the reinforcement learning algorithm continuously learns the optimal restoration and judgment structure. Through this, automatic adaptation to game types, camera placement, environmental conditions, etc. is possible.
[0049] In addition, the AI processor (21) utilizes a HybridRAG structure that combines a large language model (LLM) and a sports domain-specific 3D pose database to automatically generate 3D player models, visualizations, and review reports that match the analysis context. The generated outputs are converted into various formats such as video, images, JSON, PDF, and UI overlays, and can be linked with referee review UIs, broadcasting systems, and record and statistics servers, and through repeated use, the completeness and reliability of the review results are continuously improved. Thus, the AI processor (21) performs a core function of enhancing the objectivity and transparency of the ruling in the VAR automatic review system.
[0050] Meanwhile, the AI processor (21) that performs the functions described above may be a general-purpose processor (e.g., CPU), but may be an AI-dedicated processor for artificial intelligence learning (e.g., GPU, graphics processing unit).
[0051] The memory (25) can store various programs and data required for the operation of the AI device (20). The memory (25) can be implemented as non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), or solid-state drive (SDD). The memory (25) is accessed by the AI processor (21), and the AI processor (21) can perform reading / writing / modification / deletion / updating of data. Additionally, the memory (25) can store a neural network model (e.g., a deep learning model) generated through a learning algorithm for data classification / recognition according to one embodiment of the present specification.
[0052] Meanwhile, the AI processor (21) may include a data learning unit that learns a neural network for data classification / recognition. For example, the data learning unit may learn a deep learning model by acquiring training data to be used for learning and applying the acquired training data to a deep learning model.
[0053] The communication unit (27) can transmit the AI processing results by the AI processor (21) to an external electronic device.
[0054] Here, external electronic devices may include other terminals and / or servers (cloud servers).
[0055] Meanwhile, although the AI device (20) illustrated in FIG. 2 is described by functionally separating it into an AI processor (21), memory (25), and communication unit (27), the aforementioned components may be integrated into a single module and referred to as an AI module or an artificial intelligence (AI) model.
[0056] FIG. 3 illustrates a system for 2D image-based 3D player model restoration and automatic reading to which the present specification may be applied.
[0057] Referring to FIG. 3, the system (300) can support VAR decisions by extracting field coordinate systems and object information from a 2D image, and performing key player recognition, ID loss correction, 3D player model restoration, visualization, and automatic reading.
[0058] 2D image acquisition and field coordinate system generation module (310) It captures the entire stadium using dual-lens or multi-lens cameras and generates 2D image data suitable for analysis by performing frame synchronization and preprocessing. For example, it automatically recognizes fixed structures within the video (goals, corner flags, half lines, etc.) and can establish an XYZ coordinate system that matches the actual stadium through camera pose calculation and homography transformation. The field coordinate system generated through this process is subsequently used as a reference axis for 3D player model reconstruction, distance calculation, and rule reading, and can provide stable spatial reference regardless of camera installation environments or differences in stadium structure.
[0059] VAR request reception and analysis frame extraction module (320) The system detects a signal requesting a review from a referee or system in real time and designates frames before and after that point in time as targets for analysis. The analysis section is stored in a buffer and is subsequently utilized in the object recognition and tracking process. Through this, the system (300) does not analyze all game footage in real time but processes only the key section at the time of the request, thereby reducing computational resources and enabling high-precision reading even in a low-cost environment.
[0060] Key player recognition and 2D joint extraction module (330) It identifies key objects necessary for VAR rule decisions. For example, in the case of offside, it automatically detects the attacker closest to the goal line, the second-closest defender, and the ball. Subsequently, it extracts the 2D joint coordinates of these objects through a pose estimation model and matches them with a soccer domain-specific pose database to obtain accurate joint position and posture information.
[0061] ID loss correction and partial joint-based 3D estimation module (340) In cases where ID tracking fails due to player overlap or obscuration during a match, this module operates to infer the entire joints and pose even when only some joints are visible. To achieve this, it performs time-based interpolation using data from the preceding and succeeding frames, as well as similar sample matching based on a 3D pose database. The restored joint data undergoes a Re-ID (re-identification) process to confirm it as the same object, and is subsequently fed into the 3D player model restoration stage. This enables the acquisition of stable reading data even in complex match situations.
[0062] 3D player model restoration, visualization and automatic reading module (350)It generates a 3D player model by projecting 2D joint data onto an XYZ coordinate system and performs automatic analysis based on rules, such as offside, whether a goal has been scored, or whether a penalty kick has been awarded. For example, a 3D player model may refer to a virtual 3D model reconstructed by matching pre-acquired and learned 3D joint data and mesh data based on keypoint data of an object detected in a 2D image. The 3D model is aligned with the XYZ coordinate system of the actual stadium and can include information on the object's position, posture, and shape.
[0063] The reading results are provided with rotatable 3D visual data, distance measurements, and baseline markings, and can be output and saved in various formats (video, image, data) so that referees, spectators, and broadcasting systems can intuitively check them.
[0064] FIG. 4 illustrates a 2D image-based 3D player model restoration and automatic reading method to which the present specification may be applied.
[0065] Referring to FIG. 4, the system (300) receives a VAR review request by preprocessing 2D images collected from the stadium and linking them with a soccer domain-specific 3D pose database, and tracks an object (player, ball, etc.) at the time of the request to determine whether it is successful. If the object tracking is successful, the entire joint information is extracted and aligned, and if it fails, the system uses training data based on partial joint information to correct and align it to generate complete joint information. The aligned joint information is projected onto an XYZ coordinate system to create a 3D player model including a human mesh, and the review engine performs rule judgments such as offside and goal status based on the 3D player model, and can provide the results in the form of visual data such as a 3D view, multi-view, and distance display.
[0066] The system collects and preprocesses 2D images (S4010).
[0067] The 2D image acquisition and field coordinate system generation module of the system (300) can capture the entire stadium through a dual-lens or multi-lens camera and generate a field coordinate system. The camera module covers the left and right half-courts simultaneously and can accurately match images of the same timestamp through a frame synchronization device. The images collected in this way are used as base data for subsequent analysis. For example, the collected original images undergo processes such as noise removal, brightness and contrast adjustment, and distortion correction in the 2D image acquisition and preprocessing module. This preprocessing process is essential for increasing the accuracy of object detection and joint recognition, and enables stable analysis even under various lighting and weather conditions.
[0068] In addition, at this stage, pre-training and loading of a soccer domain-specific 3D pose database (DB) can be performed in parallel. This DB includes 3D joint coordinates and mesh shape data collected through actual matches, simulations, and motion capture, which are subsequently used as matching standards during joint extraction and 3D reconstruction.
[0069] The system receives a VAR review request (S4020).
[0070] The VAR request reception and analysis frame extraction module detects a review request event occurring from the referee or system (300) in real time. This request may be triggered for an accurate judgment on a specific situation during the match (e.g., whether a goal is scored, offside, foul, etc.).
[0071] When a request is received, the system (300) stores video of the N-second interval before and after a specific point in time of the review request in the VAR buffer. This allows for intensive processing of only the key scenes at the time of the request without analyzing the entire match video. The frame at the time of the VAR request becomes the standard for the subsequent object tracking step (S4030), and preparations for OD (object detection) and OT (object tracking) are already underway in the background. This allows the analysis step to proceed quickly immediately after the request.
[0072] The system tracks the object based on the VAR read request (S4030)
[0073] The core player recognition and 2D joint extraction module identifies players, referees, and the ball in the video at the time of the VAR request and tracks them while maintaining inter-frame IDs through a multi-object tracking algorithm. To achieve this, object detectors and trackers operate together, and the tracking status is managed in real time. During the tracking process, metadata such as position, movement path, and speed is recorded and utilized for subsequent rule readings. For example, in the case of offside calls, data is accumulated to continuously compare the relative positions of attackers and defenders.
[0074] The system determines whether object tracking was successful (S4040)
[0075] The system (300) evaluates the continuity and integrity of the object ID. For example, if the object is tracked stably without being obscured or merged, it is classified as 'Yes', and if tracking failure or discontinuity occurs, it is classified as 'No'. For example, a case where object tracking fails or is discontinuous may mean that the ID is lost.
[0076] In the case of Yes, a normal joint information extraction and matching process is performed. In the case of No, a correction procedure is performed based on incomplete joint information. This branching logic is an important judgment point that determines the accuracy of the entire system (300). While normal paths allow for fast and efficient analysis, incomplete paths may require complex correction algorithms and additional calculations.
[0077] The system extracts joint information based on the object and aligns it based on the training data (S4050).
[0078] In a normal tracking path, the key player recognition and 2D joint extraction module identifies key objects (e.g., attackers, defenders, and balls) and extracts major joint coordinates such as shoulders, knees, and ankles. This 2D joint data is linked to a field coordinate system. The extracted joint data is matched with a pre-trained soccer domain 3D pose DB, and the most similar 3D joint data and mesh shape are selected. In this process, the system (300) can improve the matching accuracy through camera viewpoint correction.
[0079] The system extracts partial joint information based on incomplete object tracking and aligns it based on training data (S4060).
[0080] In an incomplete tracking path, the ID loss correction and partial joint-based 3D estimation module extracts partial joint information when only some joints are recognizable. This data is processed only when joint omission occurs. For example, partial joint data is used to search for Top-K samples with high similarity in a 3D pose DB. Based on the pose distribution of the samples, the system (300) can probabilistically restore the entire joint and mesh. The corrected joint and mesh are linked to the existing object through the Re-ID process and are finally converted into complete joint data and passed to the next step.
[0081] The system generates a 3D athlete model based on the matched joint information (S4070).
[0082] The 3D player model restoration, visualization, and automatic reading module generates a 3D player model by projecting the restored joint data onto an XYZ field coordinate system. The 3D player model can include not only joint positions but also a human body shape mesh. During model generation, accurate spatial alignment is achieved by reflecting the relative position with fixed references such as goalposts and lines. This enables quantitative distance calculation during rule determination.
[0083] The system reads based on a 3D player model (S4080)
[0084] The 3D player model restoration, visualization, and automatic judgment module automatically determines offside, whether a goal has been scored, and whether a penalty kick has been awarded based on the 3D player model. These judgments are performed by calculating the distance between the goal line, penalty line, or baseline and the object's body part, and by comparing their relative positions. This minimizes referee subjectivity and ensures consistency in judgments. The results can be provided in the form of 3D views, multi-views, distances, and baseline displays through the 3D player model restoration, visualization, and automatic judgment module, and can be saved and transmitted as video, images, or data for use in match recording and analysis.
[0085] FIG. 5 illustrates in more detail the operation of collecting and preprocessing 2D images that can be applied to the present specification.
[0086] Referring to Fig. 5, the operation of S4010 of the 2D image acquisition and preprocessing module is illustrated in more detail.
[0087] Collect 2D images (S5010) :
[0088] The 2D image acquisition and field coordinate system generation module secures original video by capturing the entire stadium through a camera module. The camera is configured with high resolution and a high frame rate to capture even the detailed movements of fast-moving players or the ball. Additionally, in a multi-camera environment, the position and field of view of each camera are calibrated in advance to enable accurate spatiotemporal matching during the subsequent coordinate calculation process.
[0089] The collected video is precisely synchronized with frames at the same point in time through a Time Sync Controller. This allows data captured by different cameras to be compared and analyzed along the same time axis during VAR readings or 3D pose reconstruction. During the synchronization process, GPS timing or high-precision network timing protocols (NTP / PTP) are utilized to minimize frame lag to the microsecond level. Through a Frame Extractor, frames for analysis are extracted from the collected video stream at regular intervals. The extracted frames undergo processing such as lens distortion correction, noise removal, and color correction in a Preprocessing Unit to maintain uniform image quality. These preprocessed frames are then stored in Video Buffer Storage.
[0090] Based on the collected 2D image, recognize the field and calculate the coordinates (S5020) :
[0091] The 2D image acquisition and field coordinate system generation module recognizes the field of the stadium from the preprocessed 2D image and calculates the field coordinate system. To this end, a Fixed Marker Detector can detect reference points within the stadium (e.g., corner flags, goal posts, line markings). The detected reference point data is used to calculate the transformation relationship between the 2D image coordinates and the actual field coordinates through a Homography Estimator.
[0092] Subsequently, the Camera Pose Solver calculates the camera's position, orientation, and tilt to define the relationship between the captured 2D image and the actual 3D space. The projection matrix generated at this stage becomes a key parameter for converting the positions of objects, such as players and balls, into 3D. Additionally, the Field Coordinate Encoder converts the stadium coordinate system into a standard format to maintain consistency in data exchange between other modules. While the field recognition and coordinate calculation processes are primarily performed by automated algorithms, correction is possible via the Manual ROI Input Interface if recognition accuracy is low or if parts of the reference points are obscured. This enables stable coordinate calculation even in abnormal shooting situations (e.g., weather, crowd obstruction, camera shake).
[0093] Based on the calculated result, link with the 3D pose DB (S5030) :
[0094] Once the coordinate calculation is complete, the process of linking with the 3D pose DB is performed based on the data. First, the 3D pose database loader loads pre-stored standard 3D pose data. This data includes information such as the athlete's standard movements, joint coordinates, and pose changes per frame. The Camera Parameter Injector applies the projection matrix calculated in the previous step to align the 3D pose data with the coordinate system of the currently captured 2D image. Subsequently, the 3D Pose Aligner aligns the actual captured data with the pose data in the DB to achieve accurate 2D-3D mapping.
[0095] Finally, a 2D Skeleton Projector projects the 3D pose onto a 2D image, and a Pose Normalizer standardizes the joint positions and sizes to enhance comparability. A Pose Matching Indexer assigns an index to each frame, enabling the rapid retrieval and utilization of pose data at the relevant point in time when a VAR reading or analysis request is made.
[0096] FIG. 6 illustrates an object tracking method that can be applied to the present specification.
[0097] Referring to Fig. 6, the operation of S4030-S4040 of the VAR request reception and analysis frame extraction module / key player recognition and 2D joint extraction module is illustrated in more detail.
[0098] Received VAR review request (S6010) :
[0099] A VAR (Video Assistant Referee) review request is initiated by an instruction to analyze a specific situation (goal, foul, offside, etc.) during the match. The request includes the time of occurrence of the event to be reviewed, the corresponding camera ID, the match time, and the analysis conditions, and the system (300) can initialize the analysis work based on this.
[0100] Upon receiving a VAR reading request, the VAR request reception and analysis frame extraction module verifies whether the video for the corresponding time interval exists in the image storage server and whether the resolution, frame rate, and synchronization data required for analysis are available. By checking for missing or damaged video during this process, it is possible to prevent subsequent analysis failures.
[0101] Recognize and track objects (S6020) :
[0102] The VAR request reception and analysis frame extraction module recognizes target objects (players, balls, etc.) in the video within the requested time interval. For example, it utilizes a deep learning-based object detection model to generate bounding boxes for each frame, enabling the tracking of the same object across consecutive frames. Additionally, tracking algorithms (such as Kalman Filter or DeepSORT) are used to secure stable trajectory information even in the presence of occlusion or rapid movement. This trajectory data serves as a key criterion for selecting analysis frames later, and a correction procedure is performed if the tracking quality fails to meet a certain standard.
[0103] Select analysis frame (S6030) :
[0104] The VAR request reception and analysis frame extraction module selects frames from the tracked trajectory data that are directly related to the event. For example, sudden changes in position or speed, or the moments immediately before and after specific events (ball touch, contact, etc.) can serve as key selection criteria. In a multi-camera environment, timestamps can be used to align frames at the same point in time and secure images from various angles. The selected frames are transmitted to the next analysis stage in a form that includes object location coordinates, camera calibration parameters, and time information.
[0105] Based on the analysis frame, object analysis (key player recognition) (S6040) :
[0106] The key player recognition and 2D joint extraction module re-recognizes key targets (e.g., players near the offside line) in selected analysis frames. To achieve this, it combines object detection and identification algorithms to confirm that they are the same person, thereby improving accuracy even in crowded scenes. Additionally, it can perform image preprocessing, such as high-resolution magnification and noise reduction, to ensure that the body parts of key players are clearly identifiable.
[0107] 2D joint extraction based on object analysis (S6050) :
[0108] The core player recognition and 2D joint extraction module can calculate the 2D coordinates of major body joints using deep learning-based pose estimation models such as OpenPose and MediaPipe Pose. The coordinates of each joint are calculated in pixel units and are represented as the connectivity relationships (skeleton structure) between body parts. The extracted joint data serves as basic data for subsequent 3D pose reconstruction, distance calculation, and motion analysis.
[0109] Determine whether object tracking was successful (S6060) :
[0110] The core player recognition and 2D joint extraction module compares the joint extraction results with object trajectory data to determine whether tracking has been consistently maintained throughout the entire section. If joint data is missing, a discontinuous trajectory occurs, or misrecognition occurs in a specific section, the video of that section may be reanalyzed or another camera video may be used as a supplement.
[0111] FIG. 7 illustrates an object information extraction and training data-based matching method applicable to the present specification.
[0112] Referring to Fig. 7, the operation of the key player recognition and 2D joint extraction module is illustrated in more detail when object tracking is successful according to the result of the aforementioned S4040.
[0113] Recognize key players based on tracked objects (S7010) :
[0114] The key player recognition and 2D joint extraction module recognizes key players who are the primary subjects of analysis within the game scene based on the acquired tracking results. This process is performed by comprehensively analyzing the player's position, movement, and role within the game. In particular, candidate players are selected by utilizing various features such as distance from the ball, frequency of play participation, and center position within the camera frame.
[0115] For example, in the automatic recognition process, a machine learning-based Player Role Classifier is utilized to classify each player's position (forward, defender, goalkeeper, etc.), and rule-based logic is applied to ultimately determine the key players suitable for the game situation. In addition, to enhance analysis accuracy, average data across multiple frames is reflected to prevent temporary misrecognition.
[0116] In addition, depending on the situation, analysts can directly designate key players through a Manual Selection Interface rather than automatic recognition. This is necessary in VAR reviews or referee assistance situations and is an important procedure to complement the limitations of AI judgment and ensure reliability.
[0117] Extract key player's joint information (S7020) :
[0118] The key player recognition and 2D joint extraction module extracts 2D joint coordinates (keypoints) of the recognized key players. To this end, the Pose Estimation Module can detect the player's body by dividing it into more than 17 major joint points (head, shoulders, elbows, knees, ankles, etc.). This process is optimized by considering the video resolution, shooting angle, occlusion status, etc.
[0119] After joint extraction, a Multi-Object Tracker is utilized to reliably link joint positions between frames. During this process, the Tracking State Manager continuously manages the joint tracking status, preventing ID loss issues that may occur due to collisions between players, obstructed views, or frame loss.
[0120] In addition, missing joint coordinates are corrected through Background Object Detection (Background OD) and Temporal Data Estimation (TD) modules. This maintains the continuity of joint data and secures high-quality input data that can be used in the subsequent skeleton registration step.
[0121] 2D projected skeleton registration (S7030) based on joint information :
[0122] The core player recognition and 2D joint extraction module performs 2D projected skeleton registration based on the extracted joint information. First, the Pose Matcher module searches for the most similar skeleton structure by comparing the joint coordinates of the current frame with a pre-trained pose dataset. During the comparison, the distance ratio between joints, relative angles, and differences in position coordinates are used as key feature values.
[0123] Subsequently, a module that selects the Best Match determines the candidate skeleton with the highest match and aligns it to the current frame. During this process, transformations such as scaling, rotation correction, and coordinate translation are applied to minimize the discrepancy between the actual athlete's movements and the data.
[0124] The 2D skeleton data that has been matched can be transferred to various application modules such as subsequent 3D pose estimation, motion analysis, referee judgment support, and game record analysis, and can be reused as a ground truth dataset required for future model training to contribute to continuously improving the recognition accuracy of the system (300).
[0125] FIG. 8 illustrates partial joint information to which the present specification can be applied, and training data-based alignment.
[0126] Referring to Fig. 8, the operation of the ID loss correction and partial joint-based 3D estimation module is illustrated in more detail when object tracking fails according to the result of the aforementioned S4040.
[0127] ID recovery (S8010) based on the location where the ID was lost :
[0128] The ID loss correction and partial joint-based 3D estimation module detects locations where the ID of a specific object is lost during the object tracking process. ID loss may occur due to temporary occlusion of the object, missing frames within the image, or color and shape similarity with the background. The system (300) utilizes an ID Loss Region Detector to automatically detect these loss locations and can initiate an ID recovery process based on the detection results.
[0129] Subsequently, the Tracking Module and the Pose Estimation Module are linked to verify continuity with past frames. To this end, previous joint information stored in the Pose History Buffer is compared and analyzed with time-series flow data extracted by the Temporal Flow Extractor. The Flow Continuity Matcher evaluates the similarity between consecutive poses to generate a candidate set of lost IDs.
[0130] Finally, the ID Candidate Generator and Re-ID Decision Engine operate to determine the ID with the highest consistency among the candidates as the recovery ID. The Appearance Matcher and Pose Similarity Scorer operate auxiliaryly to confirm the object with a high match rate in both appearance features and joint patterns as the final recovery target.
[0131] Extract partial joint information based on the recovered ID (S8020) :
[0132] The ID loss correction and partial joint-based 3D estimation module extracts partial joint information of the corresponding object based on the recovered ID. First, the Partial Skeleton Extractor operates to select only a subset of joints that can be reliably detected within the image, rather than all joints. This is to ensure tracking stability despite occlusion or the loss of some joint information. Next, the Keypoint Index Filter is applied to remove unnecessary joint coordinates, leaving only the core joints for analysis. Subsequently, the Noisy Joint Remover filters out false positives or irregularly fluctuating joint coordinates. This enhances the reliability of the joint data and guarantees alignment accuracy in subsequent steps. Finally, the Partial Keypoint Formatter normalizes the extracted joint coordinates and converts them into a standard data structure.
[0133] 2D projected skeleton registration (S8030) based on partial joint information :
[0134] The ID loss correction and partial joint-based 3D estimation module performs alignment with a 2D-projected skeleton database based on partial joint information. First, the Partial Pose Matcher compares the input partial joints with the skeleton data in the DB, and the Pose Similarity Calculator calculates similarity by determining the position, angle, and length ratio between joints. For example, the Top-K Match Pool stores the top K candidate skeletons with the highest similarity. Subsequently, the Best Match Selector selects the most suitable skeleton among them. The selection process prevents false positives by considering not only simple similarity but also the object's motion context and temporal continuity.
[0135] Finally, the Matching Success Checker determines whether the match satisfies a predefined threshold. If the match is successful, subsequent actions can be performed based on the skeleton information, and if it fails, the process moves to the next step, S8040, to perform the restoration procedure.
[0136] If alignment fails, re-align after 2D skeleton restoration (S8040) :
[0137] This step is performed only when matching fails in S8030. First, the ID loss correction and partial joint-based 3D estimation module uses a 2D Skeleton Reconstructor to estimate and restore missing joints based on partial joint information. During the restoration process, reasonable coordinates are calculated by utilizing the distance ratio between adjacent joints, symmetry, motion patterns, etc.
[0138] The reconstructed skeleton is compared again with the projected skeleton in the database using the Pose Similarity Calculator. At this stage, the Top-K Match Pool and Best Match Selector are applied equally to re-select the optimal match. During the re-alignment process, the reliability score of the reconstructed joint coordinates is also considered to reduce the possibility of false positives caused by incorrect reconstruction.
[0139] Finally, the Matching Success Checker determines the re-matching result. If successful, the skeleton is adopted as the final matching result, and if unsuccessful, an error log is left and the process can proceed to an auxiliary tracking path or manual verification procedure. Through this, the system (300) can reliably track and match objects even in situations of ID loss and partial joint loss.
[0140] FIG. 9 illustrates a method for generating and reading a 3D player model to which the present specification can be applied.
[0141] Referring to Fig. 9, the operation of the S4070-S4080 3D player model restoration, visualization, and automatic reading modules is illustrated in more detail.
[0142] 3D athlete model restoration (S9010) based on aligned joint information :
[0143] The 3D athlete model restoration, visualization, and automatic reading module restores the 3D athlete model based on the aligned joint information generated in the preceding processing steps (S7030 / S8040, etc.). The joint information includes the coordinates of major joints of the human body, length ratios, and relative positional relationships, which are used to generate a standard 3D skeletal structure. During the restoration process, individual joint coordinates are mapped into three-dimensional space, and connection relationships between body parts are established to construct the basic frame of the entire dummy. At this time, the actual body shape, height ratio, and direction of movement of each athlete are also reflected to minimize the error with the original scene.
[0144] Visualization task (S9020) based on the restored 3D player model:
[0145] The restored 3D player models are combined with the real environment through a visualization module. First, the scene reconstruction engine aligns the models with the stadium coordinate system or the shooting environment coordinate system to complete a realistic spatial layout. Next, tactical visual aids (e.g., offside lines, distance indicators, attack and defense zone indicators, etc.) can be added. These are automatically generated by reflecting the analysis results of the tactical rules engine and support referees and analysts in intuitively assessing situations.
[0146] Finally, the visualization results are output in the form of high-quality 3D renderings and can be provided in real time by linking to a VAR (video interpretation) system or broadcast transmission equipment.
[0147] Based on the visualization task, VAR reading (S9030) :
[0148] The 3D player model restoration, visualization, and automatic review module performs VAR (Video Assistant Referee) reviews based on the visualized scene. For example, the review module can automatically analyze whether there has been a violation of the rules of the game in the visualized data and, if necessary, provide evidence footage that the referee can refer to.
[0149] The review process considers not only simple positional comparisons but also changes in movement along the time axis and the context of the event's occurrence. For example, in the case of an offside call, the position of the ball at the moment of passing and the defensive line are reviewed simultaneously. The results derived from this review are transmitted to the referee in real time to be directly reflected in match decisions, and can also be stored for post-match analysis and data archiving purposes.
[0150] Through this, the system (300) can reconstruct accurate 2D and 3D poses through ID recovery and skeleton matching processes based on partial joint information, even when the identification of the tracked object is lost or incomplete during the VAR (Video Assistant Review) review process. In addition, key scenes of the game situation can be analyzed without omission, and the possibility of incorrect judgment can be significantly reduced. In particular, the multi-frame-based continuous tracking and similarity matching procedure enables fast and stable ID recovery, making it applicable even in real-time review environments.
[0151] The system (300) performs projected skeleton matching based on learning data using only partial joint information, thereby enabling the restoration of the entire pose with high accuracy even in cases of occlusion or failure to detect some joints. In this process, the optimal matching result is selected using a Top-K matching pool and similarity calculation, and in the event of a failure in matching, a 2D skeleton restoration and re-matching procedure is performed to finally obtain pose data of high completeness. Through this, stable reading quality can be maintained even in input images of various shooting angles and resolutions.
[0152] Finally, the 3D player model restoration and visualization features recreate actual match scenes based on aligned joint information and aid the intuitive understanding of the reviewer by clearly visualizing offside lines and tactical indicators. When integrated with the VAR review module, this enables rapid and objective rulings on complex match situations, thereby enhancing match fairness and contributing to the reduction of officiating controversies.
[0153] The foregoing specification may be implemented as computer-readable code on a medium on which a program is recorded. A computer-readable medium includes all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include Hard Disk Drives (HDDs), Solid State Disks (SSDs), Silicon Disk Drives (SDDs), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, optical data storage devices, etc., and also include implementations in the form of carrier waves (e.g., transmission over the Internet). Accordingly, the above detailed description should not be interpreted restrictively in all respects and should be considered exemplary. The scope of this specification should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of this specification are included within the scope of this specification.
[0154] Furthermore, although the above description has focused on the services and embodiments, this is merely illustrative and does not limit the scope of this specification. Those skilled in the art will understand that various modifications and applications not exemplified above are possible without departing from the essential characteristics of the services and embodiments. For example, each component specifically shown in the embodiments may be modified and implemented. Differences related to such modifications and applications should be interpreted as being included within the scope of this specification as defined in the appended claims.
Claims
Claim 1 A method for restoring a 3D player model based on 2D video and performing automatic reading, comprising: a step of collecting and preprocessing 2D video of a stadium; a step of detecting a VAR (Video Assistant Referee) reading request; a step of tracking an object based on the VAR reading request, wherein the object includes a player, a referee, or a ball, and is tracked to maintain the ID of the player, the referee, or the ball between frames; and, if the object tracking is successful, a step of extracting joint information based on the object and aligning it based on training data, wherein the training data is a 3D pose database of the soccer domain; a step of generating the 3D player model based on the aligned joint information; and a step of performing a reading related to the VAR based on the 3D player model. The method comprises the step of, when the object tracking fails, extracting partial joint information based on the object and matching it based on the training data; wherein the case of the object tracking failure includes cases where the ID is lost due to overlap between players, non-visible areas, or failure to recognize partial joints occurring during a game; and the step of, when the object tracking fails, extracting partial joint information based on the object and matching it based on the training data includes the step of recognizing only some joints of the object to extract partial joint information; the step of searching for the sample with the highest similarity by comparing the partial joint information with the training data; and the step of restoring the entire joints and mesh of the object based on the sample. Claim 2 An automatic reading method according to claim 1, wherein the step of collecting and preprocessing a 2D image of the stadium comprises the step of generating a field coordinate system of the stadium based on the 2D image. Claim 3 delete Claim 4 An automatic reading method according to paragraph 2, wherein the matching step comprises the step of comparing the joint information with the 3D pose database to select the most similar skeleton. Claim 5 In paragraph 4, the automatic reading method comprises a 3D player model including a joint structure projected onto the field coordinate system and a human body mesh. Claim 6 delete Claim 7 delete
Citation Information
Patent Citations
Occluded pedestrian re-identification method based on pose estimation and background suppression
US11908222B1
System and method for officiating interference in sports, powered by artificial intelligence
US12233327B1
Techniques for inferring three-dimensional poses from two-dimensional images
US20210374993A1