Cap behavior recognition method and system based on 3D depth modeling and semantic segmentation
Patent Information
- Application Number
- CN202410279530.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-03-12
AI Technical Summary
[0003]在篮球比赛中,盖帽是一种常见的防守行为,但是否伴随犯规(如打手)往往争议较大
[0023]本公开中,利用计算机视觉以及机器学习方法,通过对篮球运动中防守球员和进攻球员在盖帽时手部区域进行语义分割,通过深度学习的方法将二维图像像素映射到三维点云区域,设置阈值对盖帽是否犯规进行判定,将数据映射三维空间,根据多只手的三维关系,并通过基于机器视觉的智能判断方法,提高了盖帽行为判断的准确性,能够减少裁判判罚争议。
Smart Images

Figure CN118155282B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of behavior recognition technology, specifically to a hat-giving behavior recognition method and system based on 3D depth modeling and semantic segmentation. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] In basketball, blocking is a common defensive action, but whether it is accompanied by a foul (such as a hand strike) is often controversial. When judging whether a foul has occurred manually, the angle and line of sight can make an inaccurate judgment, thus affecting the fairness of the ruling. Therefore, behavior recognition methods based on image recognition are applied.
[0004] The inventors discovered in their research that existing methods for identifying blocked shots commonly include: object detection algorithms to identify athletes in tracking videos; the OpenPose method to identify player arm postures to recognize the entire blocking process; and time-velocity analysis, which compares speed and time differences to determine the blocking action. All of these methods are based on two-dimensional planar visual images. Blocking action judgment based on two-dimensional planar vision is affected by occlusion angles, resulting in low accuracy. Furthermore, multiple players in the same frame occluding each other increases the difficulty of recognition, making it impossible to accurately determine whether a block is a normal block or a foul. 。 Summary of the Invention
[0005] To address the aforementioned issues, this disclosure proposes a method and system for identifying blocking behavior based on 3D depth modeling and semantic segmentation. Using visual 3D depth modeling, semantic segmentation, and deep learning, the method identifies the hand posture of players during a block. By acquiring 3D data and observing the hand postures of both players from multiple angles during the block, the accuracy of foul behavior identification is improved.
[0006] To achieve the above objectives, the present disclosure adopts the following technical solution:
[0007] One or more embodiments provide a hat-capping behavior recognition method based on 3D depth modeling and semantic segmentation, including the following steps:
[0008] Semantic segmentation is performed on each frame of the multiple capped video images from different perspectives to obtain the hand region in each frame image;
[0009] An instance segmentation method is used to distinguish the hands of different subjects, and a pixel-level mask for each hand is output.
[0010] Depth information is extracted from each frame of the image. Based on the obtained mask and depth information, the two-dimensional spatial pixels are mapped to the corresponding regions in the three-dimensional point cloud. The 3D point cloud data obtained from multi-view videos are fused to obtain a complete three-dimensional model of the hand.
[0011] Based on the 3D model of the hand, the coverage area of the hand of different subjects is determined to obtain the result of the capping behavior recognition.
[0012] One or more embodiments provide a cap-giving behavior recognition system based on 3D depth modeling and semantic segmentation, including: a camera device and a processor;
[0013] The camera device is a multi-camera system that uses multiple cameras to capture video from different set shooting angles;
[0014] The processor is configured to execute the above-described hat-capping behavior recognition method based on 3D depth modeling and semantic segmentation.
[0015] One or more embodiments provide a hat-capping behavior recognition system based on 3D depth modeling and semantic segmentation, including:
[0016] Semantic segmentation module: configured to perform semantic segmentation on each frame of the acquired capped video from multiple different viewpoints to obtain the hand region in each frame image;
[0017] Instance segmentation module: Configured to use instance segmentation methods to distinguish the hands of different subjects and output a pixel-level mask for each hand;
[0018] Mapping module: It is configured to map two-dimensional spatial pixels to corresponding regions in three-dimensional point cloud based on the obtained mask and the depth information extracted from each frame of image; and fuse the 3D point cloud data obtained from multi-view video to obtain a complete three-dimensional model of the hand.
[0019] The result recognition module is configured to determine the hand coverage area of different subjects based on the 3D hand model, and obtain the capping behavior recognition result.
[0020] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps in the above-described method for hat-making behavior recognition based on 3D depth modeling and semantic segmentation.
[0021] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the steps in the above-described method for hat-making behavior recognition based on 3D depth modeling and semantic segmentation.
[0022] Compared with the prior art, the beneficial effects of this disclosure are as follows:
[0023] In this disclosure, computer vision and machine learning methods are used to perform semantic segmentation of the hand areas of defensive and offensive players when blocking shots in basketball. Deep learning is used to map two-dimensional image pixels to three-dimensional point cloud regions. A threshold is set to determine whether a block is a foul. By mapping the data to three-dimensional space and considering the three-dimensional relationship between multiple hands, and through intelligent judgment methods based on machine vision, the accuracy of judging blocking behavior is improved, which can reduce referee disputes.
[0024] The advantages of this disclosure, as well as its additional advantages, will be described in detail in the following specific embodiments. Attached Figure Description
[0025] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute a limitation thereof.
[0026] Figure 1 This is a schematic diagram of the identification method process of Embodiment 1 of this disclosure;
[0027] Figure 2 This is a flowchart of the identification method according to Embodiment 1 of this disclosure;
[0028] Figure 3 This is a system block diagram of Embodiment 3 of this disclosure. Detailed Implementation
[0029] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.
[0030] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0031] It should be noted that the terminology used herein is for descriptive purposes only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features within those embodiments can be combined with each other. The embodiments will now be described in detail with reference to the accompanying drawings.
[0032] Based on the problems raised in the background art, this disclosure proposes a method integrating 3D depth modeling and semantic segmentation. By acquiring three-dimensional information and combining semantic segmentation and instance segmentation, it accurately identifies the position and movement of a player's hands. It utilizes the two-dimensional spatial pixels corresponding to the hand region in the image as an index to map to the corresponding region in the three-dimensional point cloud. Based on the three-dimensional relationship of multiple hands, it determines whether a blocking foul has occurred. The intelligent behavior recognition method provided by this disclosure can accurately identify whether a blocking action is a foul, assisting referees in making more accurate judgments and contributing to improving the fairness of the game. Specific embodiments are described below.
[0033] Example 1
[0034] In one or more of the technical solutions disclosed in the embodiments, such as Figures 1 to 2 As shown, a hat-blocking behavior recognition method based on 3D depth modeling and semantic segmentation includes the following steps:
[0035] Step 1: Perform semantic segmentation on each frame of the multiple video clips taken from different perspectives to obtain the hand region in each frame image;
[0036] Step 2: Use instance segmentation to distinguish the hands of different subjects and output the pixel-level mask for each hand;
[0037] Step 3: Based on the obtained mask and the depth information extracted from each frame of the image, map the two-dimensional spatial pixels to the corresponding areas in the three-dimensional point cloud; fuse the 3D point cloud data obtained from the multi-view video to obtain a complete three-dimensional model of the hand.
[0038] Step 4: Determine the hand coverage area of different subjects based on the 3D hand model to obtain the capping behavior recognition result.
[0039] In this embodiment, computer vision and machine learning methods are used to perform semantic segmentation of the hand areas of defensive and offensive players when blocking shots in basketball. Deep learning is used to map two-dimensional image pixels to three-dimensional point cloud regions. A threshold is set to determine whether a block is a foul. The data is mapped to three-dimensional space. Based on the three-dimensional relationship of multiple hands, and through intelligent judgment methods based on machine vision, the accuracy of blocking behavior judgment is improved, which can reduce referee disputes.
[0040] In step 1, a multi-camera system can be set up, using RGB cameras to capture the scene from different angles, and adjusting the angles of different cameras to ensure that the acquired images correspond accurately in three-dimensional space;
[0041] Specifically, the multi-camera system uses RGB cameras to acquire scene video; multiple synchronized high-frame-rate RGB cameras capture the scene from different angles to ensure coverage of all angles requiring monitoring. Furthermore, the cameras are calibrated to ensure that images acquired from different cameras accurately correspond in three-dimensional space.
[0042] The above steps can obtain videos of athletes blocking shots from different angles. The following steps will perform semantic segmentation on the videos, treating each frame as a separate image and annotating each hand instance at the pixel level.
[0043] In some embodiments, semantic segmentation can be achieved by training a semantic segmentation model using an FCN neural network. The FCN neural network may include ordinary convolutional layers and transposed convolutional layers, enabling dense classification of each pixel.
[0044] Specifically, the FCN neural network uses 1×1 convolutional layers and transposed convolutional layers to preserve feature information while restoring the image to its original size in order to extract the hand region in each frame of the image.
[0045] The FCN employs two convolutional operations: a 1×1 convolutional layer to extract features using standard convolution, reducing computation while preserving feature information; and a transposed convolutional layer to upsample the feature map, restoring or increasing the spatial dimensions of the image. By utilizing these two convolutional operations, the FCN restores the image to its original size while preserving feature information, completing image segmentation, identifying the hand region in the image, and outputting the two-dimensional pixel coordinates of the hand region in each frame.
[0046] After semantic segmentation in step 1, instance segmentation is performed to further analyze the hand region and distinguish different instances. Once the hand location is determined, further instance segmentation is performed to differentiate the hands of different athletes during a block.
[0047] In step 2, optionally, Mask R-CNN can be used for instance segmentation based on FCN. The Mask R-CNN instance segmentation method is used to distinguish the hands of different subjects, outputting a pixel-level mask for each hand.
[0048] Specifically, the Mask R-CNN network structure is built on top of the Region Proposal Network (RPN) and Fast R-CNN, and adds a masking branch to generate a binary mask for each hand, assigning a unique identifier to each hand instance.
[0049] Step 3 involves mapping two-dimensional spatial pixels to a three-dimensional point cloud, including the following steps:
[0050] Step 3-1: Locate the pixels with a value of 1 in the mask after instance segmentation, and determine the two-dimensional coordinates of the hand object in the image as an index;
[0051] After the instance segmentation by Mask R-CNN, the pixel-level mask is obtained. The two-dimensional coordinates of the object in the image are determined by finding the pixel with a value of 1 in the mask. This coordinate is then used as an index to map the corresponding region in the three-dimensional point cloud.
[0052] Step 3-2: Use the P2-Net DepthCNN learning module to process individual frames of the video to obtain the corresponding depth information Dt for each frame.
[0053] Step 3-3: Based on the obtained index and the depth information of each frame image, project all pixel values of the two-dimensional plane onto the three-dimensional space to create a three-dimensional point cloud corresponding to each point;
[0054] During the mapping process, it is necessary to establish a fine-grained correspondence between 2D images and 3D point clouds. Specifically, in this embodiment, the P2-Net DepthCNN learning module is used to take individual frames of the video as input and output the corresponding depth information Dt, thereby projecting all pixel values of the two-dimensional plane onto the three-dimensional space, thus creating a three-dimensional point cloud corresponding to each point.
[0055] Furthermore, the obtained 3D point cloud data is fused to obtain a complete 3D hand model for each subject, including the following steps:
[0056] Step 31: Normalize the three-dimensional spatial coordinates of the hand obtained from hand instances with different identifiers to ensure the accuracy of the hand coordinate information.
[0057] Step 32: Based on the normalized coordinates, fuse the 3D point cloud data obtained from all perspectives together to obtain a complete 3D model of the hand.
[0058] Based on multiple sets of images of the hand region obtained from various viewpoints using a camera system, 3D point clouds of the same hand instance are constructed repeatedly. Using deep learning methods, each pixel of the hand is constrained to its corresponding hand spatial region. The 3D point cloud data from all viewpoints are then fused together to obtain a complete 3D model of the hand.
[0059] Using deep learning methods, each pixel of the hand is constrained to a corresponding hand space region, as follows:
[0060] Step 321: Obtain historical images and corresponding 3D point cloud data, as well as annotation information of the hand region, and construct a training set;
[0061] Specifically, the annotation information for the hand area is the 3D coordinates of the hand area;
[0062] Step 322: Construct a deep learning model. The deep learning model includes an encoder and a decoder. Specifically, a convolutional neural network is used as the encoder and a fully connected network is used as the decoder.
[0063] Step 323: Train the deep learning model using historical images and corresponding 3D point cloud data as input and the annotation information of the hand region as output;
[0064] Step 324: Based on the output of the deep learning model and the annotation information of the hand region, calculate the prediction loss, add regularization to prevent overfitting, use the gradient descent optimization method to minimize the loss function, iteratively train and adjust the parameters of the deep learning model to obtain the trained deep learning model.
[0065] Specifically, the prediction loss calculation method uses mean squared error to calculate the error between the predicted 3D coordinates and the actual input 3D coordinates.
[0066] The trained deep learning model is used to identify the hand spatial region in the image.
[0067] Furthermore, visualizing the generated 3D point cloud using a point cloud library allows for a more intuitive determination of blocked fouls.
[0068] In basketball, when a block occurs, obvious fouls and physical contact can be directly judged by the referee. However, hand contact between offensive and defensive players is often difficult to discern with the naked eye. According to the rules for blocking fouls, the criteria for determining a blocking foul are as follows:
[0069] Before the ball leaves the offensive player's hand, if the defensive player touches the ball with one or both hands but does not or only slightly touches the offensive player's hand, it is considered a legal block.
[0070] When a defensive player makes clear and forceful contact with an offensive player's hand or arm while attempting to touch the ball, this is usually ruled a hand foul, also known as a blocking foul.
[0071] In step 4, the blocking behavior recognition is used to determine whether it is a blocking foul. Specifically, it includes the following steps:
[0072] Step 41: For the obtained 3D hand model, use radial basis functions to approximate the surface of the point cloud and extract the isosurface as the reconstructed surface;
[0073] Step 42: On the reconstructed surface, integrate over the surface area to estimate the area;
[0074] Step 43: Based on the set threshold, determine whether it is a blocking foul;
[0075] Specifically, a threshold of 3% can be set. According to the criteria for judging blocking fouls, a slight contact can be considered a legitimate block. With a threshold of 3%, a blocking foul can be judged when the area covered by the hands of the offensive and defensive players exceeds 3%, and a legitimate block can be judged when the area covered by the hands of the offensive and defensive players does not exceed 3%.
[0076] The 3D point cloud region mapped in step 3 is used to distinguish the degree of contact between the defensive and offensive players' hands. If the defensive player's hand only makes slight contact with the offensive player's hand, the mapped mask can show the range of contact; if the defensive player makes forceful contact with the offensive player's hand, the mask will show a larger contact area. A 3% threshold is used as the criterion to characterize whether the blocking action is standard.
[0077] This embodiment proposes a blocking foul determination method based on 3D depth modeling and semantic segmentation, solving the problem of determining blocking fouls in basketball confrontation scenarios. It uses 3D data combined with visual semantic segmentation and instance segmentation to distinguish the hand information of offensive and defensive players. It uses the 2D spatial pixels corresponding to the hand area in the image as an index to map to the corresponding area in the 3D point cloud. Based on the 3D relationship of multiple hands, the foul is accurately determined.
[0078] Example 2
[0079] Based on Embodiment 1, this embodiment provides a hat-covering behavior recognition system based on 3D depth modeling and semantic segmentation, including: a camera device and a processor;
[0080] The camera device is a multi-camera system that uses multiple cameras to capture video from different set shooting angles;
[0081] The processor is configured to execute the capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in Example 1.
[0082] Specifically, the multi-camera system uses RGB cameras to acquire scene video; multiple synchronized high frame rate RGB cameras are used to capture the scene from different angles to ensure coverage of all angles that need to be monitored.
[0083] Example 3
[0084] Based on Example 1, this example provides a hat-covering behavior recognition system based on 3D depth modeling and semantic segmentation, such as... Figure 3 As shown, it includes:
[0085] Semantic segmentation module: configured to perform semantic segmentation on each frame of the acquired video of the cap from multiple different viewpoints to obtain the hand region in each frame image;
[0086] Instance segmentation module: Configured to use instance segmentation methods to distinguish the hands of different subjects and output a pixel-level mask for each hand;
[0087] Mapping module: It is configured to map two-dimensional spatial pixels to corresponding regions in three-dimensional point cloud based on the obtained mask and the depth information extracted from each frame of image; and fuse the 3D point cloud data obtained from multi-view video to obtain a complete three-dimensional model of the hand.
[0088] The result recognition module is configured to determine the hand coverage area of different subjects based on the 3D hand model to obtain the capping behavior recognition result.
[0089] It should be noted that each module in this embodiment corresponds one-to-one with each step in embodiment 1, and their specific implementation process is the same, so it will not be repeated here.
[0090] Example 4
[0091] This embodiment provides an electronic device, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the processor executes the computer instructions, it completes the steps in the capping behavior recognition method based on 3D depth modeling and semantic segmentation in Embodiment 1.
[0092] Example 5
[0093] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, they complete the steps in the capping behavior recognition method based on 3D depth modeling and semantic segmentation in Embodiment 1.
[0094] The electronic devices proposed in this disclosure can be mobile terminals and non-mobile terminals. Non-mobile terminals include desktop computers, and mobile terminals include smartphones (such as Android phones, iOS phones, etc.), smart glasses, smartwatches, smart bracelets, tablets, laptops, personal digital assistants, and other mobile internet devices capable of wireless communication.
[0095] It should be understood that in this disclosure, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0096] The memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store information about the device type.
[0097] In implementation, each step of the above method can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The steps of the method disclosed herein can be directly implemented by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in mature storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0099] In the embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or units may be electrical, mechanical, or other forms.
[0100] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0101] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
[0102] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.
Claims
1. A hat-capping behavior recognition method based on 3D depth modeling and semantic segmentation, characterized in that, Includes the following steps: Semantic segmentation is performed on each frame of the multiple capped video images from different perspectives to obtain the hand region in each frame image; An instance segmentation method is used to distinguish the hands of different subjects, and a pixel-level mask for each hand is output. Extract the depth information of each frame image, and based on the obtained mask and depth information, map the two-dimensional spatial pixels to the corresponding regions in the three-dimensional point cloud; The 3D spatial coordinates of the hand obtained from hand instances with different identifiers are normalized; based on the normalized coordinates, the 3D point cloud data obtained from all perspectives are fused together to obtain a complete 3D hand model. Based on the 3D model of the hand, the coverage area of the hand of different subjects is determined to obtain the blocking behavior recognition result. This includes: for the obtained 3D model of the hand, the radial basis function is used to approximate the surface of the point cloud and the isosurface is extracted as the reconstructed surface; on the reconstructed surface, the area is estimated by integrating on the surface area; and based on the set threshold, it is determined whether it is a blocking foul.
2. The hat-capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in claim 1, characterized in that: A semantic segmentation model is trained using an FCN neural network, which includes regular convolutional layers and transposed convolutional layers. Ordinary convolutional layers perform ordinary convolution to extract features; Transposed convolutional layers upsample feature maps to restore or increase the spatial dimensions of the image.
3. The capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in claim 1, characterized in that, Mapping two-dimensional pixels to a three-dimensional point cloud includes the following steps: Find the pixels with a value of 1 in the mask after instance segmentation, and use the two-dimensional coordinates of the hand object in the image as an index; The P2-Net DepthCNN learning module is used to process individual frames of the video to obtain the corresponding depth information of each frame. Based on the obtained index and the depth information of each frame image, all pixel values of the two-dimensional plane are projected onto the three-dimensional space to create a three-dimensional point cloud corresponding to each point.
4. The hat-capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in claim 1, characterized in that: A multi-camera system is deployed, using RGB cameras to capture the scene from different angles, and the angles of different cameras are adjusted to ensure that the acquired images correspond accurately in three-dimensional space.
5. A hat-capping behavior recognition system based on 3D depth modeling and semantic segmentation, characterized in that, include: Camera device and processor; The camera device is a multi-camera system that uses multiple cameras to capture video from different set shooting angles; The processor is configured to perform the capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in any one of claims 1-4.
6. A hat-capping behavior recognition system based on 3D depth modeling and semantic segmentation, characterized in that, The method for implementing the hat-capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in any one of claims 1-4 includes: Semantic segmentation module: configured to perform semantic segmentation on each frame of the acquired video of the cap from multiple different viewpoints to obtain the hand region in each frame image; Instance segmentation module: Configured to use instance segmentation methods to distinguish the hands of different subjects and output a pixel-level mask for each hand; Mapping module: It is configured to map two-dimensional spatial pixels to corresponding regions in three-dimensional point cloud based on the obtained mask and the depth information extracted from each frame of image; and fuse the 3D point cloud data obtained from multi-view video to obtain a complete three-dimensional model of the hand. The result recognition module is configured to determine the hand coverage area of different subjects based on the 3D hand model to obtain the capping behavior recognition result.
7. An electronic device, comprising a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps in the capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in any one of claims 1-4.
8. A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the steps of the capping behavior recognition method based on 3D depth modeling and semantic segmentation as described in any one of claims 1-4.
Citation Information
Patent Citations
Three-dimensional modeling system for human hand based on posture recognition
CN115620344A
Detection of intentional contact between object and body part of player in sport
US20230033533A1