Gaming activity monitoring system and method

A distributed computing system with machine learning techniques addresses the challenge of monitoring gaming venues by efficiently processing high-resolution images for real-time player tracking and anomaly detection, improving operational efficiency and security.

JP2025183250APending Publication Date: 2025-12-16ANGEL GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025143613
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-30
Filing Date
2025-08-29
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Gaming venues face challenges in efficiently monitoring dynamic and fast-paced gaming activities due to the large size and high computational intensity of image data, which is difficult to process in real-time using existing monitoring systems, especially in environments with numerous gaming tables and players.

Method used

A distributed computing system utilizing machine learning techniques, including object detection, pose estimation, and facial recognition, is deployed to monitor gaming activity by processing high-resolution images from multiple cameras, enabling real-time anomaly detection and player tracking.

Benefits of technology

The system effectively tracks player activities and game objects, providing operational insights and enhancing security by identifying fraudulent behavior in real-time, while overcoming the constraints of space, power, and heat generation in gaming environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025183250000001_ABST
    Figure 2025183250000001_ABST
Patent Text Reader

Abstract

To provide a system, a method and a computer-readable medium for gaming monitoring.SOLUTION: In a system that monitors a gaming activity in a gaming area including a gaming table, a gaming monitoring method includes: causing an edge gaming monitoring device to receive an image captured by a camera system or image data on the image (1310); processing the image so as to determine the presence of a game object on the gaming table in the image (1320); detecting and tracking a game object including one or a plurality of players in the image (1330); determining object detection event data including meta data corresponding to the game object (1340); and transmitting the object detection event data to a gaming monitoring server (1350).SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Technical Field

[0001] The described embodiments relate generally to computer-implemented methods and systems for monitoring gaming activity within a gaming venue. The embodiments apply image processing and machine learning processes to monitor gaming activity. [Background technology]

[0002] background

[0002] Gaming venues, such as casinos, are busy environments with multiple individuals involved in various gaming activities. Gaming venues can be large spaces that accommodate many customers within various portions of the gaming venue. Some gaming venues include tables or gaming tables where various games are played by dealers or operators.

[0003]

[0003] Monitoring of a gaming environment may be performed by an individual responsible for monitoring. The dynamic nature of gaming, the significant number of individuals freely moving about the gaming environment, and the size of the gaming venue often limit the degree of monitoring that can be performed by an individual. Gaming venue operators may benefit from automated monitoring of gaming activity within the gaming venue. Data regarding gaming activity may improve operational management of the gaming venue or facilitate data analysis to determine player rankings (e.g., to award players special bonuses).

[0004]

[0004] Throughout this specification, the term "comprise" or variations thereof such as conjugations thereof will be understood to mean the inclusion of a stated element, integer or step or group of elements, integers or steps, but not the exclusion of any other element, integer or step or group of elements, integers or steps.

[0005]

[0005] In this specification, a statement that an element may be "at least one" of a list of options is to be understood as meaning that the element may be any one of the listed options or any combination of two or more of the listed options.

[0006]

[0006] Any discussion of documents, acts, materials, devices, articles or the like contained in this specification should not be construed as an admission that any or all of such matters form part of the basis of the prior art or were common general knowledge in the field relevant to the present disclosure by virtue of existing before the priority date of each claim of this application. Summary of the Invention [Means for solving the problem]

[0007] overview

[0007] Some embodiments relate to a system for monitoring gaming activity in a gaming area including a gaming table, the system including at least one camera configured to capture images of the gaming area, at least one processor configured to communicate with the at least one camera and a memory, the memory storing instructions executable by the at least one processor to configure the at least one processor to: determine a presence of a first game object on the gaming table in a first image from a series of images of the gaming area captured by the at least one camera, the series of images including one or more images; in response to determining the presence of the first game object in the first image, process the first image to estimate a pose of one or more players in the first image; and determine a first target player among the one or more players associated with the first game object based on the estimated pose.

[0008]

[0008] In some embodiments, at least one processor is further configured to execute instructions stored in the memory to determine, within the series of images, an image area of ​​the face of a first target player that is associated with the first game object.

[0009]

[0009] In some embodiments, processing the image to estimate the posture of one or more players in the first image includes identifying one or more peripheral indicator regions in the first image, each peripheral indicator region corresponding to a distal hand periphery of one or more players.

[0010] In some embodiments, each of the one or more peripheral indicator regions may correspond to a distal left hand perimeter or a distal right hand perimeter.

[0011]

[0011] In some embodiments, the at least one processor is further configured to execute instructions stored in the memory to determine a first target player associated with the first game object by estimating a distance of each peripheral indicator area from the first game object in the first image, identifying a nearest peripheral indicator area based on the estimated distances, and determining a first target player associated with the first game object based on the identified nearest peripheral indicator area.

[0012]

[0012] In some embodiments, estimating the posture of one or more players in the first image includes estimating skeletal models of the one or more players, and determining a first target player associated with the first game object based on the estimated skeletal models of the one or more players.

[0013]

[0013] In some embodiments, estimating a skeletal model of one or more players includes estimating key points in the first image associated with one or more of the wrists, elbows, shoulders, neck, nose, eyes or ears of the one or more players.

[0014]

[0014] In some embodiments, the at least one processor is further configured to execute instructions stored in the memory to determine the presence of a second game object on the gaming table in a second image from the series of images, process the second image to estimate a pose of the player in the second image in response to determining the presence of the second game object in the second image, and determine a second target player from among the players associated with the second game object based on the estimated pose.

[0015] In some embodiments, the first game object comprises any one of a game object, a token, a currency note, or a coin.

[0016] In some embodiments, the memory includes one or more pose estimation machine learning models trained to estimate the pose of one or more players in a sequence of images.

[0017]

[0017] In some embodiments, the one or more pose estimation machine learning models include one or more deep learning artificial neural networks trained to estimate the pose of one or more players in the captured image.

[0018]

[0018] In some embodiments, identifying the face of the first target player further includes determining a vector representation of an image region of the face of the first target player using a facial recognition machine learning model stored in memory.

[0019] In some embodiments, the at least one processor is further configured to estimate a game object value associated with the first game object.

[0020] In some embodiments, determining the presence of the first game object on the gaming table further includes determining a location on the gaming table area of ​​the first game object.

[0021]

[0021] In some embodiments, at least one processor is further configured to execute instructions stored in the memory to identify, within the series of images, a plurality of facial regions corresponding to the face of a first target player associated with the first game object.

[0022]

[0022] In some embodiments, the at least one processor is further configured to execute instructions stored in the memory to process the plurality of facial regions to determine facial orientation information of the face of the first target player in each of the plurality of facial regions.

[0023]

[0023] In some embodiments, the at least one processor is further configured to execute instructions stored in the memory to process facial orientation information of the first target player's face in each of the plurality of facial regions to determine the most frontal facial region corresponding to the target player.

[0024]

[0024] Some embodiments relate to a method of monitoring gaming activity in a gaming area including a gaming table, the method including providing at least one camera configured to capture images of the gaming area, at least one processor configured to communicate with the at least one camera, and a memory storing instructions executable by the at least one processor; determining, by the at least one processor, the presence of a first game object on the gaming table in a first image from a series of images of the gaming area captured by the at least one camera, the series of images including one or more images; in response to determining the presence of the first game object in the first image, processing, by the at least one processor, the first image to estimate a pose of one or more players; and determining, by the at least one processor, a first target player from the one or more players associated with the first game object based on the estimated pose.

[0025] Some embodiments further include identifying, by at least one processor, an image region of the first target player's face within the sequence of images that is associated with the first game object.

[0026]

[0026] In some embodiments, estimating the posture of one or more players in the first image includes identifying one or more peripheral indicator regions in the first image, each peripheral indicator region corresponding to an area around a distal portion of a hand of one or more players.

[0027] In some embodiments, each of the one or more peripheral indicator regions may correspond to a distal left hand perimeter or a distal right hand perimeter.

[0028]

[0028] Some embodiments further include determining a first target player associated with the first game object by at least one processor estimating a distance of each peripheral indicator area from the first game object in the first image, identifying a nearest peripheral indicator area based on the estimated distances, and determining a first target player associated with the first game object based on the identified nearest peripheral indicator area.

[0029]

[0029] In some embodiments, estimating the posture of one or more players in the first image includes estimating a skeletal model of the one or more players, and determining a first target player associated with the first game object is based on the estimated skeletal model of the one or more players.

[0030]

[0030] In some embodiments, estimating a skeletal model of one or more players includes estimating key points in the first image associated with one or more of the wrists, elbows, shoulders, neck, nose, eyes or ears of the one or more players.

[0031]

[0031] Some embodiments further include determining, by at least one processor, the presence of a second game object on the gaming table in a second image from the series of images; in response to determining the presence of the second game object in the second image, processing, by at least one processor, the second image to estimate a posture of one or more players; and based on the estimated posture, determining, by the at least one processor, a second target player from the one or more players associated with the second game object.

[0032] In some embodiments, the first game object comprises any one of a game object, a token, a currency note, or a coin.

[0033] In some embodiments, the memory includes one or more pose estimation machine learning models trained to estimate the pose of one or more players in a sequence of images.

[0034]

[0034] In some embodiments, the one or more pose estimation machine learning models include one or more deep learning artificial neural networks trained to estimate the pose of one or more players in the captured image.

[0035]

[0035] In some embodiments, identifying an image region of the first target player's face by at least one processor includes extracting a vector representation of the first target player's face using a facial recognition machine learning model stored in memory.

[0036]

[0036] The method of some embodiments further includes estimating, by the at least one processor, a game object value associated with the first game object.

[0037] In some embodiments, determining the presence of the first game object on the gaming table further includes determining a gaming table area associated with the first game object.

[0038]

[0038] The method of some embodiments further includes identifying a plurality of facial regions within the sequence of images that correspond to a face of a first target player associated with the first game object.

[0039]

[0039] The method of some embodiments further includes processing the plurality of facial regions to determine facial orientation information of the face of the first target player in each of the plurality of facial regions.

[0040]

[0040] The method of some embodiments further includes processing facial orientation information of the face of the first target player in each of the plurality of facial regions to determine the most frontal facial region corresponding to the target player.

[0041]

[0041] Some embodiments relate to a system for monitoring gaming activity in a gaming area including a gaming table, the system including at least one camera configured to capture images of the gaming area, an edge gaming monitoring computing device provided in proximity to the gaming table, the edge gaming monitoring computing device including a memory and at least one processor having access to the memory, the at least one processor configured to communicate with the at least one camera, the memory storing instructions executable by the at least one processor to configure the at least one processor to: determine a presence of a first game object on the gaming table in a first image from a series of images of the gaming area captured by the at least one camera; track the first game object in a plurality of images in the series of images in response to determining the presence of the first game object in the first image; determine object detection event data including image data extracted from the plurality of images and metadata corresponding to the first game object; and transmit the object detection event data to a gaming monitoring server.

[0042] In some embodiments, the object detection event data is determined in response to tracking of the first game object in at least two or more of the plurality of images in the sequence of images.

[0043] In some embodiments, the at least one processor is further configured to determine the presence of one or more players within the sequence of images of the gaming area.

[0044] In some embodiments, the image data extracted from the plurality of images includes image data corresponding to one or more players and image data corresponding to a first game object.

[0045]

[0045] Some embodiments relate to a system for monitoring gaming activity within a gaming area, the system including a gaming monitoring server including at least one processor configured to communicate with a memory, the memory storing instructions executable by the at least one processor to configure the at least one processor to: receive object detection event data from a table gaming monitoring computing device, the object detection event data including image data corresponding to a plurality of images and metadata corresponding to a first game object; process the image data corresponding to the plurality of images to estimate poses of the one or more players; and determine a first target player from the one or more players associated with the first game object based on the estimated poses.

[0046] In some embodiments, the at least one processor is further configured to identify, within the plurality of images, a plurality of facial regions of the first target player associated with the first game object.

[0047]

[0047] In some embodiments, the at least one processor is further configured to process multiple facial regions of the first target player to determine facial orientation information of the first target player's face in each of the multiple facial regions.

[0048]

[0048] In some embodiments, the at least one processor is further configured to process facial orientation information of the face of the first target player in each of the plurality of facial regions to determine the most frontal facial region corresponding to the target player.

[0049] In some embodiments, the most frontal face region corresponding to the target player relates to the face region that is the most information-rich face region for face recognition operations.

[0050]

[0050] Some embodiments relate to a method for monitoring gaming activity in a gaming area including a gaming table, the method including providing at least one camera configured to capture images of the gaming area, at least one processor configured to communicate with the at least one camera, and memory storing instructions executable by the at least one processor; determining a presence of a first game object on the gaming table in a first image from a series of images of the gaming area captured by the at least one camera; tracking the first game object in a plurality of images in the series of images in response to determining the presence of the first game object in the first image; determining object detection event data including image data extracted from the plurality of images and metadata corresponding to the first game object; and transmitting the object detection event data to a gaming monitoring server.

[0051] In some embodiments, the object detection event data is determined in response to tracking of the first game object in at least two or more of the plurality of images in the sequence of images.

[0052]

[0052] The method of some embodiments further includes determining the presence of one or more players within the sequence of images of the gaming area.

[0053] In some embodiments, the image data extracted from the plurality of images includes image data corresponding to one or more players and image data corresponding to a first game object.

[0054]

[0054] Some embodiments relate to a method of monitoring gaming activity in a gaming area, the method including, optionally, providing a gaming surveillance server including at least one processor configured to communicate with a memory storing instructions executable by the at least one processor; receiving, at the gaming surveillance server, object detection event data from a table gaming surveillance computing device, the object detection event data including image data corresponding to a plurality of images and metadata corresponding to a first game object; processing, by the gaming surveillance server, the image data to estimate poses of one or more players in the plurality of images; and determining, by the gaming surveillance server based on the estimated poses, a first target player from the one or more players associated with the first game object.

[0055] In some embodiments, the method includes identifying, within the plurality of images, a plurality of facial regions of the first target player associated with the first game object.

[0056]

[0056] The method of some embodiments further includes processing a plurality of facial regions of the first target player to determine facial orientation information of the first target player's face in each of the plurality of facial regions.

[0057]

[0057] The method of some embodiments further includes processing facial orientation information of the face of the first target player in each of the plurality of facial regions to determine the most frontal facial region corresponding to the target player.

[0058]

[0058] In some embodiments, the most frontal face region corresponding to the target player relates to the face region that is the most information-rich face region for face recognition operations.

[0059]

[0059] In some embodiments, the gaming monitoring server includes a gaming monitoring server located on the gaming premises or a gaming monitoring server located remotely from the gaming premises.

[0060]

[0060] In some embodiments, the gaming surveillance server includes a gaming surveillance server located on the gaming premises or a gaming surveillance server located remotely from the gaming premises.

[0061]

[0061] In some embodiments, the gaming monitoring server includes a secure data storage component for object detection event data and determined target player information.

[0062]

[0062] In some embodiments, at least one camera and edge gaming surveillance computing device may be part of a smartphone or tablet computing device.

[0063] In some embodiments, the captured image includes a depth of field image, and the determination of the presence of the first game object on the gaming table is based on the depth of field image.

[0064]

[0064] Some embodiments relate to a non-transitory computer-readable storage medium storing program code that, when executed by at least one processor, configures the at least one processor to perform any one of the methods of the embodiments. [Brief explanation of the drawings]

[0065] Brief description of the diagram [Figure 1]

[0065] FIG. 1 is a block diagram of a gaming monitoring system according to some embodiments. [Figure 2]

[0066] 1 is an image illustrating a portion of a method for pose estimation according to some embodiments. [Figure 3]

[0067] 1 is an image illustrating a portion of a method for pose estimation according to some embodiments. [Figure 4]

[0068] 2 is an image of an exemplary gaming environment monitored by the gaming monitoring system of FIG. 1. [Figure 5A]

[0069] 1 is an image illustrating a portion of a method for pose estimation according to some embodiments. [Figure 5B]

[0069] An image illustrating a portion of a method for pose estimation according to some embodiments. [Figure 5C]

[0069] An image illustrating a portion of a method for pose estimation according to some embodiments. [Figure 5D]

[0069] An image illustrating a portion of a method for pose estimation according to some embodiments. [Figure 6]

[0070] 1 is a flowchart of a method of gaming monitoring according to some embodiments. [Figure 7]

[0071] 1 is an image of an exemplary gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. [Figure 8]

[0072] 1 is another image of an exemplary gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. [Figure 9]

[0073] 1 is another image of an exemplary gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. [Figure 10]

[0074] 1 is another image of an exemplary gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. [Figure 11]

[0075] FIG. 1 is a block diagram of a gaming monitoring system according to some embodiments. [Figure 12]

[0076] FIG. 12 is a block diagram of a portion of the gaming monitoring system of FIG. 11 according to some embodiments. [Figure 13]

[0077] 12 is a flowchart of a portion of a method of gaming monitoring performed by the table gaming monitoring device of FIG. 11. [Figure 14]

[0078] 12 is a flowchart of a portion of a method of gaming monitoring performed by the on-premise gaming monitoring server of FIG. 11. [Figure 15]

[0079] 1 is an image of an exemplary gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. [Figure 16]

[0080] FIG. 1 is a schematic diagram of an example of determining the distance between two bounding boxes. [Figure 17]

[0081] 1 illustrates an exemplary computer system architecture according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0066] Detailed Description

[0082] Various table-based games are played at gaming venues. Games may include, for example, baccarat, blackjack, roulette, and craps. Such games may involve a chance event or a series of chance events with a chance outcome or an unpredictable outcome on which a player, participant, or patron may place a wager or a series of wagers. The chance event may include the drawing of a card, the dealing of a card, the throwing of a die, or the spinning of a roulette wheel. Players participate in games by placing game objects in several locations on the gaming table. Game objects may include, for example, chips, cards, or tokens, or coins or bills issued by the gaming venue. In some games, the gaming table defines zones or areas associated with specific outcomes of the game. For example, in a game of baccarat, the gaming table includes zones or regions on the table surface corresponding to the player and the banker. A bet on a specific outcome or chance event in a game may be placed by a patron by placing a game object within the respective zone or region associated with the specific outcome.

[0067]

[0083] In some games, such as blackjack, players may sit at a specific portion of a gaming table and bet on each hand of cards by placing game objects within a designated zone or area of ​​the table. However, it is also possible for a seated player to bet on the hands of other seated players at the gaming table. It is also possible for a player not seated at the gaming table to bet on one or more hands of seated players. This practice is known as backbetting. When several players participate in a game, some seated and others not seated, and each player places wagers in different zones or areas within a fast-paced gaming environment, it can be difficult to monitor each player's activities. Additionally, players may move between various gaming tables within a venue during a visit, making it more difficult to monitor each player during that visit.

[0068]

[0084] Due to the dynamic and fast-paced nature of gaming environments, monitoring and supervising gaming events using image data can be highly computationally intensive. High-resolution image data is often required to identify objects or identify events within gaming environment image data with reasonable confidence. For example, image data having resolutions of 720p (1280 x 720), 1080p (1920 x 1080), 4MP (2560 x 1920), or higher may be captured at 10 frames per second or higher. A gaming premises or venue may include numerous gaming tables or gaming environments. For example, a gaming premises may include 1000 or more gaming tables or gaming environments. Each gaming environment may include a gaming table or gaming area where gaming can occur. Each gaming environment may be monitored using one or more sensors, including cameras and / or other imaging and / or ranging sensors.

[0069]

[0085] For example, a camera capturing image data at 1080p resolution at 30 frames per second may generate image data at a rate of 2.5 Mbps (megabits per second). In a gaming premises or venue with hundreds or thousands of gaming tables or gaming environments, each gaming environment may be equipped with two cameras, and all image data may be generated at a rate of, for example, 5 Gbps. In some embodiments, the image data may also include data from a camera capturing images in an image spectrum not visible to the naked eye (e.g., the infrared spectrum). In some embodiments, the image data may also include data from a depth sensor or depth camera or 3D camera that captures depth or 3D scene information of the gaming environment. Additional image data sources may further increase the volume and rate of image data generated from the supervision of the gaming environment.

[0070]

[0086] According to some embodiments of the present disclosure, the significant rate and volume of data generated by sensors monitoring a gaming environment can be addressed by specific distributed computing architectures to efficiently process image data and derive insights from the captured data.

[0071]

[0087] Gaming environments also impose additional constraints on the deployment of distributed computing systems. For example, the placement of computing devices within a gaming environment (e.g., near or under a gaming table) for the performance of processing-power-intensive operations may generate undesirable amounts of heat, posing safety risks such as the risk of fire, or may require additional cooling equipment, requiring additional expense, space, and power. Providing cooling capacity within a gaming environment within the tight constraints of a gaming environment (including physical space, power, and security constraints) may be impractical.

[0072]

[0088] Some embodiments provide the computing power to effectively monitor large gaming environments while providing improved distribution of computing operations within a distributed computing environment deployed within a gaming premises to meet constraints imposed by the gaming environment. Some embodiments also provide a distributed monitoring system that can be scaled to cover larger premises or dynamically scaled depending on fluctuations in occupancy within the premises. Some embodiments also allow for dynamic variation in the degree of monitoring performed by the distributed monitoring system or the degree of monitoring capacity implemented. For example, additional computing monitoring capacity can be efficiently deployed by the distributed monitoring system across some or all gaming environments within a gaming premises using a distributed computing system. Some embodiments also enable scaling of the computing resources of the distributed monitoring system to monitor more than one gaming premises.

[0073]

[0089] Some embodiments relate to computer-implemented methods and systems that use machine learning techniques to monitor gaming activity within a gaming premises and assist gaming venue operators in responding expediently to gaming anomalies or fraudulent activity using distributed computing systems and / or devices.

[0074]

[0090] Some embodiments relate to computer-implemented methods and systems for monitoring gaming activity within a gaming venue. The embodiments incorporate one or more cameras positioned to capture images of a gaming area, including gaming tables. The one or more cameras are positioned to capture images of the gaming tables as well as images of players near the gaming tables participating in games. The embodiments incorporate image processing techniques, including object detection, object tracking, pose estimation, image segmentation, and facial recognition, to monitor gaming activity. The embodiments rely on machine learning techniques, including deep learning techniques, to perform various image processing tasks. Some embodiments that perform gaming monitoring tasks in real time or near real time use machine learning techniques to assist gaming venue operators in expediently responding to gaming anomalies or fraudulent activity.

[0075]

[0091] Some embodiments relate to a method of gaming monitoring. The method may include receiving, by a computing device, a sequence of images and timestamp information of the capture time of each image in the sequence of images. The sequence of images may be in the form of a video stream. Each image in the sequence of images may be an image of a gaming environment. The computing device processes a first image in the sequence of images to determine a first event trigger indicator in the first image. The computing device may be configured to identify a gaming monitoring initiation event based on the determined first event trigger indicator. In response to identifying the gaming monitoring initiation event, the computing device may initiate transmission of image data of a first image or images in the sequence of images captured subsequent to the first image to an upstream computing device. The computing device may be configured to process a second image in the sequence of images to determine a second event trigger indicator in the second image. The second image may be an image captured subsequent to the first image. The computing device may identify a gaming monitoring termination event based on the determined second trigger indicator. In response to identifying the gaming monitoring termination event, the computing device may be configured to terminate transmission of the image data.

[0076]

[0092] In some embodiments, a method of gaming monitoring may include receiving, at an upstream computing device, image data and corresponding timestamp information for an image captured within a gaming environment from a computing device. The upstream computing device may be configured to process the received image data and image regions corresponding to the detected game objects to detect game objects. The upstream computing device may be configured to process image regions corresponding to the identified game objects to determine game object attributes.

[0077]

[0093] In some embodiments, a method of gaming surveillance may include receiving, by a computing device, a series of images and timestamp information of the capture time of each image in the series of images. Each image in the series of images may be an image of a gaming environment. In embodiments, the series of images may be received as a video feed or video. The computing device may process a first image in the series of images to determine an event trigger indicator in the first image. The computing device may identify a gaming surveillance event based on the determined event trigger indicator. The computing device may transmit image data of the first image and an image closest to the first image in the series of images to an upstream computing device. Determining the event trigger indicator may include detecting a first game object in the first image that may not have been detected in images captured prior to the first image in the series of images.

[0078]

[0094] Some embodiments may relate to a computing device or edge computing device or upstream computing device configured to perform gaming monitoring according to the above methods. Some embodiments relate to a computer-executable medium storing program code that, when executed by a processor, configures the processor to perform gaming monitoring methods according to the above methods.

[0079]

[0095] FIG. 1 is a block diagram of a gaming monitoring system 100 according to some embodiments. The gaming monitoring system 100 may be configured to monitor gaming activity on a particular gaming table 140 within a gaming venue. The gaming monitoring system 100 may include a first camera 110. The gaming monitoring system 100 may optionally include a second camera 112. In some embodiments, the gaming monitoring system 100 may include two or more cameras, each capturing images of the table surface from various perspectives of the gaming table. The first camera 110 is positioned to capture images of the upper (playing) surface of the gaming table 140 and players near the gaming table participating in gaming activity. The first camera 110 is also positioned to include the faces of various players in the captured images. The second camera 112 and any other cameras may similarly be positioned at various locations around the gaming table 140 to capture images of the gaming table 140 and players participating in gaming from various angles.

[0080]

[0096] The first camera 110 and / or the second camera 112 may capture images at a resolution of, for example, 1280 x 720 pixels or greater. The first camera 110 and / or the second camera 112 may capture images at a rate of, for example, 20 frames per second. In some embodiments, the first camera 110 and / or the second camera 112 may include a See3CAM130, a UVC-compliant AR1335 sensor-based 13MP autofocus USB camera. In some embodiments, the first camera 110 and / or the second camera 112 may include an AXIS P3719-PLE network camera. Any additional (e.g., third, fourth, fifth, sixth) cameras positioned to capture images of the gaming table 140 may have similar features or operating parameters as those set forth above for the first and second cameras 110, 112.

[0081]

[0097] The first and second cameras 110, 112 (and any additional cameras included as part of the system 100) are positioned or mounted on a wall, pedestal, or rod with a substantially uninterrupted view of the table playing surface, allowing the player to view the table from the dealer's side of the table 140. toward the player side and away from the outside or edge of the table 140 (toward the table playing surface). and oriented to capture images in a viewing direction (at least partially horizontal across the viewing direction), e.g. If necessary, leave at least 1 meter (maximum approx. 2 meters) of vertical space above the table surface and Captures the image area including the vertical space above the player side and the image area below the player side.

[0082]

[0098] Gaming surveillance system 100 also includes gaming surveillance computing device 120. Gaming surveillance computing device 120 is configured to communicate with cameras 110 and 112, as well as any additional cameras present as part of system 100. For example, communication between computing device 120 and cameras 110 and 112 may be provided via a wired medium, such as a Universal Serial Bus cable. In some embodiments, communication between computing device 120 and cameras 110 and 112 may be provided via a wireless medium, such as a Wi-Fi™ network or other short-range, low-power wireless network connection. In some embodiments, cameras 110 and 120 may be connected to a computer network and configured to communicate over the computer network using, for example, the Internet Protocol.

[0083]

[0099] Computing device 120 may be positioned near the gaming table 140 being monitored. For example, computing device 120 may be located in an enclosed room or cavity beneath the gaming table. In some embodiments, computing device 120 may be positioned away from gaming table 140 but may be configured to communicate with cameras 110 and 112 over a wired or wireless communication link.

[0084]

[0100] The computing device 120 includes at least one processor 122 in communication with a memory 124 and a network interface 129. The memory 124 may include both volatile and non-volatile memory. The network interface 129 may enable communication with other devices, such as the cameras 110, 112, and over a network 130.

[0085]

[0101] The memory 124 stores executable program code for providing various computational capabilities of the gaming monitoring system 100 described herein. The memory 124 includes at least an object detection module 123, a pose estimation module 125, a game object value estimation module 126, a face recognition module 127, and a game object association module 128.

[0086]

[0102] Various modules stored in memory 124 for execution by at least one processor 122 may incorporate or have functional access to machine learning-based data processing models or computational structures to perform various tasks related to monitoring gaming activity. In particular, the software code modules of various embodiments may have access to AI models incorporating deep learning-based computational structures, including artificial neural networks (ANNs). ANNs are computational structures inspired by biological neural networks and include one or more layers of artificial neurons configured or trained to process information. Each artificial neuron includes one or more inputs and an activation function to process received inputs and generate one or more outputs. The outputs of each layer of neurons are connected to subsequent layers of neurons using links. Each link may have a predetermined numerical weighting that determines the strength of the link as information progresses through the layers of the ANN. During the training phase, the various weightings and other parameters defining the ANN are optimized to obtain a trained ANN using inputs and known outputs of the inputs. Optimization may occur through various optimization processes, including backpropagation. An ANN incorporating deep learning techniques includes several hidden layers of neurons between a first input layer and a final output layer, which allows the ANN to model complex information processing tasks, including those of object detection, pose estimation, and face recognition performed by gaming surveillance system 100.

[0087]

[0103] In some embodiments, the various modules implemented within memory 124 may incorporate one or more variations of a convolutional neural network (CNN) (a type of deep neural network) to perform various image processing operations for gaming surveillance. A CNN includes various hidden layers of neurons between an input layer and an output layer to convolve an input and generate an output through the various hidden layers of neurons.

[0088]

[0104] The object detection module 123 includes program code for detecting specific objects in images received by the computing device 120 from the cameras 110 and 112. Objects detected by the object detection module 123 may include game objects, such as chips, cash, coins, or bills placed on the gaming table 140. The object detection module 123 may be trained to determine areas or regions of the gaming table where game objects are present or may be detected. The results of the object detection process performed by the object detection module 123 may be or include information regarding the class to which each identified object belongs and the location or area of ​​the gaming table where the identified object is detected. The location of the identified object may be indicated, for example, by the image coordinates of a bounding box surrounding the detected object or an identifier of the area of ​​the gaming table in one or more images where the object was detected. The object detection results may also include, for example, a probability number associated with a confidence level of the accuracy of the identified object class. The object detection module 123 may also include program code for identifying people, human faces, or specific body parts in images. The object detection module 123 may include a game object detection neural network 151 trained to process images of a gaming table and detect game objects placed on the gaming table. The object detection module 123 may include a person detection neural network 159 trained to process images and detect one or more people or parts of one or more people (e.g., faces) in the images. The object detection module 123 may generate results (as output of the neural network 159) in the form of coordinates in the processed images that define a rectangular bounding box around each detected object. The bounding boxes may overlap for objects that may be placed next to each other or partially overlap in the image.

[0089]

[0105] The object detection module 123 may incorporate a region-based convolutional neural network (R-CNN) or one of its variants (e.g., including Fast R-CNN, or Fast R-CNN, or Mask R-CNN) to perform object detection. R-CNN may include three modules: a region proposal module, a feature extractor module, and a classifier module. The region proposal module is trained to determine one or more candidate bounding boxes around potentially detected objects in an input image. The feature extractor module processes a portion of the input image corresponding to each candidate bounding box to obtain a vector representation of the features within each candidate bounding box. In some embodiments, the vector representation generated by the feature extractor module may include 4096 elements. The classifier module processes the vector representation to identify the class of object present in each candidate bounding box. The classifier module generates a probability score representing the likelihood of the presence of each class or object in each candidate bounding box. For example, for each candidate bounding box, the classifier module may generate a probability of whether the bounding box corresponds to a person or a game object.

[0090]

[0106] Based on the probability scores generated by the classifier module and a predetermined threshold, an assessment may be made regarding the class of the object present within the bounding box. In some embodiments, the classifier may be implemented as a support vector machine. In some embodiments, the object detection module 123 may incorporate a pre-trained ResNet-based convolutional neural network (e.g., ResNet-50) for feature extraction from images to enable object detection operations.

[0091]

[0107] In some embodiments, the object detection module 123 may incorporate a You Look Only Once (YOLO) model for object detection. The YOLO model includes a single neural network trained to process an input image and directly predict bounding boxes and class labels for each bounding box. The YOLO model divides the input image into a grid of cells. Each cell in the grid is processed by the YOLO model to determine one or more bounding boxes that contain at least a portion of the cell. The YOLO model is also trained to determine a confidence level associated with each bounding box and an object class probability score for each bounding box. The YOLO model then considers each bounding box determined from each cell and its respective confidence and object class probability score to determine a final set of reduced bounding boxes around objects that have object class probability scores higher than a predetermined threshold object class probability score.

[0092]

[0108] In some embodiments, object detection module 123 implements one or more image processing techniques described in published PCT applications "System and method for machine learning driven object detection" (Publication No. WO 2019 / 068141) or "System and method for automated table game activity recognition" (Publication No. WO 2017 / 197452), the contents of which are incorporated herein by reference.

[0093]

[0109] Pose estimation module 125 includes executable program code for processing one or more images of players within a gaming environment to identify a pose of one or more players. Each identified pose may include the location of an area within the image that corresponds to a particular body part of the player. For example, the identified body part may include a left or right hand, a left or right wrist, a distal left or right hand area, or a face within the image.

[0094]

[0110] The pose estimation module 125 may be configured to identify the poses of multiple people in a single image without prior knowledge of the number of people in the image. Because gaming venues are dynamic and fast-paced environments with multiple customers passing through various parts of the venue, the ability to identify multiple people facilitates improving the monitoring capabilities of the gaming surveillance system 100. The pose estimation module 125 may include a keypoint estimation neural network trained to estimate keypoints corresponding to specific portions of one or more people in an input image. The pose estimation module 125 may include a 3D mapping neural network trained to map pixels associated with one or more people in an image to a 3D surface model of the person.

[0095]

[0111] In some embodiments, pose estimation may involve a top-down approach, where people in an image are first identified, followed by identifying the pose or various parts of the person. Object detection module 123 may be configured to identify portions or regions of the image that correspond to a single person. Pose estimation module 125 may rely on the identified portions or regions of the image that correspond to a single person and may process each identified portion or region of the image to identify the pose of the person.

[0096]

[0112] In some embodiments, pose estimation may involve a bottom-up approach, where various body parts of all people in an image are first identified, followed by a process of establishing relationships between the various parts to identify the pose of each person in the image. The object detection module 123 may be configured to identify portions or regions of the image that correspond to specific body parts of a person (e.g., face, hands, shoulders, legs, etc.). Each specific portion or region in the image that corresponds to a specific body part may be referred to as a keypoint. The pose estimation module 125 may receive information about the identified keypoints (e.g., coordinates of each keypoint and coordinates of the body part associated with each keypoint) from the object detection module 123. Based on this received information, the pose estimation module 125 may associate the identified keypoints with each other to identify the pose of one or more people in the image.

[0097]

[0113] In some embodiments, the pose estimation module 125 may incorporate the OpenPose framework for pose estimation. The OpenPose framework includes a first feedforward ANN trained to identify body part locations within an image in the form of a confidence map. The confidence map includes identifiers of parts identified within a region of the image and a confidence level in the form of a probability of confidence associated with the detection. The first feedforward ANN is also trained to determine part similarity field vectors for the identified parts. The part similarity field vectors represent the associations or similarities between the parts identified in the confidence map. The determined part similarity field vectors and the confidence map are iteratively filtered by a convolutional neural network (CNN) to remove weaker part affinities and ultimately predict the pose of one or more people in the image. The output of the pose estimation module 125 may include the identified coordinates or parts (keypoints) for each person identified in the image and an indicator of the class to which each part belongs (e.g., whether the identified part is a wrist, hand, or knee).

[0098]

[0114] 2 shows several example images conceptually illustrating various keypoints associated with various parts of the human body and determined by pose estimation module 125. Image 210 shows various keypoints corresponding to specific parts of the human body. For example, keypoints 4 and 7 correspond to the wrist, and keypoints 9 and 12 correspond to the knee. Image 220 shows various limbs or parts of a human body model based on the keypoints in image 210. For example, points 16 and 24 correspond to the part of the person between the elbow and wrist. The combination of various keypoints or limbs or parts of the human body may be referred to as the skeletal model determined by pose estimation module 125.

[0099]

[0115] In some embodiments, the pose estimation module 125 may incorporate a pose analysis or estimation framework for pose estimation. The pose estimation module 125 framework may map pixels in a 2D image corresponding to a person to a 3D model of the surface of a human body. In some embodiments, the pose estimation module 125 framework may identify pixels in an image corresponding to a person and map each pixel to a 3D surface model of a human body. The pose estimation module 125 framework may enable the identification of a more complete representation of a person's pose in an image. For example, the pose estimation module 125 framework may enable the differentiation of pixels in an image corresponding to the outer surface of a person's hand from pixels in an image corresponding to the person's palm, thereby mapping the 2D image to a 3D surface model of a human body. Another example may include the differentiation of a person's face from the back of a person's head. The pose estimation module 125 framework may utilize a trained convolutional neural network to perform this task.

[0100]

[0116] In some embodiments, the operation of the pose estimation module 125 framework may be limited to detecting only the player's hands and head to further improve the computational efficiency of the pose estimation process. In some embodiments, the open source DensePose framework may be incorporated within the pose estimation module 125 to perform pose estimation. In some embodiments, the pose estimation module 125 framework may be deployed using the Torchserve framework for deployments that enable high-throughput pose estimation operations and scalable distributed execution using multiple processors, including multiple graphics processing units.

[0101]

[0117] FIG. 3 illustrates an example of a 3D surface model of a human body in image 310, a 3D surface model of a 2D representation in image 320, and an example mapping between the 3D surface model and a 2D image-within-image 330. The 3D surface model shown in image 310 includes a human body decomposed into several separate portions. For example, portions 312 and 314 in the image correspond to a 3D surface model of the tip of a person's hand, including the palm and the outer surface of the hand. Portion 316 corresponds to one side of the head. Each portion is parameterized using a coordinate system such as that illustrated in mapping 330. The pose estimation module 125 can process the image through a pose estimation framework and identify which 3D surface portion in the 3D surface model of image 310 a pixel corresponds to. The output of the pixel mapping can include an identifier of the portion (e.g., left or right hand, or head) and the U,V coordinates (illustrated at 330) associated with the 3D surface model to which the pixel is mapped.

[0102]

[0118] The pose estimation module 125 is trained on a training dataset that includes several examples of images having individuals in cluttered environments, where various parts of the individual's body are obscured by one or more other individuals in the image. For example, an individual sitting at a table may obscure parts of the arm of an individual standing behind them. The training dataset that includes images having multiple players in a cluttered environment enables the neural network included in the pose estimation module 125 to model the cluttered gaming environment and generate accurate pose estimates in images of a gaming environment having multiple players and where the players are partially obscured or occluded by one another.

[0103]

[0119] In some embodiments, a large-scale image dataset may be used to train the neural network included in the pose estimation module 125. The large-scale image dataset may be a large-scale object detection / segmentation / captioning dataset. This dataset may include information about wrist keypoints. In some embodiments, a Common Objects in Context (COCO) dataset may be used to train the neural network included in the pose estimation module 125.

[0104]

[0120] In some embodiments, the large-scale image dataset may include distal hand-periphery keypoints within images therein. Accordingly, the large-scale image dataset may be used to train the keypoint estimation neural network 152 to identify poses of individuals within images, the identified poses including, for example, a skeletal model of the player and distal hand-periphery keypoints. The large-scale image dataset may be large and thus may include example images from a variety of environments. The training dataset may enable the training of a robust keypoint estimation neural network 152 suitable for application within cluttered, varied, and fast-paced gaming environments.

[0105]

[0121] The face recognition module 127 includes program code for processing an image to identify regions within the image that correspond to faces. The face recognition module 127 may determine regions within the image that correspond to faces based on one or more keypoints that correspond to facial features (e.g., eyes, nose, mouth, ears) determined by the pose estimation module 125. Upon identifying regions within the image that correspond to faces, the face recognition module 127 is configured to extract image features from the identified regions to enable face recognition. The face recognition module 127 may include one or more trained machine learning models to perform image processing steps for face recognition.

[0106]

[0122] In some embodiments, the face recognition module 127 may include a deep neural network, such as a convolutional neural network, to perform face recognition. In some embodiments, the face recognition module 127 may incorporate the FaceNet framework for face recognition. An embodiment incorporating the FaceNet face recognition framework may include a neural network configured to process images of faces, map facial features into Euclidean space, and obtain embedded representations of the facial features in the images. The neural network trained to obtain the embedded representations may be trained using the triplet loss training principle. According to the triplet loss training principle, the loss or error during training is calculated using two positive examples and one negative example per training dataset. A neural network trained using the triplet loss training principle generates embedded representations that are very close in Euclidean space for two different images of the same person, but far from the embedded representations of any other individuals. The embedded representations are in the form of vectors in feature space that allow comparison of the embedded representations with images of known individuals, enabling recognition of faces in the images. In some embodiments, the embedded representations may be in the form of 512-element vectors.

[0107]

[0123] In some embodiments, gaming monitoring computing device 120 is configured to communicate with gaming monitoring server 180 over network 130. Gaming monitoring server 180 may provide central computing power for system 100, enabling centralized performance of one or more monitoring operations. For example, in some embodiments, gaming monitoring server 180 may include or have access to memory 184 in communication with processor 182, memory 184 including facial feature and identifier database 189. Gaming monitoring server 180 may communicate with multiple (and possibly large numbers) gaming monitoring devices 120 deployed within a gaming venue or across multiple gaming venues.

[0108]

[0124] Facial features and identifier database 189 may include records of facial features and identifier details of various known patrons of the gaming venue. The records of facial features of known patrons may be in the form of embedded representations suitable for comparison with embedded representations generated by facial recognition module 127. Facial features database 189 may essentially include images or other information that provides a basis for comparison against facial features recognized by facial recognition module 127. Facial features database 189, or another database accessible to gaming monitoring server 180, may include or have access to information or identifier details of known patrons, such as names, official identification numbers, or addresses.

[0109]

[0125] Game object association module 128 includes program code that considers the output generated by each of object detection module 123, pose estimation module 125, game object value estimation module 126, and face recognition module 127 to identify gaming events on the gaming table and associate the gaming events with identities of people participating in or initiating the gaming events. Game object association module 127 may receive information about game objects identified by object detection module 123, game object values ​​estimated by game object estimation module 126, and facial feature information of people initiating or participating in the gaming events, through a combination of the output from pose estimation module 125 and the output from face recognition module 127. Various steps performed by game object association module 128 are identified in the flowchart of FIG. 6.

[0110]

[0126] The game object estimation module 126 is configured to process the image, where at least one game object is detected by the object detection module 123, and estimate a game object value associated with the detected game object. In some embodiments, the game object estimation module may include a height-based game object estimation sub-module 154 configured to estimate game object values ​​based on the heights of a large number of game objects and the color of the game objects using any of the game object value estimation processes described in PCT specification "System and method for automated table game activity recognition" (Publication No. WO 2017 / 197452), the contents of which are incorporated herein by reference. In some embodiments, the game object estimation module may include a trained edge pattern recognition neural network configured to estimate game object values ​​based on the determined edge pattern of each game object in a large number of game objects using any of the game object value estimation processes described in PCT specification "System and method for machine learning driven object detection" (Publication No. WO 2019 / 068141), the entire contents of which are incorporated herein by reference. The object detection module 123 may also be configured to determine a gaming table area in which a game object is detected. For example, a gaming table for the game of baccarat may include areas associated with the player, banker, or tie. The object detection module 123 may be configured with a model of the gaming table layout that may be superimposed on the image to determine in which gaming table area a game object is present.

[0111]

[0127] 4 shows a gaming environment 400 including cameras 110 and 112 for monitoring gaming activity on a gaming table 420. The cameras 110 and 112 are embedded in posts 410 and 412, respectively. The posts 410 and 412 may be positioned on opposite sides of a dealer's position on the table. The cameras 110 and 112 are positioned to look away from the dealer's side of the table and toward players participating in the game. The cameras 110 and 112 are positioned to have viewing angles through openings in the posts 410, 412 that allow them to capture images of the gaming activity on the gaming table 420 and the faces and postures of individuals participating in the gaming activity. In some embodiments, cameras 110 and 112 may be positioned at a height ranging from, for example, 45 cm to 65 cm, or 35 cm to 55 cm, or 55 cm to 75 cm, or 35 cm to 45 cm, or 45 cm to 55 cm, or 55 cm to 65 cm, or 65 cm to 75 cm, or 75 cm to 85 cm from gaming table 420. In some embodiments, cameras 110 and 112 may be positioned at a distance from a center point 425 of the gaming table, for example, 130 cm to 150 cm, or 120 cm to 140 cm, or 140 cm to 160 cm, or 160 cm to 180 cm. In some embodiments, three or more cameras may be used to monitor gaming activity. Additional cameras may be positioned on the dealer's side of gaming table 420 and used to capture images from unique angles that allow for coverage of a wider field of view above gaming table 420. In some embodiments, additional cameras may be used to capture redundant images of the gaming area to enable verification of image processing results from multiple angles. The additional cameras may be configured to communicate with gaming monitor computing device 120 wirelessly or via a wired connection.

[0112]

[0128] The pose estimation module 125 of some embodiments may be configured to process images of people in a gaming environment and identify one or more key points associated with specific parts of the people's bodies. In some embodiments, the pose estimation module 125 may be configured to determine distal hand perimeter key points of one or more players in the image. The distal hand perimeter key points may be pixels or sets of adjacent pixels in the image that correspond to portions of the player's hands that are farthest from the player's wrist. Players may have hands in various orientations, and one or more of the player's fingers may be flexed in the image. The pose estimation module 125 of some embodiments is trained to determine which pixel or sets of adjacent pixels in the image corresponds to the player's farthest hand perimeter that is visible in the image. Determining distal hand perimeter key points may be performed for both left and right hands.

[0113]

[0129] 5A, 5B, 5C, and 5D illustrate the determination of wrist keypoints and distal hand peripheral keypoints of a person in each image. In FIG. 5A , which shows image 510, keypoints 514 and 518 corresponding to the wrist and keypoints 512 and 516 corresponding to the distal hand peripheral are determined by pose estimation module 125. In FIG. 5B , which shows image 520, keypoints 524 and 528 corresponding to the wrist and keypoints 522 and 526 corresponding to the distal hand peripheral are determined by pose estimation module 125. In FIG. 5C , which shows image 530, keypoints 534 and 538 corresponding to the wrist and keypoints 532 and 536 corresponding to the distal hand peripheral are determined by pose estimation module 125. In FIG. 5D , which shows image 540, keypoints 544 and 548 corresponding to the wrist and keypoints 542 and 546 corresponding to the distal hand peripheral are determined by pose estimation module 125. 5A-5D, pose estimation module 125 identifies the distal perimeter of the hand regardless of the angle at which the image was taken or the degree to which the person's fingers are flexed or extended. Identifying the distal perimeter of the hand within the gaming environment provides an effective key point for assigning or associating game objects on a gaming table to the person or player responsible for placing the game objects on the gaming surface.

[0114]

[0130] 6 shows a flowchart of a process 600 performed by gaming surveillance computing device 120 according to some embodiments. At 610, gaming surveillance computing device 120 receives one or more images captured by camera 110. In embodiments having two or more cameras, computing device 120 may receive images from each camera simultaneously. The received images may be in the form of, for example, a time-stamped stream of images of the gaming environment.

[0115]

[0131] At 620, object detection module 123 processes the stream of images received at 610 to identify one or more game objects in a first image from the stream of images. The region of the first image from which one or more game objects are identified may be identified by one or more coordinates or a bounding box around the identified game object. If a game object is detected in the first image, game object estimation module 126 estimates a game object value associated with the detected game object at 630. Estimating the game object value of the game object at 630 is an optional step, and in some embodiments, game object association module 128 may be configured to associate a player with an identified game object without estimating the game object value of the identified game object.

[0116]

[0132] At 640, the pose estimation module processes the first image to determine the posture or pose of one or more players present in the first image. The determined pose may include one or more key points associated with various specific body parts of the player. The key points may include key points associated with distal hand regions, including the left hand region, the right hand region, or both. The pose estimation may also include estimating a unique skeletal model of the player. Identifying the unique skeletal model may enable association of facial regions in the first image with hand regions of the same player. Because gaming environments can be highly cluttered with multiple players crowded around a table and participating in a game at different times, estimating the player's skeletal model enables accurate association of faces in the image with hands or distal hand regions identified in the image. The skeletal model may include an approximation and partial mapping of various body parts of the player with key points and segments associated with the various body parts. For example, the skeletal model may include keypoints associated with one or more of the distal hand keypoints, wrists, elbows, shoulders, neck, nose, eyes, and ears.

[0117]

[0133] At 650, the game object association module 128 processes the posture or attitude determined at 640 and the game object detected at 620 to determine a target player associated with the game object. The target player is the player who likely placed the game object on the gaming table. The game object association module 128 estimates a distance between the game object identified at step 620 and each key point associated with the distal periphery of the hand. The estimated distance may be based on the coordinates of each key point associated with the distal periphery of the hand and the coordinates of the game object detected at 620 in the first image. Based on the calculated distance, the key point associated with the distal periphery of the hand closest to the game object may be considered to belong to the target player. A skeletal model of the target player may be determined based on the distal periphery of the hand closest to the game object.

[0118]

[0134] Based on the skeletal model of the target player identified in 650, a region of the first image corresponding to the target player's face may be determined at 655. For example, key points corresponding to one or more of the eyes, nose, mouth, or ears may be used to extrapolate and determine a bounding box associated with the target player's face. The skeletal model may enable association between the distal perimeter of the hand closest to the game object and the target player's head or facial region. In some embodiments, the person detection neural network 159 may determine a bounding box (e.g., in the form of image coordinates) around the face of one or more players in the image, and the skeletal model may enable association between the target player's face and the distal perimeter of the hand closest to the game object.

[0119]

[0135] In some embodiments, the facial region of the image corresponding to the target player may be obscured or the target player may be turned away (e.g., showing only the back of the target player's head). To address such situations, the object detection module 123 of some embodiments may be configured to track a region in a first image identified as a facial region corresponding to the target player across subsequent images in the stream of images received from the camera 110. The pose estimation module 125 may be configured to perform pose estimation across subsequent images in the stream of images to determine the target player's posture or posture. Determining the target player's posture or posture may include determining whether the tracked facial region of the target player corresponds to the target player's face or the back of his or her head. Based on the determination of the pose estimation module 125 across subsequent images in the stream of images, the game object association module 128 may extract one or more image regions corresponding to the target player's face. In some embodiments, the game object association module 128 may extract image regions corresponding to the target player's face from, for example, 2-3, or 3-5, or 5-7, or 7-9, or 9-11 images from the stream of images. Extracting additional image regions corresponding to the target player's face may provide additional information about the target player's facial features to improve the facial recognition process.

[0120]

[0136] At 660, based on the one or more image regions corresponding to the target player's face extracted at 655, a facial recognition module may process the one or more image regions to obtain a vector representation or embedded representation of the target player's face. The embedded representation of the target player's face captures information that encodes the player's unique facial features and allows comparison to a database of similarly encoded information of facial features.

[0121]

[0137] At 670, the gaming monitor computing device may transmit information regarding the determined gaming event and the target player to the gaming monitor server 180. The information regarding the determined gaming event may include the nature of the gaming event (e.g., the placement of game objects in an identified region of interest on a particular table). The gaming event information may include a unique identifier corresponding to the table, the region of the table where the gaming event occurred, and a timestamp at which the gaming event occurred. The timestamp may include date, time of day, and time zone information. The time of day may include hour and minute information. In some embodiments, the time of day may include hour, minute, and second information. The gaming event information may also include game object values ​​associated with the gaming event. The transmitted information regarding the target player may include an embedded representation of the face of the target player determined at 660.

[0122]

[0138] Process 600 may be implemented using a multithreading computing architecture to perform real-time or near-real-time gaming monitoring. For example, for each game object detected in 620, processor 122 may initialize a separate thread to perform the image processing tasks of steps 630-660, enabling parallel processing of the gaming monitoring process for each detected game object. If multiple game objects are detected in step 620 during the course of being monitored by process 600, a separate thread may be initiated to associate each game object with a target player. In some embodiments, a series of images captured by camera 110 may be processed separately by each thread to track the target player across the series of images. Tracking the target player across the series of images allows for the extraction of multiple images of the target player's face. The multiple images of the target player's face may be made available to face embedding generation neural network 156 in step 660 to obtain a robust embedding representation of the target player.

[0123]

[0139] 7 is an image 700 of a gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. In image 700, pose estimation module 125 has determined a skeletal model 715 associated with player 750. The determined skeletal model 715 includes various key points associated with particular body parts of player 750 (e.g., joints of player 750). The determined skeletal model 715 also includes segments connecting key points associated with particular body segments of player 750. Distal hand key points 712 correspond to the distal periphery of player 750's hands.

[0124]

[0140] Also identified in image 700 is a bounding box 710 determined by object detection module 123. Bounding box 710 identifies one or more game objects placed on tabletop surface 730. The image distance (in terms of image coordinates) between key points 712 and bounding box 710 is determined by game object association module 128 to associate game objects within bounding box 710 with a skeletal model 715 of player 750. In embodiments where multiple players are present, the distance between key points corresponding to the distal periphery of each player's hand and bounding box 710 may be determined by game object association module 128, and the distal hand key point 712 (in terms of image coordinates) closest to bounding box 710 may be determined to correspond to player 750 placing one or more game objects within bounding box 710 on tabletop surface 730. This distance may be calculated, for example, based on the (minimum) Cartesian distance between the pixel in image 700 corresponding to distal hand keypoint 712 and bounding box 710.

[0125]

[0141] 7 also shows a bounding box 718 identified by the person detection neural network 159 of the object detection module 123 around the face of the player 750. In some embodiments, the coordinates of the bounding box 718 may be determined by processing the image to extrapolate facial keypoints (keypoints corresponding to one or more of a person's eyes, nose, mouth, and ears) determined by the keypoint estimation neural network 152. The identification of the skeletal model 715 enables association of distal hand keypoints 712, corresponding to the distal periphery of the player's 750's hands, with the bounding box 718 around the player's face, thereby enabling association of one or more game objects within the bounding box 710 with the face of the player 750 within the bounding box 718.

[0126]

[0142] FIG. 8 is an image 800 of a gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. In image 800, pose estimation module 125 has determined a skeletal model 815 associated with player 850. Unlike player 750 of FIG. 7, player 850 of FIG. 8 is seated at a gaming table. The determined skeletal model includes various key points associated with specific body parts of player 850 (e.g., joints of player 850). The determined skeletal model also includes segments connecting key points associated with specific body segments of player 850. Key point 812 corresponds to the distal vicinity of player 850's hands.

[0127]

[0143] Also identified in image 800 is a bounding box 810 determined by object detection module 823. Bounding box 810 identifies one or more game objects placed on tabletop surface 830. The distance between distal hand keypoint 812 and bounding box 810 is determined by game object association module 128 to associate the game objects within bounding box 810 with a skeletal model 815 of player 850. In embodiments where multiple players are present, the distance between keypoints corresponding to the distal periphery of each player's hand and bounding box 810 may be determined by game object association module 128, and the distal hand keypoint 812 closest to bounding box 810 may be determined to correspond to a player placing one or more game objects at a location on tabletop surface 830 corresponding to bounding box 810.

[0128]

[0144] 8 also shows a bounding box 818 identified by the person detection neural network 159 of the object detection module 123 around the face of the player 850. In some embodiments, the bounding box 818 may be determined by extrapolating facial keypoints (keypoints corresponding to one or more of a person's eyes, nose, mouth, and ears) determined by the keypoint estimation neural network 152. The identification of the skeletal model 815 allows for the association of distal hand keypoints 812, corresponding to the periphery of the distal portion of the player's 850's hands, with the bounding box 818 around the player's face, thereby allowing for the association of one or more game objects within the bounding box 810 with the face of the player 850 within the bounding box 818.

[0129]

[0145] 9 is an image 900 of a gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. Image 900 includes a first player 910 and a second player 920. Pose estimation module 125 processes image 900 to identify a first skeletal model 918 of first player 910 and a second skeletal model 928 of second player 920. Each determined skeletal model includes a first distal hand point 914 associated with first player 910 and a second distal hand point 924 associated with second player 920. Object detection module 123 processes image 900 to determine bounding boxes 916 and 926 (in the form of image coordinates) around game objects placed on the game table.

[0130]

[0146] The game object association module 128 may determine the relationships between the game objects within bounding boxes 916 and 926 and the players 910 and 920 based on the minimum distance (in image coordinates) between a first distal point 914 associated with the first player 910 and a second distal point 924 associated with the second player 920. Based on the distances determined by the game object association module 128, the game object association module 128 associates the game objects within bounding box 916 with the player 910 and the game objects within bounding box 926 with the player 920. In some embodiments, the person detection neural network 159 of the object detection module 123 also processes the image 900 to determine the players' faces and to determine bounding boxes 912 and 922 around the faces of the players 910 and 920, respectively. In some embodiments, the bounding boxes 912 and 922 may be determined by extrapolating facial keypoints (keypoints corresponding to one or more of a person's eyes, nose, mouth, and ears) determined by the keypoint estimation neural network 152.

[0131]

[0147] FIG. 10 is an image 1000 of a gaming environment illustrating a portion of a method for gaming monitoring according to some embodiments. The image 1000 includes a first player 1010 and a second player 1020. Unlike players 910 and 920 of FIG. 9, the hands of players 1010 and 1020 in FIG. 10 are significantly closer to each other. The pose estimation module 125 processes the image 1000 to identify a first skeletal model 1018 of the first player 1010 and a second skeletal model 1028 of the second player 1020. Each determined skeletal model includes a first distal hand point 1014 associated with the first player 1010 and a second distal hand point 1024 associated with the second player 1020. The object detection module 123 processes the image 1000 to determine bounding boxes 1016 and 1026 around game objects placed on the game table.

[0132]

[0148] The game object association module 128 may determine the relationships between the game objects within bounding boxes 1016 and 1026 and the players 1010 and 1020 based on the minimum distance (in image coordinates) between a first distal hand point 1014 associated with the first player 1010 and a second distal hand point 1024 associated with the second player 1020. Based on the distances determined by the game object association module 128, the game object association module 128 associates the game objects within bounding box 1016 with the player 1010 and the game objects within bounding box 1026 with the player 1020. In some embodiments, the person detection neural network 159 of the object detection module 123 also processes the image 1000 to determine the players' faces and to determine bounding boxes 1012 and 1022 around the faces of the players 1010 and 1020, respectively. In some embodiments, the bounding boxes 1012 and 1022 may be determined by extrapolating facial keypoints (keypoints corresponding to one or more of a person's eyes, nose, mouth, and ears) determined by the keypoint estimation neural network 152.

[0133]

[0149] As shown in image 1000, even though the player is positioned with their limbs or body parts possibly closely overlapping each other, embodiments can associate the player with game objects based on the described image processing techniques.

[0134]

[0150] 11 is a block diagram of a gaming monitoring system 1100 according to some embodiments. The gaming monitoring system 1100 includes computing and network components located within or near a gaming environment 1120. The gaming monitoring system 1100 also includes computing and network components located remotely from the gaming environment 1120. The gaming environment 1120 may include a venue, such as a gaming venue or casino, where gaming monitoring may take place.

[0135]

[0151] The gaming surveillance system 1100 includes a camera system 1104 in communication with a computing device or edge gaming surveillance computing device 1106. In some embodiments, the edge gaming surveillance computing device 1106 may be a computing device configured to perform low-latency operations or low-latency image processing operations on image data received from the camera system 1104. The camera system 1104 and the edge gaming surveillance computing device 1106 may be designated for a particular area or region within the gaming environment 1120. A particular area or region of the gaming environment 1120 may include, for example, a particular table or a group of tables located in close proximity. The combination of the camera system 1104 and the edge gaming surveillance computing device 1106 may be repeated throughout various areas, regions, or regions of the gaming environment 1120.

[0136]

[0152] In some embodiments, the camera system 1104 and the edge gaming surveillance device 1106 may be implemented as a single smart camera or machine vision system that combines the capabilities of capturing images and processing image data by performing some or all of the processing operations of the edge gaming surveillance computing device 1106 within the single smart camera or machine vision system.

[0137]

[0153] In some embodiments, at least a portion of camera system 1104 and edge gaming surveillance computing device 1106 may be implemented using a smartphone that includes a camera for capturing images of the gaming environment. In some embodiments, camera system 1104 may include a depth-of-field camera, or a motion-sensing camera, or a time-of-flight camera, or a 3D camera, or a range-imaging camera that captures depth-related information associated with the line of sight of camera system 1104. Depth-of-field images captured using the 3D camera of camera system 1104 may include information regarding relative or absolute distances between camera system 1104 and various objects from the perspective of camera system 1104. In some embodiments, object detection module 123, or face recognition module 127, or pose estimation module 125, or game object value estimation module 126, or game object association module 128, or face orientation determination module 1289 may perform various image processing tasks based on a combination of visual image data and depth-of-field image data captured by camera system 1104. For example, the object detection module 123 may perform an image segmentation operation as part of the object detection process. The image segmentation process may include segmenting an image into separate segments, each corresponding to an object or background of potential interest. The image segmentation operation may include analyzing depth-of-field image data to perform at least a portion of the image segmentation operation based on depth-of-field information captured by the 3D camera of the camera system 1104. For example, a customer sitting at or standing near a gaming table may be distinguished from the background of the image based on depth information associated with portions of the image corresponding to the customer. Using the depth-of-field information, the image may be segmented to identify portions of the image that correspond to the gaming table or to customers sitting at or standing near the gaming table.

[0138]

[0154] Each edge gaming monitoring computing device 1106 is configured to communicate with an on-premises gaming monitoring server 1110 over a network 1108. The network 1108 may be a local computer communications network deployed within the gaming environment 1120. The network 1108 may include, for example, a local area network (LAN) deployed using a combination of routers and / or switches. The on-premises gaming monitoring server 1110 may be deployed within a secure portion of the gaming environment 1120 to receive and process information or data from each of the edge gaming monitoring devices 1106. The on-premises gaming monitoring server 1110 is configured to communicate with a remote gaming monitoring server 1130 over a network 1112. The network 1112 may include a wide area network, such as the Internet, over which the remote gaming monitoring server 1130 and the on-premises gaming monitoring server 1110 may communicate.

[0139]

[0155] A significant amount of image data is continuously captured by the camera systems 1104. Each camera system 1104 may generate image data from the captured images at a rate of, for example, 1-20 MB / s or even greater than 1 MB / s. A gaming environment 1120 may be particularly large, with 100 or more gaming tables 1102 or gaming territories or areas that need to be monitored. Accordingly, a significant amount of data may be generated at a rapid rate that may require efficient and fast processing to support the monitoring operations. The gaming monitoring system 1100 employs an approach that distributes the processing of the image data recorded by each camera system 1104 using a hierarchical combination of edge gaming monitoring devices 1106, on-premise gaming monitoring servers 1110, and remote gaming monitoring servers 1130. The hierarchical distribution of image processing and analysis operations enables the gaming monitoring system 1110 to handle a significant volume and velocity of image data while allowing needed gaming monitoring inferences to be generated in near real time in response to occurrences or events within the gaming environment 1120 that may require immediate action.

[0140]

[0156] The camera system 1104 may include at least one camera pointed in a direction that provides visibility of patron activity playing games on the gaming table 1102. In some embodiments, the camera system 1104 may include two or more cameras to provide sufficient coverage at various angles around the gaming table 1102 to adequately capture images of patron gaming activity and their faces. In some embodiments, the camera system 1104 may include at least one panoramic camera configured to capture images at a capture angle of 180 degrees or greater. Some embodiments may incorporate, for example, a Mobotix™ S16 DualFlex camera or a Jabra™ PanaCast camera. Each camera in the camera system 110 may continuously capture images at a resolution of, for example, 6144 x 2048 pixels, 3840 x 2160 pixels, 2592 x 1944 pixels, 2048 x 1536 pixels, 1920 x 1080 pixels, or 1280 x 960 pixels. These listed example image resolutions are non-limiting, and thus other imaging resolutions may be incorporated by embodiments to perform gaming activity monitoring.

[0141]

[0157] The on-premises gaming surveillance server 1110 may include a secure data storage module or component 1291 provided as part of the memory 1284. The secure data storage module 1291 may be configured to store data received from the edge computing device 1106 or data generated as part of image processing operations performed by the gaming surveillance server 1110. Because the data received from the edge computing device 1106 or data generated as part of image processing operations performed by the gaming surveillance server 1110 may pertain to sensitive personal information, the data stored in the secure data storage module 1291 may be encrypted to protect the data from information security breaches. The physical location of the on-premises gaming surveillance server 1110 may also be physically secured, such as by placement in a locked environment accessible only to authorized personnel. In some embodiments, the remote gaming surveillance server 1130 may also include an equally secure data storage module for storing data or information received from the on-premises gaming surveillance server 1110. In some embodiments, the edge computing device 1106 may be configured to automatically delete image data received from the camera system 1104 after the image data has been processed.

[0142]

[0158] Figure 12 is a block diagram of system components 1200 of a subset of the gaming surveillance system 1100 of Figure 11. Figure 12 shows components of the on-premise gaming surveillance server 1110 and the edge gaming surveillance computing device 1106 in more detail. The edge gaming surveillance computing device 1106 includes at least one processor 1222 and a network interface 1229 in communication with memory 1214. The memory 1214 includes program code for implementing the object detection module 123 and the event detection module 1223. The object detection module 123 was described with reference to the gaming surveillance computing device 120 of Figure 1. The program code for the object detection module 123 enables the edge gaming surveillance computing device 1106 to process images captured by the camera system 1104 and perform object detection operations on the captured images. As described with reference to gaming surveillance computing device 120 of FIG. 1 , object detection operations performed on captured images by object detection module 123 may include detecting game objects by game object detection neural network 151. Game objects may include chips, playing cards, or cash. Game objects may be detected by object detection module 123 when they are partially or fully within the field of view of camera system 1104. Game objects may be within the field of view of camera system 1104 when placed on gaming table 1102. Person detection neural network 159 may similarly detect image regions corresponding to people or portions thereof in images captured by camera system 1104.

[0143]

[0159] The object detection module 123 may continuously process images captured by the camera system 1104 to generate a stream of object detection data including data packets with information or data about detected objects. Each data packet in the stream of object detection data may correspond to object data in one or more image frames captured by the camera system 1104 and may have a timestamp including date and time information related to the date and time the image frame was captured by the camera system 1104. The time information may include time information in 24-hour HH:MM:SS format. Each data packet in the stream of object detection data may also include a list or set of object information related to objects detected in one or more image frames. The object information may include a class identifier or label that identifies the type of object detected. The class or label may refer, for example, to a person or part of a person (e.g., a person's face or hand), or a game object or cash. The object information may also include one or more attributes associated with the detected object (e.g., coordinates in the captured image associated with the detected object or coordinates defining a bounding box around the image region corresponding to the detected object, etc.). The one or more attributes may also include a unique identifier associated with the detected object to uniquely identify each detected object. The one or more attributes may also include an area or region identifier of the gaming table 1102 in which the detected game object was located at the time it was detected by the object detection module 123.

[0144]

[0160] The event detection module 1223 may include program code that defines logical or mathematical operations for processing object detection data generated by the object detection module 123 to identify the occurrence of an event. Based on the identified event, the event detection module 1223 may transmit data related to the identified event to the on-premise gaming surveillance server 1110 for further analysis. The transmitted data related to the identified event may include various attributes or labels associated with the detected object. The transmitted data related to the identified event may also include image data corresponding to some or all of an image captured by the camera system 1104 related to an object detected by the object detection module 123 that is related to the identified event.

[0145]

[0161] The program code of the event detection module 1223 may define multiple event triggers or triggers. Each trigger may include a set of conditions or parameters that can be evaluated to determine the occurrence or non-occurrence of an event based on data received by the event detection module 1223. The event detection module 1223 may identify an event based on object detection data and various attributes and information included in the object detection data. The detection of a new game object (not previously seen in images captured by the camera system 1104) in an image captured by the camera system 1104 is an example of an event. The condition of the detection of a new game object (not previously seen in images captured by the camera system 1104) in an image captured by the camera system 1104 may be defined, for example, as a trigger in the event detection module 1223 for the detection of an event or gaming event. The detection of a new game object may correspond, for example, to a real-world event of a customer placing a game object on the gaming table 1102 during the course of game play. The detection of a new person or a new face (not previously seen in images captured by camera system 1104) in an image captured by camera system 1104 is another example of an event. The condition of detection of a new person or face (not previously seen in images captured by camera system 1104) in an image captured by camera system 1104 may be defined as another trigger in event detection module 1223, for example, for detection of an event or gaming event.

[0146]

[0162] In some embodiments, the edge gaming monitoring computing device 1106 may be implemented using a low-computing-power computing device, such as an NVIDIA Jetson Xavier NX or NVIDIA Jetson Nano-based system-on-module. The edge gaming monitoring computing device 1106 may be implemented using a less expensive computing device, for example, that consumes less power, generates less heat, and has lower computing power requirements. The use of a less expensive computing device as the edge gaming monitoring computing device 1106 allows the gaming monitoring system 1100 to be inexpensively scaled, for example, by deploying hundreds or thousands of edge gaming monitoring computing devices 1106 within the gaming environment 1120.

[0147]

[0163] The on-premises gaming surveillance server 1110 includes at least one processor 1282 and memory 1284 in communication with a network interface 1283. The memory 1284 includes several modules or components described with reference to the gaming surveillance device 120 of FIG. 1 , including face recognition module 127, pose estimation module 125, game object value estimation module 126, game object association module 128, and facial feature identification database 189. Unlike the gaming surveillance computing device 120 of FIG. 1 , which receives data directly from the camera 110, the on-premises gaming surveillance server 1110 receives processed data from the edge gaming surveillance computing device 1106. Various code modules of the on-premises gaming surveillance server 1110 perform face recognition, pose estimation, game object value estimation, and game object association operations on the processed data received from the edge gaming surveillance computing device 1106.

[0148]

[0164] The on-premise gaming surveillance server 1110 may include a face orientation determination module 1289. The face orientation determination module 1289 includes program code for processing image data related to an image of a face captured by the camera system 1104 and detected by the object detection module 123 of the edge gaming surveillance computing device 1106. The face orientation determination module 1289 determines face orientation information in the image data corresponding to the face.

[0149]

[0165] As a customer participates in game play within the gaming environment 1120, the camera system 1104 may capture images of the customer's face from various angles due to natural movements of the customer's face. The object detection module 123 may detect faces in images captured by the camera system 1104 and may transmit image data of areas of the captured images corresponding to the customer's face to the on-premises gaming surveillance server 1110.

[0150]

[0166] Some captured images may include a more frontal, forward, or straight-on snapshot of the customer's face captured as the customer is looking directly in the direction of one or more cameras of camera system 1104. A more frontal, forward, or straight-on snapshot of the customer's face may be, for example, an image that captures at least the customer's eye region. Alternatively, a more frontal, forward, or straight-on snapshot of the customer's face may be an image that captures a larger surface area of ​​the customer's face and is therefore more information-rich for facial recognition purposes.

[0151]

[0167] Some captured images may include a more lateral snapshot of a customer captured when the customer was looking less directly toward one or more cameras of camera system 1104. For example, a more lateral snapshot of a customer may be an image that captures only one eye region of the customer. Alternatively, a lateral snapshot of a customer may be an image that captures a smaller surface area of ​​the customer's face, which may be less informative and therefore less suitable for facial recognition purposes.

[0152]

[0168] A more frontal image of the customer's face may be effective and efficient for performing facial recognition operations by the facial recognition module 127. The program code of the face orientation determination module 1289 processes image data corresponding to the face to identify landmarks in the image that correspond to particular points in the face. The particular points in the face may include, for example, points corresponding to various parts of the eyes, mouth, nose, eyebrows, and chin.

[0153]

[0169] The face orientation determination module 1289 may also determine 3D coordinates corresponding to each identified landmark. Based on the 3D coordinates of each identified landmark, the face orientation may be determined. The determined orientation may be represented in the form of a rotation matrix, or using Euler angles, or using a quaternion (a scalar component and a three-dimensional vector component), or using a multi-dimensional vector representation that encodes the detected face orientation.

[0154]

[0170] Using a representation of the detected facial orientation, a more frontal facial image of the customer may be selected from multiple images of the customer's face. Selecting the more frontal facial image improves the accuracy of the facial recognition module 127 and the efficiency of the facial recognition process. The more frontal facial image contains more significant data about the customer's unique facial features and therefore provides a more effective starting point for more precise and efficient facial recognition operations. Discarding less frontal images of the customer's face also allows the on-premises gaming surveillance server to avoid performing computationally expensive facial recognition operations on images with less information about facial features.

[0155]

[0171] In some embodiments, the face orientation determination module 1289 may include a deep neural network trained to process facial images and identify facial landmarks and the 3D coordinates associated with each facial landmark. For example, the deep neural network may be trained using the 300W-LP dataset or the 300-VW dataset. The deep neural network of the face orientation determination module 1289 may include a face alignment network (FAN). The FAN may include one or more stacked hourglass networks described in the paper "Stacked Hourglass Networks for Human Pose Estimation" by Newell et al., published by the European Conference on Computer Vision in 2016. The FAN may include one or more bottleneck networks or layers described in the paper "Binarized Convolutional Landmark Localizers for Human Pose Estimation and Face Alignment with Limited Resources" by Bulat et al., published by the International Conference on Computer Vision in 2017. In some embodiments, face orientation determination may be performed using techniques described in the paper "How far are we from solving the 2D&3D Face Alignment problem?" by Bulat et al., published by the International Conference on Computer Vision in 2017.

[0156]

[0172] In some embodiments, the on-premise gaming surveillance server 1110 may include a facial feature and identifier database 1287. Similar to the facial feature and identifier database 189 described with reference to FIG. 1 , the facial feature and identifier database 1287 may include records of facial features and identifying details of various known patrons of the gaming venue. The records of facial features of known patrons may be in the form of embedded representations suitable for comparison with embedded representations generated by the facial recognition module 127. The facial feature database 1287 may essentially include images or other information that provides a basis for comparison against facial features recognized by the facial recognition module 127. The facial feature database 1287 or another database accessible to the on-premise gaming surveillance server 1110 may include or have access to information or identifier details of known patrons, such as names, official identification numbers, or addresses. In some embodiments, the remote gaming surveillance server 1130 may include the facial feature database 1287 and may perform comparisons based on facial feature data transmitted by the on-premise gaming surveillance server 1110.

[0157]

[0173] 12 illustrates an example allocation or distribution of various processing modules and / or computing resources between the edge computing device 1106 and the on-premise server 1110. The allocation or distribution of modules between the edge computing device 1106 and the on-premise server 1110 may vary in other embodiments depending on various constraints or objectives of the gaming monitoring system 1100. For example, in some embodiments, the various processing modules (123, 1223, 125, 125, 127, 1291, 128, 189, 1289) illustrated in FIG. 12 may be deployed across the edge computing device 1106, the on-premise gaming monitoring server 1110, and the remote gaming monitoring server 1130. Various constraints or objectives governing the allocation or deployment of the various processing modules may include: constraints related to handling the amount of data generated by the camera system 1104, constraints related to memory buffer capacity across the various computing components of the gaming surveillance system 1100, constraints related to the latency requirements of the monitoring operations of the gaming surveillance system, constraints related to the physical and data security of the various computing devices of the gaming surveillance system 1100, and constraints related to the data link capacity between the various components of the gaming surveillance system 1100. Any computing devices of the gaming surveillance system 1100 remote from the on-premise server 1110 and the remote gaming surveillance server 1130 and the edge computing device 1106 may be collectively referred to as upstream computing devices.

[0158]

[0174] In some embodiments, some of the image processing operations performed by the edge computing device 1106 may be performed by the on-premises gaming surveillance server 1110. In some embodiments, some of the image processing operations performed by the on-premises gaming surveillance server 1110 may be performed by the edge computing device 1106. In some embodiments, some of the image processing operations performed by the on-premises gaming surveillance server 1110 may be performed by the remote gaming surveillance server 1130.

[0159]

[0175] 13 shows a flowchart of a method 1300 performed by the edge gaming monitoring device 1106 according to some embodiments. At 1310, the edge gaming monitoring device 1106 receives images or image data captured by the camera system 1104. The images or image data may be received at a rate of, for example, 2-10 frames / second, 10-20 frames / second, 10-30 frames / second, 10-60 frames / second, or 10-120 frames / second. The frame rate of the camera system 1104 may be configured to balance the computational time required to process each frame with the need to capture image frames in close temporal proximity in order to capture images of all or most events occurring in the gaming environment.

[0160]

[0176] At 1320, the object detection module 123 of the edge gaming monitoring device 1106 processes the image received at 1310 to perform object detection. As described with reference to the object detection module 123, object detection may include detecting one or more game objects or people within the image received at 1310. The process of object detection may include identifying image segments within the image, each image segment corresponding to an identified object. Multiple image segments may be identified within each image, and the multiple image segments may be superimposed on one another. In some embodiments, object detection may be limited to detecting game objects such as chips.

[0161]

[0177] In optional step 1330, the object detected in 1320 is tracked across multiple image frames received from camera system 1104. Tracking may occur, for example, across 2 to 10 image frames received from camera system 1104. In embodiments in which object detection in step 1320 is limited to detecting game objects, an optional step of tracking the detected game object may be performed. The game object may be initially held in a player's hand before being placed on the gaming table 1102. While being held in the player's hand, the game object may be detected in step 1320. Performing the optional step of tracking the detected object across multiple image frames provides greater certainty that the detected object is being used as part of the course of game play on the gaming table 1102. In situations in which a game object is initially detected (e.g., presented in a player's hand) but may not be tracked across multiple image frames (e.g., if the player withdraws the game object from view), event detection module 1223 may determine in step 1340 that an event associated with the initially detected game object did not occur.

[0162]

[0178] At 1340, the event detection module 1223 processes the data generated regarding the object detection in step 1320 and the object tracked in step 1330 to determine whether an event, such as a gaming event, has occurred. An event may be, for example, a player placing a game object on the gaming table 1102. The occurrence of an event may be determined based on successfully tracking the game object detected in 1320, for example, across multiple image frames. Based on the determined event, the event detection module 1223 may prepare an event data packet or event data for further analysis by the on-premise gaming surveillance server 1110.

[0163]

[0179] The event data packet may include image regions or image segments corresponding to the objects detected in step 1320 (including image regions or segments corresponding to game objects and image regions corresponding to people detected in the images captured by the camera system 1104). The event data packet may also include any specific attributes related to the objects detected by the object detection module 123. The specific attributes may include a class or label identifier associated with the detected object that identifies the category of the detected object. The specific attributes related to the game object may include a label or identifier associated with an area of ​​the gaming table 1102 where the specific game object may be detected. For example, with respect to a gaming table 1102 for the game of baccarat, the gaming table 1102 may have, for example, ten areas, each associated with one or more potential players. Each area may have three sub-areas: a banker area, a player area, and a tie area. One of the attributes related to the detected game object may include an area identifier and a sub-area identifier that indicate in which portion of the gaming table 1102 the game object was detected. In some embodiments, the event data may include coordinates defining a rectangle relating to a detected game object in the image received at 1310 or a portion of the image received at 1310. In some embodiments, the event data may include a timestamp associated with the image received at 1310 that includes the object detected at 1320. The timestamp may include date and time information, and the timestamp data may indicate the time of occurrence of the associated event.

[0164]

[0180] At 1350, the event data packet prepared at step 1340 is transmitted to the on-premise gaming monitoring server for further analysis. In some embodiments, the event data packet prepared at step 1340 may alternatively or additionally be transmitted to the remote gaming monitoring server 1130 for further analysis.

[0165]

[0181] The various steps of method 1300 may be performed in a multi-threaded computing environment to respond in parallel to various events occurring on gaming table 1102. For example, if multiple game objects are detected at 1320, a separate processing thread of steps 1330, 1340, and 1350 may be started for each game object detected.

[0166]

[0182] In some embodiments, some or all of steps 1320, 1330, and 1340 may be performed by edge computing device 1106 operating in coordination with on-premises gaming surveillance server 1110. The processing workload of steps 1320, 1330, and 1340 may be distributed across edge computing device 1106 and on-premises gaming surveillance server 1110 to meet the latency and scalability requirements of gaming surveillance system 1100.

[0167]

[0183] 14 shows a flowchart of a method 1400 performed by the on-premises gaming monitoring device 1110 according to some embodiments. In some embodiments, the method 1400 may be performed by a remote gaming monitoring server 1130 that includes the various software and hardware components described with reference to the on-premises gaming monitoring device 1110.

[0168]

[0184] At 1410, the premises gaming monitoring device 1110 receives the event data packet prepared at step 1340 of Figure 13. At optional step 1420, the game object value may be determined by the game object value estimation module 126. Step 1420 may incorporate various image processing operations described with reference to step 630 of Figure 6.

[0169]

[0185] In step 1430, image segments corresponding to people or players in the event data packet are analyzed by pose estimation module 125 to estimate posture or posture information associated with each player. Estimating posture or posture information may include semantic segmentation of the image segment to identify specific body parts of the players in the image segment. The identified specific body parts may include, for example, one or more of the head, left hand, right hand, and torso. In some embodiments, the semantic segmentation may be limited to, for example, identifying only the head, left hand, and right hand. Step 1430 may also include defining a bounding box around each identified body part based on the results of the semantic segmentation. In some embodiments, semantic segmentation to identify distinct body parts of the player may be performed using Google's Semantic Image Segmentation with DeepLab in TensorFlow implementation as described in the paper entitled "Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation" by Chen et al., published by the European Conference on Computer Vision 2018, which is incorporated herein by reference in its entirety.

[0170]

[0186] The pose estimation of step 1430 may also include determining U,V coordinates (illustrated at 330) of each identified body part in the image segment relative to the 3D surface model described with reference to FIG. 3, using the pose estimation or analysis framework described with reference to pose estimation module 125 of FIG. 1.

[0171]

[0187] After determining the bounding box around each player's hand at 1430, the hand image segment that is closest to the game object is identified at 1440. In some embodiments, the distance between the hand and the game object may be evaluated using the formula described with reference to FIG.

[0172]

[0188] At 1450, based on the identified hand image segment closest to the game object, an image region of the target player's face associated with the hand image segment closest to the game object may be determined. The image segment corresponding to the target player's face may be used to identify the target player by comparing the records in the facial features with an identification database 189.

[0173]

[0189] In some embodiments, the event data packet 1410 may include a series of additional image frames or additional image frame segments captured immediately before or after the occurrence of the event is determined in step 1340. The series of additional images may allow verification of the determination made by the on-premise gaming surveillance server 1110. In some embodiments, the target player may momentarily look away from the camera system 1104. If the image region determined in 1450 pertains to an image where the target player is looking away or where a sufficient range of the target player's face is not captured, the series of additional image frames may be analyzed via steps 1460-1470 to obtain a better, more informative, or more distinctive image segment corresponding to the target player's face. The more informative or more distinctive image segment corresponding to the target player's face may include the target player's binocular region providing an image, thereby showing a larger portion of the player's entire face.

[0174]

[0190] At 1460, the series of additional images may be analyzed by face orientation determination module 1289 to determine orientation information of the target player's face within each frame or additional image frame segment of the series of additional image frames and within the target player's facial image region determined at 1450. At 1470, based on the orientation information determined at 1460, the most frontal, or most distinctive, or most informative image segment corresponding to the target player's face is identified from among the series of additional image frames or additional image frame segments and target player's facial image region determined at 1450.

[0175]

[0191] At 1480, similar to step 660 described with reference to FIG. 6, the facial recognition module 127 processes the facial image segments of the target player identified at 1450 or 1470 to obtain a vector representation or embedded representation of the target player's face. The embedded representation of the target player's face captures information encoding the player's unique facial features and allows comparison to a database of similarly encoded information for facial features. In some embodiments, the on-premise gaming surveillance computing device 1110 may also determine the target player's identity by comparing the embedded representation of the target player's face determined at 1480 with various records in the facial features and identifier database 1287. The target player's identity may include, for example, information regarding the player's name, address, date of birth, a membership identifier assigned by the gaming venue, or any other identifier or information for uniquely identifying the player.

[0176]

[0192] At 1490, similar to step 670 described with reference to FIG. 6, the on-premise gaming monitor computing device 1110 may transmit information regarding the determined gaming event and the target player to the remote gaming monitor server 180. The information regarding the determined gaming event may include the nature of the gaming event (e.g., the placement of game objects in an identified area of ​​interest on a particular table). The gaming event information may include a unique identifier corresponding to the table, the area of ​​the table where the gaming event occurred, and a timestamp at which the gaming event occurred. The timestamp may include date, time of day, and time zone information. The time of day may include hour and minute information. In some embodiments, the time of day may include hour, minute, and second information. The gaming event information may also include game object values ​​associated with the gaming event. The transmitted information regarding the target player may include an embedded representation of the face of the target player determined at 1480.

[0177]

[0193] 14 may be performed by the on-premise gaming surveillance server 1110 operating in coordination with the remote gaming surveillance server 1130. The processing workload of the various steps of the method 1400 may be distributed across the on-premise gaming surveillance server 1110 and the remote gaming surveillance server 1130 to meet the latency and scalability requirements of the gaming surveillance system 1100.

[0178]

[0194] 15 shows an image frame 1500 illustrating some results of object detection and pose estimation operations performed by the edge gaming monitoring device 1106 and the on-premise gaming monitoring server 1110, or alternatively, the gaming monitoring computing device 120. A face bounding box 1502 bounds the face region of a target player 1501. A left hand bounding box 1504 bounds the left hand region of the target player. A game object bounding box 1506 bounds the game object. The left hand bounding box 1504 is closest to the game object bounding box 1506, and therefore, the hand region of the target player 1501 can be associated with the game object within the game object bounding box 1506. Due to the association of the player's 1501 hand region with the game object within the game object sticking box 1506, the player's 1501 face region within bounding box 1052 can be associated with the game object within the game object sticking box 1506.

[0179]

[0195] 16 shows a schematic diagram 1600 of an example of determining the distance between two bounding boxes 1602 and 1604. Bounding box 1602 may have a center point 1601 with coordinates (x1, y1). Bounding box 1602 may have a length of l1 and a width of w1. Bounding box 1604 may have a center point 1603 with coordinates (x2, y2). Bounding box 1604 may have a length of l2 and a width of w2. The length of segment 1706 may be determined using the following formula: max(|x1-x2|-(l1+l2) / 2,|y1-y2|-(w1+w2) / 2)

[0180]

[0196] The above formula may be used in step 1440 of Figure 14 or step 650 of Figure 6 to determine the distance between a bounding box around a player's hand and a game object. Bounding box 1602 may correspond to a bounding box around a player's hand (e.g., bounding box 1504 of Figure 15). Bounding box 1605 may correspond to a bounding box around a game object (e.g., bounding box 1506 of Figure 15).

[0181]

[0197] 17 illustrates an exemplary computer system 1700 according to some embodiments. In particular embodiments, one or more computer systems 1700 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 1700 provide functionality described or illustrated herein. In particular embodiments, software executing on one or more computer systems 1700 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 1700. As used herein, references to a computer system may encompass computing devices, and vice versa, where appropriate. Additionally, references to a computer system may encompass one or more computer systems, where appropriate. Computing device 120, gaming monitoring server 180, edge computing device 1106, on-premise gaming monitoring server 1110, and remote gaming monitoring server 1130 are examples of computer system 1700.

[0182]

[0198] The present disclosure contemplates any suitable number of computer systems 1700. By way of example and not limitation, computer system 1700 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a dedicated computing device, a desktop computer system, a laptop or notebook computer system, a mobile phone, a server, a tablet computer system, or a combination of two or more thereof. Where appropriate, computer system 1700 may include one or more computer systems 1700, may be single or distributed, may span multiple locations, may span multiple machines, may span multiple data centers, or may reside partially or entirely within a computing cloud that may include one or more cloud computing components within one or more networks. Where appropriate, one or more computer systems 1700 may perform one or more steps of one or more methods described or illustrated herein with little or no space or time limitations. By way of example, and not limitation, one or more computer systems 1700 may perform one or more steps of one or more methods described or illustrated herein in real time or batch mode. One or more computer systems 1700 may perform one or more steps of one or more methods described or illustrated herein at different times or at different locations, where appropriate.

[0183]

[0199] In particular embodiments, computer system 1700 includes at least one processor 1702 , memory 1704 , storage 1706 , an input / output (I / O) interface 1708 , a communication interface 1710 , and a bus 1712 .

[0184]

[0200] In particular embodiments, processor 1702 includes hardware for executing instructions (such as those making up a computer program). By way of example and not limitation, to execute instructions, processor 1702 may retrieve (or fetch) instructions from an internal register, an internal cache, memory 1704, or storage 1706, decode them, execute them, and then write one or more results to an internal register, an internal cache, memory 1704, or storage 1706. In particular embodiments, processor 1702 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 1702 including any suitable number of any suitable internal caches, where appropriate. By way of example and not limitation, processor 1702 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 1704 or storage 1706, and the instruction cache may speed retrieval of the instructions by processor 1702. Data in a data cache may be a copy of data in memory 1704 or storage 1706 for instructions executing on processor 1702 to access by subsequent instructions executing on processor 1702 or to act on results of previous instructions executed by processor 1702 or other suitable data for writing to memory 1704 or storage 1706. The data cache may expedite read or write operations by processor 1702. The TLB may expedite virtual address translation for processor 1702. In particular embodiments, processor 1702 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 1702 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 1702 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include one or more processors 1702. While this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0185]

[0201] In particular embodiments, memory 1704 includes main memory for storing instructions for processor 1702 to execute or data on which processor 1702 operates. By way of example and not limitation, computer system 1700 may load instructions from storage 1706 or another source (such as, for example, another computer system 1700) into memory 1704. Processor 1702 may then load the instructions from memory 1704 into an internal register or cache. To execute instructions, processor 1702 may retrieve the instructions from the internal register or cache and decode them. During or after execution of instructions, processor 1702 may write one or more results (which may be intermediate or final results) to an internal register or cache. Processor 1702 may then write one or more of the results to memory 1704. In particular embodiments, processor 1702 executes only instructions in one or more internal registers or internal caches or in memory 1704 (as opposed to storage 1706 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 1704 (as opposed to storage 1706 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 1702 to memory 1704. Bus 1712 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 1702 and memory 1704 to facilitate accesses to memory 1704 requested by processor 1702. In particular embodiments, memory 1704 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. This RAM may be dynamic RAM (DRAM) or static RAM (SRAM), where appropriate. Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 1704 may include one or more memories 1704, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0186]

[0202] In particular embodiments, storage 1706 includes mass storage for data or instructions. By way of example and not limitation, storage 1706 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, a universal serial bus (USB) drive, or a combination of two or more thereof. Storage 1706 may include removable or non-removable (i.e., fixed) media, where appropriate. Storage 1706 may be internal or external to computer system 1700, where appropriate. In particular embodiments, storage 1706 is non-volatile solid-state memory. In particular embodiments, storage 1706 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically erasable re-writeable ROM (EAROM), or flash memory, or a combination of two or more thereof. The present disclosure contemplates mass storage 1706 taking any suitable physical form. Storage 1706 may include, where appropriate, one or more storage control units that facilitate communication between processor 1702 and storage 1706. Where appropriate, storage 1706 may include one or more storages 1706. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.

[0187]

[0203] In particular embodiments, I / O interface 1708 includes hardware, software, or both that provide one or more interfaces for communication between computer system 1700 and one or more I / O devices. Computer system 1700 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system 1700. By way of example and not limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more thereof. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 1708. Where appropriate, I / O interface 1708 may include one or more device or software drivers that enable processor 1702 to drive one or more of these I / O devices. I / O interface 1708 may include, where appropriate, one or more I / O interfaces 1708. Although this disclosure describes and illustrates particular I / O interfaces, this disclosure contemplates any suitable I / O interface.

[0188]

[0204] In particular embodiments, communication interface 1710 includes hardware, software, or both that provide one or more interfaces for communications (e.g., packet-based communications, etc.) between computer system 1700 and one or more other computer systems 1700 or one or more networks. By way of example and not limitation, communication interface 1710 may include a network interface controller (NIC) or network adapter for communicating with a wireless adapter for communicating with a wireless network, such as Wi-Fi or a cellular network. This disclosure contemplates any suitable network and any suitable communication interface 1710. By way of example and not limitation, computer system 1700 may communicate with an ad-hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 1700 may be in communication with a wireless cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network or a 3G, 4G, or 5G cellular network, etc.) or other suitable wireless network or a combination of two or more thereof. Computer system 1700 may include any suitable communication interface 1710 for any of these networks, where appropriate. Communication interface 1710 may include one or more communication interfaces 1710, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0189]

[0205] In particular embodiments, bus 1712 includes hardware, software, or both that couple components of computer system 1700 together. By way of example, and not limitation, bus 1712 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus, or a combination of two or more thereof. Bus 1712 may include one or more buses 1712, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0190]

[0206] As used herein, a computer-readable non-transitory storage medium or media may, where appropriate, include one or more semiconductor-based or other integrated circuits (ICs) (such as field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives (FDDs), solid-state drives (SSDs), RAM drives, or any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more thereof. Computer-readable non-transitory storage media may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.

[0191]

[0207] The scope of the present disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that would be understood by a person skilled in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although the present disclosure describes and illustrates each embodiment herein as including particular components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or permutation of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that would be understood by a person skilled in the art. Furthermore, references in the appended claims to devices or systems, or parts of devices or systems, that are adapted, arranged, capable, configured, enabled, operative, or effective to perform a particular function encompass that device, system, or part, so long as the device, system, or part is so adapted, arranged, capable, configured, enabled, operative, or effective, regardless of whether the particular function is activated, turned on, or unlocked. Additionally, although this disclosure describes and illustrates particular embodiments as providing certain advantages, the particular embodiments may provide none, some, or all of these advantages.

[0192]

[0208] It will be understood by those skilled in the art that several variations and modifications may be made to the described embodiments without departing from the broad general scope of the present disclosure, and the described embodiments are therefore to be considered in all respects as illustrative.

Claims

1. 1. A system for monitoring gaming activity within a gaming area including a gaming table, comprising: at least one camera configured to capture images of the gaming area; an edge gaming surveillance computing device provided proximate to the gaming table, the edge gaming surveillance computing device including: a memory; and at least one processor having access to the memory, the at least one processor configured to communicate with the at least one camera. Including, The memory may include: determining a presence of a first game object on the gaming table in a first image from a series of images of the gaming area captured by the at least one camera; tracking the first game object in a plurality of images in the sequence of images in response to determining the presence of the first game object in the first image; determining object detection event data including image data extracted from a plurality of images and metadata corresponding to the first game object; and transmitting the object detection event data to a gaming monitoring server. storing instructions executable by said at least one processor to configure said system to:

2. The system of claim 1 , wherein the object detection event data is determined in response to tracking of the first game object in at least two or more of the plurality of images in the sequence of images.

3. The system of claim 1 or 2, wherein the at least one processor is further configured to determine the presence of one or more players within the sequence of images of the gaming area.

4. The system of claim 3 , wherein the image data extracted from the plurality of images includes image data corresponding to the one or more players and image data corresponding to the first game object.

5. 1. A system for monitoring gaming activity within a gaming area, comprising: A gaming monitoring server including at least one processor configured to communicate with a memory. Including, The memory may include: receiving object detection event data from the table gaming monitor computing device, the object detection event data including image data corresponding to a plurality of images and metadata corresponding to a first game object; processing the image data corresponding to the plurality of images to estimate a pose of one or more players; determining a first target player among the one or more players associated with the first game object based on the estimated pose; storing instructions executable by said at least one processor to configure said system to:

6. 6. The system of claim 5, wherein the instructions are executable to configure the at least one processor to identify, within the plurality of images, a plurality of facial regions of the first target player associated with the first game object.

7. 7. The system of claim 6, wherein the instructions are executable to configure the at least one processor to process the plurality of facial regions of the first target player to determine facial orientation information of the first target player's face in each of the plurality of facial regions.

8. 8. The system of claim 7, wherein the instructions are executable to configure the at least one processor to process the facial orientation information of the first target player's face in each of the plurality of facial regions to determine a front-most facial region corresponding to the target player.

9. The system of claim 8 , wherein the front-most facial region corresponding to the target player relates to a facial region that is the most information-rich facial region for facial recognition operations.

10. 1. A method of monitoring gaming activity within a gaming area including a gaming table, comprising: providing at least one camera configured to capture images of the gaming area, at least one processor configured to communicate with the at least one camera, and a memory storing instructions executable by the at least one processor; determining a presence of a first game object on the gaming table in a first image from a series of images of the gaming area captured by the at least one camera; tracking the first game object in a plurality of images in the sequence of images in response to determining the presence of the first game object in the first image; determining object detection event data including image data extracted from a plurality of images and metadata corresponding to the first game object; and transmitting the object detection event data to a gaming monitoring server. A method comprising:

11. The method of claim 10 , wherein the object detection event data is determined in response to tracking of the first game object in at least two or more of the plurality of images in the sequence of images.

12. The method of claim 10 or 11, further comprising determining the presence of one or more players within the sequence of images of the gaming area.

13. The method of claim 12 , wherein the image data extracted from the plurality of images includes image data corresponding to the one or more players and image data corresponding to the first game object.

14. 1. A method of monitoring gaming activity within a gaming area, comprising: receiving, at the gaming surveillance server, object detection event data from the gaming surveillance computing device, the object detection event data including image data corresponding to a plurality of images and metadata corresponding to a first game object; processing, by the gaming surveillance server, the image data to estimate a pose of one or more players within the plurality of images; determining, by the gaming monitoring server, a first target player among the one or more players associated with the first game object based on the estimated pose; A method comprising:

15. The method of claim 14 , further comprising identifying, within the plurality of images, a plurality of facial regions of the first target player associated with the first game object.

16. 16. The method of claim 15, further comprising processing the plurality of facial regions of the first target player to determine facial orientation information of the first target player's face in each of the plurality of facial regions.

17. 17. The method of claim 16, further comprising processing the facial orientation information of the first target player's face in each of the plurality of facial regions to determine a front-most facial region corresponding to the target player.

18. The method of claim 17 , wherein the front-most face region corresponding to the target player relates to a face region that is the most information-rich face region for face recognition operations.

19. The method of claim 10 , wherein the gaming surveillance server comprises a gaming surveillance server located on a gaming premises or a gaming surveillance server located remotely from the gaming premises.

20. The system of claim 1 , 2 or 5 , wherein the gaming surveillance server comprises a gaming surveillance server located on a gaming premises or a gaming surveillance server located remotely from the gaming premises.

21. 6. The system of claim 1, 2, or 5, wherein the gaming monitoring server includes a secure data storage component for the object detection event data and the determined target player information.

22. The system of claim 1 or 2, wherein the at least one camera and the edge gaming surveillance computing device are part of a smartphone.

23. the captured images include depth-of-field images; and The system of any one of claims 1 to 4, wherein determining the presence of a first game object on the gaming table is based on the depth of field image.