Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

228 results about "Landmark" patented technology

A landmark is a recognizable natural or artificial feature used for navigation, a feature that stands out from its near environment and is often visible from long distances. In modern use, the term can also be applied to smaller structures or features, that have become local or national symbols.

Geofenced ai landmark information system

A method is provided for delivering location-based information using artificial intelligence. The method includes automatically inferring a location of a user system at a geographic landmark based on location data, and triggering an AI assistant in response to the inferred location. The AI assistant generates information about the geographic landmark, and a graphical indication of the AI-generated information is displayed proximate to a graphical representation of the user on a map interface. Upon user selection of the graphical indication, a chat conversation with the AI assistant is initiated, the conversation including the AI-generated information about the geographic landmark
Owner:SNAP INC

Outdoor unmanned aerial vehicle visual language navigation method based on multi-modal memory map

The invention discloses an outdoor unmanned aerial vehicle visual language navigation method based on a multi-modal memory map, and aims to solve the problems of memory discretization, insufficient space connectivity and overhigh calculation load caused by dependence on global mapping in an unstructured environment in the prior art. The method comprises the following steps: extracting semantic features from an aerial view image by using a visual language model, and constructing memory nodes in combination with a local truncated symbol distance field; establishing topological connection containing six-degree-of-freedom constraints according to a visual common-view relationship and spatial trafficability among the nodes, and performing hierarchical organization on semantic landmarks in the nodes to form a multi-modal memory map; a semantic weight is calculated based on a cross attention mechanism, and a three-dimensional semantic navigation potential field is constructed by fusing geometric accessibility probability; and determining a navigation target according to the potential field distribution, executing path search in a topology layer, and performing obstacle avoidance and track generation in a measurement layer. According to the invention, the long-voyage autonomous visual language navigation of the unmanned aerial vehicle is realized under the condition that the computing resources are limited.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Method and system for dynamically planning low-altitude route of unmanned aerial vehicle

The invention provides a dynamic planning method and system for a low-altitude route of an unmanned aerial vehicle, and relates to the technical field of intelligent planning, and the method comprises the steps: receiving a new delivery point instruction, obtaining a three-dimensional coordinate, carrying out the task integration of a new delivery point and an existing delivery point, defining a geographic range through three objects, namely a communication base station, a wireless tower and a permanent landmark building, and carrying out the task integration. Performing space structure division on the geographic range, and disassembling the geographic range into uniformly distributed three-dimensional voxel grids; and generating an adjustment value according to the coordinate interval of the three-dimensional voxel grid, including the delivery point density, the linear distance with the landmark building and other structural attributes, and adjusting the task integration result by using the adjustment value to obtain the adjusted task integration result of all delivery points. According to the invention, safe, efficient and accurate route planning of a multi-delivery-point dynamic newly-added scene in a complex low-altitude environment is realized.
Owner:TONGHANG FUTURE (BEIJING) AVIATION TECH DEV GRP CO LTD

Positioning system for space-time calibration of multiple sensors in monitoring blind area

The invention belongs to the technical field of computer vision and intelligent sensing, particularly relates to a positioning system for monitoring blind area multi-sensor space-time calibration, and aims to solve the problems of monitoring blind areas and track breakage caused by non-uniform space-time reference of multi-source heterogeneous sensors. The system comprises a space-time reference generation unit, distributed sensing nodes, a joint calibration engine, a real-time positioning calculation module and a self-adaptive feedback controller. Through nanosecond time synchronization, space registration based on a static landmark and a mobile calibration carrier and a dynamic re-calibration mechanism of closed loop feedback, continuous smooth trajectory output of cross-device target tracking is realized. The system supports layered architecture and block chain evidence storage, improves expandability and security, and is suitable for smart cities and industrial automation scenes.
Owner:MINGSHANG TECH CO LTD

Unmanned aerial vehicle natural language multi-modal navigation method and system based on multi-dimensional thinking chain

The invention discloses an unmanned aerial vehicle natural language multi-modal navigation method and system based on a multi-dimensional thinking chain. The method mainly comprises the following steps: constructing a software-in-the-loop simulation environment deeply integrated with task planning software, and collecting a live-action image with a time-space stamp; secondly, the natural language task instruction is analyzed, and key landmarks and position coordinates of the key landmarks are recognized in combination with a real scene image; then the recognition result and the structured task sequence are input into a multi-dimensional thinking chain reasoning engine in parallel to be processed; and finally, integrating the output of the multi-dimensional thinking chain inference engine, generating a flight action sequence, and issuing the flight action sequence to a flight control system to complete a navigation task. According to the method, the explicit thinking chain reasoning process is embedded into each step of navigation decision and is deeply integrated with the existing task planning ecology, so that the autonomy, adaptability and task success rate of the unmanned aerial vehicle in an unknown or dynamic environment can be remarkably improved.
Owner:HUAZHONG UNIV OF SCI & TECH

SLAM (Simultaneous Localization and Mapping) optimization mapping method and device combining geometric verification and constraint of reflector

The invention discloses an SLAM (Simultaneous Localization and Mapping) optimization mapping method and device combining geometric verification and constraint of a reflector, relates to the technical field of industrial-grade mobile robot navigation, and solves the problem that in the prior art, a continuous constraint mechanism cannot be established in a mapping process to correct accumulative errors in motion. The method comprises the following steps: acquiring an original laser radar point cloud containing a reflector, screening out reflection points with the distance smaller than a preset distance, and carrying out clustering and quintuple geometric verification on the screened points by adopting a density clustering algorithm, so as to obtain a multi-dimensional geometric verification mechanism based on PCA (Principal Component Analysis); local coordinates of the center of the reflector are calculated and converted into global coordinates, the reflector is used as a stable Landmark depth to be fused into an SLAM back-end graph optimization framework, and optimization nodes including a timestamp, a robot pose, a reflector ID and an observation pose are constructed based on a global coordinate system to participate in graph optimization; and the matching problem caused by insufficient natural characteristics or environment change in a long corridor scene can be solved in a targeted manner.
Owner:ZHEJIANG MILEY ROBOT CO LTD

Waypoint prediction method fusing instruction landmark features in visual language navigation

The invention discloses a waypoint prediction method fusing instruction landmark features in visual language navigation, and belongs to the crossing field of artificial intelligence and computer vision. The method is realized through the following steps: firstly, extracting landmarks and co-occurrence landmarks thereof from a natural language instruction by utilizing a large language model, and generating a candidate landmark sequence; then, the occurrence probability of the candidate landmarks in actual observation is calculated based on a visual large model CLIP, and noise interference is dynamically corrected and suppressed through a learnable co-occurrence scoring module; secondly, fusing the corrected landmark features with the depth visual features, inputting a double-layer Transform model to model a spatial relationship, generating a waypoint probability heat map, and outputting adjacent waypoints through non-maximum suppression (NMS); and finally, topological mapping is carried out on the predicted waypoints, an optimal path is obtained through cross-modal path planning, and an intelligent agent is guided to execute low-level actions. The method does not need to depend on a predefined environment map or manually annotate data, navigation errors in a continuous environment are remarkably reduced, the generalization ability of unseen scenes is improved, and the method is suitable for open world navigation scenes such as robot navigation and automatic driving.
Owner:DALIAN UNIV OF TECH

Illuminated multi-view sensing using 3D reconstruction for in-cabin applications

Optical sensors (e.g., cameras) and (e.g., IR) illumination sources may be distributed in an environment (e.g., an interior space such as a cabin or cockpit of an ego-machine) and synchronized to generate frames of sensor data, which may be used to reconstruct 3D geometry and / or 3D pose of an occupant, operator, or other object in the environment. For example, stereo vision may be used to generate one or more depth maps from image data generated using different cameras, the depth map(s) may be transformed into a 3D point cloud, and surface reconstruction may be applied to reconstruct the 3D geometry of surface(s) in the environment. A 3D pose, one or more keypoints (e.g., facial landmarks), or some other representation of the shape of the reconstructed surface(s) may be extracted from the reconstructed surface and used in one or more downstream tasks, such as driver and / or occupant monitoring tasks.
Owner:NVIDIA CORP

Geometrically assisted visual positioning method and system

A geometric structure aided visual localization method and a system implementing the method are provided. The method includes retrieving a 3D map of initialization location data, the 3D map being constrained by visual structures and geometric structures and modeled by a Gaussian mixture model of a set of Gaussian distributions with mapped landmarks; acquiring a series of real-time image frames by a camera of a mobile system; and for each real-time image frame: extracting local features from the real-time image frame; predicting a camera pose corresponding to the real-time image frame; creating a key frame by tracking the predicted camera pose in the 3D map; identifying temporally visible landmarks with respect to the created key frame; acquiring 2D-3D correspondences between the local features and the temporally visible landmarks; and localizing the mobile system by estimating a state of the real-time image frame based on the 2D-3D correspondences.
Owner:HONG KONG UNIV OF SCI & TECH R & D CORP LTD

Real-time ir fundus image tracking in the presence of artifacts using reference landmarks

To provide a more efficient system / method for eye motion tracking.SOLUTION: A system and method for eye motion tracking. An anchor point and a plurality of auxiliary points are selected from the reference image. Individual live images in the sequence of images are then searched for a match between the anchor point and the ancillary point. First, an anchor point is found, and then the search for individual auxiliary points is limited to a search window defined by the known distance and / or orientation of the auxiliary point to be searched relative to the anchor point.SELECTED DRAWING: Figure 4
Owner:CARL ZEISS MEDITEC INC +1

Time shift determination unit and method for determining time shifts

Methods and system are herein provided for determining a time shift between images captured by means of different cameras. In one example, a time shift determination unit comprises a computing device comprising one or more processors and a memory, wherein the computing device is configured to determine a defined plurality of landmarks in each image of a first plurality of successive images of a defined environment captured by means of a first camera, each landmark being defined by an x-coordinate and a y-coordinate in each of the respective images, determine the defined plurality of landmarks in each image of a second plurality of successive images of the defined environment captured by means of a second camera, each landmark being defined by an x-coordinate and a y-coordinate in each of the respective images, for each image captured by means of the first camera and the second camera, determine a mean x-coordinate or a mean y-coordinate for the defined plurality of landmarks in the image, compile a first chart representing a progression of the mean x-coordinate of the defined plurality of landmarks over the first plurality of successive images, and a second chart representing a progression of the mean x-coordinate of the defined plurality of landmarks over the second plurality of successive images, or compile a first chart representing a progression of the mean y-coordinate of the defined plurality of landmarks over the first plurality of successive images, and a second chart representing a progression of the mean y-coordinate of the defined plurality of landmarks over the second plurality of successive images, determine a similarity measure of the first chart and the second chart, perform a defined number of shifts of the second chart with respect to the first chart, wherein with each shift the second chart is shifted by a defined number of images, and for each shift of the plurality of shifts determine a similarity measure of the first chart and the shifted second chart, and evaluate the determined similarity measures in order to determine a time shift between the first plurality of successive images and the second plurality of successive images.
Owner:HARMAN BECKER AUTOMOTIVE SYST GMBH

Driver scoring system and method using optimum path deviation

Techniques are disclosed to determine driver scoring that consider deviations in map and sensor data in driver score computations. The deviations can be based on deviations from an optimum driving path, and include the determination of vectors between a reference point and one or more landmarks, and one or more reference vectors between the reference point the landmark(s) when traveling on an optimum driving path. A difference between the vectors may then be used determine the deviations from the optimum driving path. In contrast to the conventional approaches, the use of positional deviations (e.g. determined from road markings) in the computation of driver scores allows for improved driver scoring techniques and driver characterizations.
Owner:INTEL CORP

Position estimation of an anatomical landmark by text inputs

Training framework for creating an artificial intelligence (AI) system for estimating a position of an anatomical landmark by text inputs. The training framework includes providing a context-set comprising a list of names of anatomical landmarks, and a position-list comprising position-tokens being expressions referring to relative positions. A plurality of question-prompts asking for the relative position of a landmark are generated by using varying combinations of the landmarks and position-tokens of the context-set and the position-list. A number of target-landmarks to each question-prompt are generated by inputting the question-prompts in a large language model. The answer is parsed for landmarks and the found landmarks are defined as target-landmarks. A plurality of training-datasets are formed, wherein each training-dataset comprises the landmark, the position-token from a question-prompt and the target-landmark from the answer to this question-prompt. The AI-system is trained with the training-dataset and additional spatial coordinates of a part of the landmarks of the context-set.
Owner:SIEMENS HEALTHINEERS AG

Robot localization and mapping accommodating non-unique landmarks

Robot localization or mapping can be provided without requiring the expense or complexity an “at-a-distance” sensor, such as a camera, a LIDAR sensor, or the like. Adjacency-derived landmark features can be used and non-unique landmark features can be accommodated. Uncertainty in robot pose can be tracked and compared to an adaptive threshold, and non-dock and dock-based localization behavior can be controlled based on the uncertainty, the adaptive threshold, one or more other thresholds, and the accessibility of available differently oriented landmark features, such as perpendicularly oriented straight wall segments landmark features. Available features can be sorted according to a quality metric, and path planning and navigation techniques are also included for helping obtain successful wall-following and localization observations.
Owner:IROBOT CORP

Illuminated multi-view sensing using 3D reconstruction for in-cabin applications

Optical sensors (e.g., cameras) and (e.g., IR) illumination sources may be distributed in an environment (e.g., an interior space such as a cabin or cockpit of an ego-machine) and synchronized to generate frames of sensor data, which may be used to reconstruct 3D geometry and / or 3D pose of an occupant, operator, or other object in the environment. For example, stereo vision may be used to generate one or more depth maps from image data generated using different cameras, the depth map(s) may be transformed into a 3D point cloud, and surface reconstruction may be applied to reconstruct the 3D geometry of surface(s) in the environment. A 3D pose, one or more keypoints (e.g., facial landmarks), or some other representation of the shape of the reconstructed surface(s) may be extracted from the reconstructed surface and used in one or more downstream tasks, such as driver and / or occupant monitoring tasks.
Owner:NVIDIA CORP

Household robot article searching method and system based on multi-source pre-established map

The invention discloses a household robot article searching method and system based on a multi-source pre-built map, and the method comprises the following steps: S1, controlling a household robot to move in a household scene through employing an autonomous synchronous positioning and mapping technology, obtaining environment data obtained by a multi-source sensor carried by the household robot, and storing the environment data in a database; constructing an environment map fusing the ground occupation information and the overlook visual information, wherein the environment map is used as a multi-source pre-established map for subsequent article search; s2, calculating three types of scores of each landmark in the home scene; calculating the comprehensive score of each landmark according to the three types of scores of each landmark, and selecting the first K landmarks with the highest comprehensive score to construct an associated landmark set; and S3, in combination with the multi-source pre-established map and the associated landmark set, constructing a high-low layer collaborative target exploration mechanism, and realizing layer-by-layer reasoning from region planning to action execution. According to the method, the semantic reasoning result is accurately sent to the ground, the actual navigation behavior is driven, and the accuracy of target positioning and the efficiency of path planning are improved.
Owner:HUNAN UNIV

Methods for presenting route guidance at an augmented reality headset and devices and systems thereof

An example method includes, displaying, at an augmented-reality (AR) display of an AR headset, a map extended-reality (XR) augment that includes a first set of XR landmarks. The AR headset has a first orientation in the physical environment, and the first set of XR landmarks are presented to a wearer such that the first set of XR landmarks overlays the corresponding first set of physical landmarks of the physical environment seen through the AR display. The method includes that in response to detecting a change in orientation of the AR headset to a second orientation in the physical environment, displaying the map XR augment to include a second set of XR landmarks. The second set of XR landmarks are presented to the wearer such that the second set of XR landmarks overlays the corresponding second set physical landmarks of the physical environment seen through the AR display.
Owner:META PLATFORMS TECHNOLOGIES LLC

Data coupling system and method for satellite image data and video information

The invention relates to the field of image data processing, and discloses a satellite image data and video information data coupling system and method, and the system comprises a data calibration module which is used for carrying out the space-time calibration processing of satellite image data, and obtaining a reference remote sensing image; the anchor point frame modeling module is used for carrying out background modeling processing on the video anchor point frame to obtain a monitoring background image; the ground feature fusion rendering module is used for identifying landmark ground features in the reference remote sensing image and performing feature fusion rendering on the landmark ground features to obtain a rendered remote sensing image; the lacking object compensation module is used for determining a lacking object in the reference remote sensing image and performing compensation rendering on the reference remote sensing image to obtain a compensation remote sensing image; and the image coupling generation module is used for carrying out coupling processing on the rendered remote sensing image and the compensated remote sensing image to obtain an accurate monitoring graph of the target monitoring area. According to the invention, the coupling accuracy of the satellite image data and video information data coupling system can be improved.
Owner:MINISTRY OF ECOLOGY & ENVIRONMENT CENT FOR SATELLITE APPL ON ECOLOGY ENVIRONMENT

Systems and methods for performing an autonomous aircraft visual inspection task

This disclosure provides system and method for performing an autonomous aircraft visual inspection task using an unmanned aerial vehicle (UAV). The UAV is equipped with a front-facing RGB-D camera, one Velodyne three dimensional Light Detection and Ranging with 64 channels, and one Inertial Measurement Unit. In the method of the present disclosure, the UAV takeoff from any nearby location of the aircraft and face the RGB-D camera towards the aircraft. The UAV find the nearest landmark using a template matching approach and register with the aircraft coordinate system. The UAV navigate using LiDAR and IMU measurements, whereas the inspection process uses measurements from the RGB-D camera. The UAV navigate using a proposed safe navigation around the aircraft by avoiding obstacles. The system identifies the objects of interest using a deep-learning based object detection tool and then performs the inspection. A simple measuring algorithm for simulated objects of interest is implemented.
Owner:TATA CONSULTANCY SERVICES LTD

A landmark-based positioning success determination method, chip and robot

The application discloses a landmark-based positioning success judgment method, a chip and a robot, and the positioning success judgment method comprises the following steps: 1, when the robot judges that the same landmark is consistent for multiple times, it is judged whether the concentration of the feature point corresponding to the landmark is in a preset concentration threshold range; if yes, step 2 is executed, otherwise, it is determined that the landmark positioning is successful; 2, when the robot judges that at least one reference landmark is consistent with the landmark in step 1, it is determined that the landmark in step 1 is successfully positioned; the reference landmark is different from the landmark in step 1, and the reference landmark is a landmark with uneven feature point distribution.
Owner:AMICRO SEMICONDUCTOR CO LTD

Method and system for tracking a state of a camera

The invention relates to a method for determining a state xk (7) of a camera (11) at a time tk, the state xk (7) being a realization of a state random variable Xk, wherein the state is related to a state-space model of a movement of the camera (11). The method comprises the following steps: a) receiving an image (1) of a scene of interest (8) in an indoor environment (8) captured by the camera (11) at the time tk, wherein the indoor environment (8) comprises N landmarks (9) having known positions and orientations in a world coordinate system (12), N being a natural number; b) receiving a state estimate Formula I (2) of the camera (11) at the time tk, wherein the state estimate (2) comprises an estimate of the pose of the camera; c) determining (3) positions of M features in the image (1), M being a natural number; and d) determining (6) the state xk (7) of the camera (11) at the time tk based on (i) observation zk at the time tk, the observation zk being a realization of a joint observation random variable zk, the observation zk comprising the positions of the M features and data indicative of distance between each of the M features and its corresponding object point in the scene of interest, respectively, and (ii) the state estimate Formula {circumflex over ( )}I (2), wherein the determining (6) of the state xk (7) comprises determining (4) an injective mapping estimate from at least a subset of the M features into the set of the N landmarks (9), and wherein the determining (6) of the state xk (7) is based on an observation model set up (5) based on the determined injective mapping estimate. The invention also relates to a computer program product and to an assembly.
Owner:VERITY AG

System

A system is provided.SOLUTION: A system comprising: means for registering a user's hometown; means for recommending a niche landmark near the registered hometown based on data of another application; means for giving a point when the user visits the recommended place and performs an action; means for transferring the given point to another piece of electronic money; and means for giving the point to an organization having jurisdiction over the place.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

SYSTEM AND METHOD FOR VEHICLE LOCALIZATION IN A TUNNEL

The present disclosure provides a system (104) and a method (800) for vehicle localization in a tunnel, utilizing sensor data and tunnel-specific features. The system (100) is configured to determine the lateral and longitudinal positions of the vehicle (101) in real time. The lateral position is identified by updating a light probability distribution and applying Bayesian filtering to establish the vehicle's ego lanes. Furthermore, a real-time landmark identifier (ID) for lights is generated based on normalized tunnel features, including distances, light intensity, and tunnel geometry. The longitudinal position is determined by matching the generated landmark ID with predefined map data, enabling robust localization for navigation and thus allowing the vehicle (101) to be located in tunnels.
Owner:MERCEDES BENZ GROUP AG

Indoor positioning method and system, storage medium and server

The embodiment of the invention discloses an indoor positioning method and system, a storage medium and a server, and is applied to the technical field of computer vision, three-dimensional reconstruction and indoor positioning. The method comprises the following steps: a server extracts a semantic landmark in a current view of a mobile terminal and semantic information for describing a relationship between the semantic landmark and the mobile terminal through a semantic understanding technology based on a received video stream sent by the mobile terminal; and the real-time feature vector obtained based on the semantic landmark and the relational semantic information is matched with the preset semantic landmark feature library, so that the current position of the mobile terminal can be quickly determined. In this way, the snapshot type positioning experience of scanning and knowing is achieved, the method does not depend on the initial position and the continuous movement track of the mobile terminal, and the complexity of indoor positioning is greatly reduced.
Owner:SHENZHEN SMARTCITY TECH DEV GRP CO LTD

Sensor system for mobile platform and shape-based landmark identification method

The mobile platform is configured to perform one or more tasks in a work site including a passive landmark. The mobile platform may include a first laser rangefinder and at least one processor configured to: scan the first passive landmark with the first laser rangefinder to collect a first plurality of distance measurements for a first plurality of yaw angles; fitting a first shape to the first plurality of distance measurements based on a predetermined shape of the first passive landmark; and determining the position of the geometric center of the first passive landmark relative to the first position of the first laser range finder based on the fitted first shape.
Owner:ROBUST ROBOTICS LTD

SLAM pose optimization method and three-dimensional reconstruction equipment

The invention relates to the technical field of three-dimensional reconstruction, in particular to an SLAM pose optimization method and three-dimensional reconstruction equipment. According to the method, after a three-dimensional scene is marked, an SAM segmentation model is introduced, disordered three-dimensional point cloud data is projected into an ordered two-dimensional intensity image, then the SAM segmentation model can process laser radar data, the SAM model can be used for extracting three-dimensional coordinates of mark points more accurately, and the calculation precision of the mark points is remarkably improved. In addition, a global factor graph is generated by taking the extracted semantic three-dimensional coordinates of the mark points as semantic road sign factors, the pose of the laser radar is constrained and optimized according to the global factor graph, constraint is provided in scenes such as tunnels and long corridors when geometric features are missing, and SLAM drift is remarkably inhibited.
Owner:SHENZHEN XGRIDS-INNOVATION CO LTD

Unintended use restricted use method and apparatus, product, sighting telescope apparatus and medium

The embodiment of the invention provides an unexpected use limited use method and equipment, a product, sighting telescope equipment and a medium. The use limiting method for the unexpected purpose comprises the following steps: acquiring scene image data in a current visual field range of sighting telescope equipment; detecting a target feature contained in the scene image data, and judging whether the current use is the unexpected use or not based on the matching degree of the target feature and a landmark feature representing the unexpected use in a sensitive feature library; and if the sighting telescope equipment is not used for the expected purpose, executing a use limiting strategy on the sighting telescope equipment, and performing read-only storage on an event used for the unexpected purpose.
Owner:HEFEI YINGJU INNOVATION TECHNOLOGY CO LTD

A landmark-based positioning data screening method

The application discloses a landmark-based positioning data screening method, which comprises the following steps: 1, determining the landmark with successful positioning according to the concentration of the feature points corresponding to the landmark, and marking the positioning data of the landmark with successful positioning as target positioning data; 2, when the robot traverses one target positioning data, screening the target positioning data falling within the preset displacement range of the target positioning data, performing angle difference filtering on the screened target positioning data, and then obtaining the reference positioning data group and the target positioning data existing in the reference positioning data group; 3, performing average value processing on the target positioning data in the reference positioning data group to obtain average positioning data, and selecting the target positioning data with the smallest pose difference degree relative to the average positioning data from the target positioning data in the reference positioning data group to set as optimal positioning data.
Owner:AMICRO SEMICONDUCTOR CO LTD

Scalable Real-Time Hand Tracking

PendingUS20260017806A1Image enhancementImage analysisRadiologyPoint localization
Example aspects of the present disclosure are directed to computing systems and methods for hand tracking using a machine-learned system for palm detection and key-point localization of hand landmarks. In particular, example aspects of the present disclosure are directed to a multi-model hand tracking system that performs both palm detection and hand landmark detection. Given a sequence of image frames, for example, the hand tracking system can detect one or more palms depicted in each image frame. For each palm detected within an image frame, the machine-learned system can determine a plurality of hand landmark positions of a hand associated with the palm. The system can perform key-point localization to determine precise three-dimensional coordinates for the hand landmark positions. In this manner, the machine-learned system can accurately track a hand depicted in the sequence of images using the precise three-dimensional coordinates for the hand landmark positions.
Owner:GOOGLE LLC