Similarity-based image clustering plus in-context learning cuts segmentation labeling effort while preserving annotation accuracy at scale.
Automatic object tracking and wireless camera control let one operator manage multiple video devices while reducing fatigue during long shoots.
User-selected ratio switching exchanges capability data with external devices so returned video matches the display format without distortion.
A low-power AO sensor detects and decodes visual codes before waking the AP, cutting scan time and power use in electronic devices.
Reference probabilities from image sets refine low-confidence object classification results, improving accuracy in complex scenes.
By classifying training data with uncertainty and inconsistency scores, the model update flow cuts manual labeling and limits outlier overlearning.
Automatic detection of inclusion relations between stroke sets generates hierarchical metadata, reducing manual tagging effort and improving ink data handling.
Auxiliary images from a second sensor train bulk-flow classifiers to keep sorting accuracy high while reducing sensor cost.
Machine learning tracks local and global features across comic panels to automate annotation despite changing drawing styles and character appearance.
Spatial self-alignment, multi-sampling, and temporal aggregation improve video action boundary localization with less end-to-end training.
Laser-based presence sensing on rig walkways blocks hazardous machine motion when personnel enter monitored access areas.
Measures eyelid opening distance and eye reopening duration to detect driver drowsiness more reliably across individuals and lighting conditions.
Stored facial profiles let a security system recognize authorized people, suppress false alarms, and reduce manual event review.
RFID, cameras, and presence sensors identify pallets at the dock to verify truck assignment and loading sequence without slowing operations.
A server-mediated shared virtual space keeps AR objects aligned in position and time across remote physical environments for seamless multiuser interaction.
Automated template matching groups similar document pages into batches, reducing slow manual review in large unstructured data sets.
Correlates facial recognition with MEMS-tracked item interactions to prevent misidentification, unauthorized use, and deepfake-based approval.
Shared material attributes and edit locking keep AR scenes synchronized across devices while preventing multi-user editing conflicts.
User-corrected facial map points refine 3D face meshes, improving AR cosmetics try-on accuracy and personalized product recommendations.
A gesture mask isolates on-screen content and maps gestures to search, save, or share actions with less input and better relevance.
A flow-modeled error distribution reshapes keypoint loss, improving multi-task joint training accuracy without separate models.
Correcting size and position of small handwritten character rectangles helps OCR avoid misreading commas and punctuation.
AI ranks content interest to retain primary and auxiliary video streams at different resolutions, cutting storage and transcoding overhead.
Language-based document splitting improves bulk OCR speed and accuracy, then recombines sections into a searchable PDF.
Brain waves and facial biosignals are fused to predict mouth motion for avatar speech when users cannot speak or must stay silent.
Reference feature detection and projective correction normalize folded or distorted document images for faster, more accurate content extraction.
Combines OCR, NLP, fuzzy matching, and ML refinement to improve text extraction and entity matching across noisy, inconsistent documents.
Automatic frame-by-frame privacy detection and blurring protects sensitive screen content without interrupting recording.
Multiple vehicle observations are aligned into anchor-based point clouds to correct GPS drift and produce more reliable lane line maps.
Heterogeneous graph features fuse face, device, and object data to improve similar-face verification accuracy in face scan payments.
Multiple cameras identify and track vehicles across drive-through zones, improving POS updates, lane alerts, and kitchen preparation.
Compares detected road line pairs, resolves overlapping lane conflicts, and generates more reliable boundary line maps for automated driving.
Immersive XR overlays combine video feeds and real-time facility data to speed remote decisions while preserving spatial awareness and operator safety.
Relabeling scores combine severity incoherence, user feedback, and reviewer bias to target mislabeled videos and retrain models efficiently.
Captured exterior imagery is converted into 3D building models for real-time remote viewing, navigation, and post-construction updates.
Multiple feature maps are fused into point-level vectors to improve image matching accuracy for map information updates.
Location data from a secure document and the user device are matched to score identity confidence and reduce false authentication decisions.
Hierarchical feature aggregation helps vision transformers classify images accurately with less training data and simpler architecture.
Computer vision and machine learning turn web images and videos into assistant skills, reducing manual skill development effort.
A CNN predicts document posture correction from target and reference images, cutting feature-point processing load in continuous eKYC capture.
Two imaging devices with different views combine facial recognition and rear-object detection to block tailgating at passage devices.
Combining image analysis with text signatures verifies declared image content with less training data and broader object coverage.
A shared optical transmit-receive chain combines face distance sensing and data transfer, cutting module size and stopping emission when too close.
Camera-based valve image matching identifies unlabeled fluid components on site, reducing maintenance delays and helping technicians act immediately.
Pixel weight templates and half-row offsets extract 1D signals from 2D images with uniform scaling, high resolution, and less blurring.
Automated prompt retrieval and iterative crop refinement let one vision-language model handle diverse image cropping tasks without fine-tuning.
Segmented skin tone references define lighting limits that avoid blocked shadows and blown highlights in face authentication.
Dynamic ROI exposure follows the aimer position and target distance to keep machine-readable symbols properly exposed in complex scenes.
Parallel convolution and pooling paths with different kernel sizes preserve local detail while improving multi-scale semantic segmentation.
Local and global attention windows capture long-term video context for frame-level action segmentation with lower computation and training time.