Similarity-based image clustering plus in-context learning cuts segmentation labeling effort while preserving annotation accuracy at scale.
Automatic object tracking and wireless camera control let one operator manage multiple video devices while reducing fatigue during long shoots.
User-selected ratio switching exchanges capability data with external devices so returned video matches the display format without distortion.
A low-power AO sensor detects and decodes visual codes before waking the AP, cutting scan time and power use in electronic devices.
Reference probabilities from image sets refine low-confidence object classification results, improving accuracy in complex scenes.
By classifying training data with uncertainty and inconsistency scores, the model update flow cuts manual labeling and limits outlier overlearning.
Automatic detection of inclusion relations between stroke sets generates hierarchical metadata, reducing manual tagging effort and improving ink data handling.
Auxiliary images from a second sensor train bulk-flow classifiers to keep sorting accuracy high while reducing sensor cost.
Machine learning tracks local and global features across comic panels to automate annotation despite changing drawing styles and character appearance.
Spatial self-alignment, multi-sampling, and temporal aggregation improve video action boundary localization with less end-to-end training.
Laser-based presence sensing on rig walkways blocks hazardous machine motion when personnel enter monitored access areas.
Measures eyelid opening distance and eye reopening duration to detect driver drowsiness more reliably across individuals and lighting conditions.
Stored facial profiles let a security system recognize authorized people, suppress false alarms, and reduce manual event review.
RFID, cameras, and presence sensors identify pallets at the dock to verify truck assignment and loading sequence without slowing operations.
A server-mediated shared virtual space keeps AR objects aligned in position and time across remote physical environments for seamless multiuser interaction.
Automated template matching groups similar document pages into batches, reducing slow manual review in large unstructured data sets.
Correlates facial recognition with MEMS-tracked item interactions to prevent misidentification, unauthorized use, and deepfake-based approval.
Shared material attributes and edit locking keep AR scenes synchronized across devices while preventing multi-user editing conflicts.
User-corrected facial map points refine 3D face meshes, improving AR cosmetics try-on accuracy and personalized product recommendations.
A gesture mask isolates on-screen content and maps gestures to search, save, or share actions with less input and better relevance.
A flow-modeled error distribution reshapes keypoint loss, improving multi-task joint training accuracy without separate models.
Correcting size and position of small handwritten character rectangles helps OCR avoid misreading commas and punctuation.
AI ranks content interest to retain primary and auxiliary video streams at different resolutions, cutting storage and transcoding overhead.
Language-based document splitting improves bulk OCR speed and accuracy, then recombines sections into a searchable PDF.
Brain waves and facial biosignals are fused to predict mouth motion for avatar speech when users cannot speak or must stay silent.
Reference feature detection and projective correction normalize folded or distorted document images for faster, more accurate content extraction.
Combines OCR, NLP, fuzzy matching, and ML refinement to improve text extraction and entity matching across noisy, inconsistent documents.
Automatic frame-by-frame privacy detection and blurring protects sensitive screen content without interrupting recording.
Multiple vehicle observations are aligned into anchor-based point clouds to correct GPS drift and produce more reliable lane line maps.
Heterogeneous graph features fuse face, device, and object data to improve similar-face verification accuracy in face scan payments.
Multiple cameras identify and track vehicles across drive-through zones, improving POS updates, lane alerts, and kitchen preparation.
Compares detected road line pairs, resolves overlapping lane conflicts, and generates more reliable boundary line maps for automated driving.
Immersive XR overlays combine video feeds and real-time facility data to speed remote decisions while preserving spatial awareness and operator safety.
Relabeling scores combine severity incoherence, user feedback, and reviewer bias to target mislabeled videos and retrain models efficiently.
Captured exterior imagery is converted into 3D building models for real-time remote viewing, navigation, and post-construction updates.
Multiple feature maps are fused into point-level vectors to improve image matching accuracy for map information updates.
Location data from a secure document and the user device are matched to score identity confidence and reduce false authentication decisions.
Hierarchical feature aggregation helps vision transformers classify images accurately with less training data and simpler architecture.
Computer vision and machine learning turn web images and videos into assistant skills, reducing manual skill development effort.
A CNN predicts document posture correction from target and reference images, cutting feature-point processing load in continuous eKYC capture.
Two imaging devices with different views combine facial recognition and rear-object detection to block tailgating at passage devices.
Combining image analysis with text signatures verifies declared image content with less training data and broader object coverage.
A shared optical transmit-receive chain combines face distance sensing and data transfer, cutting module size and stopping emission when too close.
Camera-based valve image matching identifies unlabeled fluid components on site, reducing maintenance delays and helping technicians act immediately.
Pixel weight templates and half-row offsets extract 1D signals from 2D images with uniform scaling, high resolution, and less blurring.
Automated prompt retrieval and iterative crop refinement let one vision-language model handle diverse image cropping tasks without fine-tuning.
Segmented skin tone references define lighting limits that avoid blocked shadows and blown highlights in face authentication.
Dynamic ROI exposure follows the aimer position and target distance to keep machine-readable symbols properly exposed in complex scenes.
Parallel convolution and pooling paths with different kernel sizes preserve local detail while improving multi-scale semantic segmentation.
Local and global attention windows capture long-term video context for frame-level action segmentation with lower computation and training time.
A neural network classifies facial images by processing color channel data and applying parameter transformations for alignment.
Redundant sensing elements sample signal amplitude differences to calculate block-to-block adjustments, correcting noise-induced variations across sensor rows.
Agricultural vehicle cameras sample images using trained machine-learning models to store data for future training.
Jointly trains multi-task dense prediction models with hardware-aware neural architecture search to reduce latency while stabilizing relative error noise.
A vehicle lighting system digitally models detected targets to generate precise illumination matching their contours.
A temporal fusion net generates hidden states from current LiDAR embeddings and previous camera data to process sensor inputs asynchronously.
Region of interest maps guide local encoding quality adjustments, reducing bandwidth consumption while maintaining object detection accuracy.
Infrared detection captures subdermal facial features to resolve measurement precision issues caused by visible light interference.
An object discriminating apparatus converts similarity scores using derived variation differences to identify objects accurately.
Automated anomaly detection monitors network resources to prevent unexpected costs from sudden usage spikes.
Segmented layer processing automates depth estimation and occlusion filling, resolving the trade-off between visual quality and conversion time.
An automated method analyzes video frames to identify clean areas for ad insertion without obstructing faces.
A weakly supervised model generates heat maps and bounding boxes from keywords to detect objects without manual annotation.
Unsupervised machine learning connects dispersed user actions into significant automation routines, resolving manual process identification bottlenecks.
A makeup support device holds pigment on a non-water-soluble sheet surface for stable skin adhesion.
Automated textile visual quality control selects specific image sub-areas for pixel evaluation.
A modifying unit transforms intermittent lines into solid marks using machine learning before classification.
Iterative depth feature learning estimates articulated object posture accurately while shifting computational load to an offline training phase.
A content moderation application detects inappropriate video streams using trained machine learning models.
A magnetic pad and image processing system track interactive tools in real time, resolving noise interference and improving tracking accuracy.
A gesture-based security system uses mobile device metadata to interpret video feeds and classify motion patterns.
Apparatus senses user movement to classify sound-producing gestures and provides audio feedback by modifying the cancellation process.
A multi-task training system adjusts shared parameters to minimize task-induced variance between gradients.
An image decoder identifies independently decodable blocks and pre-scales them based on display size.
A search engine maps images into a vector space and calculates scores using hyperplanes representing distinct query senses to rank results.
A dimple position detection device adjusts binarization levels based on reflected light intensity to locate the dimple top accurately.
A 3D vertical integration architecture processes track patterns using high bandwidth board-to-board communications.
Computing device selects machine-readable link type based on document characteristics, resolving manual selection bottlenecks that hinder automated workflows.
A face recognition method segments images into overlapping patches to build KNN graphs for calculating local and global structural similarities.
A Poisson cloning algorithm integrates facial feature templates into original images using reduced resolution processing.
A bidirectional compact deep fusion framework merges modality-specific encoder branches into a single shared structure.
Video processing segments the chest region to determine target hand positions, resolving setup delays and enabling effective guidance without dedicated sensors.
Automated machine vision system inspects rail components to detect defects like displaced anchors and missing spikes.
A configurable filter adapts to image data characteristics to perform de-blocking and de-ringing operations.
A processor determines a feature vector from an image of a load carrier to identify its identity using optical codes.
Electronic apparatus generates driving line information from vehicle image data for augmented reality navigation.
A self-supervised language model creates hierarchical similarity matrices to score variable length documents without manual labels.
Interface system merges separate touch and non-touch detection outputs via differential threshold comparisons to resolve accuracy versus complexity trade-offs.
A recognition device segments monitored images into blocks to process only event-centralized regions.
Endpoint-based compression reduces memory bandwidth for high dynamic range data by storing interpolated pixel indices instead of full bit-depth values.
Cluster video frames by scene similarity to sample representative images for visual recognition processing.
A touch sensing layer uses distinct reflectivity to identify the biometric sensing region on an electronic device display.
AI analyzes drone bridge images to identify maintenance needs and calculate optimal solutions.
Calculates check image resolution from MICR character coordinates, resolving inaccurate header metadata for reliable OCR processing.
Messaging application detects user location and notifies of popular content, eliminating manual QR code scanning.
Extracts significant points from product images to filter noise and align features for accurate authentication.
Infrared and visible light cameras capture container images to identify codes and colors, eliminating manual counting errors and reducing inspection time.
A similarity hash function identifies redundant binary data pages using feature vectors and alignment tolerance.
Neural networks combine keypoint vectors with feature maps to classify documents, reducing computational complexity while maintaining accuracy.