Dual-camera monitoring and voice commands track learner focus, detect interruptions, and deliver real-time remote guidance.
Ear images and deep learning replace earbud trial and error by predicting fit and comfort from anatomical features.
Local face-region CNNs and fusion learning improve recognition under lighting, pose, and occlusion changes by separating similar targets.
Combining 2D barcode images with 3D scene data improves decode success, object recognition, and fault detection in barcode readers.
Face and body feature fusion enables on-device pet photo clustering with strong precision and recall, even in growing offline galleries.
Cameras and AI map occupancy and light usage across open offices to automate lighting control and cut energy and carbon waste.
Image analysis defines a target sensing zone so radar can switch resolution modes, improving detection accuracy while cutting power use.
Physical markers placed beside local features automate ROI and mask generation, cutting manual labeling time while preserving training data quality.
Skeleton feature extraction turns simple keyword input into pose-based image search, reducing query preparation time while preserving search accuracy.
Image recognition detects AR glasses and sends restricted vehicle display content only to the wearer, preventing unauthorized viewing.
Large language models generate rule determination logic from input rules and case collections, avoiding manual tuning across diverse work sites.
Adaptive frame capture triggered by speech and hand detection improves AR query interpretation while reducing power use and continuous visual collection.
Interior sensors classify driver actions into vehicle operation modes and transmit concise status data to reduce distraction and processing load.
Combining low-resolution body views with high-resolution hand-object regions improves action recognition speed and accuracy for real-time vision.
Registers face features from user-selected regions in multi-person images, enabling consent-based recognition with lower privacy and storage risk.
Combining low-resolution body views with high-resolution hand-object regions helps identify actions faster without losing small-object detail.
RF sensing adds user location and direction cues to audio, improving voice command recognition in noise and soft-speech conditions.
Adjusts elevator voice and display notification speed using detected age or disability attributes for clearer in-car information delivery.
Surrounding-scene analysis grades nuisance behavior and selects alert levels that curb violations without causing discomfort in public spaces.
Two neural networks compare radar and camera-derived features to detect elements accurately while avoiding camera use in privacy-sensitive areas.
Combining moiré, hand, facial, and hologram checks improves ID card forgery detection while supporting efficient online verification.
Separating RF reflection capture from trusted remote classification helps secure pose and location imaging when wireless devices are exposed.
A treat-guided camera arm stabilizes pet images for accurate AI identification, personalized recommendations, and in-store item navigation.
Pattern-of-life analysis and surveillance training help officers anticipate evasive behavior, including encrypted communications and digital currencies.
Real-time optical scanning captures animal biometric indicators to fit saddles or shoes accurately while reducing manual measurement errors.
Skeleton and attribute height cues are compared to identify camera parameters when a person's true height is unknown.
Viewpoint-only transmission and shared 3D object data enable immersive replay while reducing network traffic and terminal processing load.
Microphone and camera cues trigger a conversational state so appliances can interpret natural speech more accurately and avoid unintended actions.
Human body region division and feature vector matching improve dress code recognition for specific body parts while keeping model complexity manageable.
A mediator-based distribution architecture spreads virtual items across multiple live streams and viewers to expand recognition beyond the original audience.
Machine learning converts industrial video frames into compact metadata so remote sites can render scenes with less bandwidth, lag, and delay.
A moving guide image shifts the capture area with the hand path, reducing blur and improving fingertip image clarity.