Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2755 results about "Object detection" patented technology

Object detection is a computer technology related to computer vision and image processing that deals with detecting instances of semantic objects of a certain class (such as humans, buildings, or cars) in digital images and videos. Well-researched domains of object detection include face detection and pedestrian detection. Object detection has applications in many areas of computer vision, including image retrieval and video surveillance.

Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Unmanned aerial vehicle image-based small object detection method for target areas

The present invention relates to the technical field of deep learning and computer vision. Disclosed is an unmanned aerial vehicle image-based small object detection method for target areas. The present invention crops images of obvious small objects in certain target areas, and annotates the small objects of different categories to form a raw training and testing dataset, so as to ensure the accuracy of data required in the early stage of the algorithm and further ensure the scientificity of the algorithm; uses the computing capability of an improved YOLOv7 detection model to collect image features of different degrees in the dataset, the improved YOLOv7 detection model using YOLOv7 as a basic model and adding to a neck network an MS-CET module, which is constituted by an improved self-attention mechanism and convolution module SPPCSP, and a BHC-FB module, which is constituted by bidirectional mixed convolution modules NConv and RPConv connected in parallel; and finally fuses different feature layers as a final judgment basis of an unmanned aerial vehicle for small object detection in the target areas, to further check the accuracy of the algorithm and criteria for dataset selection, thereby improving recognition accuracy.
Owner:CHONGQING UNIV OF TECH

Multi-stage filtering road thrown object detection method based on dynamic difference analysis

The invention relates to a multi-stage filtering road spilled object detection method based on dynamic difference analysis, which is suitable for automatic identification of unstructured foreign matters in video monitoring. The method comprises the following steps: firstly, extracting a reference image road mask, eliminating vehicle and pedestrian interference by using YOLOv8 detection, and extracting a motion candidate area through a frame difference method and background modeling; and then context expansion and super-resolution reconstruction are carried out on the candidate region, the candidate region is converted into an HSV space, multi-dimensional features such as color similarity, structural similarity and shadow determination are synthesized for screening, false detection is further removed in combination with inter-frame time sequence consistency, and finally a stable detection result is output. The method provided by the invention has the advantages of strong anti-interference capability, high adaptability, high detection precision and the like, and is suitable for the intelligent recognition task of the expressway thrown objects in a complex environment.
Owner:CCCC HUAKONG (TIANJIN) CONSTR GRP CO LTD

Laser - based targeting and object detection system

A pest control system is disclosed comprising an optical, computational, and monitoring subsystem, optionally mounted on a mobile platform. The optical system may include a neutralizing laser or multi-wavelength light source, discovery and detail cameras (optionally stereo), a beam-steering mechanism, tunable focus, and optional thermal or depth sensors. The processor, such as a GPU or FPGA, identifies insect or biological targets, adjusts laser focus by depth, and controls beam activation. A monitoring system verifies safety by detecting humans or other non-target entities using environmental and thermal cameras; if detected, laser firing is inhibited. The mobile platform may use wheels, propellers, tracks, or cables, with GPS and data links for remote control. A visible light pre-flash may induce a blink reflex before firing. In some embodiments, a scouting drone transmits target coordinates to the neutralization unit, enabling coordinated, efficient, and safe laser-based pest control.
Owner:REYNTJENS NICK

Multi-modal sensor-based detection and tracking of objects using bounding boxes

A perception system may be used to generate bounding boxes for objects in a vehicle scene. The perception system may receive images and feature maps corresponding to the received images. The perception system may correlate object queries from previous time steps with object queries from the current time step.
Owner:MOTIONAL AD LLC

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Method for Fusing Grid Maps Obtained Based on Multi-Sensors and Mobility Device Using the Method

PendingUS20260028041A1Image enhancementScene recognitionFused gridAlgorithm
A method performed by an apparatus for controlling autonomous driving of a vehicle is introduced. The method may comprise generating, based on a segmentation model processing point cloud data, a first semantic grid map, generating, based on an object detection model, a second semantic grid map, adjusting a probability regarding whether occupancy exists for an element included in each grid of the first semantic grid map and the second semantic grid map, and generating a fused grid map by determining, as a representative label, at least one label corresponding to a highest value among final probabilities of the at least one label, wherein the final probabilities are determined based on whether the at least one label matches the element, outputting, based on the fused grid map, a signal, and controlling, based on the signal, autonomous driving of the vehicle.
Owner:HYUNDAI MOTOR CO LTD +2

Federated object detection learning method based on representation enhancement and weighted aggregation under cloud-edge-terminal environment

A federated object detection learning method based on representation enhancement and weighted aggregation under cloud-edge-terminal environment comprises the steps of: 1) building a centralized federated learning framework under cloud-edge-terminal environment; 2) locally conducting representation enhancement training to strengthen model learning for few-shot category after receiving a model from the server at the client; 3) carrying out the weighted aggregation for client models in accordance with sample distribution to obtain the global model after receiving models from all clients at the server. With regard to the problem of existing federated object detection learning on low global model accuracy and weak generalization ability, the present invention can improve the accuracy and generalization ability of global object detection model.
Owner:ZHEJIANG UNIV OF TECH

Intelligent ship detection system based on convolutional neural network and radar signal processing

The invention relates to the technical field of computer vision and radar perception, in particular to a ship intelligent detection system based on a convolutional neural network and radar signal processing, which comprises a ship three-dimensional perception modeling module, a visual feature hierarchical fusion module, a target detection module and an anomaly detection module. The system emphatically utilizes deep learning methods such as a convolutional neural network and the like to realize three-dimensional space modeling and visual feature extraction of a ship by a radar in a water area environment. And a layered adaptive fusion and mutual information enhancement mechanism is adopted. The end-to-end detection model introduces a spatial hierarchy weighting strategy, so that the object detection accuracy and interpretability under the conditions of multi-target density, shielding and dynamic change in a complex water scene are improved. The system realizes continuous tracking, anomaly detection and risk early warning of ship navigation behaviors based on space-time dynamic modeling. The whole scheme has high precision, strong robustness and adaptive ability, and can meet the requirements of ship detection and intelligent management and control in complex water area environments such as smart ports and water traffic.
Owner:JIANGSU HUASHUN INTELLIGENT TECH CO LTD

Small object detection method based on improved yolov8

Disclosed in the present invention is a small object detection method based on improved YOLOv8. The small object detection method comprises: inputting, into a pre-trained small object detection model based on improved YOLOv8, a small object image to be subjected to detection for identification, so as to obtain a detection result. A training method for a small object detection model based on improved YOLOv8 comprises: acquiring a small object image dataset, and dividing same into a training set and a validation set; using a backbone network ATDeNet to replace a YOLOv8 backbone network, so as to construct a small object detection model based on improved YOLOv8; and using the training set and the validation set to train the constructed small object detection model, so as to obtain a trained small object detection model based on improved YOLOv8. The accuracy and efficiency of small object detection can be significantly improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Automatic focusing method, electronic device, and readable storage medium

PCT designated stageWO2025261246A1Imaging equipmentElectric devices
The present invention provides an automatic focusing method, an electronic device, and a readable storage medium. The automatic focusing method comprises: using a target detection model to detect a target object in an image frame acquired at an initial focal length by an optical imaging device to be focused, so as to acquire a target object detection result; on the basis of the target object detection result, determining whether a target object is present in the image frame; if yes, determining an actual object distance on the basis of the target object detection result and the initial focal length, and on the basis of the actual object distance and a mapping relationship between the object distance and an optimal imaging focal length, determining the optimal imaging focal length; and if not, using a preset search algorithm to search for a focal length until an image having the highest definition value is found, and using a focal length corresponding to the image having the highest definition value as the optimal imaging focal length. The present invention primarily employs a deep learning-based automatic focusing method, supplemented by an image definition evaluation method. Compared with traditional passive focusing methods, the present invention greatly improves the focusing efficiency. Compared with traditional active focusing methods, the present invention eliminates a need for adding an additional optical ranging component, ensuring that the imaging device has a simple structure and a low cost.
Owner:MICROPORT UROCARE(SHANGHAI) CO LTD

Small-sample target detection method and system based on aggregation variational prototype

The invention discloses a few-sample target detection method and system based on an aggregation variational prototype. The method comprises the steps of constructing a data set containing a base class and a new class, dividing the data set into a support set and a query set, generating a class prototype by utilizing a P-VAE module in combination with CLIP semantic features and a feature discriminator, realizing bidirectional fusion of the support set and the query set features by means of an MFM module, fusing the query features and the class prototype, and inputting the fused query features and the class prototype into a detection head to complete target detection. The system comprises a data set construction module, a priori variational automatic encoder P-VAE module, a mutual fusion module MFM, a feature fusion module and a target detection module. According to the scheme, by introducing semantic priori, optimizing prototype generation and feature interaction, the problems of data imbalance and insufficient new class feature representation in a few-sample scene are solved, improvement of new class detection precision is verified on PASCAL VOC, MS COCO and other data sets, and an effective solution is provided for target detection of sample scarce scenes such as medical images and rare species monitoring.
Owner:CHONGQING UNIV OF TECH

Enhanced image and video object detection using multi-stage paradigm

This disclosure describes systems, methods, and devices related to object detection in images. A device may input an image, representing an object, to a manual labeling learner system; identify, using the system, first coordinates of an upper left corner of a bounding box representing the object based on a heatmap indicative of a probability of the first coordinates representing the upper left corner; identify, using the system, second coordinates of a bottom right corner of the bounding box based on the first coordinates and a first distance regression map indicative of coordinate differences between the second coordinates and ground truth coordinates input to the machine learning model as training data; generate, using the system, adjustments to the first coordinates and the second coordinates based on a second regression map; and generate, using the system, the adjusted first and second coordinates, the bounding box.
Owner:INTEL CORP

Image object detection method, system and apparatus, and storage medium

Embodiments of the present description provide an image object detection method. The method comprises: on the basis of an image to be retrieved, an object description text, and an object retrieval condition, determining, by means of an object detection model, a target position of an object to be retrieved in said image, wherein the object description text is used for describing said object, and the object retrieval condition comprises at least one of a mask image, a pose, and a texture corresponding to said object.
Owner:ZHEJIANG DAHUA TECH CO LTD

Unmanned cross-domain positioning and acoustic fingerprint processing method and system for underwater static target

ActiveCN121899836ASuppress the cumulative drift problemAchieve highly robust target identificationNavigational calculation instrumentsNavigation by speed/acceleration measurementsSonarUncrewed vehicle
The invention relates to the technical field of underwater static target object detection and processing, in particular to an unmanned cross-domain positioning and acoustic fingerprint processing method and system for an underwater static target. Comprising the following steps: performing wide-area scanning on a task sea area, and scheduling an AUV and unmanned aerial vehicle cluster to a target area after discovering a suspicious target; the unmanned aerial vehicle cluster receives an acoustic signal of the AUV and transmits the acoustic signal to the cooperative resolving center in combination with self-positioning data so as to provide accurate coordinate guidance for the AUV; the AUV sails to a target area, parallel processing map construction and accurate positioning, target identification and dual-mode fingerprint generation are carried out, and a target task package is packaged; the ROV receives and analyzes the task packet, autonomously plans a path and sails to a target area, a sonar is started to collect data and generate real-time feature fingerprints, and the real-time feature fingerprints are matched with fingerprints in the task packet; and after matching succeeds, a specific job task is autonomously executed. According to the invention, a complete automatic solution is provided for scenes such as deep and far sea detection, emergency salvage, pipeline maintenance and the like.
Owner:SHANDONG UNIV OF SCI & TECH

Uncertainty estimation for object detection in autonomous and semi-autonomous systems and applications

In various examples, systems and methods for uncertainty estimation for object detection in autonomous and semi-autonomous systems and applications are provided. The systems and methods may use data from one or more sensors (e.g., camera(s) and / or LiDAR sensor(s) to generate a representation of features surrounding a machine. A model may be used to generate probabilities of objects being present in the representation of features and uncertainty estimates corresponding to the object presence probabilities. The uncertainty estimates may be used to identify scenes that are significantly different from the training data, detect errors in the bounding shapes for objects, and / or highlight areas where object detections may have been missed. The systems and methods may also be used to auto-label scenes associated with the representation of features, and the auto-labeled scenes may be used for training purposes.
Owner:NVIDIA CORP

Power system equipment image anomaly detection and quality diagnosis method based on multi-modal visual language model

The invention discloses an electric power system equipment image anomaly detection and quality diagnosis method based on a multi-modal visual language model. The method comprises the following steps: constructing a large-scale multi-modal data set comprising an electrical equipment image, an object detection annotation, a pairing question and answer knowledge base and an official supervision document, constructing a basic diagnosis model based on a visual language model, and carrying out instruction tuning; carrying out post-training on the model by adopting group relative strategy optimized reinforcement learning, and generating an interpretable step-by-step diagnostic reasoning chain; in the reasoning process, related knowledge is dynamically retrieved based on a retrieval enhancement generation technology of a graph structure, and the accuracy and compliance of a diagnosis decision are enhanced; and finally, generating a diagnosis report containing the exception type, the root cause and the decision suggestion. Compared with a traditional method, the method solves the three problems of data scarcity, opaque reasoning and knowledge isolation in the field of electric power detection, and has the remarkable advantages that the diagnosis process can be explained, complex multi-step reasoning is supported, and domain knowledge can be dynamically integrated.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1

Vehicle driving track prediction method and system and electronic equipment

ActiveCN121106349AVehicle drivingData mining
The embodiment of the invention provides a vehicle driving track prediction method and system and electronic equipment. The method comprises the steps that vehicle driving information corresponding to a target vehicle is determined; obtaining aerial view feature information according to the vehicle driving information; determining query feature information, semantic segmentation information and a target object detection result according to the vehicle driving information and the aerial view feature information; determining predefined anchor point information, and obtaining first driving track information corresponding to the target vehicle according to the predefined anchor point information and the query feature information; and obtaining more accurate second driving track information corresponding to the target vehicle according to the aerial view feature information, the query feature information, the semantic segmentation information, the target object detection result and the first driving track information. Therefore, based on a strategy from coarse to fine, the relatively rough first driving track information can be converted into the relatively accurate second driving track information, so that a more accurate vehicle driving track is obtained, and the accuracy of the vehicle driving track is improved.
Owner:NULLMAX INC

Robot vision object semantic understanding and posture generation method based on large model

The invention provides a robot vision object semantic understanding and posture generation method based on a large model, and belongs to the technical field of image processing. Comprising the following steps: S1, inputting an image and a 3D model; s2, object detection; s3, multi-modal feature alignment is carried out; s4, diffusion model sampling; s5, performing geometric screening; s6, carrying out NeRF (New Random Field) morphological modeling; s7, optimizing the optical flow; s8, joint loss calculation; s9, confidence coefficient analysis; and S10, outputting the attitude junction. According to the invention, through an innovative single-view rendering-optical flow optimization closed loop strategy, the calculation overhead is significantly reduced and the estimation precision is improved. According to the method, the algorithm performance in a complex scene is remarkably improved through multi-scale feature fusion and a confidence decomposition strategy.
Owner:GUANGDONG UNIV OF TECH

Dense target detection method based on comparative learning characterization and reinforcement learning decision

The invention provides a dense target detection method based on comparative learning characterization and reinforcement learning decision, which is suitable for dense animal individual detection in a livestock breeding scene. The method comprises the following steps: constructing a dense target detection model comprising a feature encoder, a reinforcement learning decision module and a detection head; the training process is divided into two stages: in the contrast learning characterization stage, an anchor sample, a positive sample and a negative sample are constructed, a first loss function is utilized to optimize a feature encoder, and an optimal feature encoder is obtained; in the reinforcement learning decision-making stage, the feature vector is used as an input state, a decision-making module dynamically selects a threshold lowering, maintaining or improving action, a detection head outputs a detection result in combination with the feature and the action, and parameters of the decision-making module are optimized through a reward function. During detection, after features of a to-be-detected image are extracted through the feature encoder, the optimized decision module selects the optimal action, and finally a detection result is output. The method has relatively high robustness and accuracy.
Owner:XIANGTAN UNIV

Training a model to identify items based on image data and load curve data

A smart shopping cart includes internally facing cameras and an integrated scale to identify objects that are placed in the cart. To avoid unnecessary processing of images that are irrelevant, and thereby save battery life, the cart uses the scale to detect when an object is placed in the cart. The cart obtains images from a cache and sends those to an object detection machine learning model. The cart captures and sends a load curve as input to the trained model for object detection. Labeled load data and labeled image data are used by a model training system to train the machine learning model to identify an item when it is added to the shopping cart. The shopping cart also uses weight data and the image data from a timeframe associated with the addition of the item to the cart as inputs.
Owner:MAPLEBEAR INC

Computer-readable recording medium having stored therein fraud detection program, information processing apparatus, and information processing system

A computer-readable recording medium having stored therein a fraud detection program causing a computer to execute a process including obtaining a result of object detection by inputting a target image group including a self-checkout-apparatus in an imaging range, into a model trained using a target image and an annotation, and performing fraud detection at the self-checkout-apparatus based on information about an item registered thereto and the result. The target image is identified by calculating statistical information of a position of a detection region of an object in each image in a first group based on positions by inputting the first group into the model, obtaining a position in each image in a second group using the model, and identifying the target image in which a region having an appearance probability equal to or less than a threshold is present, from the second group, based on the statistical information.
Owner:FUJITSU LTD

Systems and methods for object tracking

Systems and methods for object tracking are described. One or more aspects of the systems and methods include receiving a video depicting an object; generating object tracking information for the object using a student network, wherein the student network is trained in a second training phase based on a teacher network using an object tracking training set and a knowledge distillation loss that is based on an output of the student network and the teacher network, and wherein the teacher network is trained in a first training phase using an object detection training set that is augmented with object tracking supervision data; and transmitting the object tracking information in response to receiving the video.
Owner:ADOBE INC

Vehicle throwing object detecting and positioning method and system based on multi-channel spatial-temporal feature fusion

The invention provides a vehicle throwing object detection and positioning method and system based on multi-channel spatial-temporal feature fusion, and the method comprises the steps: obtaining a continuous time sequence image frame sequence of a road scene, and generating an optical flow image sequence through an optical flow algorithm; outputting a detection frame of each vehicle in each standardized image through a deep neural network target detection model, and generating a region of interest according to the detection frames; a fusion feature vector is generated for each region of interest, a space-time fusion feature sequence is constructed, and a thrown object classification result and a positioning result of each region of interest are generated based on the space-time fusion feature sequence and the multi-branch full-connection network; and based on the classification result and the positioning result of the thrown object, inputting the obtained coordinates of the suspected area of the thrown object into a spherical camera for tracking. According to the method, the thrown object can be efficiently and accurately detected and positioned automatically, the accuracy and real-time performance of thrown object detection are improved, the false alarm rate is reduced, and the precision degree of positioning the position of the thrown object is improved.
Owner:HANGZHOU URBAN CONSTR & INVESTMENT GRP CO LTD

Multi-degree-of-freedom mechanical arm obstacle avoidance path planning method based on industrial vision

The invention discloses a multi-degree-of-freedom mechanical arm obstacle avoidance path planning method based on industrial vision, and relates to the technical field of intelligent manufacturing, and the method comprises the steps: collecting and preprocessing an initial environment image of a current working scene of a mechanical arm, obtaining a standardized working scene image data set, carrying out dynamic object detection, and obtaining an environment perception evaluation report; identifying an unobserved area of the working scene based on the environment perception evaluation report, performing collision path calculation, and generating a visual angle adjustment action instruction; the safety path sequence is issued to a mechanical arm joint and executed, images in front of the mechanical arm are continuously collected during execution, and if an unpredicted sudden obstacle is detected, the four-dimensional risk map is dynamically updated, and path re-planning is conducted; and when the tail end of the mechanical arm successfully completes the task, obstacle avoidance path planning data of the mechanical arm are recorded and stored persistently, and an obstacle avoidance path planning record is generated. According to the method, the real-time problem of path planning of the mechanical arm is solved through construction of the four-dimensional dynamic risk map.
Owner:KUNSHAN GANYUAN KANGSHENG TECHNOLOGY CO LTD

Fine-tuning-based cross-scenario object detection method and system, device, and medium

The present application relates to a fine-tuning-based cross-scenario object detection method and system, a device, and a medium. The method comprises: improving a backbone network of a YOLOv8 algorithm by means of Transformer modules, and inputting an image in a first scenario into an optimized YOLOv8 model for training to obtain a first object detection model; embedding a LORA model into each Transformer module in the first object detection model, training the model on the basis of an image in a second scenario, fixing parameters of the first object detection model, and training model parameters of each LORA model; and on the basis of the trained model parameters of each LORA model, fine-tuning a weight parameter generated by each Transformer module in the first object detection model, to obtain a second object detection model applicable to the second scenario. The present application enables object detection models to adapt to different actual application scenarios within a short period of time, thereby improving the accuracy and efficiency of cross-scenario object detection and recognition.
Owner:E SURFING VISION TECHNOLOGY CO LTD

Multimedia object tracking and merging

In multimedia object tracking and merging of tracked objects, an object is tracked through frames of multimedia content until a frame appears in which the tracked object is not detected. A first track is designated as one or more consecutive frames in which the tracked object is detected, the first track ending at the first frame. Tracking continues to try to detect the tracked object in a second frame subsequent to the first frame. If the tracked object is not again detected, information about the first track is output. If the tracked object is detected subsequently, a second track of consecutive tracked object detection is designated. The tracked objects in the two tracks are then compared with the aid of trained data models, and a matching score is determined to reflect the degree of match. If the matching score meets or exceeds a first threshold, the compared tracks are merged using the same identifier assigned to both tracks. If the matching score does not exceed a second threshold that is less than the first threshold, the tracks may be discarded as showing no match. If the matching score falls between the first and second thresholds, an indication is output for further analysis of the compared tracked objects.
Owner:GETAC TECH CORP +1

Three-dimensional chat thread visualization and interaction in augmented reality

A system and method for contextual three-dimensional messaging in augmented reality (AR) environments is disclosed. The system receives chat messages with specified real-world destinations and stores them associated with those locations. When a user wearing an AR device enters a destination location, the system detects their presence using techniques like GPS, Wi-Fi positioning, or computer vision. It then generates a 3D visual representation of the message and determines an appropriate spatial position within the physical environment based on environmental analysis and object detection. The 3D message is displayed at the determined position in the AR view. The system can analyze message content to identify topics and match them to detected real-world objects for contextual placement. Users can interact with displayed messages through gestures or voice commands to reply, forward, delete, or reposition messages. This enables immersive, location-aware messaging experiences that seamlessly blend digital content with the physical world.
Owner:SNAP INC

Object detection using visual language models via latent feature adaptation with synthetic data

Systems and techniques are described herein for adapting a pretrained machine learning model. For instance, a process can include encoding a training image into a first feature vector, the training image including a first object located at a first location; generating a second feature vector based on a set of sinusoidal functions using a set of weights; combining the first feature vector with a second feature vector to generate a combined feature vector; processing the combined feature vector using a visual language model to obtain a second location for the first object; and adjusting the set of weights based on a comparison between the first location and the second location.
Owner:QUALCOMM TECHNOLOGIES INC