Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1224 results about "Ground truth" patented technology

Ground truth is a term used in various fields to refer to information provided by direct observation (i.e. empirical evidence) as opposed to information provided by inference.

System and method for causality-augmented generative intelligence to discover non-obvious insights from heterogeneous data sources

The present invention provides a system and method for causality-augmented generative intelligence capable of autonomously discovering non-obvious actionable insights from heterogeneous and multimodal data sources. The system integrates a data ingestion unit for semantic and temporal harmonization of structured and unstructured datasets, a causal inference processor for constructing a dynamically evolving directed causal knowledge representation using perturbation-based validation, a latent representation processor that combines multimodal semantic embeddings with causal parameters to generate fused latent vectors, and a generative insight processor utilizing causally constrained generative reasoning to synthesize hypotheses anchored to verified cause-effect dependencies. A validation processor performs counterfactual assessment and observational verification to ensure retention of only those insights that remain consistent with causal ground truth.
Owner:MIA MD TOFAYEL GONEE MANIK

Method and system for cross-domain predictive modeling using bedrock based foundation models and blockchain-anchored data

The present invention relates to a system and method for cross-domain predictive modeling using Bedrock-based foundation models and blockchain-anchored data. The invention integrates large-scale foundation model reasoning with distributed ledger-based data provenance to enable verifiable, secure, and explainable predictive analytics across heterogeneous domains such as finance, healthcare, logistics, and environmental systems. The system comprises a data ingestion unit for receiving and normalizing multi-domain datasets, a blockchain anchoring unit for generating cryptographic hashes and recording data provenance into a distributed ledger, a cross-domain harmonization processor for aligning heterogeneous feature representations into a unified latent space, a foundation model processor configured to execute Bedrock-based predictive inference with adaptive domain contextualization, a verification processor for validating predictions against blockchain-anchored ground truths, and a governance processor for maintaining immutable audit trails of model evolution.
Owner:VAYYASI NAVEEN KUMAR

Communication device and method for determining channel state information report based on artificial intelligence / machine learning

Communication devices and methods for determining channel state information (CSI) report based on artificial intelligence (AI) / machine learning (ML) are provided. The method for determining CSI report based on AI / MI performed by a communication device includes determining, by the communication device, one or more CSI reports according to an AI / ML based CSI feedback, wherein each of the one or more CSI reports contains an output of an auto-encoder, a compression ratio, a rank indicator, quantization levels, a ground truth of an enhanced CSI feedback, and / or an ML model monitoring outcome, and determining, by the communication device, priority rules for the CSI reports according to the AI / ML based CSI feedback.
Owner:SHENZHEN TCL NEW-TECH CO LTD

Method and system for performing end-to-end evaluation of a large language model (LLM)

A method and a Large Language Model (LLM) evaluation system provides an end-to-end evaluation of LLM, which includes evaluating both input prompts and output prompt responses, wherein the evaluation includes assessing a plurality of input and output characteristics that encompasses both quality and quantity. Each of the plurality of input characteristics are assigned with a corresponding normalized score by employing one or more statistical techniques to derive a composite health score for the input prompts. Evaluation further comprises evaluating output prompt responses in both absence and presence of the ground truth. Upon evaluating both input prompts and output prompt responses, a final aggregated health score for the LLM is computed by a scorer module employing threshold based statistical techniques that considers input prompt health and output prompt response health, wherein the aggregated health score is generated based on the granular scores of each characteristic.
Owner:LTIMINDTREE LTD

Methods for Automatically Generating a Training Dataset for Training an Optical Recognition Model for Reading Street Signs

Various embodiments include methods for generating image datasets for training an artificial intelligence machine learning (AI / ML) optical character recognition (OCR) model. Image processing may be performed on a plurality of roadway images to identify street signs within the images and generate a dataset of sign images categorized into sign variants of the same shape, color, pictogram, and characters. An OCR model may process sign images to obtain OCR results for images of each sign variant. An aggregation process may be performed on the OCR results for all sign images within each sign variant to identify a ground truth OCR result for each sign variant. The ground truth OCR result may be used to automatically label all sign images of each sign variant to produce an OCR model training dataset. The produced training dataset may then be used to retrain the initial AI / ML OCR model and / or train other AI / ML OCR models.
Owner:QUALCOMM INC

Distillation-trained machine learning models for efficient trajectory prediction

The described aspects and implementations enable training and deploying of accurate one-shot models capable of predicting trajectories of vehicles and other objects in driving environments. The disclosed techniques include, in one implementation, obtaining training data that includes a training input representative of a driving environment of a vehicle and one or more ground truth trajectories associated with a forecasted motion of the vehicle within the driving environment. The one or more ground truth trajectories are generated by a teacher model using the training input. The techniques further include training, using the training data, a student model to predict one or more trajectories of the vehicle and / or objects in the driving environment of the vehicle.
Owner:WAYMO LLC

Information retrieval in machine learning question answering systems

Evaluating and improving information retrieval in question-answering systems is an area of importance in machine learning growth. Retrieval components in a retrieval-augmented generation (RAG) question answering system enable machine learning models to provide more accurate and reliable answers to questions. Systems for retriever evaluation involve processing queries in comparison to reference documents. The system first retrieves documents deemed relevant, then generates a first answer based on them. A second answer is generated using a set of documents that includes ground truth documents known to be relevant to the query. By analyzing semantic overlap between these responses, a quantitative evaluation of the retrieval component is obtained. This evaluation then informs automatic modifications to retrieval parameters, enhancing future document selection and response accuracy.
Owner:THOMSON REUTERS ENTERPRISE CENTRE GMBH

Circadian rhythm-based data augmentation for occupant state analysis

In various examples, circadian rhythm-based data augmentation for drowsiness detection systems and applications are provided. Embodiments described herein may produce an estimated circadian rhythm for a test subject and / or vehicle driver or other machine operator or occupant, and use the pattern of that circadian rhythm to correct, confirm, calibrate, or otherwise augment drowsiness assessments derived from video image data. The position of a person in the context of their process C circadian cycle may be used as indication of their level of drowsiness. An estimated process C circadian cycle may be used to generate more accurate ground truth training data for training machine learning models, and may be used by real-time, in-vehicle drowsiness detection systems that infer driver drowsiness levels based on captured images. In various embodiments, a circadian rhythm drowsiness estimate may be used to correct, calibrate, augment, and / or replace a drowsiness score predicted by a machine learning model.
Owner:NVIDIA CORP

Multi-modality reinforcement learning in logic-rich scene generation

Generating high-quality images of logic-rich three-dimensional (3D) scenes from natural language text prompts is challenging, because the task involves complex reasoning and spatial understanding. A reinforcement learning framework utilizing a ground truth data set can be implemented to train a policy network. The policy network can learn optimal parameters to refine a text prompt to obtain a modified text prompt. The modified text prompt can be used to obtain a three-dimensional scene, and the three-dimensional scene can be rendered and projected to obtain a rendered image. The framework involves an action agent for text modification, a generation agent to produce rendered images, and a reward agent to evaluate the rendered images. The loss function used in training the policy network optimizes visual accuracy and quality of the rendered images and semantic alignment between the rendered images and the text prompt.
Owner:INTEL CORP

Automatic test performance benchmark construction method and device and storage medium

The invention relates to the field of industrial visual inspection, and provides an automatic test performance benchmark construction method which comprises the following steps: dynamically generating ROI (Region of Interest) parameters; performing pixel cutting on the target area of the production line according to the dynamically generated ROI parameters to generate a high-speed image data stream, and aligning the image acquisition moment with the mechanical motion phase of the production line; selecting a matched visual detection model based on target features in the high-speed image data stream, and outputting a defect detection result by the matched visual detection model; according to defect distribution parameters of a preset test scene, the simulation defect features are embedded into the high-speed image data stream, and a mixed test data stream is generated; performing multi-device time sequence alignment on the mixed test data stream and the defect detection result, and packaging into a standardized data set containing a unified time reference; and analyzing the standardized data set, comparing the simulated defect truth value label with a defect detection result, calculating a multi-dimensional performance index of the visual inspection system, and generating a test report containing parameter optimization suggestions.
Owner:SHENZHEN GEYUAN TECH CO LTD

Hazard detection in autonomous and semi-autonomous systems and applications

Embodiments relate to hazard detection in autonomous and semi-autonomous systems and applications. A transformer may use sampled image and LiDAR features to extract and decode a representation of whether there is a hazard at the 3D location corresponding to each initial transformer query, the shape of the hazard, and / or its class. These detections may be provided to one or more control components of an autonomous vehicle, which may use the detections to navigate, plan, or otherwise perform one or more operations (e.g., obstacle avoidance, lane keeping, lane changing, merging, splitting, etc.). Some embodiments employ an automated approach to derive ground truth data from sensor data collected by data collection vehicle(s), such as data representing detected static scene points, navigable space boundaries, or detected hazard objects. Accordingly, hazards such as road debris and other obstacles may be detected and ground truth data may be generated for a variety of sensing tasks.
Owner:NVIDIA CORP

Three-dimensional (3D) head pose prediction for automotive systems and applications

In various examples, head pose prediction for automotive occupant sensing systems and applications is presented. The systems and methods described herein provide for a machine learning model trained using a dataset that comprises ground truth head pose data computed using a registered head model of a training subject. While operating a vehicle, one or more cameras and a depth sensor capture synchronized images of the training subject. To compute a ground truth 3D head pose, angular deviations between a 3D point cloud and the registered head model may be computed to obtain a 3D ground truth head pose measurement. Using an extrinsic calibration transform, the head pose measurement may be mapped into the sensor coordinate frame. Training samples may be produced for training the machine learning model that comprise an optical image frame and the head pose measurement transposed into the frame of reference for that optical image frame.
Owner:NVIDIA CORP

Optical guidance SAR (Synthetic Aperture Radar) target detection method based on frequency domain enhancement and dynamic mask

The invention provides an optically guided SAR target detection method based on frequency domain enhancement and dynamic masks, which comprises the following steps: acquiring an SAR image and a corresponding optical image, and taking a pre-trained optical detection model as a teacher model and a to-be-trained SAR model as a student model; optical and SAR images are respectively input into corresponding models to generate feature maps, SAR features are decomposed into low-frequency global and high-frequency detail components through wavelet transform, and noise is suppressed and target features are enhanced through multi-scale convolution and a self-attention mechanism; generating a target area mask through a dynamic mask module based on the ground truth value; inputting the enhanced SAR features and the optical features into an optical detection head of a teacher model, and calculating loss by using a cross-detection-head distillation strategy; and through combination of detection loss and distillation loss, the student model is subjected to back propagation training until convergence, so that the SAR target detection precision and real-time performance in a complex scene can be effectively improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Uncertainty decomposition for in-context learning of large language models

Methods and systems for prompting a Large Language Model (LLM) with a set of text data outside pre-inference trained categories and a test prompt for an initial parameter which has a known ground truth, calculating an uncertainty of an LLM's output, selecting another LLM model parameter and calculating the total uncertainty of the LLM's output with the other LLM model parameter. The methods and systems further include prompting the LLM with another test prompt, with the initial LLM parameter and the other LLM parameter, and calculating the total uncertainty of the LLM's output for initial LLM model parameter and the other LLM model parameter, decomposing the total uncertainty of the LLM into Aleatoric Uncertainty (AU) and Epistemic Uncertainty (EU) components, and rating the total uncertainty of the LLM, using the decomposed total uncertainty as a metric.
Owner:NEC LABORATORIES AMERICA INC

System and method for embedding uncertainty estimation into deep-neural-network-based autonomous driving perception frameworks

Embodiments of this disclosure can provide a system and method for training a perception model to perform an autonomous driving task. During operation, the system can obtain labeled training data comprising images captured by multiple cameras mounted at different locations on a vehicle, and the perception model can generate, in parallel, a prediction output associated with the task and a confidence score based on the labeled training data. The confidence score can indicate a level of uncertainty associated with the prediction output. The system can generate an uncertainty-weighted prediction based on ground truth indicated by the labeled training data, the prediction output, and the confidence score; compute a loss function based on the uncertainty-weighted prediction; and update the perception model based on the loss function.
Owner:BLACK SESAME TECH INC

Weakly-supervised referring expression segmentation

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modifies parameters of a fused feature extractor. In particular, the disclosed systems generate inferred masks from digital images and digital text prompts using a fused feature extractor. Furthermore, the disclosed systems identify a subset of the inferred masks that satisfy a validity threshold. Moreover, the disclosed systems generate an augmented training set by combining the subset of the inferred masks with a training set that includes the ground truth masks. Further, the disclosed systems generate object mask predictions from the augmented training set and determine ground truth and pseudo measures of loss by comparing the object mask predictions with the inferred masks and the ground truth masks. From the ground truth and pseudo measures of loss, the disclosed systems modify parameters of the fused feature extractors.
Owner:ADOBE INC

Blood vessel continuous segmentation method based on graph network

The invention discloses a blood vessel continuous segmentation method based on a graph network, and relates to the technical field of blood vessel continuous segmentation, multi-scale texture features based on coronary artery influence and topological structure features of blood vessels are fused, and correlation among different features is enhanced through an attention mechanism; segmenting the fused multi-scale texture features influenced by the coronary artery and topological structure features of the blood vessel, extracting multi-scale features, and reducing the resolution; gradually recovering the resolution by using a decoder to obtain a segmentation result, and optimizing the network weight by using an error between the segmentation result and a real label; and applying the trained segmentation network to test data to obtain a three-dimensional segmentation result of the coronary artery, and evaluating the accuracy of the segmentation result by comparing the difference between a predicted value and a true value to obtain connection constraint loss. The problem that a three-dimensional blood vessel structure is difficult to extract by a general medical segmentation model is solved, so that a blood vessel segmentation result is more continuous and more accurate.
Owner:FUDAN UNIVERSITY

AI assisted calibration of IMUS

Provided herein are methods for Inertial Measurement Unit (IMU) sensor compensation utilizing deep neural network (DNN) technology. In some embodiments, the methods involve receiving sensor values from gyroscopes, accelerometers, and non-motion sensors, then loading these values into a deep learning algorithm alongside true sensor values. Unlike traditional calibration techniques that rely on complex parametric models, this approach may utilize a neural network to directly process raw sensor data and enhance output signal accuracy. In some embodiments, the training process employs ground truth data for approximately 70% of inputs, with the remaining 30% dedicated to sensor compensation. The technique may be implemented in standalone IMUs, augmented IMUs, and navigation systems, offering improved accuracy and reduced computational overhead across various technological applications.
Owner:AIM DESIGN LLC

Systems and methods for automated inspection of vehicles for body damage

There is provided a method of automatically detecting that a target image is deepfake, comprising: receiving authentic images depicting a vehicle with actual damage, receiving the target image depicting potential damage to the vehicle, feeding the target image into a machine learning (ML) model, obtaining a candidate set of human-readable text describing the potential damage to the vehicle, feeding the authentic images into the ML model, obtaining from the ML model, a ground truth set of human-readable text describing the actual damage to the vehicle depicted in the authentic images, computing a similarity metric indicating a difference between the potential damage described in the candidate set of human-readable text and the actual damage described in the ground truth set of human-readable text, and in response to the difference being above a threshold or meeting a requirement indicating a significant difference, detecting that the target image is likely deepfake.
Owner:UVEYE LTD

Visible light-thermal infrared image semantic segmentation method and system driven by plug-and-play prompt

The invention relates to the technical field of multi-modal image fusion perception and scene understanding, in particular to a plug-and-play prompt-driven visible light-thermal infrared image semantic segmentation method and system, and the method comprises the steps: respectively extracting the feature representation of an input visible light image and a thermal infrared image through a dual-branch LoRA fine-tuning image encoder; converting a segmentation mask generated by the existing visible light-thermal infrared image semantic segmentation model into unified prompt information, including bounding box prompt or point prompt; using a prompt encoder to encode prompt information into prompt embedding; the prompt-based mask decoder fuses the image feature representation and the prompt embedded feature through a spatial channel cross attention mechanism, and generates a final segmentation mask through a classification head; in the training stage, parameterized random disturbance is added to the prompt of the true value mask, and error distribution of the prediction prompt is simulated. According to the method, high-precision semantic segmentation adaptive to various segmentation models can be realized without retraining, and the segmentation robustness in a complex scene is remarkably improved.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Depth completion method based on geometric perception and channel attention mechanism

The invention discloses a depth completion method based on geometric perception and a channel attention mechanism, and relates to the technical field of computer vision and deep learning. Comprising the following steps: generating an initial dense depth map by using a nonlinear propagation model; pointNet + + is adopted to extract global 3D geometric features; performing preliminary fusion on the initial dense depth map and the image features by using U-Net; performing weighted fusion on the global 3D geometric features and the preliminary fusion features through a multi-modal fusion module based on a channel attention mechanism to generate optimized fusion features; residual learning is carried out by using the fusion features to correct the initial dense depth map; and optimizing the complementation result by adopting CSPN + + in combination with the original sparse truth value in the sparse depth map. According to the method, the GAC-Net structure is integrally optimized, the scene geometric perception capability is improved, the 3D global features are fully utilized, and the adaptability to the complex environment is improved.
Owner:ZHEJIANG UNIV OF SCI & TECH

Neural network training using ground truth data augmented with map information for autonomous machine applications

In various examples, training sensor data generated by one or more sensors of autonomous machines may be localized to high definition (HD) map data to augment and / or generate ground truth data—e.g., automatically, in embodiments. The ground truth data may be associated with the training sensor data for training one or more deep neural networks (DNNs) to compute outputs corresponding to autonomous machine operations—such as object or feature detection, road feature detection and classification, wait condition identification and classification, etc. As a result, the HD map data may be leveraged during training such that the DNNs—in deployment—may aid autonomous machines in navigating environments safely without relying on HD map data to do so.
Owner:NVIDIA CORP

Quality matrix for evaluating ai agent performance

A variety of metrics are described for evaluating the performance of artificial intelligence agents, e.g., in the context of user requests and generative model responses within a specific domain, such as physiological monitoring or associated health and wellness coaching, that provides a ground truth for responses to requests. These metrics may be used, e.g., to determine whether and how to deliver responses to a user, as well as for evaluating the performance of underlying generative models, agents, and so forth. In another aspect, a quality matrix may be provided for an agent that compares expected to actual behavior for different classes of user requests.
Owner:WHOOP INC

Generation of synthetic data for image registration training

A method includes generating, using at least one processing device, a ground truth optical flow map for displacement of pixels within a reference image based on a motion model. The motion model determines 3D coordinates of pixels within the reference image based on estimated depths of the pixels within the reference image. The method also includes performing, using the at least one processing device, 3D to 2D reprojection of the pixels within the reference image based on the motion model to generate a reprojected image view corresponding to a shifted camera perspective. The method further includes generating, using the at least one processing device, an occlusion mask for the reference image. The occlusion mask corresponds to occluded pixels within the reprojected image view. In addition, the method includes performing, using the at least one processing device, occlusion region inpainting of the occluded pixels to generate an inpainted reprojected image view.
Owner:SAMSUNG ELECTRONICS CO LTD

Enhanced image and video object detection using multi-stage paradigm

This disclosure describes systems, methods, and devices related to object detection in images. A device may input an image, representing an object, to a manual labeling learner system; identify, using the system, first coordinates of an upper left corner of a bounding box representing the object based on a heatmap indicative of a probability of the first coordinates representing the upper left corner; identify, using the system, second coordinates of a bottom right corner of the bounding box based on the first coordinates and a first distance regression map indicative of coordinate differences between the second coordinates and ground truth coordinates input to the machine learning model as training data; generate, using the system, adjustments to the first coordinates and the second coordinates based on a second regression map; and generate, using the system, the adjusted first and second coordinates, the bounding box.
Owner:INTEL CORP

Circadian rhythm-based training data correction for drowsiness detection systems and applications

In various examples, circadian rhythm-based data augmentation for drowsiness detection systems and applications are provided. Embodiments described herein may produce an estimated circadian rhythm for a test subject and / or vehicle driver or other machine operator or occupant, and use the pattern of that circadian rhythm to correct, confirm, calibrate, or otherwise augment drowsiness assessments derived from video image data. The position of a person in the context of their process C circadian cycle may be used as indication of their level of drowsiness. An estimated process C circadian cycle may be used to generate more accurate ground truth training data for training machine learning models, and may be used by real-time, in-vehicle drowsiness detection systems that infer driver drowsiness levels based on captured images. In various embodiments, a circadian rhythm drowsiness estimate may be used to correct, calibrate, augment, and / or replace a drowsiness score predicted by a machine learning model.
Owner:NVIDIA CORP

View synthesis using camera poses learned from a video

View synthesis is a computer graphics process that generates a new image of a scene from a novel (previously unseen) viewpoint of the scene. Typically, the graphics process relies on a machine learning model that has been trained with ground truth pose information. Since ground truth pose information is not readily available, some solutions rely on a Structure-from-Motion (SfM) library COLMAP to generate pose information for a given image. However, this pre-processing step is not only time-consuming but also can fail due to its sensitivity to feature extraction errors and difficulties in handling texture-less or repetitive regions. The present disclosure provides view synthesis from learned camera poses without relying on SfM pre-processing.
Owner:NVIDIA CORP

Depth estimation using active sensing and structured light

A depth map for a scene may be generated by projecting, by a light projector, an illumination pattern onto a scene; capturing, by a camera, a camera image of the scene; generating, by a computing device, a plurality of ground truth depth values for sample pixels of the camera image based at least in part on the illumination pattern; and estimating a depth map for the scene based at least in part on the camera image and the ground truth depth values for sample pixels.
Owner:QUALCOMM INC

Voice attribute conversion using speech to speech

There is provided a computer-implemented method of training a speech-to-speech (S2S) machine learning (ML) model for adapting voice attribute(s) of speech, comprising: creating an S2S training dataset of S2S records, wherein an S2S record comprises: a first audio content comprising speech having first voice attribute(s), and a ground truth label of a second audio content comprising speech having second voice attribute(s), wherein the first audio content and the second audio content have the same lexical content and are time-synchronized, wherein duration of phones of the second audio content are controlled in response to segment-level durations defined by the segment-level start and end time stamps and training the S2S ML model using the S2S training dataset, wherein the S2S ML model is fed an input of a source audio content with source voice attribute(s) and generates an outcome of the source audio content with target voice attribute(s).
Owner:MEANING TEAM INC

Lightweight change detection system on low-resolution video stream

Systems and methods are provided for change detection in low-resolution video streams, which can be used for applications such as high resolution video restoration and processing. The techniques effectively detect changes by leveraging a large receptive field and lightweight computation, which are achieved by working with low-resolution images. In particular, the techniques include extracting features from a change detection model and a semantic segmentation model, and integrating the extracted feature outputs from the models to produce a robust change detection map. A pre-processing phase can be employed to optimize the input for each model, ensuring minimal complexity and enhanced performance. The change detection model can be implemented as a deep neural network, and methods are provided for generating ground truth (GT) data, which semantically guides the change detection neural network to perform change detection inpainting during training.
Owner:INTEL CORP