Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

167 results about "Synthetic data generation" patented technology

Synthetic data generation has become a surrogate technique for tackling the problem of bulk data needed in training deep learning algorithms. Areas such as computer vision have greatly benefited from advances in deep learning and now generating synthetic data is serving as a good starting point for researchers who are trying to bridge the data gap.

AI-based system and method for automated API discovery and action workflow generation

A system and a method for automatically discovering and managing actions in an application is disclosed. The system includes a data ingestion layer for receiving application data from multiple sources, a scanning and systematic traversal engine for interacting with UI elements and capturing network calls, an action mapping and generation module for correlating UI actions with API calls and categorizing actions, an AI-driven icon and description generator for creating visual representations and textual descriptions of actions, a user interface for displaying and modifying discovered actions, and a continuous monitoring component for triggering re-scanning based on coverage metrics, error detection, or version updates. The system employs synthetic data generation and AI-driven exploration to uncover hidden or undocumented APIs, enabling comprehensive mapping of an application's capabilities at the API level.
Owner:ADOPT AI INC

Synthetic data generation utilizing generative artifical intelligence and scalable data generation tools

Systems, methods, devices, and computer readable storage media described herein provide techniques for generating synthetic data utilizing generative artificial intelligence (AI) models and scalable data generation tools (SDGTs). A prompt comprising a domain is provided to an AI model. A parameter associated with the domain that specifies a boundary for synthetic values in a column of data is received from the AI model. An argument comprising the parameter is provided to an SDGT to generate scaled data based on the parameter. The scaled data comprises a column of synthetic values wherein each synthetic value is within the boundary specified by the parameter. In an aspect, an emulated workload is caused to utilize synthetic data comprising the scaled data to generate a performance benchmark. In a further aspect, a lightweight model is trained to generate synthetic sentences based on training data generated by the generative AI model based on the prompt.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Determining lighting and composition parameters using machine learning models for synthetic data generation

Approaches presented herein provide for the determination of realistic lighting parameters for a scene represented in an image. Realistic lighting parameters can allow for the insertion of one or more virtual objects into a scene image, where the lighting or shading applied to the virtual object(s) can be consistent with those for other objects in the scene. A machine learning model such as a discriminator or diffusion model can be used to analyze a composed image generated by a differential renderer, for example, in which at least one virtual object has been inserted into a scene image and had lighting effects applied in accordance with a set of lighting parameters. A loss value can be determined based on the results of this machine learning model, which can be used to optimize the lighting parameters and / or adjust the weights or parameters of a model used to generate the lighting parameters. Once fine-tuned or optimized, the lighting parameters can represent an accurate light map for the scene or environment that can be used to generate composed images.
Owner:NVIDIA CORP

Anti-fact fair synthesis data generation method and device based on causal reasoning

The invention provides an anti-fact fair data synthesis method and device based on causal reasoning, and aims to generate high-quality synthesis data meeting the fairness requirement by mining the causal relationship between observable features. The synthesis method comprises the following steps: extracting observable features, sensitive features and labels from original data, extracting potential features through a variational automatic codec, and constructing a causal relationship graph; designing a generator according to a topological sequence of the causal relationship graph, connecting a causal path, inputting the potential features and the related features into the generator in sequence, and constructing a data generation process conforming to a causal structure; introducing a discriminator to carry out adversarial training on a generation result and original data, and optimizing generator parameter distribution; finally, synthetic data meeting fairness requirements are generated. According to the method, effective regulation and control on the influence of sensitive characteristics and strict constraint on a causal structure are realized, the generated data has higher fairness and interpretability, and the method can be applied to the fields with higher fairness requirements, such as finance, medical treatment and education.
Owner:JINAN UNIVERSITY

Generation of synthetic data for image registration training

A method includes generating, using at least one processing device, a ground truth optical flow map for displacement of pixels within a reference image based on a motion model. The motion model determines 3D coordinates of pixels within the reference image based on estimated depths of the pixels within the reference image. The method also includes performing, using the at least one processing device, 3D to 2D reprojection of the pixels within the reference image based on the motion model to generate a reprojected image view corresponding to a shifted camera perspective. The method further includes generating, using the at least one processing device, an occlusion mask for the reference image. The occlusion mask corresponds to occluded pixels within the reprojected image view. In addition, the method includes performing, using the at least one processing device, occlusion region inpainting of the occluded pixels to generate an inpainted reprojected image view.
Owner:SAMSUNG ELECTRONICS CO LTD

Adaptive optimization techniques for accelerated battery charging protocols

Disclosed are methods and systems for implementing optimization techniques for accelerated battery charging protocols. A state estimator model configured as a synthetic data generator generates a synthetic dataset from an input data. The synthetic dataset comprises one or more states of a battery. A surrogate model generates a set of features comprising exotic parameters representing internal attributes of the battery by applying time-series analysis to the synthetic dataset. A reduced order ML model determines a minimum charging time based on an optimization control parameter and a subset of features derived from the set of features while maintaining the optimization control parameter in an acceptable range. One or more charging protocols are generated based on the minimum charging time and the optimization control parameter. In some embodiments, an optimal charging protocol is further selected from the one or more charging protocols.
Owner:FACTORIAL INC

Bayesian graph-based retrieval-augmented generation with synthetic feedback loop (BG-rag-SFL)

An advanced AI system, known as Bayesian Graph-Based Retrieval-Augmented Generation with Synthetic Feedback Loop (BG-RAG-SFL), combines Bayesian evaluation, graph-based retrieval, and synthetic data feedback to create a continuously improving AI platform. The present invention integrates multiple LLMs, optimizing their performance while managing complexities across different models. Key features include a knowledge graph-based RAG system, a Bayesian evaluation network, a secondary ground-truth graph for verification, synthetic data generation for ongoing improvement, and a multi-agent verification system. The system also functions as an AI operating system capable of acting as a virtual user with screen I / O control and managing multiple computers as an intelligent process automation system.
Owner:ZON GLOBAL IP INC

Synthetic data generation for modality-agnostic zero-shot foundation model for medical images

One or more systems, devices, computer program products and / or computer-implemented methods of use provided herein relate to assessing certainty of artificial intelligence models used for detection or segmentation of pathologies. Accordingly, a system can comprise a memory that can store computer executable components. The system can further comprise a processor that can execute at least one of the computer executable components. The computer executable components can comprise a synthetic data generation component that generates biologically-inspired synthetic data that approximates a task-specific data manifold of a medical image from a radiomic features perspective; an artificial intelligence component that uses an artificial intelligence model to learn relevant representations of the synthetic data for an at least one image task; and a training component that utilizes the relevant representations and the artificial intelligence model to generate a task-specific model for the at least one image analysis task.
Owner:GE PRECISION HEALTHCARE LLC

Synthetic data generation for retrieval evaluation and fine-tuning

In various examples, a technique for generating synthetic data includes inputting a first prompt that includes (i) a first portion of content and (ii) a plurality of user personas into a first machine learning model and generating, via the first machine learning model and based on the first prompt, a plurality of points of interest associated with the user personas and the first portion of content. The technique also includes inputting a second prompt that includes mappings between the points of interest and a plurality of question types into a second machine learning model and generating, via the second machine learning model and based on the second prompt, a plurality of questions associated with the user personas and the first portion of content. The technique further includes retrieving a second portion of content based at least on the plurality of questions and a third prompt.
Owner:NVIDIA CORP

Bayesian graph-based retrieval-augmented generation with synthetic feedback loop (BG-RAG-SFL)

An advanced AI system, known as Bayesian Graph-Based Retrieval-Augmented Generation with Synthetic Feedback Loop (BG-RAG-SFL), combines Bayesian evaluation, graph-based retrieval, and synthetic data feedback to create a continuously improving AI platform. The present invention integrates multiple LLMs, optimizing their performance while managing complexities across different models. Key features include a knowledge graph-based RAG system, a Bayesian evaluation network, a secondary ground-truth graph for verification, synthetic data generation for ongoing improvement, and a multi-agent verification system. The system also functions as an AI operating system capable of acting as a virtual user with screen I / O control and managing multiple computers as an intelligent process automation system.
Owner:ZON GLOBAL IP INC

Synthetic data generation method for training artificial intelligence model and client apparatus

Proposed is a method of generating synthetic data, which includes receiving, by a client apparatus, original data including personal information, acquiring, by the client apparatus, seed data based on the original data, transmitting, by the client apparatus, the seed data to a server, receiving, by the client apparatus, first candidate synthetic data which is generated based on the seed data from the server, validating, by the client apparatus, the first candidate synthetic data based on a similarity between the first candidate synthetic data and the original data, and storing, by the client apparatus, the first candidate synthetic data as a member of candidate synthetic dataset if the first candidate synthetic data is valid synthetic data from the validation result.
Owner:CUBIG CORP

Systems and methods for generating and utilizing synthetic data

Systems and methods for generating and utilizing synthetic data include receiving a set of real network traffic data; generating synthetic data from the received set of real network traffic data based on patterns learned from the set of real network traffic data; and utilizing the synthetic data for any of training a machine learning model, testing a machine learning model, and configuring a customer cloud environment. The systems are adapted to generate a large amount of synthetic data from a limited set of real network traffic data. The produced synthetic data is altered in one or more ways to anonymize sensitive information present in the real data. Therefore, the systems are adapted to generate a large amount of synthetic data which accurately resembles real network traffic data while complying with data privacy practices.
Owner:ZSCALER INC

Generating synthetic training data including document images with key-value pairs

Automated techniques are for generating a large volume of diverse training data that can be used for training machine learning models to extract KV pairs from document images. Given a single input document image and associated annotation data, a large number of diverse synthetic training datapoints are automatically generated by a synthetic data generation system, each datapoint including a synthetic document image and associated annotation data. The generated synthetic training datapoints can be used to train and improve the performance of ML models for extracting KV pairs from document images. In certain implementations, multiple synthetic datapoints are generated by varying the values associated with a key for a content item within the input document image.
Owner:ORACLE INT CORP

Machine-learned architecture for structured synthetic data generation

Techniques may generate realistic synthetic data by programmatically generating a configuration file object type and relationship data. This configuration file may be used to retrieve source data matching the object type(s) and / or specific records indicated by the configuration file. The techniques may detect and anonymize private / proprietary information and may determine statistical characteristic(s) of the source data. A batch of prompt(s) may be generated using the source data, the statistical characteristic(s), and the configuration file and may be transmitted to one or more instances of a transformer-based machine-learned model. Sets of synthetic data received from the model instance(s) may be de-duplicated, checked for similarity to the source data (e.g., via embedding the synthetic data and the source data), and may be used to generate synthetic object(s) using the relationship(s) and / or other data indicated by the configuration file. These synthetic object(s) may then be deployed in a software environment.
Owner:SALESFORCE INC

Synthesizing high resolution 3D shapes from lower resolution representations for synthetic data generation systems and applications

In various examples, a deep three-dimensional (3D) conditional generative model is implemented that can synthesize high resolution 3D shapes using simple guides—such as coarse voxels, point clouds, etc.—by marrying implicit and explicit 3D representations into a hybrid 3D representation. The present approach may directly optimize for the reconstructed surface, allowing for the synthesis of finer geometric details with fewer artifacts. The systems and methods described herein may use a deformable tetrahedral grid that encodes a discretized signed distance function (SDF) and a differentiable marching tetrahedral layer that converts the implicit SDF representation to an explicit surface mesh representation. This combination allows joint optimization of the surface geometry and topology as well as generation of the hierarchy of subdivisions using reconstruction and adversarial losses defined explicitly on the surface mesh.
Owner:NVIDIA CORP

Competitive intelligence analysis system and method based on differential privacy

ActiveCN121563600ADigital data protectionCommerceCompetitive intelligenceMarket dynamics
The invention relates to a differential privacy-based competitive intelligence analysis system and method. The differential privacy-based competitive intelligence analysis system comprises a differential privacy processing and budget management module, a differential privacy feature engineering module, a privacy protection analysis and prediction module and a differential privacy synthetic data generation module. According to the design of the invention, an end-to-end advanced analysis framework with mathematics provable privacy protection capability is constructed, so that accurate, timely and reliable insight of market dynamics and competitor strategies is realized on the premise of strictly protecting enterprise sensitive commercial confidentials.
Owner:HUIZHOU UNIV +2

Breathing lung sound auxiliary identification method and system for clinical nursing

The invention relates to the technical field of respiratory lung sound recognition, in particular to a respiratory lung sound auxiliary recognition method and system for clinical nursing. The invention provides a respiratory lung sound auxiliary identification method and system for clinical nursing, solves the technical problems of data scarcity, complex noise interference and insufficient cross-device generalization ability in lung sound signals through generative data enhancement and cross-modal migration, and combines self-supervised comparative learning and causal feature discovery to identify the respiratory lung sound. Multi-source signal feature fusion and noise robustness improvement are realized, a real-time noise environment is dynamically adapted, the limitation that a traditional method depends on annotation data and feature extraction is easily influenced by hybrid noise is broken through, a full-link closed-loop system from synthetic data generation, cross-modal alignment to causal-driven decision is constructed, and multi-source signal feature fusion and noise robustness improvement are realized. And the accuracy and clinical applicability of clinical lung sound analysis are effectively improved.
Owner:GENERAL HOSPITAL OF THE NORTHERN WAR ZONE OF THE CHINESE PEOPLES LIBERATION ARMY

Synthetic data generation system and method

A computer-implemented method for generating synthetic data is provided. The method includes receiving user input specifying domain-specific requirements for synthetic data generation and selecting a scenario type. The scenario type is one of Seedless, Seeded, or combination of Seeded and Knowledge Base (KB). The method defines a structured schema based on the user input. The structured schema includes data fields, relationships between data fields, and distributional targets. Based on the structured schema, the method generates an initial set of synthetic data samples using a neural template-driven generation model trained on domain-specific data. The method applies adversarial contrastive sampling. This involves training a discriminator neural network to distinguish between the initial set of synthetic samples and real data samples. The discriminator neural network is used to identify generated samples similar to real data. A contrastive set of samples dissimilar to those identified as similar by the discriminator is generated. The method integrates the initial set of synthetic data samples with the contrastive set to create the synthetic data.
Owner:FUTURE AGI INC

Industrial part small sample target detection method based on synthetic data

The invention discloses an industrial part small sample target detection method based on synthetic data, and the method consists of a synthetic data generation module and an enhanced target detection model named YOLO-DC, and comprises the steps: separating a training process of the target detection model from dependence on large-scale real labeled data; and training the YOLO-DC model only by using virtual data generated by the synthetic data generation module. The trained model can be directly deployed in a real physical environment to accurately detect industrial parts, so that a visual perception task is completed. According to the method, a user can quickly deploy a detection system adaptive to new parts without collecting and marking real images, so that the development and deployment period of an industrial visual system is shortened, the data cost is remarkably saved, and meanwhile, the flexibility and the intelligent level of a manufacturing system are greatly improved.
Owner:YANTAI ZHONGKELANDE CNC TECH CO LTD

Synthetic data generation for machine learning model simulation

A method and system for synthetic data generation are provided that receive a schema configuration file in a synthetic data set request from a client application, create a set of worker processes to generate the synthetic data set based on the schema configuration file, upload the generated synthetic data to an analytics platform, and enable the client application to utilize the generated synthetic data in prediction models for the analytics platform.
Owner:SALESFORCE INC

Steel defect detection method based on synthetic data and CNN

The invention discloses a steel defect detection method based on synthetic data and CNN, and the method comprises the steps: S1, setting an environment for generating the synthetic data, and making preparation for the generation of the synthetic data with steel surface defects; s2, setting a shader of the steel model of the synthetic data, and generating the synthetic data with steel surface defects; s3, preprocessing the synthesized data to generate a training set and a verification set; s4, inputting the training set and the verification set into a ResNet-50 model, and training and testing the ResNet-50 model to obtain a steel defect detection model; and S5, performing steel defect detection through the steel defect detection model. Artificially synthesized data is used for training, a large amount of high-quality training data is not needed, and the problem that an automatic visual inspection system lacks original data is solved. Compared with a traditional machine learning method, the steel defect detection model is adopted for steel defect detection, the performance is higher, and the defect detection accuracy and efficiency are improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Data quality assessment and transformation in a privacy preserving federated system

A computer-implemented method for automatically transforming client data to a common data normalization schema associated with a collaborative multi-client federated learning system while preserving data privacy. The method may include automatically generating a local data ontology based on the client data associated with a client, and automatically generating synthetic data based on the client data and the local data ontology. The method may also include automatically computing an inference risk score comprising determining a privacy risk associated with sharing the synthetic data, and automatically computing a task utility score comprising determining a utility of the synthetic data. The method may further include generating a global data ontology using ontology matching algorithms on the synthetic data associated with each local data ontology. The method may also include automatically recommending and implementing data transformations to the client data.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Omnidirectional human body activity recognition method and device based on HDMR, computer and medium

The invention specifically discloses an omni-directional human body activity recognition method and device based on an HDMR, a computer and a medium, and relates to the technical field of radar signal processing. According to the method, an HDMR-based synthetic data generation algorithm is utilized, data expansion can be carried out on collected samples in a radar zero-degree observation angle direction and a small number of samples in other non-zero-degree observation angle directions, and high-quality training data containing all observation angle directions are generated; the objective of the invention is to solve the problem of angle sensitivity of a monostatic radar in omni-directional human body activity recognition. In addition, according to the invention, the dynamic time warping distance DTWD is adopted to measure the similarity between the synthetic sample and the real sample, so that the quality of the synthetic sample is evaluated. And finally, inputting the synthesized sample data in different observation angle directions into a CNN classifier based on ResNet50 for training so as to realize omnidirectional human body activity recognition.
Owner:BEIHANG UNIV

Enhancing physical reasoning in vision-language models using procedural synthetic data generation

One embodiment sets forth a technique for fine-tuning a machine learning model to perform physical reasoning. According to some embodiments, the method can include the steps of obtaining simulation annotations that describe interactions among simulated objects within a physics-based environment and one or more question templates, each question template defining a different parameterized reasoning query; generating, based on the simulation annotations and the one or more question templates, a plurality of question-answer pairs that represent physical reasoning examples; formatting the question-answer pairs into natural-language data compatible with the machine learning model; and fine-tuning the machine learning model based on the natural-language data.
Owner:AUTODESK INC

Inline Nested Data Loss Protection (DLP)

The disclosure presents systems and methods for hierarchical classification of input data across a plurality of categories. A machine learning model processes various data formats, starting with dimensional reduction using tokenization techniques, such as Bert-tiny tokenization, to create model-readable representations. The system predicts super-categories, sub-categories, and granular categories through selective activation of sub-layers tied to identified super-categories, optimizing computational efficiency. Label smoothing during training mitigates overconfidence in predictions, while softmax normalization refines inference outputs. Synthetic data generation using Large Language Models (LLMs) supplements training datasets, and an automated data labeling pipeline efficiently generates hierarchical labels. Modifications to the model, such as stop word removal and file size limitations, further reduce latency. Inference analyzes logits to predict hierarchical paths, providing detailed classifications with clear outputs. The method is adaptable for multimodal formats, ensuring scalable and accurate predictions across diverse data types while minimizing computational costs and improving reliability.
Owner:ZSCALER INC

Synthetic data generation method and device and storage medium

The embodiment of the invention provides a synthetic data generation method and device and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: in response to a data synthesis instruction initiated in a target page, acquiring scene data under multiple view angles, and performing three-dimensional Gaussian reconstruction on the scene data under the multiple view angles to obtain three-dimensional scene data; in response to an object editing instruction initiated in the target page, determining at least one scene object; determining data to be recorded according to the three-dimensional scene data and the at least one scene object; after a sensor is configured for a recording object, recording processing is carried out on the to-be-recorded data according to the configured sensor, and synthetic data of the recording object is obtained. According to the method, the efficiency, diversity, fidelity and automation degree of the generated synthetic data can be improved.
Owner:SZ ZHUOYU TECH CO LTD

Synthetic data generation

Various embodiments described herein provide for systems, methods, devices, instructions, and like for generating synthetic data. According to various embodiments, synthetic data generation comprises receiving input specifying one or more source tables and join key columns, and generating synthetic data that preserves statistical similarity and referential integrity among columns of the source data.
Owner:SNOWFLAKE INC

Boundary detection for synthetic data generation

A computer-implemented method can determine a linguistic boundary condition for synthetic data generation in a multi-class classification problem. The method includes analyzing empirical labelled data using linguistic and vector representation techniques. The method further includes deriving a synthetic boundary conditional (SBC) model based on the analysis of the empirical labelled data and identifying a boundary location for performant synthetic data generation using the SBC model.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

LLM report interpretation method and system based on synthetic data

The invention discloses an LLM report interpretation method and system based on synthetic data, and relates to the technical field of natural language processing and large language model application, in particular to the LLM report interpretation method and system based on the synthetic data. The objective of the invention is to solve the problems of insufficient accuracy, universality and user adaptability in report interpretation in the prior art. Firstly, specialized question and answer pairs are generated through synthetic data, and a knowledge base of a system is enriched; secondly, answered question and answer pairs are stored through a cache mechanism, repeated calculation is avoided, and the query response speed is increased; and finally, by utilizing a context-enhanced generation model, when the user query is processed, an answer more conforming to the characteristics of the report field can be generated, so that the interpretation accuracy and specialty are improved.
Owner:HARBIN INST OF TECH

System(s) and / or method(s) for forecasting using generated synthetic data

One or more methods and / or systems for generating synthetic data are provided. Profiles are generated for wireless communication sites from first data gathered for a first time period. Measures of similarity of second wireless communication sites to a first wireless communication site are calculated based on the generated profiles. Second wireless communication sites are selected based upon the measures of similarity. Weightings are generated for the selected second wireless communication sites based upon the measures of similarity of the selected second wireless communication sites. Second data is gathered for a second time period from the selected second wireless communication sites. Synthetic data is generated based upon the gathered second data and the generated weightings for the selected second wireless communication sites. The generated synthetic data is for the first wireless communication site for the first time period.
Owner:VERIZON PATENT & LICENSING INC