Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

112 results about "Synthetic data generation" patented technology

Synthetic data generation has become a surrogate technique for tackling the problem of bulk data needed in training deep learning algorithms. Areas such as computer vision have greatly benefited from advances in deep learning and now generating synthetic data is serving as a good starting point for researchers who are trying to bridge the data gap.

Determining lighting and composition parameters using machine learning models for synthetic data generation

Approaches presented herein provide for the determination of realistic lighting parameters for a scene represented in an image. Realistic lighting parameters can allow for the insertion of one or more virtual objects into a scene image, where the lighting or shading applied to the virtual object(s) can be consistent with those for other objects in the scene. A machine learning model such as a discriminator or diffusion model can be used to analyze a composed image generated by a differential renderer, for example, in which at least one virtual object has been inserted into a scene image and had lighting effects applied in accordance with a set of lighting parameters. A loss value can be determined based on the results of this machine learning model, which can be used to optimize the lighting parameters and / or adjust the weights or parameters of a model used to generate the lighting parameters. Once fine-tuned or optimized, the lighting parameters can represent an accurate light map for the scene or environment that can be used to generate composed images.
Owner:NVIDIA CORP

Synthetic data generation for modality-agnostic zero-shot foundation model for medical images

One or more systems, devices, computer program products and / or computer-implemented methods of use provided herein relate to assessing certainty of artificial intelligence models used for detection or segmentation of pathologies. Accordingly, a system can comprise a memory that can store computer executable components. The system can further comprise a processor that can execute at least one of the computer executable components. The computer executable components can comprise a synthetic data generation component that generates biologically-inspired synthetic data that approximates a task-specific data manifold of a medical image from a radiomic features perspective; an artificial intelligence component that uses an artificial intelligence model to learn relevant representations of the synthetic data for an at least one image task; and a training component that utilizes the relevant representations and the artificial intelligence model to generate a task-specific model for the at least one image analysis task.
Owner:GE PRECISION HEALTHCARE LLC

Synthetic data generation for retrieval evaluation and fine-tuning

In various examples, a technique for generating synthetic data includes inputting a first prompt that includes (i) a first portion of content and (ii) a plurality of user personas into a first machine learning model and generating, via the first machine learning model and based on the first prompt, a plurality of points of interest associated with the user personas and the first portion of content. The technique also includes inputting a second prompt that includes mappings between the points of interest and a plurality of question types into a second machine learning model and generating, via the second machine learning model and based on the second prompt, a plurality of questions associated with the user personas and the first portion of content. The technique further includes retrieving a second portion of content based at least on the plurality of questions and a third prompt.
Owner:NVIDIA CORP

Machine-learned architecture for structured synthetic data generation

Techniques may generate realistic synthetic data by programmatically generating a configuration file object type and relationship data. This configuration file may be used to retrieve source data matching the object type(s) and / or specific records indicated by the configuration file. The techniques may detect and anonymize private / proprietary information and may determine statistical characteristic(s) of the source data. A batch of prompt(s) may be generated using the source data, the statistical characteristic(s), and the configuration file and may be transmitted to one or more instances of a transformer-based machine-learned model. Sets of synthetic data received from the model instance(s) may be de-duplicated, checked for similarity to the source data (e.g., via embedding the synthetic data and the source data), and may be used to generate synthetic object(s) using the relationship(s) and / or other data indicated by the configuration file. These synthetic object(s) may then be deployed in a software environment.
Owner:SALESFORCE INC

Synthesizing high resolution 3D shapes from lower resolution representations for synthetic data generation systems and applications

In various examples, a deep three-dimensional (3D) conditional generative model is implemented that can synthesize high resolution 3D shapes using simple guides—such as coarse voxels, point clouds, etc.—by marrying implicit and explicit 3D representations into a hybrid 3D representation. The present approach may directly optimize for the reconstructed surface, allowing for the synthesis of finer geometric details with fewer artifacts. The systems and methods described herein may use a deformable tetrahedral grid that encodes a discretized signed distance function (SDF) and a differentiable marching tetrahedral layer that converts the implicit SDF representation to an explicit surface mesh representation. This combination allows joint optimization of the surface geometry and topology as well as generation of the hierarchy of subdivisions using reconstruction and adversarial losses defined explicitly on the surface mesh.
Owner:NVIDIA CORP

Competitive intelligence analysis system and method based on differential privacy

ActiveCN121563600ADigital data protectionCommerceCompetitive intelligenceMarket dynamics
The invention relates to a differential privacy-based competitive intelligence analysis system and method. The differential privacy-based competitive intelligence analysis system comprises a differential privacy processing and budget management module, a differential privacy feature engineering module, a privacy protection analysis and prediction module and a differential privacy synthetic data generation module. According to the design of the invention, an end-to-end advanced analysis framework with mathematics provable privacy protection capability is constructed, so that accurate, timely and reliable insight of market dynamics and competitor strategies is realized on the premise of strictly protecting enterprise sensitive commercial confidentials.
Owner:HUIZHOU UNIV +2

Synthetic data generation system and method

A computer-implemented method for generating synthetic data is provided. The method includes receiving user input specifying domain-specific requirements for synthetic data generation and selecting a scenario type. The scenario type is one of Seedless, Seeded, or combination of Seeded and Knowledge Base (KB). The method defines a structured schema based on the user input. The structured schema includes data fields, relationships between data fields, and distributional targets. Based on the structured schema, the method generates an initial set of synthetic data samples using a neural template-driven generation model trained on domain-specific data. The method applies adversarial contrastive sampling. This involves training a discriminator neural network to distinguish between the initial set of synthetic samples and real data samples. The discriminator neural network is used to identify generated samples similar to real data. A contrastive set of samples dissimilar to those identified as similar by the discriminator is generated. The method integrates the initial set of synthetic data samples with the contrastive set to create the synthetic data.
Owner:FUTURE AGI INC

Industrial part small sample target detection method based on synthetic data

The invention discloses an industrial part small sample target detection method based on synthetic data, and the method consists of a synthetic data generation module and an enhanced target detection model named YOLO-DC, and comprises the steps: separating a training process of the target detection model from dependence on large-scale real labeled data; and training the YOLO-DC model only by using virtual data generated by the synthetic data generation module. The trained model can be directly deployed in a real physical environment to accurately detect industrial parts, so that a visual perception task is completed. According to the method, a user can quickly deploy a detection system adaptive to new parts without collecting and marking real images, so that the development and deployment period of an industrial visual system is shortened, the data cost is remarkably saved, and meanwhile, the flexibility and the intelligent level of a manufacturing system are greatly improved.
Owner:YANTAI ZHONGKELANDE CNC TECH CO LTD

Enhancing physical reasoning in vision-language models using procedural synthetic data generation

One embodiment sets forth a technique for fine-tuning a machine learning model to perform physical reasoning. According to some embodiments, the method can include the steps of obtaining simulation annotations that describe interactions among simulated objects within a physics-based environment and one or more question templates, each question template defining a different parameterized reasoning query; generating, based on the simulation annotations and the one or more question templates, a plurality of question-answer pairs that represent physical reasoning examples; formatting the question-answer pairs into natural-language data compatible with the machine learning model; and fine-tuning the machine learning model based on the natural-language data.
Owner:AUTODESK INC

Inline Nested Data Loss Protection (DLP)

The disclosure presents systems and methods for hierarchical classification of input data across a plurality of categories. A machine learning model processes various data formats, starting with dimensional reduction using tokenization techniques, such as Bert-tiny tokenization, to create model-readable representations. The system predicts super-categories, sub-categories, and granular categories through selective activation of sub-layers tied to identified super-categories, optimizing computational efficiency. Label smoothing during training mitigates overconfidence in predictions, while softmax normalization refines inference outputs. Synthetic data generation using Large Language Models (LLMs) supplements training datasets, and an automated data labeling pipeline efficiently generates hierarchical labels. Modifications to the model, such as stop word removal and file size limitations, further reduce latency. Inference analyzes logits to predict hierarchical paths, providing detailed classifications with clear outputs. The method is adaptable for multimodal formats, ensuring scalable and accurate predictions across diverse data types while minimizing computational costs and improving reliability.
Owner:ZSCALER INC

Synthetic data generation method and device and storage medium

The embodiment of the invention provides a synthetic data generation method and device and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: in response to a data synthesis instruction initiated in a target page, acquiring scene data under multiple view angles, and performing three-dimensional Gaussian reconstruction on the scene data under the multiple view angles to obtain three-dimensional scene data; in response to an object editing instruction initiated in the target page, determining at least one scene object; determining data to be recorded according to the three-dimensional scene data and the at least one scene object; after a sensor is configured for a recording object, recording processing is carried out on the to-be-recorded data according to the configured sensor, and synthetic data of the recording object is obtained. According to the method, the efficiency, diversity, fidelity and automation degree of the generated synthetic data can be improved.
Owner:SZ ZHUOYU TECH CO LTD

Synthetic data generation

Various embodiments described herein provide for systems, methods, devices, instructions, and like for generating synthetic data. According to various embodiments, synthetic data generation comprises receiving input specifying one or more source tables and join key columns, and generating synthetic data that preserves statistical similarity and referential integrity among columns of the source data.
Owner:SNOWFLAKE INC

Intelligent Apparatus for Next-GEN Carbon Credits Platform Leveraging Synthetic Data Token Generation Powered by Responsible AI

Intelligent methods, processes, and systems are disclosed for generating synthetic eco-crypto tokens for carbon credits and may include the steps of identifying the amount of carbon credits saved for a product or service purchased by a customer by using data centric artificial intelligence (AI) working with trained data sets. The trained data sets may be loaded to knowledge graphs and a unique green reward identification for a customer may be combined to then generate a new synthetic data which may be used to generate an eco-crypto token. Responsible AI may be used to generate the eco-crypto tokens in a safe, trustworthy, and ethical fashion for use in other transactions for products and services.
Owner:BANK OF AMERICA CORP

Roadbed settlement time sequence prediction and three-dimensional deformation monitoring system and method thereof

The invention discloses a roadbed settlement time sequence prediction and three-dimensional deformation monitoring system and method, and belongs to the technical field of earth surface deformation monitoring and prediction.The system comprises a data acquisition module, a spatial-temporal feature construction module, a prediction model fusion module, a deformation field reconstruction module, a synthetic data generation module and a monitoring and early warning module; according to the method, an InSAR satellite remote sensing image and foundation micro-deformation sensor data are fused, a time-space coupled Transform-GRU prediction model is constructed, a pixel-level settlement rate field algorithm is created for the first time, the earth surface deformation monitoring precision is improved to 0.5 mm every year, historical settlement images are synthesized through a generative adversarial network, training data are enhanced, and the earth surface deformation monitoring precision is improved to 0.5 mm every year. Accurate reconstruction and long-term prediction of the three-dimensional deformation field of the roadbed are realized, the problem of insufficient generalization ability under the conditions of insufficient three-dimensional deformation reconstruction, long-term prediction precision and data in the prior art is effectively solved, and powerful support is provided for safe operation of traffic infrastructures.
Owner:CCCC TUNNEL ENG CO LTD +1

Large-scale bird flock automatic annotation synthetic data generation method and system

The invention relates to the technical field of ecological monitoring and computer vision, and particularly discloses a large-scale bird flock automatic labeling synthetic data generation method and system. The technical problem to be solved is to overcome the defects that in the prior art, the acquisition cost of large-scale bird flock counting task annotation data is high, the quality is difficult to guarantee, and the authenticity and scale of synthetic data are limited. According to the technical scheme, the method comprises the following steps: based on acquired real small-scale bird flock three-dimensional trajectory data, performing group scale expansion by using a collective motion model such as an inertial spin model, and generating a large-scale bird flock three-dimensional motion trajectory; a pre-established three-dimensional bird model and a flight animation are endowed with the motion trail, rendering is performed in a virtual rendering engine, and an initial synthetic image and a corresponding pixel-level precise labeling mask including a binary mask of the whole bird flock and an instance segmentation mask of each individual are generated; and performing visual reality enhancement processing on the initial synthetic image, migrating the real bird texture to the synthetic bird through style migration, and fusing the mask and the real background image to generate a final synthetic image with high reality, and strictly maintaining the integrity of the annotated mask at the same time. The method is mainly used for automatically generating precisely labeled and highly realistic large-scale bird flock images, provides efficient data support for bird flock automatic counting model training, and can be applied to ecological monitoring and protection.
Owner:SUN YAT SEN UNIV

Financial data privacy extraction method based on generative artificial intelligence and storage medium

The invention discloses a financial data privacy extraction method based on generative artificial intelligence and a storage medium, and belongs to the technical field of computers. The method comprises the steps of multi-modal bank data asset atlas construction, intelligent data privacy grading, federated generation model construction, cross-domain data joint reasoning, synthetic data generation application, data authority dynamic management, full-link early warning, privacy risk assessment, data delivery and the like. The financial data extraction efficiency and precision are improved, the data security and privacy are guaranteed, the system is universal and extensible, the financial data value is released, the system is suitable for financial institution data management, and the financial data extraction efficiency and precision are improved. A computer program for implementing the method is stored in the storage medium.
Owner:CHINA CONSTRUCTION BANK

Synthetic data generation for machine learning models

Techniques for generating synthetic data for a machine learning (ML) model are described. A system includes a language model that processes a task and a corresponding set of example inputs to generate another input, referred to herein as machine-generated data. The machine-generated data is processed using an ML model for which data is being generated to determine a model output, and the model output is analyzed to determine whether it corresponds to a target output. If the model output corresponds to the target output, the machine-generated data is added to the set of examples, and one of the original example inputs is deleted to generate an updated set of example inputs. The updated set may be used in various training techniques.
Owner:AMAZON TECH INC

Synthetic data generation of image training data

Systems and methods to generate synthetic images for use in a machine learning training set. The process begins with accessing a database of real-world 3-D images of equipment in a power grid, the 3-D images of equipment include 3-D measurements to create a dimensionally accurate and photorealistic model of the equipment. Optionally, the 3-D images could be aged or weathered using imaging editing software. Next, a database of real-world photographs of scenes in which the equipment is installed is accessed. Optionally, the identical scenes can be captured at different times of day, different times of the year, and at different perspectives. Next, using image editing software, the 3-D images of equipment is inserted into at least one of the scenes to form a synthetic image based on a combination of the equipment and the scene in which each of the equipment and the scene were previously captured independently of each other.
Owner:FLORIDA POWER & LIGHT CO

Multi-modal synthetic data generation method and system oriented to health-care large model

The invention discloses a multi-modal synthetic data generation method and system for a Kangbao large model, and relates to the field of Kangbao large models. The method specifically comprises the following steps: analyzing multi-modal health care data to obtain a semi-structured document; constructing a health-care knowledge graph by adopting an automatic knowledge graph construction technology, and extracting question and answer pair data from a semi-structured document; the question-answer pair data comprises complete question-answer pairs and question contexts; obtaining brief and detailed question and answer pairs according to the context of the question; evaluating brief and detailed question and answer pairs through a multi-dimensional generation quality evaluation system; generating feedback information to optimize the brief and detailed question and answer pairs to obtain context question and answer pairs; optimizing the complete question and answer pair by adopting instruction fine tuning; and based on the optimized complete question and answer pairs, brief version question and answer pairs and detailed version question and answer pairs, training to obtain a Kangbao model. According to the method, multi-modal synthetic data generation oriented to the Kangbao large model is realized.
Owner:CHINA UNICOM (BEIJING) IND INTERNET CO LTD

Task-driven synthetic data generation method and system based on meta-background constraint and intention perception

The invention relates to the technical field of artificial intelligence and data engineering, and particularly provides a task-driven synthetic data generation method and system based on meta-background constraint and intention awareness, and the method comprises the steps: obtaining a historical data pool of downstream tasks, carrying out the vectorization processing of samples in the data pool, and constructing a sample distribution space in a feature space; calculating a corresponding gap attribute, calculating a generation priority based on the gap attribute, carrying out reality and compliance detection on the element background, inputting the element background into the generation model if the detection is passed, and triggering a failure write-back mechanism to correct the element background if the detection is not passed; a hierarchical heterogeneous multi-view review system is constructed, output synthetic data is subjected to multi-dimensional review, an arbitration mechanism is triggered based on a judgment result, an arbitration result is returned to a corresponding generation link, system parameters are updated, and an iterative closed loop is formed. And the effect of adapting to dynamically changing downstream task requirements is achieved.
Owner:CHINA UNIV OF GEOSCIENCES (BEIJING)

Synthetic video simulations of predicted building equipment performance and future building equipment issue identification therefrom

An example method for predicting and visualizing future building equipment performance comprises: receiving actual performance and operating condition data associated with building equipment at a site, the actual performance and operating condition data including actual performance and operating condition data received from a building management system associated with the site; predicting future operating performance and operating condition data for the building equipment at the site during a future time period; generating synthetic data indicative of the predicted operating performance and operating condition data; generating, via the synthetic data, a synthetic video simulation of the future operating performance and operating condition data; displaying, via a display of a computing device in the building management system, the synthetic video simulation of the future operating performance and operating condition data; and responsive to identification of a future building equipment issue, generating an electronic ticket to remediate a future building equipment issue.
Owner:HONEYWELL INTERNATIONAL INC

Synthetic data generation for machine learning-based post-processing

A method includes obtaining a ground truth image and generating multiple image frames using the ground truth image, a modeled optical blur, and a modeled global motion. The method also includes generating multiple mosaic image frames using the image frames and a color filter array and generating multiple raw input image frames using the mosaic image frames and a noise model associated with at least one imaging sensor. The method further includes providing the raw input image frames to a multi-frame processing pipeline in order to generate synthetic training data. In addition, the method includes training a machine learning-based image processing engine using the ground truth image and the synthetic training data.
Owner:SAMSUNG ELECTRONICS CO LTD

Synthetic data generation method and device, storage medium and computer program product

The invention provides a synthetic data generation method and device, a storage medium and a computer program product, and relates to the technical field of artificial intelligence. Performing hierarchical progressive rendering processing on the scene description parameters to obtain scene multi-modal sensing data; the hierarchical gradual rendering processing comprises the step of performing (n + 1) th layer rendering processing on a rendering result under the condition that the rendering result of the nth layer rendering processing passes value evaluation; the precision of the (n + 1) th layer of rendering processing is higher than that of the nth layer of rendering processing; n is a positive integer greater than or equal to 1; and training synthetic data based on the scene multi-modal perception data generation model. According to the invention, the invalid consumption of high-fidelity rendering can be reduced, the rendering and processing efficiency of high-value samples is prevented from being dragged, and the utilization efficiency of the overall computing power of the system is improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Synthetic data generation for ai use cases

PendingUS20260252858A1Software engineeringData mining
Generative AI (GenAI) applications make use of a prompt to guide the output of the Large Language Model (LLM). The prompt is generally composed of a static template with placeholders for input data, which are populated when the LLM is invoked. A solution is presented herein for generating synthetic data directly from the use case prompt template, reducing the dependency on actual data collection processes across a wide range of task domains. In some example embodiments, two main steps are performed. First, the LLM is asked to generate personas. Using the personas, the LLM is asked to generate input data for the placeholders of a prompt template. The responses for multiple personas are aggregated and de-duplicated, resulting in a synthetic set of values for the placeholder of the prompt template.
Owner:SAP SE

Training data generation device and training data generation method

PCT designated stageWO2025258052A1Machine learningData setEngineering
A training data generation device comprising: a secondary training dataset generation unit that generates a secondary training dataset from a primary training dataset; and a tertiary training data generation unit that generates a tertiary training dataset from the secondary training dataset. The secondary training dataset generation unit includes: a probability distribution generation unit that generates, from the primary training dataset, a latent data space corresponding to the primary training dataset, and generates a probability distribution in the latent space; a gain function generation unit that generates a gain function used for extracting primary input data expected to have a large effect on improvement of model performance through use of the latent data space and the probability distribution; and a synthesis data generation unit that extracts primary input data as the source for generation of synthesis data on the basis of a score of the gain function and generates the secondary training dataset having secondary input data generated from the extracted primary input data as an element.
Owner:NT T INC

Systems and methods for generating synthetic data and optimizing large language model performance

Provided are systems and methods and computer-implemented systems for generating synthetic data and optimizing performance of large language models, by integrating semi-automatic synthetic data generation, prompt optimization and continual learning into a unified system, without full fine-tuning of the model, to improve performance of the large language model. The cold start problem of the large language model is solved through a multi-stage process, including: generating a diversified parameter set and natural language sentences; constructing synthetic user queries using feature descriptions, parameter sentences and constraint instructions; bootstrapping with a small amount of samples with the help of a compiler, and iteratively evaluating instruction-example combinations to optimize the large language model module; and combining real-time data and synthetic data to complete continual validation and optimization.
Owner:HSBC SOFTWARE DEV (GUANGDONG) LTD

Robot dexterous operation synthetic data generation method and system based on reinforcement learning

The invention discloses a robot dexterous operation synthetic data generation method and system based on reinforcement learning, and the method comprises the steps: carrying out the preprocessing of motion capture data containing hand joint angles and object poses, and converting the posture parameters of a hand parameterized model into a unified control space; generating an action instruction through an accumulated residual action control strategy; performing strategy training on the action instruction by utilizing a reinforcement learning strategy, wherein an observation state space comprises a hand joint state, relative position information and an object pose; in the strategy execution process, a hand object contact event is detected, and contact force information containing multi-coordinate system representation is generated; operation track data is generated based on the original motion capture data, synthetic operation data containing complete contact force information is generated through a hybrid strategy, and data quality screening is executed; the invention provides a technical scheme capable of efficiently generating robot flexible operation synthetic data containing contact force information with high quality.
Owner:ZHEJIANG LINGQIAO INTELLIGENT TECHNOLOGY CO LTD

Training data processing device and training data processing method

PCT designated stageWO2026038360A1Machine learningOriginal dataEngineering
The present invention presumes the existence of public data that is similar to raw data for training. This training data processing device comprises a synthetic data generation unit, a public data acquisition unit, and a synthetic data comparison unit. The synthetic data generation unit generates synthetic data that is pseudo training data from raw data. The public data acquisition unit acquires public data that is similar to the raw data. The synthetic data comparison unit compares the distribution of the public data and the synthetic data and eliminates missing synthetic data that is synthetic data that is missing from the distribution of the public data from the synthetic data.
Owner:NT T INC

Systems and method for optimizing medical procedures

The present disclosure relates to techniques for efficient performance of medical procedures. In some aspects, these improvements are achieved through improved intra- and inter-processor data transfer and processing, improved processor prioritization and parallelization, improved infrastructure management, and / or the implementation of an omnischeduler for efficient equipment management and / or task orchestration. These improvements also provide for improved application-layer capabilities, including improved digital twins and improved generation of synthetic data for training, for example, classification models and / or control policies.
Owner:INTUITIVE SURGICAL OPERATIONS INC

Machine-learning-based OKR generation

A method for training a machine learning model using positive and negative synthetic data is implemented via a computing system including a processor. The method includes generating synthetic data using a generative pre-trained transformer bidirectional language model and self-supervising the generated synthetic data based on positive traits including rule-based criteria and / or model-based criteria. The method also includes generating a set of positive synthetic data labels with gradient scale rating based on the self-supervised synthetic data, synthesizing a set of negative synthetic data labels by self-supervising the positive synthetic data labels, and training a machine learning model using the set of positive synthetic data labels and the set of negative synthetic data labels. Another method further includes utilizing the trained machine learning model to generate Objectives and Key Results (OKRs) within the context of an enterprise application.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC