Semi-generative artificial intelligence
The semi-generative AI modeling approach addresses the precision and verification challenges of generative AI by using cross-domain diffusion models with metadata constraints, enhancing the accuracy and fidelity of generated images for object detection systems.
Patent Information
- Application Number
- JP2025035602
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-03-06
- Publication Date
- 2025-12-05
AI Technical Summary
Existing generative AI models lack precision and layers of verification for the integrity of generated content, particularly in large-scale object detection systems, due to the difficulty in creating diverse and realistic visual datasets with high precision and specificity.
A semi-generative AI modeling approach using cross-domain diffusion models, incorporating metadata from various data sources like real images, 3D CAD data, and alternative signal modalities, to control output through dual fusion techniques, ensuring accuracy and fidelity by transforming images between domains with metadata constraints.
Enhances the accuracy and fidelity of generated images by controlling output precision and specificity, enabling better training of object detection models with synthetic data that closely resembles real-world conditions.
Smart Images

Figure 2025178111000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to artificial intelligence systems, and more particularly to controlling randomized noise in generative AI models. [Background technology]
[0002] Object detection systems implementing artificial intelligence are trained to accurately detect target objects using large, realistic visual datasets with high precision, specificity, and diversity. High accuracy is achieved by reducing bias from human labelers. High specificity is achieved by capturing different images of the target object in various environmental conditions. High diversity is achieved by including various images of the target object from various viewing angles and viewpoints. However, developing large, realistic visual datasets with high precision, specificity, and diversity is difficult, especially using traditional methods that manually take photographs of target objects using a camera and have a human operator label each image with a ground truth target object class. These limitations have slowed the development and deployment of large-scale object detection systems. Summary of the Invention [Means for solving the problem]
[0003] An exemplary embodiment provides a method for cross-domain semi-generative artificial intelligence modeling. The method includes receiving a source image of an object in a first domain and diffusing the source image through a source diffusion model to generate a first Gaussian distribution in the first domain. An embedding is generated from metadata that provides constraints for image reconstruction. The embedding is fed to a double diffusion implicit bridge. The first Gaussian distribution is sampled and mapped from the first Gaussian distribution to a second Gaussian distribution in a second domain via the double diffusion implicit bridge. The second Gaussian distribution is then de-diffused through a target diffusion model to generate a target image of the object in the second domain according to the metadata.
[0004] Another exemplary embodiment provides a system for cross-domain semi-generative artificial intelligence modeling, comprising: a storage device that stores program instructions; and one or more processors operatively coupled to the storage device that execute the program instructions to cause the system to receive a source image of an object in a first domain, diffuse the source image through a source diffusion model to generate a first Gaussian distribution in the first domain, generate an embedding from metadata, where the metadata provides constraints for image reconstruction, feed the embedding to a double diffusion implicit bridge, extract samples from the first Gaussian distribution, map the samples from the first Gaussian distribution to a second Gaussian distribution in a second domain via the double diffusion implicit bridge, and de-diffuse the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain according to the metadata.
[0005] Another exemplary embodiment provides a computer program product for cross-domain semi-generative artificial intelligence modeling, comprising a computer-readable storage medium having program instructions embodied therein to perform the following: receiving a source image of an object in a first domain, diffusing the source image through a source diffusion model to generate a first Gaussian distribution in the first domain, generating an embedding from metadata, the metadata providing constraints for image reconstruction, feeding the embedding to a double diffusion implicit bridge, extracting samples from the first Gaussian distribution, mapping the samples from the first Gaussian distribution to a second Gaussian distribution in a second domain via the double diffusion implicit bridge, and de-diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain according to the metadata.
[0006] The forms and functions may be achieved independently in various embodiments of the present disclosure or may be combined in yet other embodiments, further details of which can be seen with reference to the following description and drawings.
[0007] The novel features believed characteristic of the exemplary embodiments are set forth in the appended claims. However, the exemplary embodiments, as well as their preferred modes of use, further objects and features, will best be understood by reference to the following detailed description of exemplary embodiments of the present disclosure, when considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram of a semi-generative system in accordance with an illustrative embodiment; [Figure 2] FIG. 1 illustrates a diffusion model in which exemplary embodiments may be implemented. [Figure 3] FIG. 1 illustrates an example of inter-domain image transformation according to an exemplary embodiment. [Figure 4]FIG. 1 illustrates semi-generative AI image generation using metadata constraints according to an exemplary embodiment. [Figure 5] FIG. 1 illustrates an example of image decoding from a 3D CAD model to a photorealistic model according to an exemplary embodiment. [Figure 6] 10 is a flowchart illustrating a process for cross-domain semi-generative artificial intelligence modeling in accordance with an illustrative embodiment; [Figure 7] 1 is a block diagram of a data processing system in accordance with an illustrative embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0009] The illustrative embodiments recognize and take into account that generative artificial intelligence (AI) models lack the precision and layers of verification to verify the integrity of the generated content.
[0010] Exemplary embodiments provide a method for semi-generative AI modeling that adds an additional layer of accuracy and fidelity to content generated from generative AI models. By utilizing metadata from design, production, and inspection data in manufacturing and production environments, exemplary embodiments control the output of generative AI models via dual fusion techniques.
[0011] 1 is a block diagram of a semi-generative system illustrated in accordance with an exemplary embodiment. The semi-generative system 100 utilizes data from a variety of sources, including real images and videos 102 of artifacts, three-dimensional (3D) computer-aided design (CAD) data 104, two-dimensional (2D) CAD drawings 106, and alternative signal modalities 108.
[0012] Real images and videos 102 of the artifact serve as ground truth data for real-domain dataset generation. 3D CAD data 104 provides another domain for image data. The 3D CAD data 104 can be derived from engineering product data management (PDM) and is used for synthetic-domain dataset generation.
[0013] 2D CAD drawings 106 provide production and design engineering metadata (i.e., any data that is not a 3D object). Alternative signal modalities 108 provide additional metadata, such as textual data, audio data, or graphical representations that describe the real images and videos 102. The metadata provides constraints to ensure that the AI-generated images conform to desired characteristics. When transforming image data from one domain to another, constraints exert control over the output domain (e.g., a particular type of object viewed from a particular angle). Such constraints help ensure that the output image is as close as possible to the original input image in every respect.
[0014] Data and metadata from real images and videos 102 of artifacts, three-dimensional (3D) computer-aided design (CAD) data 104, two-dimensional (2D) CAD drawings 106, and alternative signal modalities 108 are fed into an unsupervised / semi-supervised model generator 110, which comprises an end-to-end object detection system that identifies objects using structured and unstructured data. Images are cropped and transformed into a latent space and then clustered by visual similarity. Clusters are then aligned and labeled with features from provided examples.
[0015] The data of the different modalities are then fed by the unsupervised / semi-supervised model generator 110 to a dual fusion pipeline 116 .
[0016] The AI design parser 112 reads and interprets information from production references 114 (e.g., engineering drawings) regarding limitations on the target object to be generated to match the real images and videos 102, and provides additional constraints. The AI design parser 112 can provide information that may be missing from the alternative signal modalities 108. The AI design parser 112 can cross-reference text, geometry, engineering specifications, and mathematics with the real images and videos 102.
[0017] The dual fusion pipeline 116 comprises a source diffusion model 118 trained on a first domain of data 120 (e.g., 3D CAD). The source diffusion model 118 generates a first Gaussian distribution 126 for the first domain 120 from the image data using a source Schrödinger bridge 136. (See Figures 2 and 3.) The diffusion Schrödinger bridge attempts to find a stochastic process that connects two probability distributions in the most plausible way within given constraints.
[0018] The dual fusion pipeline 116 also includes a target diffusion model 122 trained in a second domain of data 124 (e.g., the real world). The target diffusion model 122 generates image data for the second domain 124 from a second Gaussian distribution 128 using a target Schrödinger bridge 138 for de-diffusion.
[0019] The double diffusion implicit bridge 130 provides a transformation between a first Gaussian distribution 126 and a second Gaussian distribution 128 by mapping from a first domain 120 to a second domain 124. (See FIG. 4.) The double diffusion implicit bridge 130 is a back-to-back concatenation from the source to the latent space and from the latent space to the target Schrödinger bridge 138. The double diffusion implicit bridge 130 can incorporate embeddings 132 generated from metadata provided by the 2D CAD 106, alternative signal modalities 108, and the AI design parser 112. The metadata represented by the embeddings 132 provides constraints when transforming between Gaussian distributions in different domains.
[0020] The dual fusion pipeline 116 takes advantage of images with better FID (Fréchet Inception Distance) scores than other models. FID is a measure of the quality of images produced by a generative model. FID compares the distribution of generated images to the distribution of a set of real images (which serve as ground truth). A lower FID score indicates greater similarity between the generated and real images. Diffusion models can produce lower FID scores than other types of models, such as generative adversarial networks (GANs).
[0021] The output from the dual fusion pipeline 116 is fed to a generative AI realistic synthetic data generator 134, which generates synthetic data for use by the object detection model.
[0022] The semi-generative system 100 may be implemented in software, hardware, firmware, or a combination thereof. When software is used, the operations performed by the semi-generative system 100 may be implemented in program code configured to run on hardware, such as a processor unit. When firmware is used, the operations performed by the semi-generative system 100 may be implemented in program code and data and stored in persistent memory running on a processor unit. When hardware is employed, the hardware may include circuitry that operates to perform the operations in the semi-generative system 100.
[0023] In illustrative examples, the hardware may take the form of at least one of a circuit system, an integrated circuit, an application-specific integrated circuit (ASIC), a programmable logic device, or other suitable type of hardware configured to perform a plurality of operations. Using a programmable logic device, the device may be configured to perform a plurality of operations. The device may be later reconfigured or may be permanently configured to perform a plurality of operations. Programmable logic devices include, for example, programmable logic arrays, programmable array logic, field programmable logic arrays, field programmable gate arrays, and other suitable hardware devices. Additionally, processes may be implemented in organic components integrated with inorganic components, or may be comprised entirely of organic components, excluding humans. For example, processes may be implemented as organic semiconductor circuits.
[0024] Computer system 150 is a physical hardware system that includes one or more data processing systems. When multiple data processing systems are present in computer system 150, the data processing systems communicate with each other using a communication medium. The communication medium may be a network. The data processing systems may be selected from at least one of a computer, a server computer, a mobile device such as a tablet computer, or other suitable data processing system.
[0025] As shown, computer system 150 includes several processor units 152 that can execute program code 154 that implements processes in the illustrative example. As used herein, a processor unit of several processor units 152 is a hardware device, composed of hardware circuits, such as those on integrated circuits, that process in response to instructions and program code that cause a computer to operate. When several processor units 152 execute program code 154 for a process, several processor units 152 are one or more processor units that may be on the same computer or different computers. In other words, processes can be distributed among processor units on the same or different computers of a computer system. Furthermore, several processor units 152 can be processor units of the same or different types. For example, several processor units can be selected from at least one of a single-core processor, a dual-core processor, a multi-processor core, a general-purpose central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), or some other type of processor unit.
[0026] 2 illustrates a diagram of a diffusion model in which exemplary embodiments can be implemented. Diffusion model 200 is an example of the first and second diffusion models 118, 122 of FIG.
[0027] Diffusion model 200 is trained by gradually adding noise to an image in stages (i.e., X1, X2, etc.) and then reconstructing (dediffusing) that image from each stage. As diffusion model 200 learns, it can add more and more noise to the original image and successfully reconstruct that original image. In the final stage of training (Z), diffusion model 200 generates a random Gaussian distribution of noise from which the original image can be reconstructed.
[0028] For diffusion processes in generative modeling, the goal is to generate a given distribution P prior Given the data distribution P data The goal is to design an algorithm to convert
number
[0029] X t is interpreted as a noise process, and the data distribution P data Disturbing. P prior is the invariant distribution of the noise process. This process adds noise gradually, and if it is done in small steps over a sufficiently long time step for T>0, then X t is the mean of 0 and
number
[0030] For despreading, in principle, the process is prior First sample from X t By reversing the dynamics of P data This time-reversal operation results in another diffusion process with explicit drift and diffusion matrices, i.e.,
number
[0031] In the formula, P t is X t is the density of . The process is then required to follow a discretization of the Markov chain and the diffusion process.
[0032] The main limitation of generative modeling is that the initial forward dynamics are prior The problem is that we need a large number of step sizes so that the step size is close to , and small enough so that the neural network approximation holds.
[0033] By adding a Schrödinger bridge, it is possible to significantly reduce the number of step sizes required to define score-based generative modeling.
[0034] A distribution that describes the process of adding noise to data
number
number
number
number
[0035] 3 illustrates a diagram showing an example of inter-domain image transformation according to an exemplary embodiment. This inter-image transformation relies on a source diffusion model 302 and a target diffusion model 304 that are trained independently in separate domains A and B, respectively. In this example, domain A contains CAD image data, and domain B contains real-world images. By training the source diffusion model 302 and the target diffusion model 304 independently in each domain, the P of each domain can be calculated. prior It is possible to identify
[0036] Diffusion by source diffusion model 302 generates Gaussian distribution 306, which represents the latent encoding of domain A. These latent encodings of the source image with source diffusion model 302 are then fed to target diffusion model 304 to construct the target image. In this way, Gaussian distribution 306 in domain A is transformed into Gaussian distribution 308 in domain B, which contains the latent encoding of domain B. Target diffusion model 304 then performs de-diffusion on Gaussian distribution 308 to construct the target image. Thus, in this example, a CAD image in domain A is transformed into a photorealistic image in domain B.
[0037] The process of encoding a source image and decoding it to produce a target image is defined via ordinary differential equations (ODEs), and therefore the process is cycle-consistent up to the discretization error of the ODS solver.
[0038] 4 illustrates a diagram showing semi-generative AI image generation using metadata constraints according to an example embodiment. In this example, a source diffusion model 402 generates a Gaussian distribution 408 from a source 3D CAD image 406 in the CAD domain. This first Gaussian distribution 408 is transformed into a second Gaussian distribution 412 in the real-world domain by a double-diffusion implicit bridge 414.
[0039] In the second half of the process (from the latent space to the second domain), more metadata 416 is added to aid in back-diffusion to the real-world domain. Examples of added metadata may include data clustering, prompt embedding (e.g., black color), material (e.g., steel), and 2D drawing information. The metadata 416 serves as constraints on the Schrödinger bridge of the target diffusion model 404, which generates a photorealistic image 410 of the real-world domain from a Gaussian distribution 412. The constraints provide controls to ensure that the resulting target image is not only the same type of image / object as the source image, but also that it is presented facing the same direction, the same angle, etc. Such precision and specificity are important for applications such as generating synthetic training data for object detection model training. Thus, the metadata helps ensure that the output is as close as possible to the original input.
[0040] FIG. 5 illustrates an example of image decoding from a 3D CAD model to a photorealistic model according to an example embodiment.
[0041] 6 depicts a flowchart illustrating a process for cross-domain semi-generative artificial intelligence modeling in accordance with an illustrative embodiment. Process 600 can be implemented in semi-generative system 100 of FIG.
[0042] Process 600 begins by receiving a source image of an object in a first domain (operation 602). The first domain may include three-dimensional CAD (computer-aided design) image data, and the source image may include a CAD image. Process 600 diffuses the source image through a source diffusion model to generate a first Gaussian distribution for the first domain (operation 604).
[0043] Next, process 600 generates embeddings from the metadata, which provide constraints for image reconstruction (operation 606). The metadata may include data clustering, prompt embeddings, segmentation masks, two-dimensional drawing information, text descriptions of the target object, audio descriptions of the target object, graphical representations of the target object, materials, or background. The metadata may also include information provided by an artificial intelligence design parser that cross-references text and specifications against images and videos.
[0044] The padding is provided to a double-diffusion implicit bridge (operation 608).
[0045] Process 600 extracts samples from a first Gaussian distribution (operation 610) and maps the samples from the first Gaussian distribution to a second Gaussian distribution in a second domain via a double diffusion implicit bridge (operation 612). The second domain may include real-world image data.
[0046] Next, process 600 de-diffuses the second Gaussian through the target diffusion model to generate a target image of the object in the second domain according to the metadata (operation 614). In the case of the real-world image data domain, the target image generated from the de-diffusion of the second Gaussian comprises a photorealistic image.
[0047] The process 600 then ends.
[0048] Referring now to Figure 7, a block diagram of a data processing system is shown in accordance with an illustrative embodiment. Data processing system 700 may be used to implement computer system 150 in Figure 1. In this illustrative example, data processing system 700 includes a communications framework 702 that provides communications between a processor unit 704, a memory 706, persistent storage 708, a communications unit 710, an input / output (I / O) unit 712, and a display 714. In this example, communications framework 702 takes the form of a bus system.
[0049] Processor unit 704 is responsible for executing instructions for software that may be loaded into memory 706. Processor unit 704 may be multiple processors, a multi-processor core, or some other type of processor, depending on the particular implementation. In one example, processor unit 704 comprises one or more conventional general-purpose central processing units (CPUs). In an alternative embodiment, processor unit 704 comprises one or more graphical processing units (GPUs).
[0050] Memory 706 and persistent storage 708 are examples of storage device(s) 716. A storage device is any hardware that can store information, such as, but not limited to, data, program code in a functional form, or other suitable information, on a temporary, persistent, or both temporary and persistent basis. Storage device 716, in these illustrative examples, may also be referred to as a computer-readable storage device. Memory 706, in these examples, may be, for example, a random access memory or any other suitable volatile or non-volatile storage device. Persistent storage 708 may take various forms depending on the particular implementation.
[0051] For example, persistent storage 708 may include one or more components or devices. For example, persistent storage 708 may be a hard drive, a flash memory, a rewritable optical disk, a rewritable magnetic tape, or some combination thereof. The medium used by persistent storage 708 may be removable. For example, a removable hard drive may be used for persistent storage 708. Communications unit 710, in these illustrative examples, provides for communication with other data processing systems or devices. In these examples, communications unit 710 is a network interface card.
[0052] Input / output unit 712 allows for the input and output of data with other devices that may be connected to data processing system 700. For example, input / output unit 712 may provide a connection for user input via at least one of a keyboard, a mouse, or some other suitable input device. Additionally, input / output unit 712 may send output to a printer. Display 714 provides a mechanism for displaying information to a user.
[0053] Instructions for at least one of the operating system, applications, or programs may be located in storage device 716, which is in communication with processor unit 704 through communications framework 702. The processes of the different embodiments may be performed by processor unit 704 using computer-implemented instructions, which may be located in a memory, such as memory 706.
[0054] These instructions are referred to as program code, computer usable program code, or computer readable program code, which may be read and executed by a processor in processor unit 704. The program code in different embodiments may be embodied on different physical or computer readable storage media, such as memory 706 or persistent storage 708.
[0055] Program code 718 is located in a functional form on computer readable media 720 that is selectively removable and may be loaded onto or transferred to data processing system 700 for execution by processor unit 704. Program code 718 and computer readable media 720 form computer program product 722 in these illustrative examples. In one example, computer readable media 720 may be computer readable storage media 724 or computer readable signal media 726.
[0056] In these illustrative examples, computer-readable storage medium 724 is not a medium that propagates or transmits program code 718, but rather a physical or tangible storage device used to store program code 718. As used herein, computer-readable storage medium 724 should not be construed to be a transitory signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted through wires. As used herein, computer-readable medium should not be construed to be a transitory signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted through wires.
[0057] Alternatively, program code 718 may be transferred to data processing system 700 using computer readable signal medium 726. Computer readable signal medium 726 may be, for example, a propagated data signal containing program code 718. For example, computer readable signal medium 726 may be at least one of an electromagnetic signal, an optical signal, or any other suitable type of signal. These signals may be transmitted over at least one of a communications link, such as a wireless communications link, an optical fiber cable, a coaxial cable, a wire, or any other suitable type of communications link.
[0058] The different components illustrated for data processing system 700 are not meant to provide architectural limitations to the manner in which different embodiments may be implemented. The different illustrative embodiments may be implemented in a data processing system including components in addition to or instead of those illustrated for data processing system 700. Other components illustrated in FIG. 7 may vary from the illustrated illustrative example. The different embodiments may be implemented using any hardware device or system capable of running program code 718.
[0059] As used herein, the phrase "at least one of," when used in conjunction with a list of items, means that different combinations of one or more of the listed items can be used, and that only one of each item in the list may be required. In other words, "at least one of" means that any combination and number of items can be used from the list, but not all of the items in the list may be required. An item can be a specific object, thing, or category.
[0060] For example, without limitation, "at least one of item A, item B, or item C" may include item A, item A and item B, or item B. This example may also include item A, item B, and item C, or item B and item C. Of course, any combination of these items can be present. In some illustrative examples, "at least one of" may be, for example, without limitation, two items A, one item B, and ten items C, four items B and seven items C, or other suitable combinations.
[0061] As used herein, "plurality," when used with reference to an item, means one or more items. For example, "plurality of different types of networks" is one or more different types of networks. In illustrative examples, a "set" used with a reference item means one or more items. For example, a set of metrics is one or more of the metrics.
[0062] The descriptions of different exemplary embodiments are presented for purposes of illustration and description and are not intended to be exhaustive or limited to the disclosed forms of embodiments. Various exemplary examples describe components that perform actions or operations. In the exemplary examples, a component may be configured to perform the described actions or operations. For example, a component may have a structural configuration or design that provides the component with the ability to perform the actions or operations described as being performed by the component in the exemplary examples. Furthermore, to the extent that the terms "include," "including," "has," "contain," and variations thereof are used herein, such terms are intended to be inclusive, similar to the open transitional term "comprise," without excluding any additional or other elements.
[0063] Numerous modifications and variations will be apparent to those skilled in the art. Additionally, various exemplary embodiments may provide different configurations than other preferred embodiments. The selected embodiment or embodiments have been chosen and described in order to best explain the principles and practical applications of the embodiments and to make the disclosure of the various embodiments understandable to those skilled in the art, along with various modifications suitable for the particular use envisioned. [Explanation of symbols]
[0064] 100 Semi-Generating Systems 102 Real images and videos of artifacts 104 Three-dimensional (3D) computer-aided design (CAD) data 106 2D CAD drawings 108 Alternative Signal Modalities 110 Unsupervised / Semi-supervised Model Generator 112 AI Design Parser 114 Production Reference 116 Double Fusion Pipeline 118 Source-Diffusion Model 120 First Domain 122 Target Dispersion Model 124 Second Domain 126 First Gaussian Distribution 128 Second Gaussian Distribution 130 Double Diffusion Implicit Bridge 132 Embed 134 Generative AI Realistic Synthetic Data Generator 136 Source Schrodinger Bridge 138 Target Schrodinger Bridge 150 Computer Systems 152 processor units 154 Program Code 200 Diffusion Model 302 Source diffusion model 304 Target Dispersion Model 306 Gaussian distribution in domain A 308 Gaussian distribution in domain B 402 Source diffusion model 404 Target Diffusion Model 406 Source 3D CAD Images in CAD Domain 408 First Gaussian Distribution 410 Photorealistic Images in the Real-World Domain 412 Gaussian Distribution 414 Double Diffusion Implicit Bridge 416 Metadata 600 processes 700 Data Processing System 702 Communication Framework 704 Processor Unit 706 memory 708 Persistent Storage 710 Communication Unit 712 Input / Output Unit 714 Display 716 Storage Devices 718 Program Code 720 Computer-Readable Medium 722 Computer Program Products 724 Computer-readable storage medium 726 Computer-Readable Signal Media
Claims
1. 1. A computer-implemented method for cross-domain semi-generative artificial intelligence modeling, said method comprising: It runs using several processors, receiving a source image of an object in a first domain; diffusing the source image through a source diffusion model to generate a first Gaussian distribution in the first domain; generating an embedding from metadata, the metadata providing constraints for image reconstruction; providing the embedding to a double-diffusion implicit bridge; Drawing a sample from the first Gaussian distribution; mapping the samples from the first Gaussian distribution to a second Gaussian distribution in a second domain via the double-diffusion implicit bridge; de-diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain according to the metadata; 11. A computer-implemented method comprising:
2. The method of claim 1 , wherein the first domain comprises three-dimensional CAD (computer-aided design) image data.
3. The method of claim 1 , wherein the second domain comprises real-world image data.
4. The method of claim 3 , wherein the target image comprises a photorealistic image.
5. The metadata includes: clustering of data, Prompt embedding, segmentation mask, 2D drawing information, a text description of the target object, A phonetic description of the target object, a graph representation of the target object, materials, or background The method of claim 1 , comprising at least one of:
6. The method of claim 1 , wherein the metadata includes information provided by an artificial intelligence design parser that cross-references text and specifications against images and videos.
7. The method of claim 1 , wherein the source diffusion model and the target diffusion model comprise a Schrodinger bridge.
8. 1. A system for cross-domain semi-generative artificial intelligence modeling, said system comprising: a storage device for storing program instructions; one or more processors operatively connected to the storage device; Equipped with The one or more processors execute the program instructions to cause the system to: receiving a source image of an object in a first domain; diffusing the source image through a source diffusion model to generate a first Gaussian distribution in the first domain; generating an embedding from metadata, the metadata providing constraints for image reconstruction; providing said embedding to a double-diffusion implicit bridge; Drawing a sample from the first Gaussian distribution; mapping the samples from the first Gaussian distribution to a second Gaussian distribution in a second domain via the double diffusion implicit bridge; de-diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain according to the metadata; A system that allows the following to be performed.
9. The system of claim 8 , wherein the first domain comprises three-dimensional CAD (computer-aided design) image data.
10. The system of claim 8 , wherein the second domain comprises real-world image data.
11. The system of claim 10 , wherein the target image comprises a photorealistic image.
12. The metadata includes: clustering of data, Prompt embedding, segmentation mask, 2D drawing information, a text description of the target object, A phonetic description of the target object, a graph representation of the target object, materials, or background The system of claim 8, comprising at least one of:
13. The system of claim 8 , wherein the metadata includes information provided by an artificial intelligence design parser that cross-references text and specifications against images and videos.
14. The system of claim 8 , wherein the source diffusion model and the target diffusion model comprise a Schrodinger bridge.
15. 1. A computer program for cross-domain semi-generative artificial intelligence modeling, the computer program comprising: receiving a source image of an object in a first domain; diffusing the source image through a source diffusion model to generate a first Gaussian distribution in the first domain; generating an embedding from metadata, the metadata providing constraints for image reconstruction; providing said embedding to a double-diffusion implicit bridge; Drawing a sample from the first Gaussian distribution; mapping the samples from the first Gaussian distribution to a second Gaussian distribution in a second domain via the double diffusion implicit bridge; de-diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain according to the metadata; A computer program comprising program instructions for executing the
16. 16. The computer program of claim 15, wherein the first domain comprises three-dimensional CAD (computer-aided design) image data.
17. The computer program of claim 15 , wherein the second domain comprises real-world image data.
18. The metadata includes: clustering of data, Prompt embedding, segmentation mask, 2D drawing information, a text description of the target object, A phonetic description of the target object, a graph representation of the target object, materials, or background 16. The computer program of claim 15, comprising at least one of:
19. 16. The computer program of claim 15, wherein the metadata includes information provided by an artificial intelligence design parser that cross-references text and specifications against images and videos.
20. 16. The computer program of claim 15, wherein the source diffusion model and the target diffusion model comprise a Schrodinger bridge.