Structural Formula Diversification for Scalable AI Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional structural formula generation programs lack diversity in generating training data for artificial intelligence models, leading to low performance and requiring extensive time to produce additional data, and struggle with deriving coordinate and class information from chemical structural formulas.
Innovation Solution
A device and method for generating training data by diversifying setpoints of objects in structural formulas using a processor and memory to create various formats, allowing for large-scale production of annotated structural formulas with adjustable features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a conventional structural formula generation program is used, then the structural formula can be generated in a fixed format, but it is difficult to diversify features such as thickness, color, interval of line segments, font, size, and color of characters
Solution Approach 1:
The patent applies parameter changes by introducing multiple feature variables (thickness, color, interval of line segments, font, size) and assigning setpoints to these variables. This allows the generation program to produce structural formulas with diverse visual characteristics while maintaining a consistent underlying generation logic, thereby increasing adaptability without proportionally increasing program complexity.
Solution Approach 2:
The patent implements dynamics by making the generation program adjustable through feature variables and setpoints. Instead of a fixed generation process, the system can dynamically modify characteristics of structural formulas (such as line thickness, color, spacing, font size) based on input parameters, enabling versatile output formats from a single program.
2Productivity
If the conventional generation program generates only one image at a time, then the processing is simple, but it takes a lot of time to process data to generate additional data
Solution Approach 1:
The patent applies preliminary action by pre-defining multiple setpoints for each feature variable before the generation process. This allows the system to quickly switch between different visual formats without performing complex calculations during generation, thereby increasing productivity and reducing the time required to generate additional training data.
Solution Approach 2:
By pre-establishing parameter sets (setpoints) for various features, the system can rapidly generate diverse structural formulas by simply changing which parameter set is applied. This approach dramatically increases generation speed and productivity while minimizing the time loss associated with creating additional training data.
3Reliability
If training data is generated using only an existing structural formula generation program, then the generation process is straightforward, but the training data does not reflect the diversity of actual structural formula data
Solution Approach 1:
The patent directly addresses this contradiction by introducing multiple feature variables (thickness, color, interval of line segments, font, size, color of characters) with predefined setpoints. This enables the generation of training data that reflects the visual diversity of actual structural formulas, thereby improving AI model performance and reliability without compromising the straightforwardness of the generation process.
Solution Approach 2:
The patent makes the generation program multi-functional by enabling it to produce structural formulas with various visual characteristics through adjustable feature variables and setpoints. This single program can generate diverse training data reflecting real-world variability, improving both the versatility of the output and the reliability of AI model training.
4Measurement precision
If coordinate information and class information of atoms and bonding regions are derived by hand, then the information can be accurately obtained, but the process is difficult and time-consuming
Solution Approach 1:
The patent applies self-service by designing the generation program to automatically output coordinate information and class information of atoms and bonding regions along with the structural formulas. This eliminates the need for manual derivation, maintaining measurement precision through programmatic accuracy while dramatically reducing the time and effort required to obtain this essential information.
Data Source
Figure 1
Figure 2
Figure 3(a)~6(b)
AI summary
A training data generation device according to an embodiment of the present invention comprises a processor and at least one memory electrically connected to the processor. The at least one memory stores: one or more feature variables determining the features of objects expressing structural formulas; and a first setpoint set determined in advance for each of the one or more feature variables. The processor: loads first chemical formula data, including information about at least one piece of node and information about at least one piece of edge, and the one or more feature variables; sets setpoints of the one or more feature variables as the first setpoint set; generates a first structural formula on the basis of the first chemical formula by using an object set as the first setpoint set that is applied; acquires, in response to an input changing the setpoints of one or more feature variables, a second setpoint set in which the setpoints of the corresponding feature variables are changed; generates a second structural formula by using an object set as the second setpoint set that is applied; and generates the training data including the generated second structural formula.