Aroma reconstruction method, device, medium and product

By acquiring chemical and sensory data, determining the sensory contribution weight values, and using natural language processing technology combined with a multi-objective optimization genetic algorithm to generate Pareto optimal formulas, this solves the problem in existing technologies where aroma reconstruction is difficult to accurately reproduce sensory experience and balance costs, thus achieving efficient aroma formula generation.

CN122117151APending Publication Date: 2026-05-29SHANGHAI HEYUN FLAVORS & FRAGRANCES CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI HEYUN FLAVORS & FRAGRANCES CO LTD
Filing Date
2026-01-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing aroma reconstruction technologies struggle to consistently produce aroma formulas with optimal overall performance, fail to accurately reproduce sensory experiences, and are difficult to balance reproduction accuracy with cost.

Method used

By acquiring chemical analysis data and olfactory sensory data of the target sample, and combining GC-O intensity value and flavor dilution factor value, the sensory contribution weight value is determined. Natural language processing technology is used to convert the odor description text into a digital aroma feature vector. A multi-objective optimization genetic algorithm is used to perform iterative calculations in the raw material database to generate a Pareto optimal formulation set.

Benefits of technology

It achieves stable output of aroma formulas with optimal comprehensive indicators under multidimensional constraints, improves the calculation accuracy and restoration realism of aroma reconstruction models, and provides intuitive sensory style prediction and formula development efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117151A_ABST
    Figure CN122117151A_ABST
Patent Text Reader

Abstract

The application provides a fragrance reconstruction method, device, medium and product, relates to the technical field of perfume manufacturing, and the method comprises the following steps: obtaining target sample chemical analysis data and olfactory sensory data; determining a sensory contribution weight value according to a GC-O intensity value and a flavor dilution factor value; calculating a sensory contribution value by combining an olfactory threshold, a relative concentration and the sensory contribution weight value; converting an odor description text into a digital fragrance feature vector by using a natural language processing technology, and synthesizing a fragrance profile vector by weighting; based on a raw material database, iteratively calculating a raw material combination by using a multi-objective optimization genetic algorithm according to constraint conditions and optimization objectives of maximizing a predicted fragrance profile vector and the fragrance profile vector similarity and minimizing a cost, to obtain a Pareto optimal formula set. The scheme solves the technical problem of how to stably output a fragrance formula with optimal comprehensive indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fragrance manufacturing technology, specifically to a method, equipment, medium, and product for aroma reconstruction. Background Technology

[0002] Aroma reconstruction technology is a core R&D component in the fragrance and flavor and daily chemical industries. It aims to analyze the aroma composition of natural plants or specific samples and then use artificial means to recreate the formula in order to meet the market demand for high-quality, natural-smelling fragrance products.

[0003] In existing technologies, aroma reconstruction mainly employs a combination of instrumental analysis and human experience. Typically, gas chromatography-mass spectrometry (GC-MS) is used to detect the volatile substances in the target sample, obtaining the types of each chemical component and the proportion of chromatographic peak areas. Based on this chemical composition data, perfumers select corresponding fragrance monomers from a raw material library and perform preliminary formulation based on the detected relative contents of the substances. Subsequently, through multiple rounds of manual olfaction and fine-tuning, they attempt to simulate the target sample in terms of chemical composition, thereby completing the formulation design.

[0004] However, in practical applications, the chemical concentration of a substance is often not linearly correlated with its perceived olfactory intensity. Simulations based solely on the physical proportions of chemical detection data easily overlook the decisive role of trace high-threshold components or trace key components in the overall aroma. This results in reconstructed formulas that, while close to the target in chemical spectra, deviate from the actual sensory experience. Furthermore, when faced with complex synergistic or inhibitory effects in fragrance combinations, and when multiple objectives such as sensory fidelity and cost control need to be considered simultaneously, iterative trial-and-error methods dominated by human experience are insufficient to efficiently explore a vast array of possible combinations and cannot consistently output standardized formulas with optimal overall performance. Summary of the Invention

[0005] This application provides an aroma reconstruction method, apparatus, medium, and product to solve the technical problem of how to stably output aroma formulations with optimal comprehensive indicators in the prior art.

[0006] In a first aspect, this application provides an aroma reconstruction method, comprising: Acquire chemical analysis data and olfactory sensory data of the target sample, wherein the chemical analysis data includes the aroma components of the target sample and the relative concentration of each aroma component, and the olfactory sensory data includes the odor description text, GC-O intensity value and flavor dilution factor value of each aroma component. Based on the GC-O intensity value and the flavor dilution factor value, the sensory contribution weight value of each aroma component is determined; Obtain the olfactory threshold of each of the aroma components, and calculate the sensory contribution value of each of the aroma components based on the relative concentration, the olfactory threshold, and the sensory contribution weight value. Natural language processing techniques are used to convert the odor description text of each aroma component into a digital aroma feature vector; Based on the digital aroma feature vectors, the sensory contribution values ​​are weighted and synthesized to generate the aroma profile vector of the target sample. Based on a pre-set raw material database, a multi-objective optimization genetic algorithm is used to iteratively calculate the raw material combination according to the optimization objective and pre-set constraints to obtain a Pareto optimal formula set, and the Pareto optimal formula set is sent to the user. The optimization objective is to maximize the similarity between the predicted aroma profile vector corresponding to the raw material combination and the aroma profile vector, and to minimize the cost of the raw material combination.

[0007] Optionally, determining the sensory contribution weight value of each aroma component based on the GC-O intensity value and the flavor dilution factor value specifically includes: The target GC-O intensity value and target flavor dilution factor value of the target aroma component are normalized respectively, wherein the target aroma component is any one of the aroma components; By using a preset weighted fusion function, the normalized target GC-O intensity value and the normalized target flavor dilution factor value are weighted and calculated to obtain the basic sensory contribution weight value of the target aroma component. The sensory correction coefficients corresponding to the target aroma components are matched in a preset correction coefficient table, and the basic weight values ​​are adjusted and calculated using the sensory correction coefficients to obtain the sensory contribution weight values ​​of the target aroma components.

[0008] Optionally, after obtaining the Pareto optimal formula set, the method further includes: In response to the final target recipe selected by the user from the Pareto optimal recipe set, the final predicted aroma profile vector corresponding to the final target recipe is calculated; Calculate the cosine similarity between the final predicted aroma profile vector and each standard descriptor vector in the preset standard sensory descriptor vector library, where each standard descriptor vector corresponds to a standard descriptor. Select a target standard descriptor from the standard descriptor vector that is greater than the cosine similarity threshold, and generate predicted sensory description information based on the target standard descriptor. The predicted sensory description information is output to the user.

[0009] Optionally, the sensory contribution value of each aroma component is calculated based on the relative concentration, the olfactory threshold, and the sensory contribution weight value, specifically using the following formula: SC i =(C i / OT i )*W i ; Among them, SC i C represents the sensory contribution value of the i-th aroma component. i OT represents the relative concentration of the i-th aroma component. i W represents the olfactory threshold of the i-th aroma component. i The sensory contribution weight value of the i-th aroma component; The aroma profile vector of the target sample is generated by weighting and synthesizing the sensory contribution values ​​based on the digital aroma feature vectors, specifically using the following formula: P target =Σ(SC i *V i ); Among them, P target V represents the aroma profile vector. i This represents the digital aroma feature vector of the i-th aroma component.

[0010] Optionally, the constraints include at least one of the following: The formula normalization constraint is that the sum of the mass percentages of each raw material in the raw material combination equals 100%. The mass percentage of each raw material in the raw material combination is less than or equal to the regulatory safety constraints of the safety limit in the preset fragrance safety standard database; The key feature constraint in the raw material combination is that the key mass percentage is greater than a preset threshold, wherein the key mass percentage is the mass percentage corresponding to the key feature component; The non-negative constraint is that the mass percentage of each raw material in the raw material combination is greater than or equal to 0.

[0011] Optionally, the method further includes: The aroma components are sorted in descending order of their sensory contribution weight values, and the aroma components that rank at the top of the preset order are identified as the key feature components.

[0012] Optionally, the method further includes: Obtain the olfactory description text, unit cost, standard flavor dilution factor value, standard GC-O intensity value, and olfactory threshold of several monomer raw materials; Using the natural language processing technology, the olfactory description text of each of the monomer raw materials is converted into a digital aroma feature vector; The raw material database is formed by associating and storing the digital aroma feature vectors, the unit cost, the standard flavor dilution factor value, the standard GC-O intensity value, and the olfactory threshold.

[0013] In a second aspect, embodiments of this application provide an aroma reconstruction device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, which includes computer instructions, and the one or more processors call the computer instructions to cause the aroma reconstruction device to perform the method described in the first aspect and any possible implementation thereof.

[0014] Thirdly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on an aroma reconstruction device, cause the aroma reconstruction device to perform the method described in the first aspect and any possible implementation thereof.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on an aroma reconstruction device, cause the aroma reconstruction device to perform the method described in the first aspect and any possible implementation thereof.

[0016] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By adopting the above technical solution, this invention obtains chemical analysis data of the target sample and olfactory sensory data including odor description text, GC-O intensity value, and flavor dilution factor value, enabling the characterization of aroma from both physicochemical and sensory dimensions. The sensory contribution value is calculated by combining olfactory threshold and relative concentration, and natural language processing technology is introduced to convert the odor description text into a digital aroma feature vector, realizing the transformation of unstructured subjective sensory descriptions into calculable aroma profile vectors. Based on this, a multi-objective optimization genetic algorithm iteratively calculates in a pre-set raw material database, simultaneously balancing the conflicting objectives of maximizing the similarity between the predicted aroma profile vector and the target sample aroma profile vector, and minimizing the raw material combination cost. The resulting Pareto optimal formulation set effectively solves the problems of inaccurate sensory experience reproduction relying solely on chemical data and the difficulty in balancing reproduction accuracy and cost in single-objective optimization, thus achieving stable output of formulations with optimal comprehensive indicators under multi-dimensional constraints.

[0017] 2. By adopting the above technical solution, in determining the sensory contribution weight values, the target GC-O intensity value and the target flavor dilution factor value are first normalized to eliminate the scale differences of data with different dimensions and ensure data consistency. Then, the basic sensory contribution weight values ​​are calculated using a preset weighted fusion function, preliminarily quantifying the sensory importance of each aroma component. More importantly, the basic weight values ​​are adjusted using sensory correction coefficients matched in a preset correction coefficient table. This step fully considers the nonlinear sensory performance or synergistic effects of specific aroma components in complex systems. This progressive weight correction mechanism allows the final determined sensory contribution weight values ​​to more accurately reflect the actual contribution of the target aroma components in the overall aroma profile, effectively improving the calculation accuracy and fidelity of the subsequent aroma reconstruction model.

[0018] 3. By adopting the above technical solution, after generating the Pareto optimal formula set, the system can respond to the user's selection, lock in the final target formula, and calculate its corresponding final predicted aroma profile vector. This vector is compared with a preset standard sensory descriptor vector library, and the cosine similarity between it and each standard descriptor vector is calculated, thus mapping the abstract mathematical vector back to a concrete sensory concept. By selecting target standard descriptors with a similarity greater than the cosine similarity threshold to generate predicted sensory description information, this solution can transform the digital characteristics of the formula into standardized language that users can understand. This process allows users to intuitively predict the sensory style of the formula before actual blending, which not only verifies the accuracy of the mathematical model but also provides an intuitive and standardized decision-making basis for formula selection and confirmation, improving the efficiency and certainty of formula development.

[0019] 4. By adopting the above technical solution, multiple constraints are introduced into the iterative process of the multi-objective optimization genetic algorithm, ensuring the engineering usability and compliance of the generated formula. Formula normalization and non-negativity constraints guarantee the physical rationality of the raw material combination; regulatory safety constraints, by comparing with the fragrance safety standard database, directly eliminate formula combinations with potential safety hazards, ensuring product compliance; key feature constraints force the mass percentage of key feature components to be maintained at a specific level, preventing the algorithm from losing the core aroma characteristics of the product in pursuit of low cost. These constraints work together to limit the algorithm's search space, ensuring that the final output formula not only theoretically meets the optimization objectives of similarity and cost, but also directly conforms to the actual standards, regulatory requirements, and sensory quality baselines of industrial production, eliminating the need for cumbersome post-processing corrections. Attached Figure Description

[0020] Figure 1 This is a schematic flowchart of an aroma reconstruction method in an embodiment of this application; Figure 2This is another schematic flowchart of the aroma reconstruction method in the embodiments of this application; Figure 3 This is a schematic diagram of the physical device structure of an aroma reconstruction device in the embodiments of this application. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0022] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0023] In the description of the embodiments of this application, the term "multiple" means two or more. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first," "second," or "third" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized. It should be noted that all data collection in this scheme is conducted after obtaining user consent.

[0024] This application provides an aroma reconstruction method, referencing... Figure 1 , Figure 1 This is a flowchart of an aroma reconstruction method provided in an embodiment of this application. The method includes: Step S101: Obtain chemical analysis data and olfactory sensory data of the target sample. The chemical analysis data includes the aroma components of the target sample and the relative concentration of each aroma component. The olfactory sensory data includes the odor description text, GC-O intensity value and flavor dilution factor value of each aroma component. Among them, chemical analysis data refers to the qualitative identification results and quantitative values ​​of the chemical composition of the target sample obtained by instrumental detection methods, such as a list containing retention time, mass spectrum and compound name obtained after gas chromatography-mass spectrometry (GC-MS) acquisition and processing; aroma components refer to single chemical substances present in the sample that are volatile and can stimulate olfactory receptors to produce odor perception, such as isoamyl acetate, linalool or phenylethanol identified by mass spectrometry library search; relative concentration refers to the proportion of a specific aroma component in the total sample or total detected substances, such as the percentage of a component in the total peak area calculated by the chromatographic peak area normalization method; olfactory sensory data refers to the data recorded by the olfactory evaluator during the gas chromatography-olfactory measurement process regarding each ellipse. The sensory characteristics information of the components include a comprehensive record sheet containing odor feature descriptions, intensity scores, and dilution tolerance; odor description text refers to the words or phrases used by the odor assessor to qualitatively characterize the smelled odor, such as recorded text information like "grassy," "roasted nut," or "citrus-like"; GC-O intensity value refers to the quantitative score given by the odor assessor to the instantaneous burst or intensity of a specific aroma component according to a preset intensity evaluation standard, such as a value of 3 or 4 recorded under a five-point scale; flavor dilution factor value refers to the maximum dilution factor at which a specific aroma component can still be perceived by smell after multiple stages of solvent dilution in aroma extract dilution analysis, such as the value of 128, the aroma threshold potential of a representative substance determined in a stepwise dilution experiment. It should be noted that the dual-channel complementary extraction unit refers to a sample pretreatment system containing a first channel (SBSE unit) and a second channel (DTD unit); the enrichment component refers to the physical carrier used in this unit to carry or adsorb volatile substances of the target sample, specifically including a stirring rod in the first channel with an adsorption coating such as polydimethylsiloxane, and a thermal desorption tube in the second channel for directly loading solid plant samples; the thermal desorption-cold trap focusing injection system refers to the interface device used to receive the above-mentioned enrichment component and introduce the components into the chromatographic column without thermal discrimination through heating desorption and low-temperature cold trap focusing; the mass spectrometer detector refers to the analytical instrument used to obtain qualitative and relative quantitative information of substances; the olfactory detection port refers to the sensory evaluation device for olfactory judges to perceive the odor of the airflow in real time; and the aroma extract dilution analysis method refers to the method of determining the contribution of key aroma components by gradient dilution of samples.

[0025] Specifically, this step is performed during the data foundation construction phase of aroma reconstruction, covering the entire process from sample pretreatment to signal acquisition and correlation. To obtain chemical analysis and olfactory sensory data of the target sample, the target sample is first enriched across the entire spectrum using a dual-channel complementary extraction unit: on the one hand, the first channel (SBSE unit) uses a stir bar as a carrier to adsorb and extract volatile and heat-sensitive top notes; on the other hand, the second channel (DTD unit) uses a thermal desorption tube filled with a solid sample to prepare for the extraction of high-boiling-point, non-volatile body and base notes. After enrichment, the enrichment component carrying the aroma substances (i.e., the stir bar that has completed the adsorption process or the thermal desorption tube loaded with the sample) is placed in a thermal desorption-cold trap focusing injection system. The system heats up according to a preset program to desorb the aroma components from the stir bar coating or the matrix in the thermal desorption tube, and then, after being focused by the cold trap, enters the chromatographic column for separation. The separated chromatographic eluent is simultaneously fed into two detection channels via a splitter to generate different data: the eluent enters the mass spectrometer detector, where the instrument records the mass spectrum and retention index to identify the chemical structure of the aroma components and calculates the relative concentration of each aroma component based on the detected peak area; simultaneously, the eluent enters the olfactory detection port, where an olfactory evaluator records the odor description text and GC-O intensity value corresponding to each eluent peak in real time. Subsequently, specialized software is used to precisely align the total ion chromatogram and the olfactory signal graph on the time axis, establishing the correspondence between chemical signals and sensory signals; finally, aroma extract dilution analysis is applied, and the extract obtained through the above enrichment components is serially diluted and the olfactory process is repeated to determine the dilution factor of each aroma component at the sensory threshold, thereby obtaining the flavor dilution factor value.

[0026] The aforementioned scheme for acquiring chemical analysis and olfactory sensory data effectively combines the advantages of stir bar adsorption extraction (SBSE) for high-capacity enrichment of low-boiling-point, volatile top notes with the ability of direct thermal desorption (DTD) for direct extraction of high-boiling-point, heat-sensitive, or non-volatile body and base notes by employing a dual-channel complementary extraction technique. This overcomes the extraction discrimination or capacity limitations inherent in conventional single pretreatment methods, achieving full-spectrum capture of aroma components from the target sample without blind spots. Furthermore, the thermal desorption-cold trap focusing injection system eliminates the interference of solvent peaks on early eluting components through solvent-free injection and utilizes low-temperature secondary focusing in the cold trap. The effect compresses the desorbed broad peaks back into extremely narrow bands at the column head, significantly improving peak quality and greatly enhancing the detection sensitivity for trace key aroma compounds. Furthermore, through the integrated application of simultaneous mass spectrometry and olfactory dual-channel detection and aroma extract dilution analysis (AEDA), precise alignment and quantitative correlation between chemical structural information and sensory olfactory signals are achieved. This enables the accurate screening of key aroma components with low concentrations but significant odor contributions from complex mixtures, eliminating background interference from high concentrations of odorless substances. This provides comprehensive, accurate, and sensory-evidence-based data support for the subsequent construction of highly reproducible aroma reconstruction formulations.

[0027] Step S102: Determine the sensory contribution weight value of each aroma component based on the GC-O intensity value and the flavor dilution factor value. Specifically, this step aims to reduce and fuse multi-dimensional heterogeneous sensory data to provide a unified evaluation standard for subsequent screening of key aroma components. Since single evaluation indicators often have limitations—for example, a simple GC-O intensity value primarily reflects the instantaneous subjective perception intensity at a specific concentration, easily affected by olfactory saturation or individual differences; while a simple flavor dilution factor value reflects a substance's dilution tolerance and aroma threshold potential, it may ignore the saturation effect of high-concentration, low-threshold substances in the actual system—it is necessary to combine the two. In this step, the system reads the GC-O intensity value and flavor dilution factor value data corresponding to each aroma component. The GC-O intensity value is used as an explicit variable characterizing the sensory impact of a substance, and the flavor dilution factor value is used as a implicit variable characterizing the aroma potential of a substance. Through preset logical rules or data fusion strategies, the data from these two dimensions are comprehensively calculated or correlated. This process transforms raw sensory data of different dimensions and physical meanings into a comprehensive index that can simultaneously reflect the intensity and persistence of a substance's aroma, namely the sensory contribution weight value, through numerical transformation. This quantifies the relative importance of each aroma component in reconstructing the target aroma, ensuring that the final calculated weight value includes both the direct intensity information of the substance in the original solution and its dilution resistance characteristics.

[0028] Step S103: Obtain the olfactory threshold of each of the aroma components, and calculate the sensory contribution value of each of the aroma components based on the relative concentration, the olfactory threshold and the sensory contribution weight value. Among them, the olfactory threshold refers to the minimum concentration of a substance that can be perceived by the human olfactory organs in a specific medium. For example, the olfactory detection threshold of vanillin measured in water is 0.02 μg / L, or the threshold of ethyl acetate measured in ethanol solution is 5 mg / L. The relative concentration refers to the proportion of a specific aroma component in the target sample to the total content of the detected substances, or the relative content value calculated based on the internal standard method. For example, the percentage of a certain ester compound in the total volatile substances calculated by the gas chromatography peak area normalization method is 15.5%. The sensory contribution weight value refers to the dimensionless coefficient used to characterize the importance of aroma components, which is calculated based on the GC-O intensity value and flavor dilution factor value in the previous step. For example, the value obtained after normalization is 0.85. The sensory contribution value refers to the comprehensive quantitative index reflecting the actual contribution of the component to the overall aroma after combining the physicochemical content of the substance, the physiological perception threshold, and the sensory evaluation weight. For example, the value obtained after correction by aroma activity value (OAV) is 500 or 1200.

[0029] Step S103 is executed after determining the basic physicochemical properties (relative concentration) of each aroma component and the sensory importance coefficient (sensory contribution weight value) calculated in the previous step. It aims to combine the objective content of a substance with its physiological perception limits and introduce weight corrections to obtain a quantitative indicator that better reflects real sensory experience. Specifically, concentration alone cannot determine the aroma contribution of a substance because different substances have vastly different olfactory thresholds; a low-concentration, low-threshold substance may contribute more than a high-concentration, high-threshold substance. In this step, the calculation system first obtains the olfactory threshold of each aroma component in its corresponding matrix by consulting a standard olfactory threshold database or through sensory group experiments. Then, the system executes the calculation logic, first calculating the ratio of relative concentration to olfactory threshold, i.e., the aroma activity value (OAV), which reflects the multiple by which the substance concentration exceeds its threshold; then, it uses the sensory contribution weight values ​​obtained in the previous step to weight and correct this ratio. This calculation process corrects the shortcomings of traditional aroma activity values, which only consider concentration and threshold and ignore dynamic changes in olfactory intensity and dilution resistance. The final output sensory contribution value can accurately quantify the core role of each aroma component in the overall flavor system. For example, components with high concentration but weak odor are given a lower contribution value, while components with trace amounts but characteristic aroma and dilution resistance are given a higher contribution value.

[0030] Optionally, the sensory contribution value can be calculated using the following formula: SC i =(C i / OTi )*W i ; Among them, SC i C represents the sensory contribution value of the i-th aroma component. i OT represents the relative concentration of the i-th aroma component. i W represents the olfactory threshold of the i-th aroma component. i The sensory contribution weight value of the i-th aroma component; Step S104: Use natural language processing technology to convert the odor description text of each aroma component into a digital aroma feature vector; Natural language processing (NLP) technology refers to the technical means of semantic analysis, understanding, and conversion of human natural language text using computer science and artificial intelligence algorithms, such as using Word2Vec, BERT, or TF-IDF algorithms to extract features from text data; odor description text refers to the qualitative word sequence used by olfactory evaluators to record the sensory characteristics of aroma components during olfactory evaluation, such as strings like "having a strong toasted bread aroma," "having a fresh, freshly cut grass aroma," or "sweet fruit aroma similar to ripe pineapple"; digital aroma feature vectors refer to numerical arrays with mathematical operation capabilities generated after mapping unstructured text descriptions to a high-dimensional numerical space, such as mapping "apple aroma" to a vector containing 128 floating-point numbers [0.12, -0.56, 0.89, ...]; aroma components refer to volatile monomers in a sample that have been identified and have specific chemical structures and odor characteristics, such as hexanal, ethyl butyrate, or rose oxide.

[0031] Step S104 is executed after obtaining the qualitative aroma description text of each aroma component. It aims to address the problem that textual descriptions cannot directly participate in subsequent mathematical modeling and numerical calculations, transforming semantic information into a computer-processable mathematical form. Specifically, human descriptions of aromas typically rely on empirical adjectives; this textual data is unstructured and cannot be directly used for quantitative synthesis or similarity calculation. In this step, the processing unit calls a pre-trained natural language processing model or constructs a dedicated aroma semantic space model to perform preprocessing operations such as word segmentation and stop word removal on each aroma description text. Subsequently, using word embedding technology or semantic encoding algorithms, the processed text is converted into a fixed-dimensional real-valued vector, i.e., a digitized aroma feature vector. The direction and position of this vector in multi-dimensional space represent the semantic features of the aroma. For example, in vector space, the vector distance between "lemon" and "citrus" is closer than the vector distance between "lemon" and "coffee." Through this transformation, the originally discrete language symbols become mathematical entities that can be added, subtracted, multiplied, and divided, providing a mathematical basis for the subsequent weighted synthesis of the characteristics of different aroma components.

[0032] Step S105: Based on each of the digital aroma feature vectors, the sensory contribution values ​​of each of the digital feature vectors are weighted and synthesized to generate the aroma profile vector of the target sample. Among them, the digital aroma feature vector refers to a high-dimensional numerical array representing the semantic features of aroma obtained by converting the aroma description text through natural language processing technology, such as the dimension vector representing the "floral" feature; the sensory contribution value refers to a scalar value representing the contribution of a specific aroma component to the overall flavor, calculated by comprehensively considering concentration, threshold and sensory weight, such as the value 25.6 or 108.2; the aroma profile vector refers to the comprehensive high-dimensional numerical vector that represents the overall flavor characteristics of the target sample, such as a final vector coordinate that can characterize the overall perception result of "rich strawberry jam flavor".

[0033] Step S105 is executed after the digital mapping of aroma semantics (obtaining digital aroma feature vectors) and the quantitative calculation of the contribution of each component (obtaining sensory contribution values). Its aim is to reconstruct the overall digital aroma profile of the target sample through mathematical synthesis. Specifically, the feature vector of a single aroma component only represents a local odor attribute, while the overall aroma of the target sample is a comprehensive result of the combined effects of all components. In this step, the calculation unit uses the sensory contribution value of each aroma component as a scalar weight and multiplies it with the corresponding digital aroma feature vector. Geometrically, this process involves lengthening or shortening the modulus of the feature vector according to the magnitude of the contribution, ensuring that key aroma components with high contributions dominate the synthesized vector. Subsequently, the system performs vector addition on all weighted feature vectors to generate a final synthesized vector, i.e., the aroma profile vector. The aroma profile vector accurately locates the overall flavor attributes of the target sample in a multidimensional semantic space. It contains information on "what it tastes like" (determined by the vector direction) and also implies information on "which components it is mainly composed of" (determined by the weighted synthesis process). This enables the transformation of complex mixed aromas into a unique digital fingerprint that can be recognized and compared by a computer.

[0034] Optionally, the aroma profile vector can be calculated using the following formula; P target =Σ(SC i *V i ); Among them, P target V represents the aroma profile vector. i This represents the digital aroma feature vector of the i-th aroma component.

[0035] Specifically, assuming the target sample contains n aroma components, the system will perform a full-weighted synthesis calculation on all these aroma components. To illustrate the calculation logic, three representative components are selected for explanation: component A as the main aroma, component B as the modifying aroma, and component C as a trace background aroma. The calculation processing unit first retrieves the digital aroma feature vectors corresponding to each component generated in step S104. Since this digital aroma feature vector is a high-dimensional vector (assumed to be m-dimensional) generated based on a word embedding model, it is represented here in the form of "first dimension, second dimension, ..., last dimension". The digital aroma feature vector of component A is [0.80, 0.10, ..., 0.05], the digital aroma feature vector of component B is [0.20, 0.90, ..., 0.10], and the digital aroma feature vector of component C is [0.05, 0.05, ..., 0.80]. At the same time, the system retrieves the sensory contribution values ​​of each component calculated in step S103, where the sensory contribution value of component A is 100, the sensory contribution value of component B is 50, and the sensory contribution value of component C is 10.

[0036] Next, the computational processing unit performs linear weighted operations according to the formula. The system first performs scalar multiplication for each component: multiplying the sensory contribution value of component A (100) with its digitized aroma feature vector to obtain a weighted vector [80.0, 10.0, ..., 5.0]; multiplying the sensory contribution value of component B (50) with its digitized aroma feature vector to obtain a weighted vector [10.0, 45.0, ..., 5.0]; and multiplying the sensory contribution value of component C (10) with its digitized aroma feature vector to obtain a weighted vector [0.5, 0.5, ..., 8.0]. Subsequently, the system performs vector addition, summing the weighted vectors of all components in the sample along each corresponding dimension. For the first dimension, calculate 80.0 + 10.0 + 0.5 to get 90.5; for the second dimension, calculate 10.0 + 45.0 + 0.5 to get 55.5; for the omitted dimensions, add them up in the same way; for the last dimension, calculate 5.0 + 5.0 + 8.0 to get 18.0.

[0037] Finally, the system outputs the aroma profile vector of the target sample as [90.5, 55.5, ..., 18.0]. This aroma profile vector is an m-dimensional vector with the same dimensions as the original digitized aroma feature vector, which integrates information from all components in the sample in mathematical space. The first dimension (90.5), with its higher value, reflects the dominant role of component A in the overall flavor, while the last dimension (18.0), which retains the unique background characteristics of trace component C, preserves this unique feature. Through this multi-dimensional cumulative calculation, the aroma profile vector of the sample can accurately record complete sensory information from the main aroma to the subtle nuances, achieving a digital holographic mapping of the flavor characteristics of complex mixtures.

[0038] Step S106: Based on the preset raw material database, a multi-objective optimization genetic algorithm is used to iteratively calculate the raw material combination according to the optimization objective and preset constraints to obtain the Pareto optimal formula set, and the Pareto optimal formula set is sent to the user. The optimization objective is to maximize the similarity between the predicted aroma profile vector corresponding to the raw material combination and the aroma profile vector, and to minimize the cost of the raw material combination. The pre-defined raw material database refers to a digital set storing detailed information on N available flavorings (monomers or natural extracts). Each record includes the raw material name, digital aroma feature vector, and unit price price_i. For example, the database records the digital aroma feature vector, olfactory threshold, standard GC-O intensity value, standard flavor dilution factor value, and unit price of the i-th flavoring. The raw material combination refers to the decision variable vector X=[x_1, x_2, ..., x_N] in the multi-objective optimization process. This vector consists of the mass percentage x_i of the N available flavorings, used to represent the specific component composition of a candidate formulation. For example, X=[10%, 20%, ..., 0%] represents the proportion of each component. The predicted aroma profile vector refers to the simulated aroma profile P calculated based on the current formulation vector X using the same calculation method as the aroma profile vector. sim (X) represents the theoretical sensory performance of the recipe; the optimization objective refers to the two mathematical functions that the algorithm needs to optimize simultaneously, specifically including maximizing sensory consistency (i.e., maximizing P). sim (X) and the target aroma profile P target The similarity between the two (cosine similarity is used in this method) and minimizing the total cost; the preset constraints refer to the set of mathematical boundaries or logical restrictions that the decision variable vector X must satisfy when searching the solution space, in order to ensure the physical and regulatory feasibility of the generated recipe; the Pareto optimal recipe set refers to a set of non-dominated solutions obtained by the multi-objective evolutionary algorithm, and each recipe in the set represents an optimal trade-off between sensory consistency and cost.

[0039] Specifically, this step aims to address the "reconstruction fault" problem and is executed after the aroma profile vector of the target sample is generated. The computational processing unit employs a multi-objective evolutionary algorithm (such as NSGA-II) to iteratively search the decision variable vector X within a solution space that satisfies preset constraints. During the evaluation of each generation of the population, the system first, based on the current raw material combination vector X, calls the digitized aroma feature vectors, olfactory thresholds, standard GC-O intensity values, standard flavor dilution factor values, and unit prices of each raw material in vector X from the database. It then uses the same computational process as for the aroma profile vector of the target sample to calculate the predicted aroma profile vector P. sim (X); Subsequently, the system performs a dual calculation based on the optimization objective function: on the one hand, it calculates the similarity value between the predicted aroma profile vector and the aroma profile vector of the target sample using metrics such as cosine similarity; on the other hand, it accumulates the costs of each component to obtain the total cost Cost(X). Based on the above two indicators, the system performs non-dominated sorting to screen for the Pareto front. In this computational example, it is assumed that the algorithm generates three candidate formulation vectors that satisfy the constraints for the aroma profile vector of the target sample: the predicted profile P of formulation X1. sim Formula X1 has a similarity of 0.98 with the aroma profile vector of the target sample, with a total cost of 300 yuan; formula X2 has a similarity of 0.90, with a total cost of 100 yuan; and formula X3 has a similarity of 0.85, with a total cost of 350 yuan. The system found that formula X3 has lower sensory consistency (0.85 < 0.90) and higher cost (350 > 100) compared to formula X2, therefore formula X3 was determined to be a dominated solution and eliminated. While formula X1 has a higher cost, it is optimal in sensory consistency, and formula X2, although slightly less consistent, has a significant cost advantage; the two do not dominate each other. Finally, the system outputs a Pareto-optimal formula set containing formulas X1 and X2, which demonstrates the optimal solutions under different trade-offs.

[0040] Optionally, the preset constraints mentioned in step S106 may, in specific implementations, include at least one of the following: formulation normalization constraints, regulatory safety constraints, key feature constraints, and non-negativity constraints.

[0041] The formula normalization constraint is that the sum of the mass percentages of each raw material in the raw material combination equals 100%. The mass percentage of each raw material in the raw material combination is less than or equal to the regulatory safety constraints of the safety limit in the preset fragrance safety standard database; The key feature constraint in the raw material combination is that the key mass percentage is greater than a preset threshold, wherein the key mass percentage is the mass percentage corresponding to the key feature component; The non-negative constraint is that the mass percentage of each raw material in the raw material combination is greater than or equal to 0.

[0042] First, it needs to be clarified that in the mathematical model definition of this scheme, the variable x_i specifically refers to the mass percentage of the i-th spice ingredient in the raw material combination vector X. It is the basic decision variable constituting the solution space of the formula. The following section provides a detailed explanation of each constraint condition in conjunction with the definition of x_i: Specifically, the formula normalization constraint refers to strictly limiting the sum of all components x_i in the raw material combination vector, that is, requiring that the sum of the mass percentages of all usable spices in the formula must be equal to 100%. This constraint not only conforms to the objective laws of physical formulation, ensuring that the generated formula is a complete percentage scheme that can be directly used for production input, but also defines a clear hyperplane search space for the multi-objective evolutionary algorithm, preventing the algorithm from generating invalid formula data with overflow or deficiency in total content.

[0043] The aforementioned regulatory safety constraints aim to ensure that the generated fragrance formulations comply with industry safety standards. Specifically, this is manifested in setting an upper limit on the mass percentage x_i of each individual fragrance component. The system will call a preset fragrance safety standard database (e.g., a database established based on the standards of the International Federation of Fragrance and Perfume Associations), retrieve the safety limit corresponding to the i-th ingredient in the formulation, and enforce that the mass percentage of each ingredient in the ingredient combination must be less than or equal to the preset safety limit. This constraint ensures that the intelligently reconstructed formulation not only has a similar odor but also fully meets the requirements of practical applications in terms of compliance and safety.

[0044] The key feature constraints are designed to ensure that the reconstructed formula retains the aroma or core structure of the target sample. The system pre-identifies the key aroma-producing components (i.e., the k-th ingredient) that play a decisive role in the overall aroma and sets corresponding mass percentage thresholds. During optimization, the algorithm must ensure that the mass percentage x_k corresponding to these key components in the ingredient combination is strictly greater than the preset threshold. This constraint prevents the algorithm from reducing certain expensive but crucial trace components to an imperceptible level simply to pursue mathematically low cost or overall contour similarity, thus ensuring the effective restoration of characteristic aromas.

[0045] The non-negativity constraint is a mathematical boundary condition set based on the physical foundation of matter, requiring that the mass percentage x_i of each raw material in the raw material combination must be greater than or equal to 0. This constraint mathematically excludes meaningless solutions with negative mass, ensuring that the value of each dimension has actual physical meaning when the algorithm searches the solution space. That is, there are only cases of "adding" or "not adding" a raw material, and there is no impossible case of "removing a raw material with negative content".

[0046] Optionally, the key feature components can be determined by: sorting the aroma components in descending order of their sensory contribution weight values, and determining the aroma components that rank first in the preset order as the key feature components. Specifically, the system executes sorting and filtering logic, arranging all aroma components in descending order according to their calculated sensory contribution weight values. Based on a pre-set cutoff number (i.e., the "preset quantity," such as selecting the top 5 or top 10 components), the top-ranked aroma components are marked as key feature components. In this way, the algorithm automatically identifies the core skeletal components that decisively support the overall aroma, thus providing a clear target for setting the constraint condition of "key quality percentage greater than a preset threshold" in subsequent steps. This maximizes the accuracy of the reconstructed formula in preserving the core aroma characteristics of the target sample from an algorithmic perspective.

[0047] The following is a more detailed description of the process of the method provided in this implementation.

[0048] Optionally, steps S10201-S10203 are more specific steps than step S102. Step S10201: Normalize the target GC-O intensity value and target flavor dilution factor value of the target aroma component, wherein the target aroma component is any one of the aroma components. Specifically, since target GC-O intensity values ​​are typically obtained based on a linear scale (e.g., a score of 1 to 10), while target flavor dilution factor values ​​are typically obtained based on an exponential scale (e.g., 2 to the power of n, i.e., 4, 8, 16, etc.), their physical dimensions and numerical magnitudes differ significantly. Direct calculation would lead to the larger numerical value dominating the calculation results. Therefore, the system needs to perform dimensionless processing on these two indicators for each target aroma component. Taking hexyl acetate as an example, if its target GC-O intensity value measured in the experiment is 7 (range 0-10), and its target flavor dilution factor value is 128 (maximum dilution factor of 1024), the system will map the intensity value of 7 to 0.7 (assuming linear normalization). At the same time, for FD values ​​that exhibit exponential changes, the system usually first takes its logarithm or linearizes it according to the dilution order, then maps it to the 0-1 interval, converting 128 into the corresponding normalized value. For example, for linalool, if its intensity value is 4 and its FD value is 16, the system will also convert it into normalized values ​​such as 0.4 and 0.15 based on the maximum and minimum values ​​of the full sample data, thereby eliminating the influence of data dimensions and providing standardized input data for subsequent weighted calculations.

[0049] Step S10202: Using a preset weighted fusion function, the normalized target GC-O intensity value and the normalized target flavor dilution factor value are weighted to obtain the basic sensory contribution weight value of the target aroma component. The preset weighted fusion function refers to a predefined linear or nonlinear combination formula containing weight coefficients w1 and w2, such as W=w1×A+w2×B; w1 represents the importance ratio coefficient of the normalized target GC-O intensity value; w2 represents the importance ratio coefficient of the normalized target flavor dilution factor value; the basic sensory contribution weight value refers to the comprehensive value used to initially characterize the importance of the target aroma components after combining information from two dimensions: direct olfactory intensity and dilution tolerance.

[0050] Specifically, step S10202 is the data fusion operation performed after data standardization. Before this step, the system needs to determine the values ​​of the weighting coefficients w1 and w2. The determination of these two coefficients typically employs either the Analytic Hierarchy Process (AHP) or Principal Component Analysis (PCA). If PCA is used, the system calculates the variance contribution rate of the two indicators (i.e., the proportion of information carried by the fluctuation of the indicator data in the total information) through statistical analysis of a large amount of historical flavor formulation data, and uses this as the objective weighting coefficients w1 and w2. For example, based on historical data statistical analysis, if the fluctuation of the intensity value data provides 55% of the effective information (i.e., the variance contribution rate is 55%) when distinguishing aroma characteristics, while the FD value provides 45% of the effective information, then w1 = 0.55 and w2 = 0.45 respectively. After determining the weights, the system performs a weighted calculation: For citral with a normalized intensity value of 0.8 and a normalized FD value of 0.3, with w1=0.6 and w2=0.4, its basic sensory contribution weight is calculated as 0.8×0.6+0.3×0.4=0.6. Similarly, for vanillin with a normalized intensity value of 0.5 and a normalized FD value of 0.9, with the same weight settings, its basic sensory contribution weight is calculated as 0.5×0.6+0.9×0.4=0.66. In this way, the system obtains a comprehensive evaluation index that takes into account both aroma intensity and potential.

[0051] Step S10203: Match the sensory correction coefficient corresponding to the target aroma component in the preset correction coefficient table, and use the sensory correction coefficient to adjust and calculate the basic weight value to obtain the sensory contribution weight value of the target aroma component. Among them, the preset correction coefficient table refers to a structured dataset stored in the database that establishes a mapping relationship between the identity of flavor chemical substances and their behavioral characteristics in the mixture system; the sensory correction coefficient is a dynamic weighting factor jointly determined by the target flavor dilution factor value (FD value) and the known synergistic / inhibitory effects of the component in the mixture; the sensory contribution weight value of the target aroma component refers to the core indicator used to characterize the retention priority of the component in the formulation reconstruction after environmental difference compensation and interaction effect calibration.

[0052] Specifically, step S10203 is the final calibration step of the sensory evaluation model. First, regarding the source of the preset correction coefficient table, this table is constructed based on the physicochemical properties (such as threshold and vapor pressure) of a large number of individual aroma compounds and sensory interaction experimental data in binary or multi-component systems. Based on this, for each target aroma component, the system does not simply look up the table, but determines its sensory correction coefficient according to specific logic: this coefficient is jointly determined by the component's FD value (direct sensory importance from GC-O) and the component's known synergistic or inhibitory effects in the mixture.

[0053] For example, the ingredient "indole" often has an extremely high FD value in GC-O testing, but at high concentrations it is prone to producing an unpleasant fecal odor, classifying it as a "high FD but high risk" ingredient. Therefore, the system will adjust its sensory correction coefficient (e.g., set to 0.3) according to a preset inhibition effect logic to prevent the algorithm from deriving excessively high concentrations of formulations due to its high FD value.

[0054] For example, regarding the ingredient "geomyosin," although its FD value is extremely high, it would ruin the overall aesthetics in perfume applications if not modified in minute quantities. The system also sets its sensory correction coefficient to a low value (such as 0.1) based on the inhibition effect.

[0055] Conversely, for the ingredient "maltol", if a large number of ester components that produce a synergistic effect with it in the formula are detected, the system will increase its sensory correction coefficient (for example, set it to 1.5) according to the synergistic effect logic.

[0056] The system ultimately adjusts the base weight values ​​obtained in step S10202 using predetermined sensory correction coefficients (e.g., through multiplication). For example, if the base weight value of indole is 0.9, after correction by coefficient 0.3, the final sensory contribution weight value becomes 0.27; if the base weight value of maltol is 0.6, after correction by coefficient 1.5, the final weight value becomes 0.9. This step ensures that the calculation results accurately reflect the sensory status and safety of the substance in the final product.

[0057] Optional, see reference Figure 2After step S106, this scheme can also execute steps S107-S110; Step S107: In response to the final target formula selected by the user from the Pareto optimal formula set, calculate the final predicted aroma profile vector corresponding to the final target formula; Here, "user" refers to the formulation designer, perfumer, or related technical personnel operating this system; "Pareto optimal formulation set" refers to the set of formulation solutions output in step S106 that achieves Pareto optimality among multiple objective functions such as aroma similarity, formulation cost, and raw material complexity. Each formulation in this set cannot further optimize a single objective without weakening the performance of other objectives; "final target formulation" refers to the unique formulation scheme to be produced or verified selected by the user from the Pareto optimal formulation set based on actual application needs (such as cost constraints or raw material inventory); "final predicted aroma profile vector" refers to a multi-dimensional numerical vector calculated based on a mathematical model, used to digitally represent the comprehensive olfactory characteristics presented by the final target formulation. The dimension of this vector is consistent with the dimension of the aroma profile vector of the target sample.

[0058] Specifically, step S107 is executed after the system displays multiple optimized formulation schemes to the user. Since the Pareto optimal formulation set may contain various trade-offs such as "high similarity, high cost" or "moderate similarity, low cost," the system first receives the user's selection instruction through the human-computer interaction interface. For example, if the user selects the formulation numbered "Solution_ID_05," the system will lock it as the final target formulation. Subsequently, the system calls the exact same calculation process used to generate the aroma profile vector of the target sample to process the formulation. For example, if the final target formulation consists of linalyl acetate, geraniol, and phenylethyl alcohol, the system calculates a 10-dimensional final predicted aroma profile vector [0.1, 0.8, 0.2, 0.05, ...], where the second component, 0.8, represents an extremely high prediction intensity for the "rose fragrance" dimension in this formulation.

[0059] Step S108: Calculate the cosine similarity between the final predicted aroma profile vector and each standard descriptor vector in the preset standard sensory descriptor vector library, where each standard descriptor vector corresponds to a standard descriptor. The pre-defined standard sensory descriptor vector library refers to a standard reference set that is pre-built and stored in the system database. This library contains a large number of industry-standard aroma types (such as "sweet orange", "jasmine", and "grassy"). The standard descriptor vector refers to the digital expression of each standard aroma type in the same multi-dimensional sensory space. For example, the descriptor "jasmine" corresponds to a specific feature vector, which has high values ​​in the dimensions of "indole" and "floral". The standard descriptor refers to the aroma adjective used in human natural language communication. Cosine similarity is a mathematical index that measures the directional consistency of two vectors by calculating the cosine value of the angle between them. The value range is usually from -1 to 1 (0 to 1 in non-negative space). The closer the value is to 1, the more similar the features of the two vectors are.

[0060] Specifically, step S108 involves matching the obtained digital features of the formula (final predicted aroma profile vector) with known standards. The system iterates through each entry in the standard sensory descriptor vector library, comparing them one by one. For example, suppose the final predicted aroma profile vector obtained in step S107 is vector A, and the library contains standard descriptor vectors B_1 representing "lemon-like" and B_2 representing "vanilla-like". The system calculates the similarity between A and B_1, and between A and B_2. If vector A has extremely high values ​​in the dimensions representing "limonene" and "citral" (i.e., exhibiting significant citrus top notes), and vector B_1 (lemon-like) also has extremely high values ​​in these dimensions, and their vector directions are highly consistent, the calculation result may show that the cosine similarity between A and B_1 is 0.92. Conversely, if the main numerical distribution of vector B_2 (vanilla-like) is on the dimension representing "vanillin" (i.e., exhibiting a sweet undertone), which is completely offset from the main feature dimension of vector A (i.e., the vectors are orthogonal or have a large angle), the calculation result may show that the cosine similarity between A and B_2 is only 0.15. The system will repeat this calculation process for hundreds or thousands of standard descriptors in the library to generate a similarity list containing all the comparison results.

[0061] Step S109: Select a target standard descriptor from the standard descriptor vector that is greater than the cosine similarity threshold, and generate predicted sensory description information based on the target standard descriptor; Among them, the cosine similarity threshold represents the numerical boundary used to determine whether two vectors are statistically significantly related. This value is usually set as a fixed value that has been verified by a large amount of historical data. The target standard descriptor refers to a specific sensory label in the standard sensory descriptor vector library whose similarity calculation result with the final predicted aroma profile vector meets the screening conditions. The predicted sensory description information refers to the semantic data generated by logically integrating, sorting or formatting one or more selected target standard descriptors, which is used to clearly characterize the sensory features of the final target formula.

[0062] Specifically, the system first determines the cosine similarity threshold used for this screening based on a preset strategy. The determination method can be as follows: the system sorts all similarity values ​​obtained in step S108 in descending order, selects the highest-ranked value as the benchmark, and sets 90% of this benchmark value as the cosine similarity threshold (e.g., if the highest similarity is 0.98, then the threshold is set to 0.882); or, the system directly reads a strictly defined fixed value (e.g., 0.90) from the database configuration parameters as the threshold. After determining the threshold, the system iterates through all comparison results, removing items below the threshold and retaining only the descriptors corresponding to items above the threshold. For example, assuming the threshold is determined to be 0.90, if the similarity of "rose-like" is 0.95, the similarity of "geranium-like" is 0.91, and the similarity of "lilac-like" is 0.85, the system will determine "rose-like" and "geranium-like" as the target standard descriptors, while ignoring "lilac-like". Next, the system generates predicted sensory description information based on these target standard descriptors. For example, the system can combine the selected "lemon-like" and "bergamot-like" to generate text information that "has significant lemon and bergamot characteristics"; or, if the selected target standard descriptors are "green," "grassy," and "leaf alcohol-like," the system can weight and sort these tags according to their similarity values ​​(e.g., 0.95 > 0.92 > 0.90) to generate structured description information that "the main fragrance note is green, accompanied by grassy and leaf alcohol-like characteristics."

[0063] Step S110: Output the predicted sensory description information to the user; Specifically, step S110 is the terminal interaction stage of the entire computer-aided design process. The system provides feedback on the prediction results to the user through a graphical user interface (GUI). The output can take various forms: for example, the system can pop up a text box in the center of the screen displaying "Predicted fragrance: sweet floral and fruity, with characteristics of green apple and jasmine"; or, the system can generate a word cloud, with "green apple" and "jasmine" in the largest font, intuitively reflecting their weight in the predicted sensory description information; or, the system can generate a radar chart, mapping the predicted sensory description information onto coordinate axes such as "floral," "fruity," and "woody," and marking the specific sensory scores. Through this step, users can know the expected sensory performance of the selected formula in advance on the computer without actually mixing samples and smelling them.

[0064] Optionally, the raw material database acquisition method is steps S111-S113; Step S111: Obtain the olfactory description text, unit cost, standard flavor dilution factor value, standard GC-O intensity value and olfactory threshold of several monomer raw materials; This step primarily involves the standardized collection of basic data, aiming to establish an objective set of raw material attributes. Among these, monomeric raw materials refer to specific chemical structures or standard natural extracts that constitute the fragrance formulation, such as isoamyl acetate and linalool; olfactory description text refers to the aroma characteristics of the raw material recorded according to standard terminology in the fragrance industry, focusing on qualitative classification rather than subjective perception. For example, a raw material might be described as having "rose-like and geranium-like aroma characteristics, accompanied by a waxy undertone," rather than subjective descriptions like "pleasant" or "fresh"; unit cost refers to the standard purchase price of the raw material within the current supply chain system; specifically, the system accesses professional fragrance raw material databases or connects to enterprise ERP systems through data interfaces to batch retrieve this standardized data. For example, for the raw material "citronellol acetate," the system obtains its standard description text as "rose-like, geranium-like, fruity," along with its unit cost value, standard flavor dilution factor value, standard GC-O intensity value, and olfactory threshold, ensuring the objectivity and traceability of the data.

[0065] Step S112: Using the natural language processing technology, the olfactory description text of each monomer raw material is converted into a digital aroma feature vector. The core of this step lies in using algorithms to transform qualitative textual descriptions into quantitative mathematical vectors, thereby achieving a numerical mapping of odor features. The natural language processing techniques used here should be consistent with those used in step S104.

[0066] In this step, the processing unit invokes a pre-trained natural language processing model or constructs a dedicated flavor semantic space model to perform preprocessing operations such as word segmentation and removal of irrelevant modifiers on the olfactory description text of each raw material. Subsequently, using natural language processing techniques, such as word embedding or semantic coding algorithms, the processed text is converted into fixed-dimensional real-valued vectors, i.e., digital aroma feature vectors. The direction and position of this vector in multidimensional space represent the semantic features of the raw material's aroma. For example, in vector space, the aroma vector corresponding to "isoamyl acetate" is closer to the concept vector of "banana" than to the concept vector of "smoked." Through this transformation, the originally discrete linguistic symbols become mathematical entities that can be subjected to addition, subtraction, multiplication, and division operations, providing a foundation for subsequent algorithms to quickly retrieve, match, and calculate the mathematical distance between raw materials and target aromas in the raw material database.

[0067] Step S113: The digital aroma feature vectors, the unit cost, the standard flavor dilution factor value, the standard GC-O intensity value, and the olfactory threshold are associated and stored to form the raw material database; Specifically, the system performs a data merging operation, writing the digital aroma feature vector calculated in step S112, along with the unit cost and olfactory threshold obtained in step S111, into the same data record. For example, a record indexed by the CAS number is created in the database, containing its high-dimensional digital aroma feature vector, unit cost, standard flavor dilution factor value, standard GC-O intensity value, and olfactory threshold. When the subsequent formula generation algorithm runs, the system can directly index this raw material database, simultaneously retrieving the digital aroma feature vectors of each raw material to calculate Euclidean distance or cosine similarity, retrieving the cost data of each raw material to calculate the total formula price, and retrieving the standard flavor dilution factor value, standard GC-O intensity value, and olfactory threshold of each raw material. Combined with concentration, the system calculates and predicts the aroma profile vector, thereby achieving multi-objective optimization iteration based on objective data.

[0068] The aroma reconstruction device in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the physical device structure of the aroma reconstruction device in the embodiments of this application.

[0069] It should be noted that, Figure 3 The structure of the aroma reconstruction device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.

[0070] like Figure 3As shown, the aroma reconstruction device includes a CPU 301, which can perform various appropriate actions and processes according to a program stored in the read-only memory ROM 302 or a program loaded from the storage section 308 into the random access memory RAM 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An I / O interface 305 is also connected to the bus 304.

[0071] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0072] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by CPU 301, it performs the various functions defined in the present invention.

[0073] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0074] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0075] Specifically, the aroma reconstruction device in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the aroma reconstruction method provided in the above embodiment.

[0076] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the aroma reconstruction apparatus described in the above embodiments; or it may exist independently and not assembled into the aroma reconstruction apparatus. The storage medium carries one or more computer programs that, when executed by a processor of the aroma reconstruction apparatus, cause the aroma reconstruction apparatus to implement the aroma reconstruction method provided in the above embodiments.

[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for aroma reconstruction, characterized in that, The method includes: Acquire chemical analysis data and olfactory sensory data of the target sample, wherein the chemical analysis data includes the aroma components of the target sample and the relative concentration of each aroma component, and the olfactory sensory data includes the odor description text, GC-O intensity value and flavor dilution factor value of each aroma component. Based on the GC-O intensity value and the flavor dilution factor value, the sensory contribution weight value of each aroma component is determined; Obtain the olfactory threshold of each of the aroma components, and calculate the sensory contribution value of each of the aroma components based on the relative concentration, the olfactory threshold, and the sensory contribution weight value. Natural language processing techniques are used to convert the odor description text of each aroma component into a digital aroma feature vector; Based on the digital aroma feature vectors, the sensory contribution values ​​are weighted and synthesized to generate the aroma profile vector of the target sample. Based on a pre-set raw material database, a multi-objective optimization genetic algorithm is used to iteratively calculate the raw material combination according to the optimization objective and pre-set constraints to obtain a Pareto optimal formula set, and the Pareto optimal formula set is sent to the user. The optimization objective is to maximize the similarity between the predicted aroma profile vector corresponding to the raw material combination and the aroma profile vector, and to minimize the cost of the raw material combination.

2. The method according to claim 1, characterized in that, The step of determining the sensory contribution weight value of each aroma component based on the GC-O intensity value and the flavor dilution factor value specifically includes: The target GC-O intensity value and target flavor dilution factor value of the target aroma component are normalized respectively, wherein the target aroma component is any one of the aroma components; By using a preset weighted fusion function, the normalized target GC-O intensity value and the normalized target flavor dilution factor value are weighted and calculated to obtain the basic sensory contribution weight value of the target aroma component. The sensory correction coefficients corresponding to the target aroma components are matched in a preset correction coefficient table, and the basic weight values ​​are adjusted and calculated using the sensory correction coefficients to obtain the sensory contribution weight values ​​of the target aroma components.

3. The method according to claim 1, characterized in that, After obtaining the Pareto optimal formulation set, the method further includes: In response to the final target recipe selected by the user from the Pareto optimal recipe set, the final predicted aroma profile vector corresponding to the final target recipe is calculated; Calculate the cosine similarity between the final predicted aroma profile vector and each standard descriptor vector in the preset standard sensory descriptor vector library, where each standard descriptor vector corresponds to a standard descriptor. Select a target standard descriptor from the standard descriptor vector that is greater than the cosine similarity threshold, and generate predicted sensory description information based on the target standard descriptor. The predicted sensory description information is output to the user.

4. The method according to claim 1, characterized in that, The sensory contribution value of each aroma component is calculated based on the relative concentration, the olfactory threshold, and the sensory contribution weight value, specifically using the following formula: SC i =(C i / OT i )*W i ; Among them, SC i C represents the sensory contribution value of the i-th aroma component. i OT represents the relative concentration of the i-th aroma component. i W represents the olfactory threshold of the i-th aroma component. i This represents the sensory contribution weight value of the i-th aroma component; The aroma profile vector of the target sample is generated by weighting and synthesizing the sensory contribution values ​​based on the digital aroma feature vectors, specifically using the following formula: P target =Σ(SC i *V i ); Among them, P target V represents the aroma profile vector. i This represents the digital aroma feature vector of the i-th aroma component.

5. The method according to claim 1, characterized in that, The constraints include at least one of the following: The formula normalization constraint is that the sum of the mass percentages of each raw material in the raw material combination equals 100%. The mass percentage of each raw material in the raw material combination is less than or equal to the regulatory safety constraints of the safety limit in the preset fragrance safety standard database; The key feature constraint in the raw material combination is that the key mass percentage is greater than a preset threshold, wherein the key mass percentage is the mass percentage corresponding to the key feature component; The non-negative constraint is that the mass percentage of each raw material in the raw material combination is greater than or equal to 0.

6. The method according to claim 5, characterized in that, The method further includes: The aroma components are sorted in descending order of their sensory contribution weight values, and the aroma components that rank at the top of the preset order are identified as the key feature components.

7. The method according to claim 1, characterized in that, The method further includes: Obtain the olfactory description text, unit cost, standard flavor dilution factor value, standard GC-O intensity value, and olfactory threshold of several monomer raw materials; Using the natural language processing technology, the olfactory description text of each of the monomer raw materials is converted into a digital aroma feature vector; The raw material database is formed by associating and storing the digital aroma feature vectors, the unit cost, the standard flavor dilution factor value, the standard GC-O intensity value, and the olfactory threshold.

8. An aroma reconstruction device, characterized in that, The aroma reconstruction device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the aroma reconstruction device to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on the aroma reconstruction device, the aroma reconstruction device performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on the aroma reconstruction device, the aroma reconstruction device performs the method as described in any one of claims 1-7.