Perfume formula composition and optimization method, system and equipment based on artificial intelligence and medium

By building multi-objective reward function and using pre-trained GNN model, and optimizing spice formulas with the Q-learning algorithm, the problem of difficulty in effectively using artificial intelligence in the existing technology is solved, and efficient, safe and economical spice formula research and development is achieved.

CN120220873AInactive Publication Date: 2025-06-27DADI HANKE BIOLOGICAL TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510290945.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively use artificial intelligence for the research and development of spice formulas, and there is a lack of an intelligent method that can combine the characteristics and costs of spices.

Method used

By defining the initial actions and states, a multi-objective reward function is constructed, combined with pre-trained GNN models to predict aroma intensity, and the spice formula is optimized using the Q-learning algorithm and the multi-objective reward function.

Benefits of technology

It achieves effective control of costs while ensuring aroma strength and durability, shortens R&D cycle, reduces experimental costs, and ensures the safety and compliance of the formula.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220873A_ABST
    Figure CN120220873A_ABST
Patent Text Reader

Abstract

The invention discloses a spice formula composition and optimization method, system and device based on artificial intelligence and a medium, and the method comprises the steps: defining an initial action and an initial state, and constructing a multi-target reward function; generating and executing a new action to obtain a state corresponding to the new action; predicting the aroma intensity corresponding to the state based on a pre-trained GNN model; determining a reward value in the state through a multi-target reward function; if the reward value of the state is greater than the reward value corresponding to the initial state, keeping the state; otherwise, keeping the current state; and S15, repeatedly executing the steps S12-S15, and when a preset termination condition is met, outputting a state corresponding to the maximum reward value. The invention belongs to the field of formula synthesis. According to the method, the aroma characteristics and intensity are predicted by combining the pre-trained graph neural network (GNN) model, and the formula performance and cost are comprehensively considered by using the multi-target reward function, so that the cost is effectively controlled while the aroma intensity and durability are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of formula synthesis, and particularly to a method, system, device, and medium for composing and optimizing fragrance formulas based on artificial intelligence. Background Art

[0002] A fragrance formula refers to combining multiple fragrance raw materials in a specific proportion to create a final product with unique fragrance and characteristics. Fragrance formulas can be applied to various fields such as perfumes, food flavorings, and cosmetics. Designing an excellent fragrance formula requires not only profound chemical knowledge but also an in-depth understanding of the characteristics of various fragrances, including their aroma types (such as floral, fruity, woody, etc.), volatility, and their performance under different conditions.

[0003] With the development of modern technology, how to utilize artificial intelligence technology for fragrance formula research and development is a hot topic of discussion. Therefore, the present invention proposes a method for composing and optimizing fragrance formulas based on artificial intelligence. Summary of the Invention

[0004] The present invention provides a method, system, device, and medium for composing and optimizing fragrance formulas based on artificial intelligence, solves the technical problem of how to use artificial intelligence for fragrance formula research and development in the prior art, and realizes the technical effect of combining artificial intelligence to achieve fragrance formula research and development.

[0005] In a first aspect, the present invention provides a method for composing and optimizing fragrance formulas based on artificial intelligence, the method comprising:

[0006] S11, defining an initial action and an initial state, and constructing a multi-objective reward function, wherein the action includes formula components and ratios, environmental parameters, and performance requirements, the state is used to adjust the compound formula according to a preset rule, and the multi-objective reward function includes formula performance and formula cost;

[0007] S12, generating and executing a new action to obtain the state corresponding to the new action;

[0008] S13, predicting the aroma intensity corresponding to the state based on a pre-trained GNN model, wherein the aroma intensity belongs to the formula performance;

[0009] S14, determining the reward value in this state by the multi-objective reward function;

[0010] S15, if the reward value of this state is greater than the reward value corresponding to the current state, then maintain this state; otherwise, maintain the current state;

[0011] S16, repeatedly execute steps S12 - S15, and when a preset termination condition is satisfied, output the state corresponding to the maximum reward value.

[0012] Further, the multi-objective reward function includes:

[0013] Q = αR S + R L + R C + P

[0014] where α, β, and γ are all weights, Q is the reward value, and R S is the aroma intensity, R L is the persistence, R C is the formulation cost, and R P is the penalty value.

[0015] Further, generating and executing a new action to obtain the state corresponding to the new action includes:

[0016] For the current state, taking the action that maximizes the multi-objective reward function in the current state as the new action, including:

[0017] a = arg\{a}{}(,a)

[0018] where a is the action, s is the state, and Q(S,a) is the reward value corresponding to executing action a in state s;

[0019] Executing the new action to obtain the state corresponding to the new action.

[0020] Further, based on the Q-learning algorithm, generating and executing a new action to obtain the state corresponding to the new action further includes:

[0021] Initializing the reward value;

[0022] Judging whether to make an arbitrary action selection with a preset probability; if not, taking the action that maximizes the multi-objective reward function in the current state as the new action;

[0023] Executing the new action to obtain the state corresponding to the new action and obtaining the reward value corresponding to this state;

[0024] Updating the reward value based on the Bellman equation.

[0025] Further, adjusting the compound formulation to be generated according to a preset rule includes:

[0026] Adjusting the concentration of the compounds in the compound formulation to be generated, replacing the compounds, or increasing or decreasing the compounds according to a preset rule.

[0027] Further, adjusting the concentration of the compounds in the compound formulation to be generated according to a preset rule includes:

[0028] Setting the adjustable range of the concentration of the compounds;

[0029] Within the adjustable concentration range, adjust the concentration of the compound in preset steps.

[0030] Furthermore, the preset termination conditions include:

[0031] Reaching the maximum number of iterations; or,

[0032] The average change rate of the reward value after several iterations is less than or equal to a preset threshold.

[0033] In a second aspect, the present invention provides an artificial intelligence-based spice formula composition and optimization system, which includes:

[0034] A definition module, used to define the initial action and initial state, and construct a multi-objective reward function. Among them, the action includes the formula ingredients and ratios, environmental parameters, and performance requirements. The state is used to adjust the compound formula according to preset rules. The multi-objective reward function includes formula performance and formula cost;

[0035] An execution generation module, used to generate and execute a new action to obtain the state corresponding to the new action;

[0036] A prediction module, used to predict the aroma intensity corresponding to the state based on a pre-trained GNN model, where the aroma intensity belongs to the formula performance;

[0037] A reward value module, used to determine the reward value in this state by the multi-objective reward function;

[0038] A judgment module, used to keep the state if the reward value of this state is greater than the reward value corresponding to the current state; otherwise, keep the current state;

[0039] An output module, used to repeatedly execute steps S12 - S15 in the first aspect above. When the preset termination conditions are met, output the state corresponding to the maximum reward value.

[0040] In a third aspect, the present invention provides an electronic device, including:

[0041] A processor;

[0042] A memory for storing instructions executable by the processor;

[0043] Among them, the processor is configured to execute to implement the artificial intelligence-based spice formula composition and optimization method provided in the first aspect.

[0044] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute to implement the artificial intelligence-based spice formula composition and optimization method provided in the first aspect.

[0045] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0046] The method for composing and optimizing a spice formula based on artificial intelligence provided by the present invention predicts aroma characteristics and intensities by combining a pre-trained graph neural network (GNN) model, and uses a multi-objective reward function to comprehensively consider formula performance and cost, achieving effective cost control while ensuring aroma intensity and persistence. This method embeds a regulatory database to intercept high-risk actions in real time, ensuring the safety and compliance of the formula; its automation and intelligence features significantly shorten the R & D cycle and reduce experimental costs, thus demonstrating significant advantages in terms of efficiency, safety, and economy. This method not only accelerates the R & D process of spice formulas but also improves the market competitiveness of the final products. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of the method for composing and optimizing a spice formula based on artificial intelligence provided by the present invention;

[0049] Figure 2 It is a schematic structural diagram of the system for composing and optimizing a spice formula based on artificial intelligence provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The embodiments of the present invention solve the technical problem of how to use artificial intelligence for the R & D of spice formulas in the prior art by providing a method for composing and optimizing a spice formula based on artificial intelligence.

[0051] The technical solution of the present invention for solving the above technical problem has the following general idea:

[0052] Artificial Intelligence-Based Spice Formula Composition and Optimization Method, the method comprising: S11, defining an initial action and an initial state, and constructing a multi-objective reward function, wherein the action includes formula ingredients and proportions, environmental parameters and performance requirements, the state is used to adjust the compound formula according to preset rules, and the multi-objective reward function includes formula performance and formula cost; S12, generating and executing a new action to obtain the state corresponding to the new action; S13, predicting the aroma intensity corresponding to the state based on a pre-trained GNN model, wherein the aroma intensity belongs to the formula performance; S14, determining the reward value of the state by the multi-objective reward function; S15, if the reward value of the state is greater than the reward value corresponding to the current state, then maintain the state; otherwise, maintain the current state; S16, repeating steps S12-S15, and when a preset termination condition is met, outputting the state corresponding to the maximum reward value.

[0053] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the specification drawings and specific embodiments.

[0054] First, it should be noted that the term "and / or" appearing in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0055] The present invention provides an artificial intelligence-based spice formula composition and optimization method as Figure 1 shown, including steps S11-S16:

[0056] S11, defining an initial action and an initial state, and constructing a multi-objective reward function, wherein the action includes formula ingredients and proportions, environmental parameters and performance requirements, the state is used to adjust the compound formula according to preset rules, and the multi-objective reward function includes formula performance and formula cost.

[0057] The information at a certain moment of the state can include formula ingredients and proportions, environmental parameters and performance requirements. For example, Compound A: 20%, Compound B: 30%, Compound C: 30%, Compound D: 20%; the environmental parameters can include parameters such as the temperature and humidity of the formula, and the performance requirements can be the aroma intensity, persistence, and price of the formula to be synthesized, etc. (the above information is encoded in vector form).

[0058] Action is a professional term in the field of reinforcement. In the present invention, it refers to the adjustment of the compound formula.

[0059] Adjust the compound formula to be generated according to preset rules, including: adjusting the concentration of the compounds in the compound formula to be generated, replacing the compounds, or increasing or decreasing the compounds according to the preset rules.

[0060]

Concentration Adjustment

[0061] When the next action needs to be executed, the concentration of the compounds in the compound formula to be generated can be adjusted according to the preset rules, including:

[0062] Set the adjustable range of the compound concentration; within this adjustable range, adjust the concentration of the compound in preset steps. For example, if the concentration of a certain compound does not exceed 50% of the total proportion of the compound formula, each time the next action needs to be generated, increase the concentration by 5%.

[0063]

Compound Replacement

[0064] Based on chemical similarity (such as functional group matching), the compounds in the formula can be replaced. For example, compound B is replaced by E. For example, it can be set that any compound replacement is performed when the action is executed an even number of times, and no replacement is performed when the action is executed an odd number of times. Compound replacement can be used in combination with concentration adjustment, and the specific rules can be determined according to actual needs and are not limited here. When performing compound replacement, compounds with similar chemical properties can be preferentially selected.

[0065]

Compound Increase or Decrease

[0066] Compound increase or decrease includes adding compounds (meeting safety regulations) or deleting the existing compounds in the formula. For example, it can be set that compounds are added or deleted when the number of actions is a multiple of 5. Compound increase or decrease can also be performed simultaneously with compound replacement and concentration adjustment in one action, and the specific preset rules can be set according to actual needs and are not limited here.

[0067] The reward value is a result reward based on the state and action, usually represented in numerical form. The reward value is used to quantify the effect after the compound formula is adjusted.

[0068] The multi-objective reward function includes:

[0069] Q = αR S +R L +R C + P

[0070] Among them, α, β, and γ are all weights, Q is the reward value, and R S is the aroma intensity, R L is the persistence, R C is the formula cost, R P is the penalty value, and α, β, and γ can all be flexibly adjusted according to actual needs.

[0071] Aroma intensity R s : It can be determined by the similarity between the actual intensity of a certain state (i.e., a certain compound formula) of the prediction model and the target fragrance (the essence of the aroma intensity in the multi-objective reward function refers to the similarity of the fragrance, and the aroma intensity in the relevant field is quantifiable in the FEMA Flavor system).

[0072] Persistence: The attenuation rate of a certain state (i.e., a certain compound formula) in the air (1 - attenuation rate / maximum acceptable attenuation rate);

[0073] Formulation cost: 1 - actual cost / budget ceiling.

[0074] Penalty value R P : When the concentration exceeds the limit or a prohibited compound is used, a fixed value is directly deducted, or a prohibited formulation is generated and the reward value is set to 0.

[0075] S12, generate and execute a new action to obtain the state corresponding to the new action.

[0076] Generating and executing a new action to obtain the state corresponding to the new action includes:

[0077] For the current state, taking the action that maximizes the multi-objective reward function in the current state as the new action, including:

[0078] a = arg\{a}{}(,a)

[0079] where a is the action, s is the state, and Q(s,a) is the reward value corresponding to executing action a in state s;

[0080] Execute the new action to obtain the state corresponding to the new action.

[0081] The present invention provides a Q-learning algorithm and a DQN algorithm. For details, please refer to the following, including:

[0082] Based on the Q-learning algorithm, generating and executing a new action to obtain the state corresponding to the new action further includes:

[0083] Initializing the reward value; determining whether to select an action randomly with a preset probability; otherwise, taking the action that maximizes the multi-objective reward function in the current state as the new action, and the preset probability is ε (such as 0.1)

[0084] In other words, there may be a probability of ε to randomly execute an action, and a probability of 1 - ε to execute the new action corresponding to the current maximum reward value. The greedy strategy can avoid falling into a local optimal solution.

[0085] Execute a new action, obtain the state corresponding to the new action, and obtain the reward value corresponding to the state; update the reward value based on the Bellman equation.

[0086] Bellman equations, including:

[0087] Q(s,a)←Q(s,a)+B[r+εmaxa , Q(s , , a , )-(s,a)]Q(s,a)

[0088] ←Q(s,a)+[r+εa , max , , a , )-(s,a)

[0089] Among them, Q(s , , a , ) is in state s , Next, perform action a , The corresponding reward value, B is the learning rate, and r is the discount factor.

[0090] The Q-learning algorithm is applicable to scenarios where the state and action space are small and discrete, for example, when the number of compounds in the formula is less than 10 and the concentration of the compound is adjusted with a fixed step size.

[0091] S13, predicting the aroma intensity corresponding to the state based on the pre-trained GNN model, wherein the aroma intensity belongs to the formula performance.

[0092] DQN is an extension of Q-Learning to handle high-dimensional or continuous states.

[0093] The update process of DQN is as follows: Main Network (Online Network): Input state s, output reward value Q for each action.

[0094] Target Network: The structure is the same as the main network, and the parameters are synchronized regularly.

[0095] Using the same greedy strategy

[0096] Store the state transition (s,a,r,s') into the experience replay buffer.

[0097] Randomly extract states (e.g., 32) from the buffer and determine the corresponding reward values.

[0098] Calculate the loss function (using mean square error as the loss value):

[0099] L=1N(target-Qonline(s,a))2L=N1Σ(target-Qonline(s,a))2.

[0100] Gradient descent updates the main network:

[0101] 0online ← 0online - avL0online ← 0online - avLe.

[0102] Every fixed number of steps (e.g., every 1000 steps), copy the main network parameters to the target network:.

[0103] 0target ← 0online0target ← Gonlinee.

[0104] DQN is applicable to high - dimensional or continuous state spaces (e.g., a formula contains 100 compounds with continuously adjustable concentrations); dynamic environments (e.g., temperature and humidity changes affect fragrance characteristics). DQN can handle complex problems and avoid the explosion of the Q - table dimension. Experience replay improves the data utilization efficiency, and the target network stabilizes the training process.

[0105] S14, determine the reward value in this state by the multi - objective reward function.

[0106] As mentioned above, after each new action is executed, obtain the state corresponding to the new action, and obtain the reward value by the multi - objective reward function.

[0107] S15, if the reward value of this state is greater than the reward value corresponding to the current state, then maintain this state; otherwise, maintain the current state.

[0108] It should be noted that S14 can be executed several times. In each judgment, the newly calculated reward value is compared with the current reward value. In the first execution, the newly calculated reward value is compared with the initial reward value.

[0109] S16, repeat steps S12 - S16. When the preset termination condition is met, output the state corresponding to the maximum reward value.

[0110] The preset termination conditions include: reaching the maximum number of iterations; or, the average change rate of the reward value after several iterations is less than or equal to the preset threshold.

[0111] In summary, the present invention provides a method for composing and optimizing a spice formula based on artificial intelligence, including: defining an initial action and an initial state, and constructing a multi-objective reward function, where the action includes the formula ingredients and proportions, environmental parameters, and performance requirements, and the state is used to adjust the compound formula according to preset rules, and the multi-objective reward function includes formula performance and formula cost; generating and executing a new action to obtain the state corresponding to the new action; predicting the aroma intensity corresponding to this state based on a pre-trained GNN model, where the aroma intensity belongs to the formula performance; determining the reward value of this state by the multi-objective reward function; if the reward value of this state is greater than the reward value corresponding to the current state, then maintain this state; otherwise, maintain the current state; repeat steps S12 - S15, and when a preset termination condition is met, output the state corresponding to the maximum reward value. The method for composing and optimizing a spice formula based on artificial intelligence provided by the present invention combines a pre-trained graph neural network (GNN) model to predict aroma characteristics and intensity, and uses a multi-objective reward function to comprehensively consider formula performance and cost, achieving effective cost control while ensuring aroma intensity and persistence. This method embeds a regulatory database to intercept high-risk actions in real time to ensure the safety and compliance of the formula; its automation and intelligence features significantly shorten the R & D cycle and reduce experimental costs, thus showing significant advantages in terms of efficiency, safety, and economy. This method not only accelerates the R & D process of spice formulas but also improves the market competitiveness of the final products.

[0112] Based on the same inventive concept, the present invention provides as Figure 2 shown a system for composing and optimizing a spice formula based on artificial intelligence, the system including:

[0113] A definition module 21, used to define an initial action and an initial state, and construct a multi-objective reward function, where the action includes the formula ingredients and proportions, environmental parameters, and performance requirements, and the state is used to adjust the compound formula according to preset rules, and the multi-objective reward function includes formula performance and formula cost;

[0114] An execution generation module 22, used to generate and execute a new action to obtain the state corresponding to the new action;

[0115] A prediction module 23, used to predict the aroma intensity corresponding to this state based on a pre-trained GNN model, where the aroma intensity belongs to the formula performance;

[0116] A reward value module 24, used to determine the reward value of this state by the multi-objective reward function;

[0117] A judgment module 25, used to, if the reward value of this state is greater than the reward value corresponding to the current state, then maintain this state; otherwise, maintain the current state;

[0118] An output module 26 is configured to repeatedly execute steps S12 - S14, and when a preset termination condition is satisfied, output the state corresponding to the maximum reward value.

[0119] Based on the same inventive concept, the present invention further provides an electronic device as shown, including:

[0120] A processor;

[0121] A memory for storing instructions executable by the processor;

[0122] Wherein, the processor is configured to execute to implement the artificial intelligence - based spice formula composition and optimization method provided as described above.

[0123] Based on the same inventive concept, the present invention further provides a non - transitory computer - readable storage medium. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute and implement the artificial intelligence - based spice formula composition and optimization method provided as described above.

[0124] Since the electronic device introduced in this embodiment is the electronic device used to implement the information - processing method in the embodiments of the present invention, based on the information - processing method introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present invention will not be described in detail here. As long as the electronic device used by those skilled in the art to implement the information - processing method in the embodiments of the present invention falls within the scope of protection of the present invention.

[0125] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk memory, CD - ROM, optical memory, etc.) containing computer - usable program code.

[0126] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general - purpose computer, a special - purpose computer, an embedded processor, or other programmable data - processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data - processing devices generate for implementation in the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks

[0127] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks

[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks

[0129] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention

[0130] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations

Claims

1. A method for the composition and optimization of flavor formulas based on artificial intelligence, characterized in that: The method comprises: S11, defining an initial action and an initial state, and constructing a multi-objective reward function, wherein the action includes formula ingredients and proportions, environmental parameters, and performance requirements, the state is used to adjust the compound formula according to preset rules, and the multi-objective reward function includes formula performance and formula cost; S12, generate and execute a new action to obtain the state corresponding to the new action; S13, predicting the aroma intensity corresponding to the state based on the pre-trained GNN model, where the aroma intensity belongs to the formula performance; S14, determining a reward value in this state according to the multi-objective reward function; S15, if the reward value of the state is greater than the reward value corresponding to the current state, then maintain the state; otherwise, maintain the current state; S16, repeating steps S12-S15, and when the preset termination condition is met, outputting the state corresponding to the maximum reward value.

2. The method for spice formula composition and optimization based on artificial intelligence as claimed in claim 1, characterized in that: Multi-objective reward function, including: Q=αR S +βR L +γR C +R P Among them, α, β, and γ are weights, Q is the reward value, and R S is the aroma intensity, R L is the durability, R C is the formula cost, R P is the penalty value.

3. The method for spice formula composition and optimization based on artificial intelligence as claimed in claim 1, characterized in that: Generate and execute new actions to obtain the status corresponding to the new actions, including: For the current state, the action corresponding to maximizing the multi-objective reward function in the current state is taken as the new action, including: a=arg\underset{a}{max}Q(s,a) Among them, a is the action, s is the state, and Q(s,a) is the reward value corresponding to executing action a in state s; Execute the new action and get the state corresponding to the new action.

4. The method for spice formula composition and optimization based on artificial intelligence as claimed in claim 3, characterized in that: Based on the Q-learning algorithm, new actions are generated and executed to obtain the state corresponding to the new actions, including: Initialize reward value; Determine whether to make any action selection with the preset probability; if not, take the action corresponding to maximizing the multi-objective reward function under the current state as the new action; Execute a new action, get the state corresponding to the new action, and get the reward value corresponding to the state; Update the reward value based on the Bellman equation.

5. The method for spice formula composition and optimization based on artificial intelligence as claimed in claim 1, characterized in that: Adjust the formula of the compound to be generated according to the preset rules, including: According to preset rules, the concentration of compounds in the compound formula to be generated is adjusted, the compounds are replaced, or the compounds are increased or decreased.

6. The method for spice formula composition and optimization based on artificial intelligence as claimed in claim 5, characterized in that: Adjust the concentration of the compounds in the compound formula to be generated according to the preset rules, including: Set the adjustable range of compound concentration; In the adjustable concentration range, the concentration of the compound is adjusted with a preset step length.

7. The method for spice formula composition and optimization based on artificial intelligence as claimed in claim 1, characterized in that: Preset termination conditions, including: The maximum number of iterations is reached; or, The average change rate of the reward value after several iterations is less than or equal to the preset threshold.

8. The spice formula composition and optimization system based on artificial intelligence is characterized by: The method for spice formula composition and optimization based on artificial intelligence applied to any one of claims 1-7, the system comprising: A definition module, used to define initial actions and initial states, and to construct a multi-objective reward function, wherein the actions include recipe ingredients and proportions, environmental parameters, and performance requirements, the states are used to adjust the compound recipe according to preset rules, and the multi-objective reward function includes recipe performance and recipe cost; The execution generation module is used to generate and execute new actions and obtain the state corresponding to the new actions; A prediction module is used to predict the aroma intensity corresponding to the state based on the pre-trained GNN model, where the aroma intensity belongs to the formula performance; A reward value module, used to determine the reward value in this state according to the multi-objective reward function; A judgment module, used to maintain the state if the reward value of the state is greater than the reward value corresponding to the current state; otherwise, maintain the current state; The output module is used to repeatedly execute steps S12-S15, and when a preset termination condition is met, output the state corresponding to the maximum reward value.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute to implement the artificial intelligence-based spice formula composition and optimization method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to implement the artificial intelligence-based flavor formula composition and optimization method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Formula generation method, device and equipment and computer readable storage medium

    CN117632923A

  • Ship agent intelligent control method and device based on reinforcement learning

    CN117991793A

  • Model generation method, computer program, and information processing device

    CN119156688A

  • Intelligent aromatic simulation of food recipe

    US20230222321A1

  • KR20250019455A