Cosmetic research and development management method and device based on reinforcement learning and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-11
- Publication Date
- 2026-03-27
Smart Images

Figure CN121743375A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus and storage medium for cosmetic research and development management based on reinforcement learning. Background Technology
[0002] In the process of cosmetic research and development, it is usually necessary to access a large amount of existing data for researchers to refer to.
[0003] However, the data that needs to be accessed is often distributed across different datasets, and each dataset typically contains massive amounts of data, making it difficult for developers to efficiently and accurately obtain the data required for research and development from various datasets, and also making it difficult to automatically generate corresponding feasible research and development solutions. Summary of the Invention
[0004] To address the aforementioned technical issues, this application proposes a cosmetics R&D management method, apparatus, and storage medium based on reinforcement learning, which can efficiently and accurately acquire the data required for R&D from various datasets and improve the feasibility of automatically recommended solutions.
[0005] In a first aspect, embodiments of this application provide a cosmetic research and development management method based on reinforcement learning, including: Obtain cosmetic research and development target information, use it as the basis for querying, and perform queries on N cosmetic research and development datasets to obtain N preliminary research and development data, where N is a positive integer; Based on the N preliminary R&D data, M R&D description sets are determined, each consisting of N R&D description information, where M is a positive integer less than N. The N R&D description information corresponds one-to-one with the N preliminary R&D data, and each R&D description information is determined based on its corresponding preliminary R&D data. Reasoning analysis is performed based on each of the aforementioned R&D description sets, and the target R&D description information corresponding to each of the aforementioned R&D description sets is determined based on each of the aforementioned R&D description sets and its reasoning analysis results. Based on the target R&D description information corresponding to each of the M R&D description sets, a reinforcement learning model is invoked to generate a cosmetic R&D recommendation scheme that matches the cosmetic R&D target information.
[0006] Optionally, the reasoning analysis based on each of the R&D description sets includes: For each of the aforementioned R&D description sets, Typical R&D description information and at least one supplementary R&D description information are selected from the R&D description set. The supplementary R&D description information refers to R&D description information whose correlation with the typical R&D description information meets the first correlation condition. Based on the typical R&D description information and the at least one supplementary R&D description information, reasoning hint information and reasoning counterexample information are constructed through semantic similarity analysis; Based on the R&D description set, the inference hints, and the inference counterexamples, inference analysis is performed.
[0007] Optionally, the step of constructing reasoning hints and counterexamples based on the typical R&D description information and the at least one supplementary R&D description information through semantic similarity analysis includes: Determine the first relevant semantic relationship between the typical R&D description information and each of the supplementary R&D description information; Based on each of the first relevant semantics, prompting R&D description information is selected from the at least one supplementary R&D description information; Based on the aforementioned prompt description information, the reasoning prompt information is constructed; Based on the supplementary R&D description information other than the prompt R&D description information in the at least one supplementary R&D description information, the inference counterexample information is constructed.
[0008] Optionally, the step of filtering out the prompting R&D description information from the at least one supplementary R&D description information based on each of the first relevant semantics includes: For each of the first related semantics, a first semantic similarity is determined between the first related semantics and the other first related semantics, and the larger of the largest first semantic similarity and the preset semantic similarity threshold is taken as the second semantic similarity corresponding to the first related semantics; Determine a fourth semantic similarity between a second related semantic and each of the first related semantics, wherein the second related semantic is a first related semantic corresponding to a third semantic similarity, and the third semantic similarity is the largest second semantic similarity; The third related semantics are determined based on the first related semantics corresponding to each of the fourth semantic similarities that are not less than the third semantic similarity among all the fourth semantic similarities; Supplementary R&D description information corresponding to the third related semantic is selected from the at least one supplementary R&D description information and used as the prompt R&D description information.
[0009] Optionally, the reasoning analysis based on the R&D description set, the reasoning hints, and the reasoning counterexamples includes: Based on this R&D description set, the first inference result is obtained; Based on the reasoning hints, a portion of the information in the first reasoning result is modified to obtain the second reasoning result; Based on the difference between the second reasoning result and the first reasoning result, the R&D description set is updated; Based on the aforementioned counterexample information and the updated R&D description set, inference analysis is performed.
[0010] Optionally, updating the R&D description set based on the difference between the second inference result and the first inference result includes: Based on the difference between the second reasoning result and the first reasoning result, first difference information is generated; Self-attention calculation is performed on the first difference information to obtain the second difference information; The second difference information is embedded into the R&D description set to update the R&D description set.
[0011] Optionally, determining the target R&D description information corresponding to each R&D description set based on each R&D description set and its reasoning analysis results includes: For each R&D description set, based on its reasoning and analysis results, target R&D description information corresponding to that R&D description set is obtained from that R&D description set.
[0012] Optionally, the step of generating a cosmetics R&D recommendation scheme that matches the cosmetics R&D target information by calling a reinforcement learning model based on the target R&D description information corresponding to each of the M R&D description sets includes: Based on the target R&D description information corresponding to each of the M R&D description sets, multiple target R&D data are obtained by querying the cosmetic formula knowledge base, wherein the cosmetic formula knowledge base stores cosmetic formula records that have been experimentally verified. Based on the aforementioned multiple target R&D data, a reinforcement learning model is invoked to generate the cosmetic R&D recommendation scheme.
[0013] Secondly, embodiments of this application provide a cosmetic research and development management device based on reinforcement learning, comprising: The preliminary research and development data acquisition module is used to acquire cosmetic research and development target information, and use it as the basis for querying N cosmetic research and development datasets to obtain N preliminary research and development data, where N is a positive integer. The R&D description set determination module is used to determine M R&D description sets composed of N R&D description information based on the N preliminary R&D data, where M is a positive integer and less than N, and the N R&D description information corresponds one-to-one with the N preliminary R&D data, and each R&D description information is determined according to its corresponding preliminary R&D data; The target R&D description information determination module is used to perform reasoning analysis based on each R&D description set, and to determine the target R&D description information corresponding to each R&D description set according to each R&D description set and its reasoning analysis results. The solution generation module is used to generate a cosmetic research and development recommendation solution that matches the cosmetic research and development target information by calling a reinforcement learning model based on the target research and development description information corresponding to each of the M research and development description sets.
[0014] Thirdly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any of the above-mentioned embodiments.
[0015] In summary, the embodiments of this application have at least the following beneficial effects: Using the embodiments of this application, cosmetic R&D target information is obtained and used as the basis for querying. N preliminary R&D data are obtained by querying N cosmetic R&D datasets, where N is a positive integer. Based on the N preliminary R&D data, M R&D description sets are determined, each consisting of N R&D description information, where M is a positive integer less than N. Each of the N R&D description information corresponds one-to-one with one of the N preliminary R&D data, and each R&D description information is determined based on its corresponding preliminary R&D data. Reasoning analysis is performed based on each R&D description set, and target R&D description information corresponding to each R&D description set is determined based on each R&D description set and its reasoning analysis results. Based on the target R&D description information corresponding to each of the M R&D description sets, a reinforcement learning model is invoked to generate a cosmetic R&D recommendation scheme that matches the cosmetic R&D target information. Thus, a highly automated R&D management system can be formed. After a user proposes their cosmetic R&D target, the system can automatically and efficiently obtain the necessary R&D data from various datasets based on the cosmetic R&D target information, thereby improving the feasibility of the automatically recommended scheme. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the cosmetic R&D management method based on reinforcement learning provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of the cosmetic R&D management device based on reinforcement learning provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments / examples are only a part of the embodiments / examples of this application, and not all of the embodiments / examples. Based on the embodiments / examples in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0018] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "multiple" means two or more. In the description of this application, the term "comprising" and its variations are open-ended, meaning "including but not limited to." The term "based on" means "at least partially based on." The term "according to" means "at least partially according to." The term "one embodiment / example" means "at least one embodiment / example"; the term "another embodiment / example" means "at least one additional embodiment / example"; the term "some embodiments / examples" means "at least some embodiments / examples."
[0019] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0020] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing specific embodiments only and is not intended to limit the application. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0021] Firstly, see [the following] Figure 1 The diagram illustrates a flowchart of a reinforcement learning-based cosmetic R&D management method provided in this application embodiment. This reinforcement learning-based cosmetic R&D management method can be applied to a computer device with data processing capabilities. The method includes S101-S104, as detailed below.
[0022] S101: Obtain cosmetic research and development target information, use it as the basis for querying, and query in N cosmetic research and development datasets respectively to obtain N preliminary research and development data, where N is a positive integer.
[0023] In some examples, there may be a one-to-one correspondence between N preliminary R&D data and N cosmetic R&D datasets, where each preliminary R&D data is data that matches the cosmetic R&D target information obtained from its corresponding cosmetic R&D dataset.
[0024] In some examples, the N cosmetic R&D datasets may include at least one of the following: cosmetic raw material dataset, cosmetic historical formula dataset, cosmetic efficacy verification dataset, cosmetic compliance dataset (e.g., it may store a list of prohibited raw materials and / or a list of restricted raw materials), cosmetic supply chain and cost dataset, and cosmetic raw material compatibility dataset.
[0025] S102, based on the N preliminary R&D data, determine M R&D description sets consisting of N R&D description information, where M is a positive integer and less than N, and the N R&D description information corresponds one-to-one with the N preliminary R&D data, and each R&D description information is determined according to its corresponding preliminary R&D data.
[0026] In some examples, the corresponding R&D description information can be determined based on each preliminary R&D data. For instance, semantic compression and structured expression transformation can be performed on each preliminary R&D data to generate the corresponding R&D description information. Semantic compression can be achieved by calling a general key semantic extraction model, so that key semantic information can be extracted from each preliminary R&D data through semantic compression. Structured expression transformation can be used to convert the extracted key semantic information into structured R&D description information, so that the R&D description information has the characteristics of a structured machine-readable form, so that the computer can process the R&D description information more efficiently.
[0027] In some examples, each of the N R&D description information can correspond to one of the R&D description sets, and each R&D description set can be composed of its corresponding R&D description information. Generally speaking, each R&D description set can correspond to at least one of the N R&D description information.
[0028] In some examples, N R&D description information items can be grouped into approximately equal groups to form M R&D description sets. Thus, the (M-1) R&D description sets within the M R&D description sets each contain an equal number of R&D description information items, all of which are of the first quantity. Furthermore, the number of items in one of the M R&D description sets (excluding the (M-1) R&D description sets) can be either this first quantity or a second quantity different from it. This ensures that the number of R&D description information items in each R&D description set is approximately equal, so that when performing parallel inference analysis on each R&D description set, the amount of data analyzed in each analysis is roughly the same, thereby improving analysis efficiency.
[0029] S103, perform reasoning analysis based on each of the R&D description sets, and determine the target R&D description information corresponding to each of the R&D description sets according to each of the R&D description sets and its reasoning analysis results.
[0030] In some examples, each of the R&D description sets and preset model prompts can be input into the inference model to obtain the inference analysis results of each of the R&D description sets output by the inference model.
[0031] For example, the inference model can employ a large model. The model prompts can guide the inference model to perform inference analysis on each piece of R&D description information in the R&D description set, obtaining inference analysis results including the inference information corresponding to each piece of R&D description information. Thus, based on the inference analysis results of each R&D description set, the R&D description information in each R&D description set can be filtered to select R&D description information whose inference information meets preset requirements, thereby constituting the target R&D description information corresponding to each R&D description set. Here, the inference information corresponding to each R&D description information may include efficacy achievement inference information, safety inference information, and / or R&D feasibility inference information, etc., and correspondingly, the preset requirements may include efficacy achievement requirements, safety requirements, and / or R&D feasibility requirements.
[0032] For example, the inference model can employ a large model. The model prompts can also guide the inference model to perform relevance inference analysis on each piece of R&D description information in the R&D description set, obtaining inference analysis results that include the degree of relevance between each piece of R&D description information. Thus, typical R&D description information can be selected from the R&D description set, and based on the degree of relevance contained in the inference analysis results of the R&D description set, R&D description information whose relevance to the typical R&D description information meets a first relevance condition can be selected as the corresponding target R&D description information. Here, the first relevance condition may include a first threshold condition, which indicates that the degree of relevance must be higher than a first relevance threshold.
[0033] S104, based on the target R&D description information corresponding to each of the M R&D description sets, call the reinforcement learning model to generate a cosmetic R&D recommendation scheme that matches the cosmetic R&D target information.
[0034] In some examples, in step S104, the target R&D description information corresponding to each of the M R&D description sets can be converted into corresponding vectors (e.g., through vectorization via encoding). All the converted vectors are then used to form the initial state S of the reinforcement learning model, enabling this initial state S to indicate the target R&D description information corresponding to each of the M R&D description sets. Inputting this initial state S into the reinforcement learning model allows it to generate an action sequence A, which can include actions at different stages of the cosmetic preparation process. After generating the action sequence A, the reinforcement learning model can decode it and convert it into a structured cosmetic R&D recommendation scheme described in natural language. For example, this recommendation scheme may include: formula composition (raw material list, dosage), preparation process parameters, efficacy verification scheme, compliance risk warnings, and / or production adaptation suggestions. The final generated cosmetic R&D recommendation scheme will match the initial cosmetic R&D target information.
[0035] In some examples, reinforcement learning models can be trained in the following way.
[0036] First, we can define the core elements required for reinforcement learning, such as: agent, which is generally used as the carrier in reinforcement learning models and is mainly responsible for learning and generating action sequences for cosmetic research and development; environment, which can be used to simulate the constraints of the entire cosmetic research and development process, such as the rules corresponding to at least some of the N research and development datasets (such as raw material compatibility, compliance restrictions, and / or efficacy requirements); state, which can be a vector encoded by sample research and development description information (i.e., the sample corresponding to the target research and development description information), which can be used to indicate all constraints and available information corresponding to the sample research and development description information; actions in the action sequence, which can be used to indicate the decision-making steps in the research and development process, such as raw material selection, process parameter setting, and the solutions adopted during efficacy verification; and reward, which can be used to evaluate the quality of the action sequences generated by the agent in order to determine the learning direction of the agent. For example, we can design the following positive rewards: +10 points when the formula meets the efficacy requirements, +8 points when it is compliant, and +5 points when the process can be industrialized. We can also design the following negative penalties: -15 points when the formula has safety risks, -10 points when the efficacy does not meet the standards, and -20 points when it is non-compliant.
[0037] Secondly, the sample R&D description information and sample action sequences corresponding to historical R&D data can be used to pre-train the agent. Pre-training can be carried out in a supervised learning manner so that the agent can initially have the ability to generate action sequences that conform to basic R&D logic.
[0038] Furthermore, the pre-trained agent can be deployed to a simulated R&D environment for online interactive training. This simulated R&D environment can receive a large amount of sample R&D description information input by R&D personnel and convert it into corresponding vector inputs to the pre-trained agent, so that the pre-trained agent outputs action sequences for R&D personnel to evaluate. The evaluation results are used to determine the latest reward and feed it back to the agent.
[0039] Finally, the intelligent agent that has completed online interactive training can be deployed in the system as a reinforcement learning model for R&D personnel to use daily. In addition, when the reinforcement learning model generates each cosmetic R&D recommendation scheme after deployment, it can also simultaneously receive the evaluation of the generated cosmetic R&D recommendation scheme input by the R&D personnel. The evaluation results can be used to fine-tune the reinforcement learning model based on real scenarios, so that the reinforcement learning model can continuously learn and iterate, and improve the effectiveness and feasibility of the cosmetic R&D recommendation schemes it generates.
[0040] It should be noted that the "similarity" described in any one or more embodiments of this application can be obtained by calculating the feature information of the two targets for which the "similarity" needs to be calculated using at least one of the following methods: cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, etc. It should be understood that the implementation of this "similarity" is merely illustrative and not intended to limit the scope of this application.
[0041] For example, cosine similarity can be calculated using the following formula.
[0042] in, Represents cosine similarity. , These represent the two types of feature information for which cosine similarity needs to be calculated. Represents the vector dot product. This represents the L2 norm of a vector.
[0043] Unless otherwise specified, the explanation of "similarity" will be omitted in the following text to avoid redundancy.
[0044] In one optional implementation, the reasoning analysis based on each of the R&D description sets includes: For each of the aforementioned R&D description sets, Typical R&D description information and at least one supplementary R&D description information are selected from the R&D description set. The supplementary R&D description information refers to R&D description information whose correlation with the typical R&D description information meets the first correlation condition. Based on the typical R&D description information and the at least one supplementary R&D description information, reasoning hint information and reasoning counterexample information are constructed through semantic similarity analysis; Based on the R&D description set, the inference hints, and the inference counterexamples, inference analysis is performed.
[0045] In some examples, the first relevance condition may include a first threshold condition, which is used to indicate that the relevance must be higher than a first relevance threshold.
[0046] In some examples, typical R&D description information can be one or more typical R&D description information in the R&D description set. Each typical R&D description information in at least some typical R&D description information can correspond to at least one supplementary R&D description information. If there is another part of typical R&D description information, the other part of typical R&D description information may not have a corresponding supplementary R&D description information.
[0047] In some examples, the above-mentioned selection of typical R&D description information in the R&D description set may include: for each R&D description information in the R&D description set, calculating the semantic similarity between the R&D description information and the other R&D description information in the R&D description set, and taking the average of all semantic similarities corresponding to the R&D description information as the corresponding semantic similarity mean; selecting each R&D description information in the R&D description set whose semantic similarity mean is less than a preset mean threshold as typical R&D description information of the R&D description set.
[0048] In other examples, the above-mentioned selection of typical R&D description information in the R&D description set may include: obtaining multiple cluster centers by performing cluster analysis on the R&D description information in the R&D description set; for each R&D description information in the R&D description set, calculating the similarity between the R&D description information and each cluster center; if the highest similarity is greater than or equal to a preset cluster similarity threshold, the R&D description information can be assigned to the cluster containing the cluster center corresponding to the highest similarity; and the R&D description information contained in the clusters containing each cluster center is taken as typical R&D description information.
[0049] In some examples, the semantic similarity between each R&D description in the R&D description set (excluding typical R&D descriptions) and the typical R&D description can be calculated to characterize the corresponding relevance. Alternatively, the R&D description set and relevance analysis prompts can be input into a large language model to obtain the relevance between each R&D description in the R&D description set (excluding typical R&D descriptions) and the typical R&D description, generated and output by the large language model. The relevance analysis prompts are used to guide the large language model in analyzing the relevance between each R&D description in the R&D description set (excluding typical R&D descriptions) and the typical R&D description.
[0050] In some examples, semantic similarity analysis can be used to obtain the semantic similarity between typical R&D description information and each supplementary R&D description information. Based on supplementary R&D description information with semantic similarity less than a counterexample threshold, or based on typical R&D description information and supplementary R&D description information with semantic similarity less than a counterexample threshold, inference counterexample information can be constructed. Similarly, based on supplementary R&D description information with semantic similarity greater than a positive threshold, or based on typical R&D description information and supplementary R&D description information with semantic similarity greater than a positive threshold, inference hint information can be constructed. In this way, similar and highly compliant inference hint information, as well as similar but non-compliant inference counterexample information, can be constructed to provide more refined hints for subsequent inference analysis, improve the reliability of the inference analysis results, and reduce the probability of generating seemingly plausible counterexamples.
[0051] In some examples, the R&D description set, inference hints, inference counterexamples, and preset inference analysis hints can be input into the inference analysis model for inference analysis to obtain the inference analysis results output by the inference analysis model. The inference analysis model can be a large model, and the inference analysis hints can be used to prompt the inference analysis model to perform inference analysis on the R&D description set based on the inference hints and inference counterexamples.
[0052] In one optional implementation, the step of constructing reasoning hints and counterexamples based on the typical R&D description information and the at least one supplementary R&D description information through semantic similarity analysis includes: Determine the first relevant semantic relationship between the typical R&D description information and each of the supplementary R&D description information; Based on each of the first relevant semantics, prompting R&D description information is selected from the at least one supplementary R&D description information; Based on the aforementioned prompt description information, the reasoning prompt information is constructed; Based on the supplementary R&D description information other than the prompt R&D description information in the at least one supplementary R&D description information, the inference counterexample information is constructed.
[0053] In some examples, the first, second, and third related semantics can be used to indicate the related semantic content between two corresponding descriptive information. Specifically, the first related semantics between two descriptive information can be generated by a large model, that is, the large model analyzes the related semantic content between two descriptive information to generate the first related semantics. The first related semantics can be text described in natural language or feature information in vector form.
[0054] In some examples, features can be extracted from the selected prompt R&D description information separately, and the extracted features can be fused to form the inference prompt information. In this case, the inference prompt information can exist based on features. Alternatively, the selected prompt R&D description information can be directly combined to form the inference prompt information, in which case the inference prompt information and the selected prompt R&D description information have the same form. Similarly, the construction method of inference counterexample information is similar; it is only necessary to replace the selected prompt R&D description information with at least one supplementary R&D description information other than the prompt R&D description information. This will not be elaborated further here.
[0055] In this embodiment, a first relevant semantic can be introduced to quantify the correlation strength between R&D description information, so as to select the most relevant prompt R&D description information from a large amount of supplementary R&D description information, filter out weakly related or irrelevant content, thereby accurately classifying construction reasoning prompt information and reasoning counterexample information, and improving the reliability of reasoning analysis results.
[0056] In one optional implementation, the step of filtering out prompting R&D description information from the at least one supplementary R&D description information based on each of the first relevant semantics includes: For each first related semantic, a first semantic similarity is determined between the first related semantic and the other first related semantics. The larger of the largest first semantic similarity and a preset semantic similarity threshold is taken as the second semantic similarity corresponding to the first related semantic. In this way, the preset semantic similarity threshold can be used as the lower limit of the second semantic similarity corresponding to each first related semantic to avoid the second semantic similarity being too small. Here, the largest first semantic similarity refers to the maximum value among all the first semantic similarities corresponding to the first related semantic, and the other first related semantics refer to all the first related semantics other than the first related semantic. A fourth semantic similarity is determined between the second related semantic and each of the first related semantics, wherein the second related semantic is the first related semantic corresponding to the third semantic similarity, and the third semantic similarity is the largest second semantic similarity; wherein the largest second semantic similarity refers to the maximum value among the second semantic similarities corresponding to each of the first related semantics; The third related semantics are determined based on the first related semantics corresponding to each of the fourth semantic similarities that are not less than the third semantic similarity; wherein the third related semantics at least includes the first related semantics corresponding to each of the fourth semantic similarities that are not less than the third semantic similarity. Supplementary R&D description information corresponding to the third related semantic is selected from the at least one supplementary R&D description information and used as the prompt R&D description information.
[0057] In this embodiment, by analyzing the third semantic similarity, it is possible to further distinguish whether the first related semantic and the second related semantic have a high degree of similarity, thereby improving the standard of distinction. Ultimately, the third related semantic and the corresponding prompting R&D description information with a high degree of matching with typical R&D description information can be obtained, thereby improving the quality of the constructed reasoning prompt information.
[0058] In one optional implementation, the reasoning analysis based on the R&D description set, the reasoning hints, and the reasoning counterexamples includes: Based on this R&D description set, reasoning is performed to obtain the first reasoning result; for example, a large model can be used to reason on this R&D description set to obtain the first reasoning result. Based on the reasoning hints, a portion of the information in the first reasoning result is modified to obtain the second reasoning result; Based on the difference between the second reasoning result and the first reasoning result, the R&D description set is updated; Based on the aforementioned counterexample information and the updated R&D description set, inference analysis is performed.
[0059] In some examples, a portion of the information in the first inference result can be replaced with at least a portion of the content in the inference hints to obtain a second inference result. For example, the first inference result can be represented as a feature vector, thus allowing a portion of the information (a portion of the features) in the first inference result to be replaced with features transformed from at least a portion of the content in the inference hints, thereby obtaining the second inference result.
[0060] In some examples, inference counterexamples, the updated R&D description set, and preset inference analysis prompts can be input into the inference analysis model for inference analysis to obtain the inference analysis results output by the inference analysis model. The inference analysis model can be a large model, and the inference analysis prompts can be used to prompt the inference analysis model to perform inference analysis on the updated R&D description set based on the inference counterexamples.
[0061] In this embodiment, preliminary reasoning can be performed using only the R&D description set to obtain a first reasoning result. Then, reasoning hints are used to correct some content in the first reasoning result to obtain a second reasoning result that better fits the R&D needs. This allows the reasoning result to gradually approach the real R&D scenario, reducing the deviation that may occur in the initial deduction. The R&D description set is dynamically updated by comparing the differences between the first and second reasoning results, so that subsequent reasoning can be based on more accurate information. Finally, counterexamples are introduced for final verification, which can identify and avoid potential R&D risks (such as raw material incompatibility, non-compliance, etc.) in advance.
[0062] In one optional implementation, updating the R&D description set based on the difference between the second inference result and the first inference result includes: Based on the difference between the second reasoning result and the first reasoning result, first difference information is generated; Self-attention calculation is performed on the first difference information to obtain the second difference information; The second difference information is embedded into the R&D description set to update the R&D description set.
[0063] In some examples, the first difference information can be generated by analyzing the error between the second inference result and the first inference result. For example, the feature error between the feature information corresponding to the second inference result and the first inference result can be calculated to obtain the first difference information.
[0064] In this embodiment, the inherent features of the first difference information can be mined through self-attention computation, thereby strengthening the first difference information to obtain the second difference information, which is then embedded into the R&D description set. This optimizes the content of the R&D description set and improves the accuracy and reliability of subsequent reasoning analysis. Furthermore, since the second difference information is embedded into the R&D description set, the model used for subsequent reasoning analysis can more directly utilize the second difference information when analyzing the R&D description set, reducing redundant computation and improving the efficiency of reasoning analysis.
[0065] In one optional implementation, determining the target R&D description information corresponding to each R&D description set based on each R&D description set and its inference analysis results includes: For each R&D description set, based on its reasoning and analysis results, target R&D description information corresponding to that R&D description set is obtained from that R&D description set.
[0066] In this embodiment, the reasoning analysis results of each R&D description set can be used as a guide for screening from that R&D description set. This allows for the selection of R&D description information that matches the guide, which can then be used as the target R&D description information for that R&D description set. In this way, by screening based on the reasoning analysis results, the target R&D description information can retain only the core information that is strongly related to the cosmetic R&D target information, reducing redundant data and making the input of the subsequent reinforcement learning model more accurate.
[0067] In one optional implementation, the step of generating a cosmetics R&D recommendation scheme that matches the cosmetics R&D target information by invoking a reinforcement learning model based on the target R&D description information corresponding to each of the M R&D description sets includes: Based on the target R&D description information corresponding to each of the M R&D description sets, multiple target R&D data are obtained by querying the cosmetic formula knowledge base, wherein the cosmetic formula knowledge base stores cosmetic formula records that have been experimentally verified. Based on the aforementioned multiple target R&D data, a reinforcement learning model is invoked to generate the cosmetic R&D recommendation scheme.
[0068] In some examples, based on each of the target R&D descriptions in the M R&D description sets, a query can be performed from the cosmetics formula knowledge base to obtain at least one target R&D data corresponding to each target R&D description. In this way, the multiple target R&D data obtained can include the target R&D data corresponding to the target R&D descriptions in the M R&D description sets.
[0069] It is understood that the content related to "reinforcement learning model" in this embodiment can refer to the description of the relevant embodiments of step S104 above. In short, it is only necessary to replace "R&D description information" in the relevant embodiments with "R&D data" in this embodiment, as follows.
[0070] Each target R&D data point is converted into a corresponding vector (e.g., through encoding). All converted vectors form the initial state S of the reinforcement learning model, enabling S to indicate all target R&D data. Inputting this initial state S into the reinforcement learning model allows it to generate an action sequence A, which can include actions at different stages of the cosmetic preparation process. After generating action sequence A, the model can decode it into a structured cosmetic R&D recommendation scheme described in natural language. This recommendation scheme may include: formulation composition (raw material list, dosage), preparation process parameters, efficacy verification plan, compliance risk warnings, and / or production adaptation suggestions. The final generated recommendation scheme will match the initial cosmetic R&D target information.
[0071] Reinforcement learning models can be trained in the following way.
[0072] First, we can define the core elements required for reinforcement learning, such as: agent, which is generally used as the carrier in reinforcement learning models and is mainly responsible for learning and generating action sequences for cosmetic research and development; environment, which can be used to simulate the constraints of the entire cosmetic research and development process, such as the rules corresponding to at least some of the N research and development datasets (such as raw material compatibility, compliance restrictions, and / or efficacy requirements); state, which can be a vector encoded by sample research and development data (i.e., the sample corresponding to the target research and development data), which can be used to indicate all constraints and available information corresponding to the sample research and development data; actions in the action sequence, which can be used to indicate the decision-making steps in the research and development process, such as raw material selection, process parameter setting, and the solutions adopted during efficacy verification; and reward, which can be used to evaluate the quality of the action sequences generated by the agent in order to determine the learning direction of the agent. For example, we can design the following positive rewards: +10 points when the formula meets the efficacy requirements, +8 points when it is compliant, and +5 points when the process can be industrialized. We can also design the following negative penalties: -15 points when the formula has safety risks, -10 points when the efficacy does not meet the standards, and -20 points when it is non-compliant.
[0073] Secondly, sample R&D data and sample action sequences corresponding to historical R&D data can be used to pre-train the agent. Pre-training can be carried out in a supervised learning manner so that the agent can initially have the ability to generate action sequences that conform to basic R&D logic.
[0074] Furthermore, the pre-trained agent can be deployed to a simulated R&D environment for online interactive training. This simulated R&D environment can receive a large amount of sample R&D data input by R&D personnel and convert it into corresponding vector inputs to the pre-trained agent, so that the pre-trained agent outputs action sequences for R&D personnel to evaluate. The evaluation results are used to determine the latest reward and feed it back to the agent.
[0075] Finally, the intelligent agent that has completed online interactive training can be deployed in the system as a reinforcement learning model for R&D personnel to use daily. In addition, when the reinforcement learning model generates each cosmetic R&D recommendation scheme after deployment, it can also simultaneously receive the evaluation of the generated cosmetic R&D recommendation scheme input by the R&D personnel. The evaluation results can be used to fine-tune the reinforcement learning model based on real scenarios, so that the reinforcement learning model can continuously learn and iterate, and improve the effectiveness and feasibility of the cosmetic R&D recommendation schemes it generates.
[0076] Secondly, correspondingly, this application also provides a cosmetic R&D management device based on reinforcement learning, which can realize all the processes of the cosmetic R&D management method based on reinforcement learning provided in the above embodiments.
[0077] See Figure 2The diagram shows a structural schematic of a reinforcement learning-based cosmetic R&D management device 200 provided in an embodiment of this application. The device includes: The preliminary research and development data acquisition module 201 is used to acquire cosmetic research and development target information, and use it as the basis for querying N cosmetic research and development datasets to obtain N preliminary research and development data, where N is a positive integer. The R&D description set determination module 202 is used to determine M R&D description sets composed of N R&D description information based on the N preliminary R&D data, wherein M is a positive integer and less than N, the N R&D description information corresponds one-to-one with the N preliminary R&D data, and each R&D description information is determined according to its corresponding preliminary R&D data; The target R&D description information determination module 203 is used to perform reasoning analysis based on each R&D description set, and to determine the target R&D description information corresponding to each R&D description set according to each R&D description set and its reasoning analysis results. The scheme generation module 204 is used to generate a cosmetic research and development recommendation scheme that matches the cosmetic research and development target information by calling a reinforcement learning model based on the target research and development description information corresponding to each of the M research and development description sets.
[0078] In one optional implementation, the reasoning analysis based on each of the R&D description sets includes: For each of the aforementioned R&D description sets, Typical R&D description information and at least one supplementary R&D description information are selected from the R&D description set. The supplementary R&D description information refers to R&D description information whose correlation with the typical R&D description information meets the first correlation condition. Based on the typical R&D description information and the at least one supplementary R&D description information, reasoning hint information and reasoning counterexample information are constructed through semantic similarity analysis; Based on the R&D description set, the inference hints, and the inference counterexamples, inference analysis is performed.
[0079] In one optional implementation, the step of constructing reasoning hints and counterexamples based on the typical R&D description information and the at least one supplementary R&D description information through semantic similarity analysis includes: Determine the first relevant semantic relationship between the typical R&D description information and each of the supplementary R&D description information; Based on each of the first relevant semantics, prompting R&D description information is selected from the at least one supplementary R&D description information; Based on the aforementioned prompt description information, the reasoning prompt information is constructed; Based on the supplementary R&D description information other than the prompt R&D description information in the at least one supplementary R&D description information, the inference counterexample information is constructed.
[0080] In one optional implementation, the step of filtering out prompting R&D description information from the at least one supplementary R&D description information based on each of the first relevant semantics includes: For each of the first related semantics, a first semantic similarity is determined between the first related semantics and the other first related semantics, and the larger of the largest first semantic similarity and the preset semantic similarity threshold is taken as the second semantic similarity corresponding to the first related semantics; Determine a fourth semantic similarity between a second related semantic and each of the first related semantics, wherein the second related semantic is a first related semantic corresponding to a third semantic similarity, and the third semantic similarity is the largest second semantic similarity; The third related semantics are determined based on the first related semantics corresponding to each of the fourth semantic similarities that are not less than the third semantic similarity among all the fourth semantic similarities; Supplementary R&D description information corresponding to the third related semantic is selected from the at least one supplementary R&D description information and used as the prompt R&D description information.
[0081] In one optional implementation, the reasoning analysis based on the R&D description set, the reasoning hints, and the reasoning counterexamples includes: Based on this R&D description set, the first inference result is obtained; Based on the reasoning hints, a portion of the information in the first reasoning result is modified to obtain the second reasoning result; Based on the difference between the second reasoning result and the first reasoning result, the R&D description set is updated; Based on the aforementioned counterexample information and the updated R&D description set, inference analysis is performed.
[0082] In one optional implementation, updating the R&D description set based on the difference between the second inference result and the first inference result includes: Based on the difference between the second reasoning result and the first reasoning result, first difference information is generated; Self-attention calculation is performed on the first difference information to obtain the second difference information; The second difference information is embedded into the R&D description set to update the R&D description set.
[0083] In one optional implementation, determining the target R&D description information corresponding to each R&D description set based on each R&D description set and its inference analysis results includes: For each R&D description set, based on its reasoning and analysis results, target R&D description information corresponding to that R&D description set is obtained from that R&D description set.
[0084] In one optional implementation, the step of generating a cosmetics R&D recommendation scheme that matches the cosmetics R&D target information by invoking a reinforcement learning model based on the target R&D description information corresponding to each of the M R&D description sets includes: Based on the target R&D description information corresponding to each of the M R&D description sets, multiple target R&D data are obtained by querying the cosmetic formula knowledge base, wherein the cosmetic formula knowledge base stores cosmetic formula records that have been experimentally verified. Based on the aforementioned multiple target R&D data, a reinforcement learning model is invoked to generate the cosmetic R&D recommendation scheme.
[0085] Thirdly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any of the above-mentioned embodiments.
[0086] Fourthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method described in any of the above-described embodiments.
[0087] Fifthly, embodiments of this application provide a computer device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the steps of the method described in any of the preceding claims.
[0088] See Figure 3 The computer device in this embodiment includes a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301, such as a reinforcement learning-based cosmetic R&D management program. When the processor 301 executes the computer program, it implements the steps described in the various reinforcement learning-based cosmetic R&D management method embodiments above, for example... Figure 1 The steps S101-S104 are shown.
[0089] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 302 and executed by the processor 301 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device.
[0090] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that the schematic diagram is merely an example of a computer device and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.
[0091] The processor 301 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 301 can be any conventional processor. The processor 301 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines.
[0092] The memory 302 can be used to store the computer programs and / or modules. The processor 301 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 302 and calling the data stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0093] Wherein, if the modules / units integrated into the computer device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a non-transitory computer-readable storage medium. When the computer program is executed by the processor 301, it can implement the steps of the various method embodiments described above. Wherein, the computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0094] In summary, the embodiments of this application have at least the following beneficial effects: Using the embodiments of this application, cosmetic R&D target information is obtained and used as the basis for querying. N preliminary R&D data are obtained by querying N cosmetic R&D datasets, where N is a positive integer. Based on the N preliminary R&D data, M R&D description sets are determined, each consisting of N R&D description information, where M is a positive integer less than N. Each of the N R&D description information corresponds one-to-one with one of the N preliminary R&D data, and each R&D description information is determined based on its corresponding preliminary R&D data. Reasoning analysis is performed based on each R&D description set, and target R&D description information corresponding to each R&D description set is determined based on each R&D description set and its reasoning analysis results. Based on the target R&D description information corresponding to each of the M R&D description sets, a reinforcement learning model is invoked to generate a cosmetic R&D recommendation scheme that matches the cosmetic R&D target information. Thus, a highly automated R&D management system can be formed. After a user proposes their cosmetic R&D target, the system can automatically and efficiently obtain the necessary R&D data from various datasets based on the cosmetic R&D target information, thereby improving the feasibility of the automatically recommended scheme.
[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware platforms, or it can be implemented entirely by hardware. Based on this understanding, all or part of the technical solutions of this application that contribute to the background technology can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0096] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A cosmetic R&D management method based on reinforcement learning, characterized in that, include: Obtain cosmetic research and development target information, use it as the basis for querying, and perform queries on N cosmetic research and development datasets to obtain N preliminary research and development data, where N is a positive integer; Based on the N preliminary R&D data, M R&D description sets are determined, each consisting of N R&D description information, where M is a positive integer less than N. The N R&D description information corresponds one-to-one with the N preliminary R&D data, and each R&D description information is determined based on its corresponding preliminary R&D data. Reasoning analysis is performed based on each of the aforementioned R&D description sets, and the target R&D description information corresponding to each of the aforementioned R&D description sets is determined based on each of the aforementioned R&D description sets and its reasoning analysis results. Based on the target R&D description information corresponding to each of the M R&D description sets, a reinforcement learning model is invoked to generate a cosmetic R&D recommendation scheme that matches the cosmetic R&D target information.
2. The method according to claim 1, characterized in that, The reasoning analysis based on each of the aforementioned R&D description sets includes: For each of the aforementioned R&D description sets, Typical R&D description information and at least one supplementary R&D description information are selected from the R&D description set. The supplementary R&D description information refers to R&D description information whose correlation with the typical R&D description information meets the first correlation condition. Based on the typical R&D description information and the at least one supplementary R&D description information, reasoning hint information and reasoning counterexample information are constructed through semantic similarity analysis; Based on the R&D description set, the inference hints, and the inference counterexamples, inference analysis is performed.
3. The method according to claim 2, characterized in that, The process of constructing reasoning hints and counterexamples based on the typical R&D description information and the at least one supplementary R&D description information through semantic similarity analysis includes: Determine the first relevant semantic relationship between the typical R&D description information and each of the supplementary R&D description information; Based on each of the first relevant semantics, prompting R&D description information is selected from the at least one supplementary R&D description information; Based on the aforementioned prompt description information, the reasoning prompt information is constructed; Based on the supplementary R&D description information other than the prompt R&D description information in the at least one supplementary R&D description information, the inference counterexample information is constructed.
4. The method according to claim 3, characterized in that, The step of filtering out prompting R&D description information from the at least one supplementary R&D description information based on each of the first relevant semantics includes: For each of the first related semantics, a first semantic similarity is determined between the first related semantics and the other first related semantics, and the larger of the largest first semantic similarity and the preset semantic similarity threshold is taken as the second semantic similarity corresponding to the first related semantics; Determine a fourth semantic similarity between a second related semantic and each of the first related semantics, wherein the second related semantic is a first related semantic corresponding to a third semantic similarity, and the third semantic similarity is the largest second semantic similarity; The third related semantics are determined based on the first related semantics corresponding to each of the fourth semantic similarities that are not less than the third semantic similarity among all the fourth semantic similarities; Supplementary R&D description information corresponding to the third related semantic is selected from the at least one supplementary R&D description information and used as the prompt R&D description information.
5. The method according to claim 2, characterized in that, The reasoning analysis based on the R&D description set, the reasoning hints, and the reasoning counterexamples includes: Based on this R&D description set, the first inference result is obtained; Based on the reasoning hints, a portion of the information in the first reasoning result is modified to obtain the second reasoning result; Based on the difference between the second reasoning result and the first reasoning result, the R&D description set is updated; Based on the aforementioned counterexample information and the updated R&D description set, inference analysis is performed.
6. The method according to claim 5, characterized in that, The step of updating the R&D description set based on the difference between the second inference result and the first inference result includes: Based on the difference between the second reasoning result and the first reasoning result, first difference information is generated; Self-attention calculation is performed on the first difference information to obtain the second difference information; The second difference information is embedded into the R&D description set to update the R&D description set.
7. The method according to claim 1, characterized in that, The step of determining the target R&D description information corresponding to each R&D description set based on each R&D description set and its reasoning analysis results includes: For each R&D description set, based on its reasoning and analysis results, target R&D description information corresponding to that R&D description set is obtained from that R&D description set.
8. The method according to claim 1, characterized in that, The step of generating a cosmetics R&D recommendation scheme that matches the cosmetics R&D target information by calling a reinforcement learning model based on the target R&D description information corresponding to each of the M R&D description sets includes: Based on the target R&D description information corresponding to each of the M R&D description sets, multiple target R&D data are obtained by querying the cosmetic formula knowledge base, wherein the cosmetic formula knowledge base stores cosmetic formula records that have been experimentally verified. Based on the aforementioned multiple target R&D data, a reinforcement learning model is invoked to generate the cosmetic R&D recommendation scheme.
9. A cosmetic R&D management device based on reinforcement learning, characterized in that, include: The preliminary research and development data acquisition module is used to acquire cosmetic research and development target information, and use it as the basis for querying N cosmetic research and development datasets to obtain N preliminary research and development data, where N is a positive integer. The R&D description set determination module is used to determine M R&D description sets composed of N R&D description information based on the N preliminary R&D data, where M is a positive integer and less than N, and the N R&D description information corresponds one-to-one with the N preliminary R&D data, and each R&D description information is determined according to its corresponding preliminary R&D data; The target R&D description information determination module is used to perform reasoning analysis based on each R&D description set, and to determine the target R&D description information corresponding to each R&D description set according to each R&D description set and its reasoning analysis results. The solution generation module is used to generate a cosmetic research and development recommendation solution that matches the cosmetic research and development target information by calling a reinforcement learning model based on the target research and development description information corresponding to each of the M research and development description sets.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-8.