A method for synthesizing nutritional value data of feed raw materials based on large language model
Through the feed raw material nutritional value data synthesis method based on the big language model, the problem of insufficient effective nutrient database data in the existing technology is solved, and the efficient synthesis of feed raw material nutritional value data is achieved, providing new ideas for the precise nutrition and efficient utilization of resources in the feed industry.
Patent Information
- Application Number
- CN202510096754.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing effective nutrient database data of feed raw materials is relatively small, resulting in poor prediction accuracy of the nutritional value prediction model of feed raw materials and insufficient dynamic degree, which limits the application of machine learning algorithms in the nutritional value prediction of feed raw materials.
The nutritional value data synthesis method of feed raw materials based on large language models is adopted. By collecting and organizing professional literature in the animal husbandry field, a question-and-answer text data set, a professional vocabulary labeling data set and a instruction fine-tuning data set are constructed, and the large language model is fine-tuned so that it can identify and understand professional terms in the field of animal nutrition and feed, and through the cross entropy loss function and RMSE loss function, a model can be generated that can synthesize the nutritional value data of feed raw materials.
It overcomes the disadvantage of difficulty in synthesising nutritional value data of feed raw materials and is difficult to utilize natural language information and real information, and provides a new method for synthesis of nutritional value data of feed raw materials, providing new data and new ideas for the precise nutrition and efficient utilization of resources in the feed industry.
Smart Images

Figure CN119538986B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of animal nutrition and feed, and specifically is a method for synthesizing feed raw material nutritional value data based on a large language model. Background Art
[0002] Food security is mainly manifested in feed grain security. Under the current background of diversified feed formula, the variety of feed raw materials and the large variation of nutritional value have seriously affected their accurate use in the formula, resulting in a large waste of feed resources. Therefore, the accurate grasp of the variation law of effective nutrients in feed raw materials is the premise for achieving accurate nutrition of livestock and poultry and efficient utilization of feed resources, and is of great significance for solving the problem of feed grain safety from the "saving" end. The current publicly available effective nutrient databases of feed raw materials generally have the defect of less data. Limited by complex measurement methods and limited measurement conditions, the effective nutrient data such as feed effective energy and available amino acids obtained based on target animal tests are more scarce, resulting in the reported feed raw material effective nutrient prediction models being based on traditional linear regression algorithms, with poor model prediction accuracy and insufficient dynamic degree. At present, many machine learning algorithms are applied to various complex scenarios to achieve more accurate predictions, but the premise for machine learning algorithms to play a good role is to be supported by a large amount of basic data. The current situation that basic data is difficult to obtain also limits its application in the prediction of the nutritional value of feed raw materials. If complete nutritional value data can be generated through data synthesis based on the existing database, it will play a vital role in increasing the amount of modeling data, applying machine learning algorithms, and building more accurate prediction models in the field of animal nutrition.
[0003] The nutritional value data of feed raw materials do not exist in isolation. Each piece of data is closely linked to a series of basic information of feed raw materials, including the wet chemical determination method and determination basis (dry matter basis or feeding basis) of the nutritional value of raw materials, raw material varieties, origin, harvest season, storage technology, processing technology, anti-nutritional factor content, physical properties (thousand-grain weight, color), etc. When feed raw materials are used on animals in the form of feed formula, they are linked to the information of the formula and animals, such as the upper and lower limit ratios used in the formula, the matching taboos between raw materials, the breed, physiological stage, and breeding technology of the applied animals, etc. Most of the above information is presented in the form of natural language, and a lot of information comes from expert experience. It is difficult for traditional database tables to carry the complex relationship between these data and information. When synthesizing the nutritional value data of feed raw materials, it is necessary to fully explore and consider the above information in the form of natural language. For the missing data, there should also be scientific methods to clean or fill the data based on the related information. In addition, the synthesis of the nutritional value data of feed raw materials depends on the existing data information and needs to consider its practical significance, so randomly generated data is completely undesirable. The emergence of large language model technology represented by generative artificial intelligence provides an effective solution for scientific reasoning based on a small amount of real data, relying on related raw material formulas and animal information, and ultimately achieving accurate synthesis of feed raw material nutritional value data. Summary of the invention
[0004] The present invention provides a method for synthesizing feed raw material nutritional value data based on a large language model, which is used to solve the key technical bottlenecks of small available data scale and limited modeling accuracy faced when constructing a dynamic prediction model for effective nutrients in feed raw materials.
[0005] The present invention is achieved through the following technical solutions:
[0006] A method for synthesizing feed raw material nutritional value data based on a large language model comprises the following steps:
[0007] Step 1: Data collection for fine-tuning the large language model: collect professional literature in the field of animal husbandry (including Chinese and English textbooks, Chinese and English scientific research papers, and Chinese and English dissertations), organize and collect literature, and build a question-answering text dataset, a professional vocabulary annotation dataset, and an instruction fine-tuning dataset;
[0008] Step 2: Use the question-answering text dataset, professional vocabulary annotation dataset, and instruction fine-tuning dataset to select a lightweight open source large language model for fine-tuning. Use three large language model fine-tuning methods, namely question-answering fine-tuning, professional vocabulary annotation, and instruction fine-tuning, to fine-tune the large language model and generate a lightweight special large language model that can recognize and understand professional terms in the field of animal nutrition and feed.
[0009] Step 3: Collect at least 100 papers on the determination of the nutritional value of feed, manually extract the nutritional value data of feed raw materials and the associated basic information of raw materials, animal test information and feed formula information, and use the question-answering fine-tuning method to enable the lightweight special large language model to master the ability to extract the nutritional value data of feed raw materials and the associated basic information of raw materials, animal test information and feed formula information from text materials, use the cross entropy loss function as the loss function of text data, and use the RMSE loss function as the loss function of numerical data for model training; use the BERTScore method and the manual evaluation method to judge the accuracy of the data extracted by the lightweight special large language model;
[0010] Step 4: Use the lightweight dedicated large language model obtained in step 3 to extract feed raw material nutritional value data and related raw material basic information, animal test information and feed formula information from more public information in turn, and construct a raw material nutritional value data set for the synthesis of feed raw material nutritional value data;
[0011] Step 5: Randomly extract data from the basic information of raw materials as the real raw material composition data, and input it into the lightweight special large language model obtained in step 3 together with the animal test information and formula information. Through question-answering fine-tuning, instruct the large language model to generate the corresponding raw material nutritional value data, calculate the root mean square error value with the true value, and backpropagate the model.
[0012] As described above, in the method for synthesizing the nutritional value data of feed raw materials based on a large language model, the question-and-answer text dataset in step 1 is constructed by extracting questions and answers from professional books, journal articles, and research reports; and inviting field experts to ask questions and provide answers based on specific topics.
[0013] In the method for synthesizing the nutritional value data of feed raw materials based on a large language model as described above, in step 1, the processed text data is converted into the CSV format required for model training.
[0014] As described above, in the method for synthesizing the nutritional value data of feed raw materials based on a large language model, the construction of the professional vocabulary annotation data set in the step 2 is based on annotating important terms and professional vocabulary using the BIO annotation method, so that the open source large language model can recognize and understand these key terms.
[0015] In the above-mentioned method for synthesizing the nutritional value data of feed raw materials based on a large language model, the construction of the instruction fine-tuning dataset in step 2 is based on jointly constructing the dataset with domain experts to ensure that the task examples corresponding to the instructions can truly reflect the actual application scenarios. The dataset includes multiple types of instructions and ensures that the difficulty is diverse, covering simple to complex tasks.
[0016] As described above, a method for synthesizing the nutritional value data of feed raw materials based on a large language model, the nutritional value data of the feed raw materials in step 2 includes the digestible energy, metabolizable energy, net energy, amino acid digestibility, crude protein digestibility, calcium digestibility and phosphorus digestibility of the feed raw materials, which are used as model prediction values.
[0017] As described above, a method for synthesizing the nutritional value data of feed raw materials based on a large language model, the animal test information includes the animal house where each test is conducted, the start and end time of the test, the test animal information used in the test, and the variety, origin, processing method and special ingredients of the raw materials used in the test.
[0018] As described above, in the method for synthesizing the nutritional value data of feed raw materials based on a large language model, the basic information of raw materials in step three includes the dry and wet material basis, water content, and basic nutrients of the feed raw materials.
[0019] In the above-mentioned method for synthesizing the nutritional value data of feed raw materials based on a large language model, in steps 2, 3, and 5, Alibaba's open source Qwen2-72B is used as the basic model, and the learning rate is set to 1e-5 for model training.
[0020] The advantages of the present invention are: based on the large language model technology, the present invention utilizes professional literature knowledge in the field of animal husbandry to enable the model to master expert knowledge, and then uses feed research articles from major universities and scientific research institutions and feed nutritional value information extracted from the articles by professionals for fine-tuning, so that the model can synthesize feed nutritional value data based on a small amount of prompt information and real raw material composition data, and synthesize feed raw material nutritional value data based on animal test information, formula information and part of the real raw material composition data, which overcomes the shortcomings of difficulty in utilizing natural language information and real information when synthesizing feed raw material nutritional value data, and proposes a new method for synthesizing feed raw material nutritional value data, which provides new data and new ideas for predicting the nutritional value of feed raw materials, and solves the pain point of the lack of information synthesis methods in the feed industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0022] Figure 1 It is the overall flow chart of the present invention;
[0023] Figure 2It is an operation flow chart of step three of the present invention;
[0024] Figure 3 It is an operational flow chart of step five of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] like Figure 1 As shown, a method for synthesizing feed raw material nutritional value data based on a large language model comprises the following steps:
[0027] Step 1: Data collection for fine-tuning the large language model: collect professional literature in the field of animal husbandry (including Chinese and English textbooks, Chinese and English scientific research papers, and Chinese and English dissertations), organize and collect literature, and build a question-answering text dataset, a professional vocabulary annotation dataset, and an instruction fine-tuning dataset;
[0028] Step 2: Use the question-answering text dataset, professional vocabulary annotation dataset, and instruction fine-tuning dataset to select a lightweight open source large language model for fine-tuning. Use three large language model fine-tuning methods, namely question-answering fine-tuning, professional vocabulary annotation, and instruction fine-tuning, to fine-tune the large language model and generate a lightweight special large language model that can recognize and understand professional terms in the field of animal nutrition and feed.
[0029] like Figure 2 As shown, step three: collect at least 100 papers on the determination of feed nutritional value, manually extract feed raw material nutritional value data and related raw material basic information, animal test information and feed formula information, and use question-answer fine-tuning to enable the lightweight special large language model to master the ability to extract feed raw material nutritional value data and related raw material basic information, animal test information and feed formula information from text materials, use the cross entropy loss function as the loss function of text data, and use RMSE as the loss function of numerical data for model training; use the BERTScore method and the manual evaluation method to judge the accuracy of the data extracted by the lightweight special large language model;
[0030] Step 4: Use the lightweight dedicated large language model obtained in step 3 to extract feed raw material nutritional value data and related raw material basic information, animal test information and feed formula information from more public information in turn, and construct a raw material nutritional value data set for the synthesis of feed raw material nutritional value data;
[0031] like Figure 3 As shown, step five: randomly extract data from the basic information of raw materials as the real raw material composition data, and input them into the lightweight special large language model obtained in step three together with the animal test information and the formula information. Through question-answering fine-tuning, instruct the large language model to generate the corresponding raw material nutritional value data, and calculate the root mean square error value with the true value, and back-propagate the model; the large language model obtained through this step can obtain a lightweight special large language model for synthetic raw material nutritional value data by setting animal test related information such as test conditions, animal breed information, animal treatment, and setting formula composition information plus part of the raw material composition data.
[0032] Specifically, the question-and-answer text dataset in step 1 of this embodiment is constructed by extracting questions and answers from professional books, journal articles, and research reports; and inviting field experts to ask questions and provide answers based on specific topics (such as feed ingredients, nutritional requirements, etc.).
[0033] Specifically, in step 1 described in this embodiment, the processed text data is converted into the CSV format required for model training.
[0034] More specifically, the construction of the professional vocabulary annotation dataset in step 2 described in this embodiment is based on annotating important terms and professional vocabulary using the BIO (Begin, Inside, Outside) annotation method, so that the open source large language model can recognize and understand these key terms.
[0035] More specifically, the construction of the instruction fine-tuning dataset in step 2 of this embodiment is based on jointly building the dataset with domain experts to ensure that the task examples corresponding to the instructions can truly reflect the actual application scenarios. The dataset includes various types of instructions (such as text generation, text summarization, classification, etc.) and ensures that the difficulty is diverse, covering simple to complex tasks.
[0036] Furthermore, the feed raw material nutritional value data in step 2 of this embodiment includes digestible energy, metabolizable energy, net energy, amino acid digestibility, crude protein digestibility, calcium digestibility and phosphorus digestibility of the feed raw material, which are used as model prediction values.
[0037] Preferably, the feed raw material nutritional value data information items described in this embodiment are shown in the following table:
[0038]
[0039] Furthermore, the animal test information described in this embodiment includes the animal house where each test is conducted, the start and end time of the test, the test animal information used in the test, and the type, origin, processing method and special ingredients of the raw materials used in the test.
[0040] Furthermore, the raw material basic information in step three of this embodiment includes the dry and wet material basis, water content, and basic nutrients of the feed raw materials.
[0041] Furthermore, in steps 2, 3, and 5 described in this embodiment, Alibaba's open source Qwen2-72B is used as the basic model, and the learning rate is set to 1e-5 for model training.
[0042] Example: The doctoral thesis published by our team in 2019, "Study on the Effective Energy and Terminal Ileal Amino Acid Digestibility of Enzymatic Soy Protein (HP300) in Pigs and Its Efficient Application in Piglets", was used as an implementation case.
[0043] 1. Extract data
[0044] 1.1 Extracting Large Language Model Fine-tuning Data
[0045] Example of command data set:
[0046] Instructions: Calculate the metabolizable energy value of enzymatically hydrolyzed soy protein (ESBM).
[0047] Input: The article mentioned that the total energy value of enzymatically hydrolyzed soy protein is 19.20 MJ / kg, and the estimated formula for metabolizable energy value is ME = DE × 0.9.
[0048] Output: The metabolizable energy value of enzymatically hydrolyzed soy protein is approximately 17.28 MJ / kg.
[0049] Example of question answering dataset:
[0050] Background: Enzymatic soy protein is made from soy protein or soybean meal, which is transformed by bioengineering technology and dried at high temperature.
[0051] Question: What are the nutritional characteristics of enzymatic soy protein?
[0052] Answer: Small peptide molecules that are easily absorbed; rich in bioactive factors, such as probiotics and antioxidant ingredients; digestible and metabolizable energy value is higher than soybean meal and fish meal.
[0053] Examples of professional vocabulary annotation (BIO method):
[0054] { "sentence": "The experiment showed that the digestible metabolizable energy value of enzymatically hydrolyzed soy protein (ESBM) is higher than that of soybean meal, soybean protein concentrate and fish meal. Adding ESBM to piglet diet can significantly improve growth performance and immune level.", "tokens":["experiment", "shows", "enzymatically hydrolyzed soy protein", "(", "ESBM", ")", "of", "digestible metabolizable energy value", "higher than", "soy meal", "soy protein concentrate", "and", "fish meal", "piglet", "diet", "in", "add", "ESBM", "can", "significantly", "improve", "growth performance", "and", "immune level"], "tags": ["O", "O", "B-Ingredient", "O", "B-Abbreviation", "O", "O", "B-NutrientValue", "O","B-Ingredient", "B-Ingredient", "O", "B-Ingredient", "B-Animal", "B-Diet", "O", "O", "B-Abbreviation", "O", "O", "B-Performance", "O", "B-HealthImpact"]}
[0055] B-Ingredient: Feed raw materials (such as soybean meal, enzymatic soy protein).
[0056] B-NutrientValue: Nutritional value description (such as digestible metabolizable energy value).
[0057] B-Performance: Terms related to growth performance.
[0058] B-HealthImpact: Description of health impact.
[0059] 1.2 Raw material nutritional value dataset extraction
[0060] Extract raw material composition data (as shown in Table 1), animal test information data (as shown in Table 2), formula information data (as shown in Table 3) and nutritional value data (as shown in Table 4) respectively.
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067] 2.1 Large Language Model Fine-tuning
[0068] Use the question-answer text dataset, professional vocabulary annotation dataset, and instruction fine-tuning dataset obtained in 1.1 to select a lightweight open source large language model for fine-tuning. Use three large language model fine-tuning methods: question-answer fine-tuning, professional vocabulary annotation, and instruction fine-tuning to fine-tune the large language model to generate a lightweight, dedicated large language model that can recognize and understand professional terms in the field of animal nutrition and feed.
[0069] 2.2 Training an open source large language model capable of extracting the nutritional value of feed ingredients
[0070] Using this doctoral thesis as input and the above four tables as theoretical output, the big language model can extract these data, and then obtain an open source big language model for extracting the nutritional value of feed raw materials.
[0071] 3. Training the nutritional value data synthesis model
[0072] 3.1 Extracting Data for Data Synthesis
[0073] Randomly select raw materials from the raw material composition data, such as raw material No. 3, and randomly extract the raw material composition in No. 3 to obtain the raw material composition data as shown in Table 5.
[0074]
[0075] The animal test information related to raw material number 3 was randomly selected, and the test information selected was test number 2, and its basic information is shown in Table VI.
[0076]
[0077] By querying the formula data, the raw material with raw material number 3, under the condition of test number 2, uses formula number 3, so the corresponding formula information is extracted as shown in Table 7.
[0078]
[0079] According to the experimental information model, the energy data of the corresponding raw materials should be generated, so the energy data corresponding to the raw materials are extracted, as shown in Table 8.
[0080]
[0081] 3.2 Using Extracted Data for Data Synthesis Training of Models
[0082] After the above steps, we get the three input information required for model training, namely, raw material composition data (Table 5), test information (Table 6), and formula information (Table 7). Inputting these data into the large language model should generate the corresponding raw material energy data. The model output values are shown in Table 9.
[0083]
[0084] According to the difference between the actual output value of the model (Table 9) and the true value (Table 8), the RMSE loss function is used to calculate the error value of the model. The large language model is back-propagated through the loss function to enable the large language model to master the ability to synthesize feed nutritional value data, thereby obtaining a large language model that can synthesize feed raw material nutritional value data.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for synthesizing feed raw material nutritional value data based on a large language model, characterized by: The steps include: Step 1: Data collection for fine-tuning of large language models: collect professional literature in the field of animal husbandry, organize and collect literature, build question-answering text datasets, professional vocabulary annotation datasets, and instruction fine-tuning datasets; Step 2: Use the question-answering text dataset, professional vocabulary annotation dataset, and instruction fine-tuning dataset to select a lightweight open source large language model for fine-tuning. Use three large language model fine-tuning methods, namely question-answering fine-tuning, professional vocabulary annotation, and instruction fine-tuning, to fine-tune the large language model and generate a lightweight special large language model that can recognize and understand professional terms in the field of animal nutrition and feed. Step 3: Collect at least 100 papers on the determination of feed nutritional value, manually extract feed raw material nutritional value data and related raw material basic information, animal test information and feed formula information, and use question-answering fine-tuning to enable the lightweight dedicated large language model to master the ability to extract feed raw material nutritional value data and related raw material basic information, animal test information and feed formula information from text materials, and use the cross entropy loss function as the loss function for text data and RMSE as the loss function for numerical data for model training; Step 4: Use the lightweight dedicated large language model obtained in step 3 to extract feed raw material nutritional value data and related raw material basic information, animal test information and feed formula information from more public information in turn, and construct a raw material nutritional value data set for the synthesis of feed raw material nutritional value data; Step 5: Randomly extract data from the basic information of raw materials as the real raw material composition data, and input it into the lightweight special large language model obtained in step 3 together with the animal test information and formula information. Through question-answering fine-tuning, instruct the large language model to generate the corresponding raw material nutritional value data, and calculate the root mean square error value with the real value, and perform back propagation on the model; The feed raw material nutritional value data in step 2 includes digestible energy, metabolizable energy, net energy, amino acid digestibility, crude protein digestibility, calcium digestibility and phosphorus digestibility of the feed raw material, which are used as model prediction values; The feed raw material composition data include: gross energy, crude protein, crude fat, crude fiber, acid detergent fiber, neutral detergent fiber, ash, moisture, carbohydrates, starch, 18 kinds of amino acids, vitamins and minerals; The animal experiment information includes the animal house where each experiment was conducted, the start and end time of the experiment, the information of the experimental animals used in the experiment, and the species, origin, processing method and special ingredients of the raw materials used in the experiment; The raw material basic information in step three includes the dry and wet material basis, water content, and basic nutrients of the feed raw materials.
2. The method for synthesizing feed raw material nutritional value data based on a large language model according to claim 1, characterized in that: In the step 1, the question-answering text dataset is constructed by extracting questions and answers from professional books, journal articles, and research reports; and inviting domain experts to ask questions and provide answers based on specific topics.
3. The method for synthesizing feed raw material nutritional value data based on a large language model according to claim 1, characterized in that: In the step 1, the processed text data is converted into the CSV format required for model training.
4. The method for synthesizing feed raw material nutritional value data based on a large language model according to claim 1, characterized in that: The construction of the professional vocabulary annotation dataset in the step 2 is based on annotating important terms and professional vocabulary using the BIO annotation method, so that the open source large language model can recognize and understand these key terms.
5. The method for synthesizing feed raw material nutritional value data based on a large language model according to claim 1, characterized in that: The construction of the instruction fine-tuning dataset in step 2 is based on jointly building the dataset with domain experts to ensure that the task examples corresponding to the instructions can truly reflect the actual application scenarios. The dataset includes multiple types of instructions and ensures that the difficulty is diverse, covering simple to complex tasks.
6. The method for synthesizing feed raw material nutritional value data based on a large language model according to claim 1, characterized in that: In steps 2, 3, and 5, Alibaba's open source Qwen2-72B is used as the basic model, and the learning rate is set to 1e-5 for model training.
Citation Information
Patent Citations
Academic question and answer model training method and device, answer generation method and device and related products
CN119311813A