Method and system for quantitatively assessing dairy product similarity
By obtaining the nutrient content values, weights, and fluctuation ranges of dairy products, and combining them with weighted deviation values and machine learning models, the problem of low assessment accuracy in existing technologies has been solved, achieving a more accurate assessment of dairy product similarity.
Patent Information
- Application Number
- CN202511225536.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing methods for assessing the similarity of dairy products use fixed standard values, which fail to reflect changes in the nutritional composition of reference dairy products and do not distinguish the importance of different nutrients, resulting in low assessment accuracy.
By obtaining the nutrient content values, importance weights, and content fluctuation ranges of the reference dairy products, a weighted deviation value is calculated. Combined with the cumulative density distribution curve and machine learning model, the final similarity score is determined.
It improves the accuracy of dairy product similarity assessment, comprehensively reflects the overall similarity between the dairy product to be assessed and the reference dairy product, and provides a scientific basis for judgment.
Smart Images

Figure CN120724178B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a dairy product similarity quantitative evaluation method and system. BACKGROUND
[0002] In the field of dairy products, the development of products often needs to take the nutritional composition of a certain reference dairy product as the target. In order to measure the closeness between the dairy product to be evaluated and the reference dairy product, and provide a basis for product optimization, the industry needs to evaluate the similarity of the nutritional ingredients of the two.
[0003] The prior art first obtains the content values of multiple nutrients in the dairy product to be evaluated through analysis and detection. Then, the content values are compared one by one with a set of pre-set standard values of the nutritional ingredients of the reference dairy product, and the deviation of the content of each nutrient from the corresponding standard value is calculated to determine the similarity between the dairy product to be evaluated and the reference dairy product. However, the content of the nutritional ingredients in the reference dairy product is not constant, and using a fixed standard value as the basis for judgment cannot reflect this objective situation. Secondly, the prior art does not consider the influence of different nutrients on the similarity of the dairy product when comparing the similarity, resulting in low accuracy of the final evaluation. SUMMARY
[0004] The present application provides a dairy product similarity quantitative evaluation method, system, electronic device, storage medium and computer program product, which solves the defects of fixed evaluation standard and not distinguishing the importance of different nutrients in the prior art, thereby improving the evaluation accuracy of the similarity of dairy products.
[0005] The present application provides a dairy product similarity quantitative evaluation method, comprising the following steps:
[0006] obtaining the content value of each nutrient in the dairy product to be evaluated, the influence weight for representing the importance degree of each nutrient, and the content fluctuation range of each nutrient in the reference dairy product;
[0007] For any nutrient, a weighted deviation value is obtained according to the content value, the content fluctuation range and the influence weight; the weighted deviation value is used to represent the content deviation degree of the dairy product to be evaluated and the reference dairy product on the nutrient;
[0008] According to all the weighted deviation values, the final similarity score of the dairy product to be evaluated and the reference dairy product is determined.
[0009] The application provides a method for quantitatively evaluating the similarity of dairy products, comprising the following steps: obtaining a reference nutrient data sample of a reference dairy product sample and a comparison nutrient data sample of a comparison dairy product sample; the reference nutrient data sample and the comparison nutrient data sample each comprise a content value of at least one nutrient substance;
[0010] According to the reference nutrient data sample and the comparison nutrient data sample, a classification model for distinguishing sample sources is trained;
[0011] The feature importance score of each nutrient substance is extracted from the classification model;
[0012] For any nutrient substance, the feature importance score is taken as the influence weight.
[0013] The application provides a method for quantitatively evaluating the similarity of dairy products, further comprising the following steps:
[0014] For any nutrient substance, a cumulative density distribution curve is constructed based on the content value distribution of the nutrient substance in the reference nutrient data sample;
[0015] The slope of each data point on the cumulative density distribution curve is determined;
[0016] The content value corresponding to the first data point is taken as the lower limit value of the content fluctuation range, and the content value corresponding to the second data point is taken as the upper limit value of the content fluctuation range;
[0017] The first data point is a data point from the low end of the cumulative density distribution curve, for which the slope is greater than a first preset threshold value for the first time; and the second data point is a data point after the first data point, for which the slope is less than a second preset threshold value for the first time.
[0018] The application provides a method for quantitatively evaluating the similarity of dairy products, wherein the weighted deviation value is obtained according to the content value, the content fluctuation range and the influence weight, and the method comprises the following steps:
[0019] A content deviation distance is obtained according to the size comparison relationship between the content value and the content fluctuation range;
[0020] The influence weight is multiplied by the content deviation distance to obtain the weighted deviation value.
[0021] The application provides a method for quantitatively evaluating the similarity of dairy products, wherein the content fluctuation range comprises a lower limit value and an upper limit value; and the content deviation distance is obtained according to the size comparison relationship between the content value and the content fluctuation range, and the method comprises the following steps:
[0022] when the content value is between the lower limit value and the upper limit value, the content deviation distance is zero;
[0023] when the content value is less than the lower limit value, the content deviation distance is the absolute value of the difference between the content value and the lower limit value;
[0024] when the content value is greater than the upper limit value, the content deviation distance is the absolute value of the difference between the content value and the upper limit value.
[0025] According to the milk product similarity quantitative evaluation method provided by the application, the final similarity score of the to-be-evaluated milk product and the reference milk product is determined according to all the weighted deviation values, and the method comprises the following steps:
[0026] performing an aggregation operation on all the weighted deviation values to obtain a comprehensive deviation value;
[0027] mapping the comprehensive deviation value to the final similarity score through a preset score function, wherein the final similarity score decreases with the increase of the comprehensive deviation value.
[0028] The application further provides a milk product similarity quantitative evaluation system, comprising the following modules:
[0029] an acquisition module, configured to acquire the content value of each nutrient substance in a to-be-evaluated milk product, an influence weight for representing the importance degree of each nutrient substance, and a content fluctuation range of each nutrient substance in a reference milk product;
[0030] a first processing module, configured to, for any nutrient substance, obtain a weighted deviation value according to the content value, the content fluctuation range and the influence weight; the weighted deviation value is used to represent the content deviation degree of the to-be-evaluated milk product and the reference milk product in the nutrient substance;
[0031] a second processing module, configured to determine a final similarity score of the to-be-evaluated milk product and the reference milk product according to all the weighted deviation values.
[0032] The application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the milk product similarity quantitative evaluation method according to any of the above when executing the program.
[0033] The application further provides a non-transitory computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the milk product similarity quantitative evaluation method according to any of the above.
[0034] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements any of the above-mentioned dairy product similarity quantification evaluation methods.
[0035] To sum up, the one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0036] By obtaining the content value of the dairy product to be evaluated, the influence weight representing the importance degree, and the content fluctuation range of the reference dairy product, the inputs of the three key dimensions of actual value, importance, and target range are prepared for similarity evaluation, overcoming the one-sidedness problem caused by the single evaluation dimension in the traditional method. By calculating the weighted deviation value according to the content value, the content fluctuation range, and the influence weight, the content deviation degree of the nutrient substance is combined with its own importance, so that the difference measurement of a single nutrient substance is no longer a simple numerical distance, but a comprehensive index that can reflect its key influence. By determining the final similarity score according to all the weighted deviation values, the comprehensive deviation in all nutrient dimensions is aggregated into a direct quantitative score, and the score is integrated with the weight and the range in its calculation basis, which can more comprehensively reflect the overall similarity degree between the dairy product to be evaluated and the reference dairy product, thereby significantly improving the accuracy of the evaluation result. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0038] Figure 1 is one of the flowcharts of the dairy product similarity quantification evaluation method provided by the present application.
[0039] Figure 2 is the second flowchart of the dairy product similarity quantification evaluation method provided by the present application.
[0040] Figure 3 is the third flowchart of the dairy product similarity quantification evaluation method provided by the present application.
[0041] Figure 4 is the fourth flowchart of the dairy product similarity quantification evaluation method provided by the present application.
[0042] Figure 5 is the fifth flowchart of the dairy product similarity quantification evaluation method provided by the present application.
[0043] Figure 6 is the sixth flowchart of the method for quantitatively evaluating the similarity of dairy products provided by the present application.
[0044] Figure 7 is a structural schematic diagram of the system for quantitatively evaluating the similarity of dairy products provided by the present application.
[0045] Figure 8 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0046] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. According to the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0047] It should be noted that in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement “comprises a” does not exclude the presence of another identical element in the process, method, article or device comprising the element. The terms “up”, “down” and the like indicate the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the indicated system or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0048] The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by “first”, “second” and the like are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, “and / or” means at least one of the connected objects, and the character “ / ” generally represents a “or” relationship between the front and rear associated objects.
[0049] The following will be described with reference to the drawings Figures 1 to 8The application provides a method, a system, an electronic device, a storage medium and a computer program product for quantitatively evaluating the similarity of dairy products.
[0050] Referring to Figure 1 , Figure 1 is one of the flowcharts of the method for quantitatively evaluating the similarity of dairy products provided by the application, as shown in Figure 1 , comprising steps 101 to 103:
[0051] Step 101: obtaining the content value of each nutrient in the dairy product to be evaluated, the influence weight for representing the importance degree of each nutrient, and the content fluctuation range of each nutrient in the reference dairy product.
[0052] Specifically, in the application, the dairy product to be evaluated is preferably infant formula milk powder, and the reference dairy product is preferably breast milk. Secondly, the dairy product to be evaluated can also be functional dairy products such as children's or elderly formula milk powder, and plant-based dairy products, and the reference dairy product can also be specific stages (such as colostrum, mature milk) or cow milk, goat milk, or a theoretical formula standard of specific nutritional requirements published by an authoritative agency.
[0053] In step 101, first, the content value of each nutrient in the dairy product to be evaluated is obtained. The content value is the direct basis for subsequent comparison and can be obtained by measuring the actual dairy product sample through standard detection means. At the same time, the influence weight for representing the importance degree of each nutrient also needs to be obtained. The influence weight is a quantitative index that reflects the key degree of a certain nutrient in distinguishing dairy products of different sources. The higher the influence weight of a nutrient, the greater the influence of the matching degree of its content on the overall similarity. In addition, the content fluctuation range of each nutrient in the reference dairy product also needs to be obtained. The content fluctuation range is not a single value but an interval, which represents a concentration range commonly found in a large number of reference dairy product samples and provides a target interval in line with biological reality for evaluation.
[0054] In a specific implementation scenario, for a dairy product to be evaluated, the content of multiple key nutrients such as amino acids and high-abundance proteins contained therein can be determined by laboratory analysis techniques such as liquid chromatography. These measurement results are the content values of each nutrient in the dairy product to be evaluated. For each nutrient, for example, lysine, the corresponding influence weight and content fluctuation range can be queried and retrieved from a pre-constructed database or parameter table. These influence weights and content fluctuation ranges are determined in advance based on a large amount of data analysis of reference dairy products.
[0055] Step 102: For any nutrient, a weighted deviation value is obtained according to the content value, the content fluctuation range, and the impact weight; the weighted deviation value is used to represent the content deviation degree of the to-be-evaluated dairy product and the reference dairy product on the nutrient.
[0056] After obtaining the basic data required for evaluation, the matching degree of a single nutrient needs to be quantified. Simply judging whether the content value falls within the content fluctuation range cannot distinguish the deviation impact of different importance levels of nutrients. Therefore, a calculation method that can comprehensively consider the content deviation and the importance level is needed to more accurately measure the difference in a single nutrient dimension.
[0057] In step 102, the weighted deviation value is a comprehensive index calculated according to the three input data obtained in the preceding steps. This weighted deviation value is used to represent the content deviation degree of the to-be-evaluated dairy product and the reference dairy product on the nutrient. A higher weighted deviation value means that the content of the to-be-evaluated dairy product on the nutrient deviates more significantly or critically from the ideal state of the reference dairy product.
[0058] In specific implementation, for any nutrient in the to-be-evaluated dairy product, such as tryptophan, a preset calculation process is performed. This process analyzes the deviation relationship between the content value and the corresponding content fluctuation range, and combines the impact weight of the substance to calculate. Through this calculation, a weighted deviation value of tryptophan can be obtained. The same calculation process is repeated for each nutrient in the analysis list, generating a corresponding weighted deviation value for each nutrient.
[0059] By performing this step, a standardized deviation measure can be generated for each nutrient. The weighted deviation value not only reflects the difference in content, but also takes into account the importance of the nutrient, making the deviation evaluation of a single nutrient more accurate and further making the final similarity score more accurate.
[0060] Step 103: According to all the weighted deviation values, determine the final similarity score of the to-be-evaluated dairy product and the reference dairy product.
[0061] After calculating the corresponding weighted deviation value for each nutrient, a set of discrete deviation indicators is obtained. Although this set of data reflects the differences in each single dimension, it cannot provide a direct conclusion about the overall similarity between the to-be-evaluated dairy product and the reference dairy product. In order to facilitate the comparison of different products or evaluate the overall formula level, it is necessary to integrate these multi-dimensional deviation information into a single final evaluation result.
[0062] In step 103, the final similarity score is a comprehensive numerical value. It is obtained by a comprehensive calculation that aggregates the deviation information represented by the weighted deviation values of all individual nutrients into a final score that can represent the overall similarity level. The score can be designed as an easily understandable scale, for example, 0 to 100 points.
[0063] In a specific embodiment, all the weighted deviation values calculated for all selected key nutrients in the dairy product to be evaluated in the preceding steps are first aggregated. Then, the entire set of weighted deviation values is input as a whole to apply a pre-set determination rule or algorithm. The process processes these values and outputs a single value, which is the final similarity score of the dairy product to be evaluated.
[0064] Through the above steps, the overall closeness of the dairy product to be evaluated to the reference dairy product in terms of key nutrient composition can be accurately and comprehensively reflected, thereby providing direct and effective basis for formula optimization and quality control of the dairy product to be evaluated.
[0065] In a possible embodiment, the method further comprises:
[0066] Step 201: obtaining a reference nutrient data sample of a reference dairy product sample and a comparison nutrient data sample of a comparison dairy product sample; the reference nutrient data sample and the comparison nutrient data sample each include a content value of at least one nutrient.
[0067] Step 202: training a classification model for distinguishing sample sources according to the reference nutrient data sample and the comparison nutrient data sample.
[0068] Step 203: extracting a feature importance score of each nutrient from the classification model.
[0069] Step 204: for any nutrient, taking the feature importance score as an influence weight.
[0070] Specifically, first, a reference nutritional data sample of a reference dairy sample and a to-be-compared nutritional data sample of a to-be-compared dairy sample are obtained. The reference nutritional data sample and the to-be-compared nutritional data sample are the basis for subsequent training of a machine learning model. Both of them include the content value of at least one nutrient. In a specific embodiment, the reference dairy is breast milk, and the to-be-compared dairy is infant formula. Researchers will collect and integrate large-scale nutritional detection data, for example, collect 2000 breast milk samples with wide representation to form a reference nutritional data sample. At the same time, detailed detection data of 22 mainstream infant formulas on the market are collected to form a to-be-compared nutritional data sample. Both data samples contain accurate content values of a series of common key nutrients, such as various amino acids and high-abundance proteins. These data are arranged in a structured format, and each data sample is labeled with its source (breast milk or infant formula). By performing this step, a high-quality, high-dimensional labeled data set can be constructed, which provides a solid data foundation for subsequent machine learning to discover the inherent difference rules between the two types of dairy.
[0071] After obtaining the data set, a tool is needed to automatically learn and quantify these complex differences. Therefore, the next step is to train a classification model for distinguishing the sample source according to the reference nutritional data sample and the to-be-compared nutritional data sample. The classification model is an algorithm model, and its goal is to learn a decision rule so that when a new nutritional data sample is given, it can accurately determine whether the sample is from the reference dairy or the to-be-compared dairy. In a preferred embodiment, the random forest algorithm is selected to build this classification model. In the training process, the labeled data set prepared in the previous step is input into the random forest model. The model makes judgments by building a large number of decision trees and combining the classification results of all decision trees. The training process will continuously adjust the parameters inside the model to make its accuracy in distinguishing the reference nutritional data sample and the to-be-compared nutritional data sample reach the highest. The effect of this step is to obtain a well-trained classification model. The model internalizes the complex nonlinear relationship of distinguishing the nutritional ingredient spectrum of the two types of dairy, so as to objectively identify the sample source.
[0072] After the model is trained, it is like a "black box" itself, although the classification is accurate, but further explore the basis for its decision-making. In order to know which nutrients are most critical to distinguish between the two types of dairy products, the feature importance score of each nutrient needs to be extracted from the classification model. Feature importance score is a numerical value that quantifies the contribution of each input feature (i.e. each nutrient) to the classification model to make accurate predictions. The higher the score, the more critical the nutrient plays in the process of distinguishing. In practice, after the random forest model is trained, its internal algorithm can be directly called to calculate the contribution of each feature. The calculation is usually based on the average degree of reducing impurity of a feature in all decision trees in the model. After calculation, the model outputs a list, in which each nutrient corresponds to a specific feature importance score. By performing this step, an objective ranking of the importance of all nutrients can be generated, providing a direct numerical basis for scientifically assigning weights.
[0073] Finally, for any nutrient, the feature importance score is used as the impact weight. This is an assignment step that confirms the objective quantitative score calculated in the previous step as the impact weight used in the final evaluation model. In practice, the name of each nutrient is bound to its corresponding feature importance score and stored in a parameter table or database. For example, if the feature importance score of lysine is calculated as 0.15, then 0.15 is determined as the impact weight of lysine. When evaluating any to-be-evaluated dairy product later, the impact weight corresponding to each nutrient can be directly queried from the parameter table. Ensuring that the impact weight used in the subsequent similarity calculation is objectively determined based on large-scale real data, thus fundamentally solving the technical problem of subjective weight setting and lack of basis in traditional evaluation methods.
[0074] In one possible implementation, the method further comprises:
[0075] Step 301: For any nutrient, based on the content value distribution of the nutrient in the reference nutrient data sample, a cumulative density distribution curve is constructed;
[0076] Step 302: Determine the slope of each data point on the cumulative density distribution curve;
[0077] Step 303: The content value corresponding to the first data point is taken as the lower limit value of the content fluctuation range, and the content value corresponding to the second data point is taken as the upper limit value of the content fluctuation range; wherein the first data point is the data point from the low end of the cumulative density distribution curve, whose slope is first greater than the first preset threshold; the second data point is the data point after the first data point, whose slope is first less than the second preset threshold.
[0078] Specifically, for any nutrient, a cumulative density distribution curve is constructed based on the distribution of the nutrient's content values in the reference nutrient data samples. The cumulative density distribution curve is a curve that represents how the cumulative percentage of samples with content values below a certain content value changes as the content value increases. In one embodiment, for a certain amino acid in the reference dairy product (e.g., breast milk), first the content values of the amino acid in all 2000 reference nutrient data samples are obtained. Then, the content value range of the amino acid is divided into 100 equal-width intervals. By counting the number of samples falling into each interval, the frequency distribution of the content values can be obtained. Based on the frequency distribution, a cumulative density distribution curve can be plotted, with the content value of the nutrient on the horizontal axis and the cumulative percentage of samples with content values less than or equal to the content value on the vertical axis. Through this step, the discrete and large amount of original data points can be transformed into a continuous curve that can intuitively reflect the data distribution density and trend, providing a necessary basis for subsequent quantitative analysis.
[0079] After obtaining the cumulative density distribution curve, a quantitative method is needed to analyze the shape of the curve, particularly to identify the region where the data points are most dense. The steepness of the curve is a key indicator of the density of the data. Therefore, the next step is to determine the slope of each data point on the cumulative density distribution curve. The slope here, in a physical sense, represents the density or probability density of the data samples near a certain specific content value. The greater the slope, the more concentrated the data points are at that point; the smaller the slope, the more sparse the data points are at that point. In a specific implementation, for each data point (or the center point of each interval) on the cumulative density distribution curve generated in the previous step, the local slope can be calculated by numerical differentiation. For example, the slope of a point can be approximated by calculating the ratio of the difference in the vertical coordinate to the difference in the horizontal coordinate between the point and its adjacent points. Repeating this process for all points on the curve yields a set of slope values corresponding one-to-one with the content values.
[0080] After the slope of all points is calculated, the core range of data concentration can be accurately defined according to the change rule of the slope. The final goal is to take the content value corresponding to the first data point as the lower limit value of the content fluctuation range, and take the content value corresponding to the second data point as the upper limit value of the content fluctuation range. The boundary points here are determined according to the comparison relationship between the slope and the preset threshold. Among them, the first data point is the data point from the low end of the cumulative density distribution curve, whose slope is greater than the first preset threshold for the first time; and the second data point is the data point after the first data point, whose slope is less than the second preset threshold for the first time. In a preferred embodiment, the first preset threshold and the second preset threshold can be determined according to experience or optimization experiment, for example, the first preset threshold is set to 1, and the second preset threshold is set to -1. In specific operation, the program starts from the lowest content value, and checks the slope corresponding to each point point by point. When the first data point whose slope is greater than 1 is found, the content value corresponding to the data point is recorded, and the content value is taken as the lower limit value of the content fluctuation range. Subsequently, the program continues to check from the point to the high content value direction, and when the first data point whose slope is less than -1 is found, the content value corresponding to the data point is recorded, and the content value is taken as the upper limit value of the content fluctuation range. Thus, the content fluctuation range of the nutrient is obtained. This step avoids the limitation of artificially setting a fixed target value by determining the typical content interval of each nutrient in the reference dairy product.
[0081] In a possible implementation, step 102 specifically includes:
[0082] Step 401: obtaining the content deviation distance according to the size comparison relationship between the content value and the content fluctuation range.
[0083] Step 402: multiplying the influence weight and the content deviation distance to obtain the weighted deviation value.
[0084] Specifically, the content deviation distance is a non-negative value, which is used to directly measure the absolute size of the content value of a single nutrient deviating from the ideal target interval. When the content value is exactly within the target interval, the deviation is considered to be zero.
[0085] In specific implementation, for a specific nutrient in the dairy product to be evaluated, for example, lysine, its content value is first obtained, and its corresponding content fluctuation range (including a lower limit value and an upper limit value) is retrieved from the preset parameters. Subsequently, the program compares the content value with the range. If the content value is between the lower limit value and the upper limit value (including the boundary), the content deviation distance is determined to be 0. If the content value is less than the lower limit value, the content deviation distance is determined to be the absolute value of the difference between the lower limit value and the content value. If the content value is greater than the upper limit value, the content deviation distance is determined to be the absolute value of the difference between the content value and the upper limit value.
[0086] After obtaining the content deviation distance, the value only reflects the absolute size of the deviation, and the importance of the nutrient itself has not yet been embodied. In order to enable the evaluation to focus on the matching degree of the key nutrients, it is necessary to combine the deviation size with the importance degree. Therefore, it is necessary to multiply the impact weight and the content deviation distance to obtain the weighted deviation value. The calculation here is the core link of generating the weighted deviation value. It enlarges or reduces the relative importance of a nutrient by a simple multiplication operation, and its impact on the content deviation.
[0087] In a specific embodiment, continuing to take lysine as an example, the program will obtain the content deviation distance of lysine calculated in the previous step. At the same time, the impact weight of lysine is obtained from the preset parameter table. Then, the two values are directly multiplied, and the operation result is the weighted deviation value of lysine. For example, if the content deviation distance of lysine is 5 units, and the impact weight of lysine is 0.15, then the weighted deviation value of lysine is 5 multiplied by 0.15, which is equal to 0.75. This calculation process will be performed once for each selected key nutrient.
[0088] The technical effect of this step is that it successfully integrates the information of deviation degree and importance degree in two different dimensions into a single index. It makes a small deviation of a high-impact-weight nutrient produce a more significant weighted deviation value than a large deviation of a low-impact-weight nutrient, so that the subsequent comprehensive evaluation is more sensitive to changes in key nutrients, significantly improving the accuracy of the evaluation method.
[0089] In one possible implementation, step 401 specifically includes:
[0090] Step 501: When the content value is between the lower limit value and the upper limit value, the content deviation distance is zero.
[0091] Step 502: When the content value is less than the lower limit value, the content deviation distance is the absolute value of the difference between the content value and the lower limit value.
[0092] Step 503: When the content value is greater than the upper limit value, the content deviation distance is the absolute value of the difference between the content value and the upper limit value.
[0093] Specifically, the implementation of the method is a conditional judgment process. For a given nutrient, its content value is first obtained, as well as the fluctuation range of the content value which is bounded by the lower limit value and the upper limit value. The first step of the process is to determine whether the content value is within the ideal interval. Specifically, when the content value is between the lower limit value and the upper limit value, the content deviation distance is zero. This includes the case where the content value is exactly equal to the lower limit value or the upper limit value. For example, if the fluctuation range of a nutrient is [10, 20], and the content value of the nutrient in the dairy product to be evaluated is 15, since 15 is between 10 and 20, the content deviation distance is determined to be 0. This rule embodies the recognition of the content value falling within the natural fluctuation range of the reference dairy product, regardless of any deviation.
[0094] If the first condition is not met, the content value falls outside the ideal interval, and it is necessary to determine the deviation direction and calculate the deviation size. The process enters the second judgment branch. Specifically, when the content value is less than the lower limit value, the content deviation distance is the absolute value of the difference between the content value and the lower limit value. Continuing to use the above example, if the measured content value is 8, since 8 is less than the lower limit value 10, the content deviation distance is calculated as |8-10|, and the result is 2. This calculation accurately quantifies the degree of content deficiency.
[0095] If the first two conditions are not met, there is only one last possibility. The process will execute the third calculation rule. Specifically, when the content value is greater than the upper limit value, the content deviation distance is the absolute value of the difference between the content value and the upper limit value. Again using the example, if the measured content value is 25, since 25 is greater than the upper limit value 20, the content deviation distance is calculated as |25-20|, and the result is 5. This calculation also accurately quantifies the degree of content over-standard. Using the absolute value ensures that the content deviation distance is always a non-negative number, only representing the magnitude of the deviation.
[0096] In one possible implementation, step 103 specifically comprises:
[0097] Step 601: Perform an aggregation operation on all weighted deviation values to obtain a comprehensive deviation value.
[0098] Step 602: Map the comprehensive deviation value to a final similarity score by a pre-set scoring function, wherein the final similarity score decreases as the comprehensive deviation value increases.
[0099] Specifically, the first step of the embodiment is to aggregate all the weighted deviation values to obtain a comprehensive deviation value. The comprehensive deviation value is a single summary value, which aims to comprehensively reflect the overall deviation degree of the to-be-evaluated dairy product in all the nutritional substances under study. The aggregation operation is a calculation operation of combining multiple values into one value. In a specific embodiment, the aggregation operation can be a summation operation. The calculation process collects each weighted deviation value calculated for all key nutritional substances (such as all amino acids and high-abundance proteins) in the previous step, and then adds all the values together. The sum obtained is the comprehensive deviation value of the to-be-evaluated dairy product.
[0100] After obtaining the comprehensive deviation value, the value itself is not intuitive, and the larger the value represents the worse the similarity, which does not conform to the conventional scoring logic. In order to obtain a final result that is easy to understand and compare, a step of conversion is needed. Therefore, the next step of the embodiment is to map the comprehensive deviation value to the final similarity score by a preset scoring function. The key here is that the final similarity score decreases as the comprehensive deviation value increases. The preset scoring function is a mathematical conversion formula, which functions to convert the non-intuitive comprehensive deviation value into a standardized and easy-to-interpret score. In a preferred embodiment, the preset scoring function can be a linear conversion function, for example: final similarity score = 100 - K x comprehensive deviation value. Wherein, K is a preset scaling coefficient, used to adjust the sensitivity of the score to ensure that the final score falls within a suitable interval (such as 0-100 points). For example, if the calculated comprehensive deviation value is 2.5 and the set scaling coefficient K is 20, then the final similarity score is 100 - (20 x 2.5) = 100 - 50 = 50 points.
[0101] By sequentially performing the two steps, a set of complex, multi-dimensional weighted deviation values can be first aggregated into a comprehensive deviation value that can represent the overall deviation level, and then converted into a standardized final similarity score through an inverse mapping function. The higher the final similarity score, the closer the product is to the nutritional profile of the reference dairy product, thereby providing a clear, reliable and highly sensitive final judgment basis for product development and quality control.
[0102] Reference Figure 7 , Figure 7 is a structural schematic diagram of a dairy product similarity quantitative evaluation system provided by the present application, the system comprising:
[0103] The acquisition module is configured to acquire the content value of each nutritional substance in the to-be-evaluated dairy product, the influence weight for representing the importance degree of each nutritional substance, and the content fluctuation range of each nutritional substance in the reference dairy product.
[0104] The first processing module is configured to obtain a weighted deviation value for any nutrient according to the content value, the content fluctuation range, and the influence weight; and the weighted deviation value is used to represent a content deviation degree of the to-be-evaluated dairy product and the reference dairy product on the nutrient.
[0105] The second processing module is configured to determine a final similarity score of the to-be-evaluated dairy product and the reference dairy product according to all the weighted deviation values.
[0106] In a possible implementation, the obtaining module is further configured to:
[0107] obtain a reference nutritional data sample of a reference dairy product sample and a to-be-compared nutritional data sample of a to-be-compared dairy product sample; the reference nutritional data sample and the to-be-compared nutritional data sample each include a content value of at least one nutrient;
[0108] train a classification model for distinguishing sample sources according to the reference nutritional data sample and the to-be-compared nutritional data sample;
[0109] extract a feature importance score of each nutrient from the classification model;
[0110] for any nutrient, the feature importance score is taken as the influence weight.
[0111] In a possible implementation, the obtaining module is further configured to:
[0112] for any nutrient, a cumulative density distribution curve is constructed based on a content value distribution of the nutrient in the reference nutritional data sample;
[0113] a slope of each data point on the cumulative density distribution curve is determined;
[0114] a content value corresponding to a first data point is taken as a lower limit value of the content fluctuation range, and a content value corresponding to a second data point is taken as an upper limit value of the content fluctuation range;
[0115] wherein the first data point is a data point starting from a low end of the cumulative density distribution curve, and the slope of the data point is first greater than a first preset threshold; and the second data point is a data point after the first data point, and the slope of the data point is first less than a second preset threshold.
[0116] In a possible implementation, the first processing module is further configured to:
[0117] obtain a content deviation distance according to a size comparison relationship between the content value and the content fluctuation range;
[0118] multiply the influence weight and the content deviation distance to obtain the weighted deviation value.
[0119] In a possible implementation, the first processing module is further configured to:
[0120] When the content value is between the lower limit value and the upper limit value, the content deviation distance is zero;
[0121] When the content value is less than the lower limit value, the content deviation distance is the absolute value of the difference between the content value and the lower limit value;
[0122] When the content value is greater than the upper limit value, the content deviation distance is the absolute value of the difference between the content value and the upper limit value.
[0123] In a possible implementation, the second processing module is further configured to:
[0124] perform an aggregation operation on all the weighted deviation values to obtain a comprehensive deviation value;
[0125] map the comprehensive deviation value to a final similarity score by using a preset scoring function, wherein the final similarity score decreases as the comprehensive deviation value increases.
[0126] It should be noted that the dairy product similarity quantification evaluation system provided in the present application can execute the dairy product similarity quantification evaluation method of any of the above embodiments when actually running, and thus the present embodiment will not be described in detail.
[0127] Figure 8 is a structural schematic diagram of an electronic device provided in the present application, as shown in Figure 8 The electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 complete mutual communication through the communications bus 840. The processor 810 can invoke a logical instruction in the memory 830 to execute a dairy product similarity quantification evaluation method, which includes: obtaining a content value of each nutrient substance in a to-be-evaluated dairy product, an influence weight used to represent an importance degree of each nutrient substance, and a content fluctuation range of each nutrient substance in a reference dairy product; for any nutrient substance, obtaining a weighted deviation value according to the content value, the content fluctuation range, and the influence weight; the weighted deviation value is used to represent a content deviation degree of the to-be-evaluated dairy product and the reference dairy product on the nutrient substance; and determining a final similarity score of the to-be-evaluated dairy product and the reference dairy product according to all the weighted deviation values.
[0128] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. According to such an understanding, the technical solutions of the present application or parts of the present application that essentially contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the method of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0129] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, when the program instructions are executed by a computer, the computer can execute the dairy product similarity quantification evaluation method provided by the above-mentioned embodiments, and the method comprises the following steps: obtaining the content value of each nutrient in the dairy product to be evaluated, the influence weight for representing the importance degree of each nutrient, and the content fluctuation range of each nutrient in the reference dairy product; for any nutrient, according to the content value, the content fluctuation range and the influence weight, a weighted deviation value is obtained; the weighted deviation value is used to represent the content deviation degree of the dairy product to be evaluated and the reference dairy product on the nutrient; according to all the weighted deviation values, the final similarity score of the dairy product to be evaluated and the reference dairy product is determined.
[0130] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the dairy product similarity quantification evaluation method provided by the above-mentioned embodiments, and the method comprises the following steps: obtaining the content value of each nutrient in the dairy product to be evaluated, the influence weight for representing the importance degree of each nutrient, and the content fluctuation range of each nutrient in the reference dairy product; for any nutrient, according to the content value, the content fluctuation range and the influence weight, a weighted deviation value is obtained; the weighted deviation value is used to represent the content deviation degree of the dairy product to be evaluated and the reference dairy product on the nutrient; according to all the weighted deviation values, the final similarity score of the dairy product to be evaluated and the reference dairy product is determined.
[0131] The system embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0132] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. According to such understanding, the above technical solutions can be embodied in the form of software product, and the computer software product can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the method of each embodiment or some parts of the embodiment.
[0133] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for quantitatively assessing the similarity of dairy products, characterized in that, include: Obtain the content value of each nutrient in the dairy product to be evaluated, the influence weight used to characterize the importance of each nutrient, and the fluctuation range of the content of each nutrient in the dairy product as a reference. For any of the nutrients, a weighted deviation value is obtained based on the content value, the content fluctuation range, and the influence weight; the weighted deviation value is used to characterize the degree of deviation between the content of the dairy product to be evaluated and the reference dairy product in that nutrient. The step of obtaining the weighted deviation value based on the content value, the content fluctuation range, and the influence weight includes: The content deviation distance is obtained by comparing the content value with the content fluctuation range; the content fluctuation range includes a lower limit and an upper limit; obtaining the content deviation distance by comparing the content value with the content fluctuation range includes: when the content value is between the lower limit and the upper limit, the content deviation distance is zero; when the content value is less than the lower limit, the content deviation distance is the absolute value of the difference between the content value and the lower limit; when the content value is greater than the upper limit, the content deviation distance is the absolute value of the difference between the content value and the upper limit. The weighted deviation value is obtained by multiplying the influence weight by the content deviation distance; Based on all the weighted deviation values, the final similarity score between the dairy product to be evaluated and the reference dairy product is determined.
2. The method for quantitatively evaluating the similarity of dairy products according to claim 1, characterized in that, Also includes: Obtain the reference nutrient data sample of the reference dairy product sample and the comparison nutrient data sample of the dairy product sample to be compared; Both the reference nutrient data sample and the nutrient data sample to be compared include the content value of at least one nutrient. Based on the reference nutrient data sample and the nutrient data sample to be compared, train a classification model to distinguish the source of the samples; Extract the feature importance score for each nutrient from the classification model; For any of the nutrients, the feature importance score is used as the influence weight.
3. The method for quantitatively evaluating the similarity of dairy products according to claim 2, characterized in that, Also includes: For any of the nutrients, a cumulative density distribution curve is constructed based on the content distribution of the nutrients in the reference nutrient data sample; Determine the slope of each data point on the cumulative density distribution curve; The content value corresponding to the first data point is taken as the lower limit of the content fluctuation range, and the content value corresponding to the second data point is taken as the upper limit of the content fluctuation range. Wherein, the first data point is the data point where the slope of the cumulative density distribution curve first exceeds a first preset threshold, starting from the low end of the curve; the second data point is the data point where the slope first falls below a second preset threshold, following the first data point.
4. The method for quantitatively evaluating the similarity of dairy products according to claim 1, characterized in that, The step of determining the final similarity score between the dairy product to be evaluated and the reference dairy product based on all the weighted deviation values includes: Aggregate all the weighted deviation values to obtain the comprehensive deviation value; The overall deviation value is mapped to the final similarity score using a preset scoring function, wherein the final similarity score decreases as the overall deviation value increases.
5. A quantitative evaluation system for the similarity of dairy products, characterized in that, include: The acquisition module is used to acquire the content value of each nutrient in the dairy product to be evaluated, the influence weight used to characterize the importance of each nutrient, and the fluctuation range of the content of each nutrient in the dairy product. The first processing module is used to obtain a weighted deviation value for any of the nutrients based on the content value, the content fluctuation range, and the influence weight; the weighted deviation value is used to characterize the degree of deviation between the content of the dairy product to be evaluated and the reference dairy product in the nutrient. The step of obtaining a weighted deviation value based on the content value, the content fluctuation range, and the influence weight includes: obtaining a content deviation distance based on a comparison between the content value and the content fluctuation range; the content fluctuation range includes a lower limit and an upper limit; obtaining the content deviation distance based on a comparison between the content value and the content fluctuation range includes: when the content value is between the lower limit and the upper limit, the content deviation distance is zero; when the content value is less than the lower limit, the content deviation distance is the absolute value of the difference between the content value and the lower limit; when the content value is greater than the upper limit, the content deviation distance is the absolute value of the difference between the content value and the upper limit; multiplying the influence weight by the content deviation distance yields the weighted deviation value. The second processing module is used to determine the final similarity score between the dairy product to be evaluated and the reference dairy product based on all the weighted deviation values.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the dairy product similarity quantitative evaluation method as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the dairy product similarity quantitative evaluation method as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the dairy product similarity quantitative evaluation method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Similarity evaluation method of human milk replacement fat
CN108132338A
Comprehensive similarity measurement and evaluation method based on data driving
CN116662825A