CIELAB-based quantitative classification modeling method for fruit color of eggplant and its application
The quantitative classification modeling method for eggplant fruit color based on CIELAB solves the problems of strong subjectivity and weak distinguishing ability of single parameters in eggplant fruit color identification. It realizes the accurate transformation of fruit color from qualitative description to quantitative scoring, provides high-precision fruit color classification rules, and supports eggplant germplasm resource identification and genetic breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF VEGETABLES GUANGDONG PROV ACAD OF AGRI SCI
- Filing Date
- 2026-04-10
- Publication Date
- 2026-06-02
AI Technical Summary
Current technologies for identifying eggplant fruit color mainly rely on manual visual inspection, which suffers from high subjectivity, low efficiency, and weak ability to distinguish between single parameters, making it difficult to meet the needs of modern eggplant breeding for high-precision phenotypic data.
A CIELAB-based quantitative classification modeling method for eggplant fruit color was adopted. By collecting CIELAB color space parameters of the fruit peel, calculating color difference derived parameters, screening key parameters, constructing a weighted comprehensive scoring function, dynamically optimizing to determine the classification threshold, and establishing fruit color classification rules, objective, accurate, and high-throughput quantitative classification of fruit color was achieved.
This study achieved objective and quantitative identification of eggplant fruit color, eliminated subjective errors, significantly improved the accuracy of fruit color differentiation, constructed standardized classification rules, ensured the stability and universality of the model, and provided a reliable phenotypic analysis tool for eggplant germplasm resource identification and genetic breeding.
Smart Images

Figure CN122135358A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-throughput and precise identification of plant phenotypes, specifically to a quantitative classification modeling method and application of eggplant fruit color based on CIELAB. Background Technology
[0002] Eggplant (Solanum melongena L.) belongs to the genus Solanum in the family Solanaceae and plays an important role in the production of Solanaceae vegetables. Eggplant, with its fruit as the product, is one of the world's top ten healthiest vegetables recommended by the WHO. Eggplant skin colors (hereinafter collectively referred to as fruit color) are rich and diverse, including purplish-red, purplish-black, green, and white. Eggplant consumption exhibits obvious regional characteristics: the characteristic eggplant type in South China is the club-shaped purplish-red eggplant, whose commercial fruit color is purplish-red, significantly different from the purplish-black eggplant mainly cultivated in Central China, Southwest China, and Northeast China. Therefore, fruit color is an important breeding target in the eggplant breeding process. Accurate and objective phenotypic identification of purplish-red / purplish-black eggplant fruit color is a prerequisite for conducting analysis of fruit color genetic patterns, QTL mapping, and candidate gene mining.
[0003] The purple color of eggplant fruit is determined by the accumulation of anthocyanins in the peel and the retention of chlorophyll in the pulp beneath, forming a continuous phenotypic spectrum from purplish-red to purplish-black. Currently, eggplant fruit color identification mainly relies on manual visual observation. This method is highly subjective, inefficient, and easily affected by factors such as the observer's experience and lighting conditions, making it difficult to meet the demands of modern breeding for large-scale, high-precision phenotypic data. Although some studies have used colorimeters to measure single... , , Parameters characterize fruit color, but eggplant fruit color is controlled by multiple genes and exhibits continuous variation. A single parameter cannot fully cover the characteristics of fruit color such as hue, saturation, and brightness, making it difficult to effectively distinguish the subtle differences between purplish-red and purplish-black fruit colors, and also failing to achieve stable classification results.
[0004] Currently, there is no dedicated multi-parameter comprehensive scoring and dynamic threshold classification system for eggplant's purple-red / purple-black fruit color, which restricts the in-depth development of eggplant fruit color genetic research and the improvement of molecular breeding efficiency. Therefore, developing an objective, accurate, and reproducible quantitative classification method for eggplant purple-red / purple-black fruit color has become an urgent technical problem to be solved in eggplant fruit color genetic research and quality breeding. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, one of the objectives of this invention is to provide a quantitative classification modeling method for eggplant fruit color based on CIELAB, which can eliminate the subjectivity and variability of manual visual inspection, overcome the weakness of single parameter discrimination, and achieve objective, accurate, and high-throughput quantitative classification of eggplant purple-red and purple-black fruit colors.
[0006] To achieve one of the objectives of this invention, the following solution is adopted: The CIELAB-based quantitative classification modeling method for eggplant fruit color includes the following steps: Step S1: Collect CIELAB color space parameters of the pericarp of the separated eggplant population. , , The values were determined, and a baseline dataset of known phenotypic traits, including purple-red and purple-black fruits, was created through manual visual inspection. Step S2, based on the benchmark dataset , , The color difference derived parameters are calculated, and key parameters that show significant differences in color between purplish-red and purplish-black fruits are selected from the derived parameters. Step S3: Assign weights based on the statistical discrimination ability of the key parameters in the benchmark dataset, and construct a weighted comprehensive scoring function; Step S4: Based on the comprehensive score distribution of the benchmark dataset calculated by the weighted comprehensive scoring function, a classification threshold is determined by dynamically optimizing the threshold coefficient, and a classification system based on the comprehensive score CS is established. i = 0 is the threshold for fruit color classification rules; Step S5: After confirming the accuracy of the fruit color classification rule through internal back-judgment of the training set and external verification of the test set, apply the fruit color classification rule to the fruit color classification prediction of unknown samples.
[0007] Furthermore, the key parameters include the hue angle h° and the red-to-blue ratio. Ratio, relative saturation and brightness The color characteristics of the fruit peel are comprehensively characterized from four dimensions: hue, red-blue ratio, relative saturation, and brightness.
[0008] Furthermore, the calculation formula for the weighted comprehensive scoring function is as follows: Where: CS i P represents the overall score of the i-th individual plant; ij B is the actual measured value of the j-th parameter of the i-th individual plant; j The midpoint reference value for the j-th parameter; SD j W represents the standard deviation of the j-th parameter in the known sample. j Let be the weight of the j-th parameter.
[0009] Furthermore, the method for determining the classification threshold includes: Calculate the average score M of the known purple-red group. R The average score of the purple-black group is M. Band the standard deviation (SD) of all known sample scores total Set the purple-red threshold T R = M R -k ×SD total Purple-black threshold T B = M B +k ×SD total Where k is the threshold coefficient, which must satisfy 0 ≤ k < (M) R -M B ) / (2×SD total ); Iterate through the k values and select the T value corresponding to the k value that maximizes the accuracy of the known samples and minimizes the median. R and T B As a classification threshold.
[0010] Furthermore, the evaluation metric for internal back-judgment in the training set is classification accuracy; the evaluation metrics for external validation in the test set include accuracy, precision, recall, and F1 score.
[0011] Furthermore, the calculation formula for the color difference derived parameter includes: chromaticity Hue angle and according to , Adjust the sign to 0°~360°; relative saturation Red to blue ratio .
[0012] Furthermore, the weight allocation method is as follows: using the ratio of the inter-group mean difference to the standard deviation of each parameter as the initial allocation basis, the total weight is set to 35 and then allocated proportionally, resulting in a final weight of 8.0 for the hue angle h° and a red-blue ratio. The final weight of the ratio is 9.0, relative saturation. The final weight is 9.5, brightness The final weight is 8.5.
[0013] Furthermore, the fruit color classification rule is as follows: if the comprehensive score CS i If the score is > 0, it is judged as a purple-red fruit; if the overall score CS i If the score is less than 0, it is judged as a purple-black fruit; if the overall score is CS i = 0, then it is classified as intermediate type.
[0014] Furthermore, the step size for traversing the k value is 0.1, and the step size is encrypted to 0.01 in the region sensitive to accuracy changes.
[0015] The second objective of this invention is to provide the application of the above-mentioned CIELAB-based quantitative classification and modeling method for eggplant fruit color in eggplant germplasm resource identification, genetic breeding, or fruit color-related gene mining.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention eliminates subjective errors and achieves objective quantitative identification. This invention collects CIELAB color space parameters of eggplant peel. , , The value is obtained by replacing manual visual judgment with instrument-quantified measurement data, and a known phenotypic benchmark dataset is established on this basis, which ensures the objectivity and repeatability of phenotypic information from the source of data.
[0017] 2. This invention achieves multi-parameter synergistic characterization, significantly improving the accuracy of fruit color differentiation. Based on the original CIELAB parameters, this invention calculates color difference-derived parameters and selects key parameters that show significant differences between purplish-red and purplish-black fruit colors. It comprehensively characterizes the peel color features from multiple dimensions, overcoming the limitation of single parameters in distinguishing subtle differences and significantly improving the ability to differentiate between purplish-red and purplish-black fruit colors.
[0018] 3. This invention achieves objective quantification and multi-parameter fusion through weighted comprehensive scoring. Based on the objective statistical distinguishing ability of each key parameter in known samples, this invention assigns weights to construct a weighted comprehensive scoring function, fusing multi-dimensional color difference parameters into a one-dimensional comprehensive score, thus realizing a precise transformation of eggplant fruit color from qualitative description to quantitative scoring.
[0019] 4. This invention establishes standardized classification rules through dynamic threshold optimization. Based on the comprehensive score distribution, this invention determines the classification boundary through dynamic optimization of the threshold coefficient k, and finally establishes a classification rule based on the comprehensive score CS. i The standardized fruit color classification rule with a threshold of 0 eliminates the subjective arbitrariness of manually defining boundaries, ensuring the uniformity and stability of classification standards.
[0020] 5. This invention employs dual internal and external validation to ensure the model's generalization ability and practical value. The accuracy of the model is verified through internal back-judgment on the training set, and its generalization ability is evaluated through external validation on an independent test set. Once accuracy is confirmed, it is applied to predict unknown samples. This validation system ensures the stability and universality of this classification rule across eggplant resources with different genetic backgrounds, providing a reliable and standardized high-throughput phenotypic analysis tool for eggplant germplasm resource identification, genetic breeding, and fruit color-related gene mining. Attached Figure Description
[0021] Figure 1 This is a flowchart of the quantitative classification and modeling method for eggplant fruit color based on CIELAB in an embodiment of the present invention; Figure 2This is a schematic diagram of the phenotypic characteristics of the male parent, female parent, and F1 generation fruit in an embodiment of the present invention; from left to right, they are the male parent "05-4", the female parent "05-1", and F1. The male parent "05-4" is a white-fleshed eggplant with purple-red fruit, and the female parent "05-1" is a green-fleshed eggplant with purple-black fruit. The F1 generation obtained by hybridization of the parents is a green-fleshed eggplant with purple-black fruit. The F2 generation segregating population is obtained by strict self-pollination of individual F1 plants. Figure 3 This is a comparison chart of the normalized model weights and objective discrimination contributions of four key parameters in an embodiment of the present invention; it shows the hue angle h° and the red-blue ratio. Ratio, relative saturation Brightness The comparison between the normalized model weights of these four key parameters and the objective distinguishing contribution (F-value); Figure 4 This is a diagram illustrating the impact of the threshold coefficient k on classification performance in an embodiment of the present invention; the left y-axis represents the model classification accuracy (%), and the right y-axis represents the proportion of intermediate cases (%). As the value of k is continuously adjusted, the model classification accuracy and the proportion of intermediate cases also show certain changing patterns; Figure 5 This is a schematic diagram of the classification results of the comprehensive scoring fruit color classification model applied to 89 known phenotypic individual plants in the F2 population in an embodiment of the present invention. It shows the comprehensive score distribution and discrimination results of known purple-red and purple-black samples, which intuitively reflects the classification accuracy and stability of the model. Figure 6 This is a schematic diagram of the phenotypic images of 89 known phenotypic individual plants from the F2 population of the training set in this embodiment of the invention; Figure 7 This is a schematic diagram of phenotypic images of 76 independent eggplant varieties used for external validation of the test set in an embodiment of the present invention; Figure 8 This is a schematic diagram of the frequency distribution of 232 F2 single plants in different comprehensive scoring intervals in an embodiment of the present invention; among them, the scoring intervals of purple-red fruit and purple-black fruit are basically separated, showing a clear bimodal distribution, which is consistent with the quality trait characteristics controlled by the major gene, and can provide a reliable phenotypic basis for subsequent genetic analysis. Detailed Implementation
[0022] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.
[0023] This invention provides a CIELAB-based quantitative classification and modeling method for eggplant fruit color, which overcomes the shortcomings of existing eggplant fruit color identification methods, such as strong subjectivity, low efficiency, and weak single-parameter discrimination ability. It realizes the accurate transformation of eggplant fruit color from qualitative visual description to quantitative scoring, improves the objectivity, accuracy, and reproducibility of fruit color phenotypic identification, and provides a reliable and standardized phenotypic analysis means for eggplant germplasm resource identification, genetic breeding, and fruit color-related gene mining.
[0024] like Figure 1 As shown in the figure, the eggplant fruit color quantitative classification modeling method based on CIELAB according to an embodiment of the present invention includes the following steps: Step S1: Collect CIELAB color space parameters of the pericarp of the separated eggplant population. The values were determined, and a baseline dataset containing known phenotypic traits of purple-red and purple-black fruits was created through manual visual inspection.
[0025] In this embodiment, step S1 is the material preparation and phenotypic data collection step, the specific contents of which are as follows: Using the F2 segregating population of eggplant as material, fruit peel color difference data of each individual plant were collected during the commercial fruiting period; multiple measurement points were evenly selected at the equatorial region of the fruit using a handheld spectrophotometer to determine the CIELAB color space parameters. , , The average value was taken as the peel color difference data of the single plant. At the same time, several experienced observers independently judged the peel color of each single plant under standard lighting conditions, and divided the fruit color into purple-red type, purple-black type and difficult to judge type. Among them, the single plants that were clearly classified by all observers were established as the "known phenotype" benchmark dataset, and the remaining single plants were marked as "unknown phenotype" and left for model prediction.
[0026] Step S2, based on the benchmark dataset The color difference derived parameters are calculated, and key parameters that show significant differences in color between purplish-red and purplish-black fruits are selected from the derived parameters.
[0027] In this embodiment, step S2 is the calculation of color difference derived parameters and the screening of key parameters, the details of which are as follows: Based on the original Values for calculating hue angle h° and red-blue ratio Ratio, relative saturation Derived parameters; based on known phenotypic samples, the differences between the two groups of each parameter were tested for significance, and the parameters that were highly significant (p<0.001) between purple-red and purple-black fruit colors and were representative were selected as model input variables to determine the key parameters that can comprehensively characterize the fruit peel color features.
[0028] Step S3: Assign weights based on the statistical discrimination ability of the key parameters in the benchmark dataset, and construct a weighted comprehensive scoring function.
[0029] In this embodiment, step S3 is the parameter weight allocation and comprehensive score calculation step, the specific content of which is as follows: Based on all known phenotypic samples, the ratio of "between-group mean difference / pooled standard deviation" for each parameter is calculated as a quantitative indicator of discriminative ability, and initial weights are assigned accordingly. The total weight is set to a value that is easy to calculate (here, 35), and the final weight W for each parameter is obtained by proportionally allocating the weights. j The midpoint of the means of the purple-red and purple-black sample groups is set as the classification baseline value B for each parameter. j And calculate the standard deviation (SD) of each parameter in the known sample. j Used for standardization; the comprehensive score CS for each individual plant i i Calculated using the following formula: In the formula: P ij B is the actual measured value of the j-th parameter of the i-th individual plant; j The midpoint reference value for the j-th parameter; SD j W represents the standard deviation of the j-th parameter in the known sample. j The weight of the j-th parameter is determined by its "|difference value| / SD". j "Ratio allocation; the model design makes purplish-red fruit color tend to get higher positive scores, and purplish-black fruit color tend to get lower negative scores."
[0030] Step S4: Based on the comprehensive score distribution of the benchmark dataset calculated by the weighted comprehensive scoring function, a classification threshold is determined by dynamically optimizing the threshold coefficient, and a classification system based on the comprehensive score CS is established. i = 0 is the fruit color classification rule with a threshold of 0.
[0031] In this embodiment, step S4 is the classification threshold optimization and determination step, the specific content of which is as follows: Based on the comprehensive score distribution of known phenotypic samples, the classification boundary is determined by dynamic optimization using a threshold coefficient k; the average score M of the known purple-red group is calculated. R The average score of the purple and black group is M. B and the standard deviation (SD) of all known sample scores. total Set the purple-red threshold T R = M R -k ×SD total Purple-black threshold T B = M B +k ×SD total Where k is the threshold coefficient, which must satisfy 0 ≤ k < (MR -M B ) / (2×SD total Within the effective range, iterate through k values with different step lengths, calculate the corresponding classification threshold for each candidate k value, and statistically analyze the classification accuracy and intermediate proportion of known samples. Select the k value that maximizes accuracy and minimizes intermediate values as the optimal threshold coefficient to determine the final classification rule.
[0032] Step S5: After confirming the accuracy of the fruit color classification rule through internal back-judgment of the training set and external verification of the test set, apply the fruit color classification rule to the fruit color classification prediction of unknown samples.
[0033] In this embodiment, step S5 is the model verification and application step, the specific content of which is as follows: The optimized model was applied to known phenotypic samples within the training set for back-judgment testing to evaluate classification accuracy. It was then further validated externally using a completely independent test set of varietal resource datasets. The model's generalization ability and universality were comprehensively evaluated using metrics such as accuracy, precision, recall, and F1 score. The validated model was then applied to predict fruit color classification in the entire F2 population (including known and unknown samples), counting the number of purple-red and purple-black plants. The classification effect was verified using a comprehensive score distribution map, and genetic ratio analysis was performed.
[0034] Through the above steps, a classification model for eggplant purple-red / purple-black fruit color based on quantitative measurement of color difference parameters is constructed, realizing objective, efficient and accurate identification of eggplant fruit color.
[0035] The CIELAB-based quantitative classification and modeling method for eggplant fruit color in this invention uses the F2 population of eggplant fruit color separation as material and collects CIELAB color parameters of the fruit peel. A benchmark dataset was constructed by combining manual visual phenotypic identification; and excellent phase angles (h°) and red-blue ratios were selected by calculating multiple color difference-derived indices. Ratio, relative saturation and brightness As key parameters for fruit color classification, each parameter is assigned a corresponding weight based on its ability to distinguish between purplish-red and purplish-black fruit colors, thus constructing a weighted comprehensive scoring function. When the optimal classification threshold coefficient k=0.932, a comprehensive scoring CS function is established. i= 0 is used as the threshold for the purple-red / purple-black dividing line in a quantitative classification model for eggplant fruit color. Testing showed that this model achieved a 97.75% accuracy rate in fruit color classification for the training set and a 93.4% accuracy rate for eggplant varietal resources in the test set, demonstrating good stability and universality. Applying this model to fruit color classification and genetic segregation analysis in the F2 population enables rapid, accurate, and objective quantitative identification of eggplant fruit color. This invention eliminates the subjectivity and variability of traditional manual visual fruit color identification, constructing a standardized, high-throughput quantitative classification system for eggplant fruit color. It provides a reliable phenotypic detection method for eggplant germplasm resource identification, genetic analysis, and the localization and cloning of fruit color-related genes.
[0036] The CIELAB-based quantitative classification and modeling method for eggplant fruit color in this invention has the following advantages: 1. This invention achieves objective quantification and eliminates subjective errors. Based on instrument measurement data using the CIELAB color space, this invention replaces manual visual judgment and combines it with a multi-parameter comprehensive scoring model to achieve a precise transformation of eggplant fruit color from qualitative description to quantitative scoring. The identification results are objective, repeatable, and unaffected by factors such as observer experience and lighting conditions.
[0037] 2. This invention achieves multi-parameter synergy, significantly improving fruit color resolution. This invention filters hue angle h° and red-to-blue ratio... Ratio, relative saturation and brightness Four core parameters comprehensively cover key factors such as hue, saturation, and brightness of fruit color. Compared with a single parameter, multi-parameter collaborative classification significantly improves the accuracy of distinguishing between purplish-red and purplish-black fruit colors.
[0038] 3. This invention has strong generalization and a wide range of applications. External validation has confirmed that the model still has high accuracy in eggplant germplasm resources with heterogeneous genetic backgrounds, and it can be applied to the classification of purple-red / purple-black fruit color in different eggplant varieties, with a wide range of application scenarios.
[0039] 4. This invention supports genetic research and improves breeding efficiency. This invention provides reliable phenotypic data for eggplant fruit color genetic breeding and candidate gene mining, accelerates the process of eggplant fruit color genetic analysis, and provides core technical support for eggplant molecular breeding.
[0040] Experimental Example 1: Construction of an eggplant fruit color classification model based on the F2 segregating population.
[0041] 1. Plant materials and phenotypic observation: Using a purified inbred line "05-1" (green flesh) of purple-black eggplant, bred through multiple generations of strict bagging and self-pollination, as the female parent and a purified inbred line "05-4" (white flesh) of purple-red fruit as the male parent, F1 (green flesh) was obtained through hybridization of the parents and single-fruit seed collection. Strict self-pollination of individual F1 plants yielded the F2 segregating population (E2730). Figure 2 As shown. All E2730 population materials (232 plants in total) were raised in July and transplanted in August of the same year, with uniform routine field management. During the commercial fruiting period, phenotypic data were collected from representative fruits of each individual plant, using a strategy combining manual visual observation and instrumental quantitative measurement.
[0042] (1) Manual visual observation and establishment of benchmark dataset: To obtain reliable phenotypic classification data, three experienced observers independently visually assessed the fruit peel color of each individual plant under natural lighting conditions. Based on the hue and gloss of the fruit surface, the fruit color was divided into three categories: purplish-red (bright red with a strong gloss), purplish-black (deep purplish-black with a weaker gloss), and indeterminate (color between the two or atypical phenotype). Ultimately, 89 plants were unanimously classified (purplish-red or purplish-black) by all observers and established as the "known phenotype" baseline dataset for subsequent classification model construction and internal validation. The remaining 143 plants, due to disagreements among observers or atypical phenotypes, were marked as "unknown phenotypes" and left for model prediction.
[0043] (2) Quantitative measurement of color difference parameters by instrument: Three measurement points were evenly selected at the equatorial region of each fruit using a handheld spectrophotometer (CM-700d). CIELAB color space parameters were measured, and the average value was taken as the peel color difference data for that single fruit. The value represents brightness (ranging from 0 to 100, where 0 is pure black and 100 is pure white). The value represents the red-green axis (positive value indicates redness, negative value indicates greenness). The value represents the yellow-blue axis (positive value indicates yellowish, negative value indicates bluish).
[0044] 2. Calculation of color difference derived parameters: To comprehensively quantify the color characteristics of eggplant peel, the original values measured in the CIELAB color space were used ( The following series of commonly used derived parameters were calculated: (1) Chroma ): Characterizes the purity and vividness of color. (2) Hue angle (h°): Characterizes the hue attribute of a color (such as red, yellow, blue, etc.). The calculation result needs to be based on... The symbol is adjusted to the 0°~360° quadrant range.
[0045] (3) Relative saturation ): Characterizes the vividness of a color at a given brightness, reflecting the proportion of color components relative to brightness.
[0046] (4) Red-Blue Ratio: Used to reflect the relative proportions of red-green and yellow-blue components.
[0047] Because of the sample All values are negative (leaning towards blue). This ratio can intuitively reflect the relative strength of the red and blue components, hence it is simply called the "red-blue ratio" and is used to quantify the hue difference between purplish-red and purplish-black samples.
[0048] 3. Color difference parameter selection and model input variable determination: Based on the phenotypic characteristics of eggplant fruit color and related research, representative parameters reflecting the peel color in terms of hue tendency, red-blue component ratio, color vividness, and brightness were initially selected as candidate parameters. Using 89 known phenotypic samples (24 purple-red plants and 65 purple-black plants) as a basis, a significance test (independent samples t-test) was conducted on the candidate parameters to assess the differences between the two groups. The results showed that the hue angle h°, red-blue ratio, and other parameters were significantly different. Ratio, relative saturation and brightness The differences between purplish-red and purplish-black fruit colors were extremely significant (p<0.001), indicating that these four parameters have good ability to distinguish categories across different dimensions. Therefore, these four parameters were selected as input variables for the subsequent classification model, and the fruit peel color characteristics were comprehensively characterized from four dimensions: hue, red-blue ratio, relative saturation, and brightness.
[0049] 4. Parameter weight allocation and discrimination ability analysis: The objective discriminative power of each parameter is the fundamental basis for weight allocation. Based on all 89 known phenotypic samples (24 purple-red plants and 65 purple-black plants), this embodiment calculated two types of statistical indicators in parallel to comprehensively evaluate the classification potential of each parameter. First, the "difference / standard deviation" ratio (i.e., the ratio of the difference between group means to the pooled standard deviation) of each parameter was calculated. This indicator directly reflects the initial discriminative power of the parameter after standardization. Second, a one-way ANOVA was performed on each parameter to calculate its univariate F-value (F = MS).between / MS within The F-values of each parameter are normalized to a percentage based on the maximum value, and this is called the "F-value contribution (%)".
[0050] The model weights are determined directly based on the aforementioned quantitative evaluation results. Using the "difference / standard deviation" ratio as the initial allocation basis, the relative weight proportions of each parameter are calculated. To obtain better numerical discrimination, the total weight is set to 35, and the weights are allocated according to this ratio to obtain the final weight (W) for each parameter. j To ensure the scoring system is symmetrical and unbiased, the midpoint of the means of the purple-red and purple-black sample groups is set as the classification baseline value (B) for each parameter. j Simultaneously, the standard deviation (SD) of each parameter in the known sample is calculated. j ), used for subsequent standardization to eliminate the influence of dimensions.
[0051] The discriminative power indices and weightings for each parameter are shown in Table 1. From the perspective of the "difference / standard deviation" ratio, the relative saturation... The red-blue ratio has the strongest distinguishing ability, with a ratio of 1.788, followed by the red-blue ratio. Ratio (1.682), Brightness (1.597) and hue angle h° (1.500). This order indicates that parameters related to color saturation and red-blue ratio are most sensitive to distinguishing between magenta / magenta fruit colors, while lightness and hue angle are relatively weaker. The model weights were determined based on the "difference / standard deviation" ratio. To ensure that the subsequent comprehensive score falls within an easily observable range (e.g., -70 to 70), the total weight was set to 35, and the absolute weights (W) of each parameter were finally determined. j The classification baseline value B for each parameter. j (i.e., the midpoint between the means of the purple-red and purple-black groups), the difference between the means of the groups, and the standard deviation SD. j and the final weight W j The specific values are shown in Table 1. This weight allocation reflects the differences in the objective distinguishing ability of each parameter, and ensures the practicality of the model through rounding and adjustment of the total weight.
[0052] Table 1. Discrimination ability indicators and weight allocation of key color difference parameters To further examine the rationality of the weight allocation from a statistical perspective, a one-way ANOVA was performed on each parameter, and its F-value was calculated as a measure of objective discrimination ability. The parameter with the strongest discrimination ability was used as the basis for the analysis. The F-value is used as a benchmark (normalized to 100%). The objective F-value contribution of each parameter is compared with the normalized model weights, such as... Figure 3 As shown. The results show that, It holds an absolutely central position in both objective discrimination ability and model weight (both are 100%). This dual verification result is consistent with its key role in the physiological mechanism by which anthocyanin accumulation directly affects the color saturation of fruit. The objective discrimination ability of the ratio parameter (F-value contribution of 69.1%) is slightly lower than the weight assigned by the model (94.7% after normalization), indicating that the model strengthens the contribution of the red-blue ratio to classification to some extent. The model weights of the two parameters, h° and h° (89.5% and 84.2%, respectively), are significantly higher than their objective univariate discrimination power (F-value contributions of 54.8% and 45.2%, respectively). This result indicates that in the comprehensive scoring model, the weight allocation of each parameter is not entirely determined by its univariate discrimination power, but rather takes into account the complementarity between parameters. While h° is less effective at distinguishing between two types of fruit color on its own, it provides color information dimensions (brightness and hue angle) that differ from relative saturation and red-blue ratio. These dimensions play an important auxiliary role in multivariate comprehensive scoring, helping to improve the overall robustness and classification accuracy of the model.
[0053] The overall score CS for each individual plant i i Calculated using the following standardized formula: In the formula: P ij B is the actual measured value of the j-th parameter of the i-th individual plant; j The midpoint reference value for the j-th parameter; SD j W represents the standard deviation of the j-th parameter in the known sample. j The weight of the j-th parameter is determined by its "|difference value| / SD". j "Ratio allocation; the model design makes purplish-red fruit color tend to get higher positive scores, and purplish-black fruit color tend to get lower negative scores."
[0054] 5. Threshold coefficient k optimization and optimal threshold determination: After calculating the weighted comprehensive scores, a threshold coefficient k is introduced for dynamic optimization to scientifically delineate the classification boundary between purple-red and purple-black fruits. The purple-red threshold T is defined as follows: R and purple-black threshold T B as follows: T R = M R - k × SD total T B = M B + k × SD total Among them, M R and M B These are the known average composite scores for the purple-red group and the purple-black group, respectively; SD total is the standard deviation of the scores for all known samples; k is the threshold coefficient, used to control the strictness of classification. The larger the value of k, the better the T... R and T B The closer to the middle, the stricter the classification, and the fewer intermediate types; conversely, the further away from the middle, the more lenient the classification. This ensures a reasonable classification range (i.e., T). R > T B k must satisfy 0 ≤ k < .
[0055] Based on 89 known phenotypic samples (24 purple-red plants and 65 purple-black plants), k values were initially screened with a step size of 0.1 (0.5, 0.6, 0.7, 0.8, 0.9), and the step size was increased to 0.01 (0.93, 0.932) in areas sensitive to accuracy changes. For each candidate k value, the corresponding T was calculated. R and The known samples were divided into purple and red (CS) i ≥ T R ), Purple-black (CS) i ≤ T B ) and intermediate type (T) B < CS i < T R The data is categorized into three classes, and the classification accuracy (number of correctly classified samples / total number of known samples) and the proportion of intermediate cases are calculated. Taking into account both accuracy and the proportion of intermediate cases, the optimal threshold coefficient (k) is selected as the value that maximizes accuracy and minimizes the number of intermediate cases.
[0056] Given that the average score of the purple-red group is M R =16.858, average score M for the purple-black group B =-16.858, total standard deviation SD total =18.085, therefore the maximum allowable value of k is calculated to be k. max =0.932. Within the effective range (k < 0.932), the threshold coefficient k was optimized through iteration. The classification threshold, accuracy, and proportion of intermediate types under different k values are shown in Table 2. As the k value increases, the purple-red threshold (T... R ) and purple-black threshold (T B The classification accuracy gradually shifts towards the middle. Simultaneously, the classification accuracy gradually increases, while the proportion of intermediate cases gradually decreases. Figure 4 As shown. When k=0.932, the accuracy reaches its highest value of 97.75%, the number of intermediate individuals is 0, and at this time both the purple-red threshold and the purple-black threshold approach 0 (T). R ≈ 0.003, T BThe threshold value is approximately -0.003, with a width of only 0.006 in the middle interval. Considering both classification accuracy and the proportion of intermediate individuals, k=0.932 is selected as the optimal threshold coefficient. Under this threshold, the overall accuracy reaches a maximum of 97.75%, and there are no intermediate individuals, which can best reflect the actual situation of fruit color separation in the F2 population.
[0057] Table 2 Classification results for different threshold coefficients k K value Purple-red threshold (TR) Purple-black threshold (TB) Accuracy (%) intermediate quantity intermediate proportion (%) 0.50 7.131 -7.131 85.39 11 12.36 0.60 5.287 -5.287 87.64 9 10.11 0.70 4.199 -4.199 92.13 7 7.87 0.80 2.390 -2.390 93.26 6 6.74 0.90 0.582 -0.582 94.38 4 4.49 0.93 0.039 -0.039 96.63 1 1.12 0.932 0.003 -0.003 97.75 0 0 6. Fruit color classification rules: Based on the optimal threshold coefficient k = 0.932, a fruit color classification rule is constructed. Since T... R ≈ 0.003, T B ≈ -0.003, both are close to 0. To simplify the classification process and improve the convenience of practical application, this invention uses a comprehensive score of 0 as the dividing value for fruit color classification and establishes the following judgment criteria: 1. If the overall score is CS i If the value is greater than 0, it is judged as a purple-red fruit; 2. If the overall score is CS i If < 0, it is judged as a purple-black fruit; 3. If the overall score is CS i = 0, then it is classified as intermediate (it rarely occurs in practice and can be handled according to the situation).
[0058] Experiment Example 2: Classification Model Validation Results.
[0059] 1. Internal validation: Backtesting based on F2 known phenotypic individual plants To evaluate the accuracy and stability of the comprehensive scoring classification model, the optimal discriminant model (k=0.932) was applied to 89 F2 individual plants with known phenotypes for back-judgment testing. Figure 5 As shown in the figure. The overall classification accuracy of the model was 97.75% (87 / 89), indicating good discrimination performance. Two misclassifications occurred, both classifying purplish-red fruit as purplish-black fruit (numbers 98 and 218), indicating a one-way error pattern in the model. Specifically, the model tends to classify a few marginal purplish-red samples into the purplish-black category, while no purplish-black samples were misclassified. This bias may be related to the fact that the minimum score for the purplish-red group (-0.686) was slightly lower than the classification threshold of 0. Phenotypic images, peel color difference parameters, and detailed classification results for 89 known phenotypic individual plants are shown below. Figure 6 As shown in Table 3, the model accurately classified the vast majority of samples, with misclassified samples concentrated near the scoring boundary. Overall, the model exhibits high discrimination stability and can be used for predicting subsequent unknown samples.
[0060] Table 3 Color difference parameters, comprehensive scores, and classification results of 89 known phenotypic samples based on the F2 population Continued from the table above Continued from the table above Continued from the table above Continued from the table above Continued from the table above 2. External validation: Generalization ability test based on independent variety resources: To test the generalization ability and universality of the constructed model, the optimal discriminant model (k=0.932, threshold ≈ 0) was applied to a completely independent varietal resource dataset. This dataset contained 76 different eggplant varieties, including 38 varieties with purple-red fruit and 38 varieties with purple-black fruit. Their genetic background, cultivation environment, and fruit morphology were unrelated to the F2 population used in the model construction. To comprehensively evaluate the model's classification performance, the following metrics were used for quantitative evaluation: (1) Accuracy: The proportion of correctly classified samples to the total number of samples, reflecting the overall classification effect of the model.
[0061] (2) Precision: For a certain category, the proportion of samples that the model predicts as belonging to that category is actually that category, reflecting the accuracy of the model's prediction.
[0062] Wherein, TP (True Positive) represents the number of positive class samples correctly predicted as positive by the model; FP (False Positive) represents the number of negative class samples incorrectly predicted as positive by the model.
[0063] (3) Recall: For a certain category, the proportion of samples that actually belong to that category that are correctly predicted by the model, reflecting the model's ability to identify the category.
[0064] FN (False Negative) is the number of positive class samples that the model incorrectly predicts as negative.
[0065] (4) F1 score: the harmonic mean of precision and recall, which comprehensively measures the classification performance of the model.
[0066] The values of the above indicators are all in the range of 0 to 1. The higher the value, the better the model performance.
[0067] The validation results are shown in Table 4. The overall classification accuracy of the model reached 93.4% (71 / 76), performing well on the independent variety set. The confusion matrix shows that all five misclassifications involved classifying deep purplish-red varieties as purplish-black, while no purplish-black varieties were misclassified. The combined scores of these five misclassified varieties ranged from -1.45 to -6.17, all close to the classification threshold of 0, but their actual fruit color, as determined by manual analysis, was a deeper purplish-red. The precision for purplish-red varieties was 100% (33 / 33), the recall was 86.8% (33 / 38), and the F1 score was 92.9%; the precision for purplish-black varieties was 88.4% (38 / 43), the recall was 100% (38 / 38), and the F1 score was 93.8%. The F1 scores for both types of varieties exceeded 92%, indicating that the model has excellent comprehensive discrimination ability for both purplish-red and purplish-black fruit colors. The above results suggest that samples with difficult-to-define purple-red or purple-black fruit color near the threshold are all dark purple-red fruit samples. Such single plants or strains should be avoided as much as possible when performing BSA pooling or GWAS association analysis.
[0068] Phenotypic images and detailed classification information of 76 independent eggplant varieties used for external validation are shown below. Figure 7 As shown in Table 5, this model is not an overfitting result specific to the F2 population. The classification rules constructed based on the color difference parameter have good stability and universality, and can be effectively applied to the classification and identification of eggplant fruit color germplasm resources in a wider range.
[0069] Table 4. Confusion matrix of model validation results on an independent eggplant variety set (n=76) Actual phenotype / Predicted phenotype Purple-red type Purple-black type total Recall rate <![CDATA[F1 value]]> Purple-red type (n=38) 33 5 38 86.8% 92.9% Purple-black type (n=38) 0 38 38 100% 93.8% total 33 43 76 Accuracy 100% 88.4% Note: This validation set contains only samples with clearly defined phenotypes: purple-red and purple-black. The overall accuracy is 93.4%, and the macro-average F1 score is 93.4%.
[0070] Table 5. Information and classification results of 76 independent eggplant varieties used for external validation. Fruit color (actual measurement) flesh color <![CDATA[Comprehensive Score (CS i )]]> Predicting fruit color Purple-black green -19.83 Purple-black Purple-black green -17.60 Purple-black Purple-black green -23.87 Purple-black Purple-black green -20.88 Purple-black Purple-black green -16.69 Purple-black Purple-black green -23.51 Purple-black Purple-black green -19.63 Purple-black Purple-black green -19.18 Purple-black Purple-black green -23.13 Purple-black Purple-black green -18.99 Purple-black Purple-black green -18.83 Purple-black Purple-black green -17.47 Purple-black Purple-black green -22.07 Purple-black Purple-black green -24.52 Purple-black Purple-black green -22.58 Purple-black Purple-black green -18.74 Purple-black Purple-black green -22.91 Purple-black Purple-black green -19.97 Purple-black Purple-black green -20.40 Purple-black Purple-black green -23.39 Purple-black Purple-black green -21.23 Purple-black Purple-black green -21.45 Purple-black Purple-black green -22.05 Purple-black Purple-black green -19.94 Purple-black Purple-black green -23.25 Purple-black Purple-black green -18.67 Purple-black Purple-black green -20.93 Purple-black Purple-black green -17.72 Purple-black Purple-black green -16.91 Purple-black Purple-black green -17.27 Purple-black Purple-black green -19.59 Purple-black Purple-black green -21.06 Purple-black Purple-black green -22.14 Purple-black Purple-black green -22.65 Purple-black Purple-black green -13.73 Purple-black Purple-black green -24.39 Purple-black Purple-black green -22.21 Purple-black Purple-black green -7.44 Purple-black Purple-red white -1.68 Purple-black Purple-red white 0.30 Purple-red Purple-red white 2.80 Purple-red Purple-red white 7.04 Purple-red Purple-red white 89.47 Purple-red Purple-red white 28.84 Purple-red Purple-red white 68.34 Purple-red Purple-red green 51.83 Purple-red Purple-red white 2.88 Purple-red Purple-red white 182.03 Purple-red Purple-red white 48.34 Purple-red Purple-red white 10.22 Purple-red Purple-red white 16.73 Purple-red Purple-red white 6.43 Purple-red Purple-red white 2.93 Purple-red Purple-red white 19.61 Purple-red Purple-red white -1.45 Purple-black Purple-red white 11.71 Purple-red Purple-red white 9.66 Purple-red Purple-red white 10.69 Purple-red Purple-red white 61.37 Purple-red Purple-red white 22.71 Purple-red Purple-red white -5.49 Purple-black Purple-red white -6.17 Purple-black Purple-red white 11.71 Purple-red Purple-red white 31.77 Purple-red Purple-red white 58.64 Purple-red Purple-red white 28.87 Purple-red Purple-red white 56.45 Purple-red Purple-red white 2.25 Purple-red Purple-red green 65.11 Purple-red Purple-red white -5.80 Purple-black Purple-red white 135.39 Purple-red Purple-red white 447.94 Purple-red Purple-red green 7.43 Purple-red Purple-red white 76.50 Purple-red Purple-red green 41.91 Purple-red Purple-red green 29.36 Purple-red Experimental Example 3: Prediction of fruit color classification and genetic ratio analysis of F2 population.
[0071] The optimized comprehensive scoring fruit color classification model was applied to all 232 F2 population individuals (including 89 known phenotypes and 143 unknown phenotypes), with a comprehensive score of 0 as the cutoff for fruit color determination. The results showed that 68 plants produced purple-red fruit, and 164 plants produced purple-black fruit; no intermediate-type individuals were found in the population. The comprehensive score distribution is shown in the figure below. Figure 8As shown, the fruit color distribution characteristics of the F2 population are clearly displayed: the scores of the purple-black fruit population are all in the negative range (<0), concentrated in the -30~0 range, with the main peak at -20~-15 (39 plants); the scores of the purple-red fruit population are mostly in the positive range (>0), concentrated in the 0~70 range, with the main peaks at 0~5 (13 plants) and 10~25 (11 plants each). The score ranges of purple-red and purple-black fruits are basically separated, showing a clear bimodal distribution, which is consistent with the quality trait characteristics controlled by the major gene.
[0072] Based on the above classification results, the segregation ratio of purple-black to purple-red in the F2 population was 164:68, approximately 2.41:1. A chi-square test showed that this ratio was not significantly different from the theoretical segregation ratio of 3:1 controlled by a single gene (χ²). 2 =2.30, df=1, p > 0.05), suggesting that the purple-black fruit color may be controlled by a single dominant major-effect gene, while the purple-red fruit color is a recessive homozygote.
[0073] In summary, this invention establishes a complete quantitative classification model for eggplant purple-red / purple-black fruit color based on CIELAB color parameters through standardized experimental material construction, systematic color difference data collection, standardized phenotypic dataset establishment, scientific parameter selection and weight allocation, dynamic threshold optimization, and multi-dimensional model validation. This specific implementation details all the technical details and key parameters of model construction, parameter optimization, classification determination, and validation. Those skilled in the art can repeat the implementation by following the steps described in this specification, ensuring the full disclosure, stability, reliability, and strong practicality of the technical solution of this invention. It can provide complete technical support for efficient, accurate, and quantitative classification of eggplant fruit color and subsequent genetic research.
[0074] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A quantitative classification modeling method for eggplant fruit color based on CIELAB, characterized in that, Includes the following steps: Step S1: Collect CIELAB color space parameters of the pericarp of the separated eggplant population. The values were determined, and a baseline dataset of known phenotypic traits, including purple-red and purple-black fruits, was created through manual visual inspection. Step S2, based on the benchmark dataset The color difference derived parameters are calculated, and key parameters that show significant differences in color between purplish-red and purplish-black fruits are selected from the derived parameters. Step S3: Assign weights based on the statistical discrimination ability of the key parameters in the benchmark dataset, and construct a weighted comprehensive scoring function; Step S4: Based on the comprehensive score distribution of the benchmark dataset calculated by the weighted comprehensive scoring function, a classification threshold is determined by dynamically optimizing the threshold coefficient, and a classification system based on the comprehensive score CS is established. i = 0 is the threshold for fruit color classification rules; Step S5: After confirming the accuracy of the fruit color classification rule through internal back-judgment of the training set and external verification of the test set, apply the fruit color classification rule to the fruit color classification prediction of unknown samples.
2. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 1, characterized in that, The key parameters include hue angle h° and red-to-blue ratio. Ratio, relative saturation and brightness The color characteristics of the fruit peel are comprehensively characterized from four dimensions: hue, red-blue ratio, relative saturation, and brightness.
3. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 1, characterized in that, The formula for calculating the weighted composite scoring function is as follows: Where: CS i P represents the overall score of the i-th individual plant; ij B is the actual measured value of the j-th parameter of the i-th individual plant; j The midpoint reference value for the j-th parameter; SD j W represents the standard deviation of the j-th parameter in the known sample. j Let be the weight of the j-th parameter.
4. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 1, characterized in that, The method for determining the classification threshold includes: Calculate the average score M of the known purple-red group. R The average score of the purple-black group is M. B and the standard deviation (SD) of all known sample scores total Set the purple-red threshold T R = M R -k ×SD total Purple-black threshold T B = M B +k ×SD total Where k is the threshold coefficient, which must satisfy 0 ≤ k < (M) R -M B ) / (2×SD total ); Iterate through the k values and select the T value corresponding to the k value that maximizes the accuracy of the known samples and minimizes the median. R and T B As a classification threshold.
5. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 1, characterized in that, The evaluation metric for internal back-judgment in the training set is classification accuracy; the evaluation metrics for external validation in the test set include accuracy, precision, recall, and F1 score.
6. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 2, characterized in that, The calculation formula for the color difference derived parameters includes: chromaticity Hue angle and according to , Adjust the sign to 0°~360°; relative saturation Red to blue ratio .
7. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 3, characterized in that, The weight allocation method is as follows: the ratio of the inter-group mean difference to the standard deviation of each parameter is used as the initial allocation basis. The total weight is set to 35 and then allocated proportionally, resulting in a final weight of 8.0 for the hue angle h° and a red-blue ratio. The final weight of the ratio is 9.0, relative saturation. The final weight is 9.5, brightness The final weight is 8.
5.
8. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 1, characterized in that, The fruit color classification rule is as follows: if the comprehensive score CS i If the score is > 0, it is judged as a purple-red fruit; if the overall score CS i If the score is less than 0, it is judged as a purple-black fruit; if the overall score is CS i = 0, then it is classified as intermediate type.
9. The method for quantitative classification and modeling of eggplant fruit color based on CIELAB according to claim 4, characterized in that, The step size for traversing the k value is 0.1, and the step size is increased to 0.01 in areas sensitive to accuracy changes.
10. The application of the CIELAB-based quantitative classification modeling method for eggplant fruit color according to any one of claims 1-9 in eggplant germplasm resource identification, genetic breeding, or fruit color-related gene mining.