Hami melon high-quality variety evaluation method and system based on big data

Through integrated learning methods and parameter optimization technology, a high-quality variety evaluation model for cantaloupe melon is constructed, which solves the problems of nonlinear relationships and feature complexity in cantaloupe melon variety evaluation, and achieves a more accurate and efficient variety quality evaluation.

CN120409901AInactive Publication Date: 2025-08-01哈密瓜鲜果农业科技发展有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510460452.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing high-quality varieties of cantaloupe melons, the advantages and disadvantages of cantaloupe melon varieties are affected by multiple dimensions and there are complex nonlinear relationships between factors. It is impossible to effectively handle the interaction between the characteristics of cantaloupe melon varieties, resulting in inaccurate and unreliable evaluation results, and the single parameter configuration is difficult to adapt to all characteristics, and it is impossible to accurately distinguish the quality of different cantaloupe melon varieties.

Method used

Using an integrated learning method, a weak evaluator is trained through a support vector machine, the weight is adjusted according to the error rate, the qualitative coefficient and threshold coefficient are introduced to update the data weight, and the optimal model parameters are found through parameter optimization, and the high-quality variety evaluation model of Hami melon is constructed.

Benefits of technology

It improves the accuracy, stability and reliability of the evaluation of high-quality varieties of cantaloupe melons, can more accurately identify and distinguish the advantages and disadvantages of cantaloupe melon varieties, and has better classification accuracy and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409901A_ABST
    Figure CN120409901A_ABST
Patent Text Reader

Abstract

The invention discloses a cantaloupe high-quality variety evaluation method and system based on big data, and belongs to the technical field of variety evaluation, and the method comprises the steps of cantaloupe data collection, cantaloupe data preprocessing, cantaloupe high-quality variety evaluation model construction, cantaloupe high-quality variety evaluation model parameter optimization and cantaloupe high-quality variety evaluation. According to the scheme, on the basis of an ensemble learning method, a rate change regulation factor is calculated according to an average error rate and the weight of a weak evaluator, a balance coefficient is introduced to update the data weight, the balance coefficient, a peak threshold coefficient and a valley threshold coefficient are updated, and weighted integration is performed on the weak evaluator; and calculating a trend value and a fluctuation parameter of each search position to obtain a corresponding guide value of each search position on each dimension, updating each dimension of each search position to find an optimal model parameter, and more accurately carrying out distinguishing evaluation on the Hami melon variety quality, so that the accuracy of the Hami melon variety quality is improved. And the accuracy, stability and reliability of high-quality variety evaluation of the Hami melons are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of variety evaluation, and specifically refers to a method and system for evaluating high-quality cantaloupe varieties based on big data. Background Art

[0002] The method for evaluating high-quality cantaloupe varieties is a method that uses big data analysis technology. By collecting and analyzing multi-dimensional factors of cantaloupe, it mines potential laws in historical data, identifies key factors affecting cantaloupe quality, evaluates the advantages and disadvantages of different varieties, helps optimize variety selection, improves the yield and quality of cantaloupe, and promotes the sustainable development of agricultural planting. However, in the existing methods for evaluating high-quality cantaloupe varieties, the advantages and disadvantages of cantaloupe varieties are affected by multiple dimensions and there are complex non-linear relationships between factors, and the interaction between cantaloupe variety characteristics cannot be effectively processed, resulting in inaccurate and unreliable evaluation results; in the existing methods for evaluating high-quality cantaloupe varieties, the variety evaluation of cantaloupe needs to determine the final variety quality according to multiple characteristics, and a single parameter configuration is difficult to adapt to all characteristics, resulting in the problem that the quality of different cantaloupe varieties cannot be accurately distinguished. Summary of the Invention

[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a method and system for evaluating high-quality cantaloupe varieties based on big data. Aiming at the problem that in the existing methods for evaluating high-quality cantaloupe varieties, the advantages and disadvantages of cantaloupe varieties are affected by multiple dimensions and there are complex non-linear relationships between factors, and the interaction between cantaloupe variety characteristics cannot be effectively processed, resulting in inaccurate and unreliable evaluation results, this solution is based on an ensemble learning method. The rate change regulation factor is calculated according to the average error rate and the weight of the weak evaluator, and the calibration coefficient is introduced to update the data weight, and the calibration coefficient, peak threshold coefficient and valley threshold coefficient are updated. The weak evaluators are weighted and integrated to obtain a high-quality cantaloupe variety evaluation model, which can evaluate the advantages and disadvantages of cantaloupe varieties more accurately and reliably, and effectively process the non-linear relationship between cantaloupe variety characteristics, thereby improving the accuracy, stability and reliability of high-quality cantaloupe variety evaluation; aiming at the problem that in the existing methods for evaluating high-quality cantaloupe varieties, the variety evaluation of cantaloupe needs to determine the final variety quality according to multiple characteristics, and a single parameter configuration is difficult to adapt to all characteristics, resulting in the problem that the quality of different cantaloupe varieties cannot be accurately distinguished, this solution calculates the tendency value and fluctuation parameter of each search position according to the fitness value, obtains the guiding value corresponding to each search position in each dimension, and updates each dimension of each search position according to the guiding value and the fitness value to find the optimal model parameters, so that the established high-quality cantaloupe variety evaluation model has better classification accuracy and generalization ability, can adapt to different variety characteristics, and thus can more accurately distinguish and evaluate the quality of cantaloupe varieties.

[0004] The technical solution adopted by the present invention is as follows: The method for evaluating high-quality varieties of Hami melons based on big data provided by the present invention includes the following steps:

[0005] Step S1: Collection of Hami melon data;

[0006] Step S2: Preprocessing of Hami melon data;

[0007] Step S3: Construction of an evaluation model for high-quality varieties of Hami melons;

[0008] Step S4: Optimization of parameters of the evaluation model for high-quality varieties of Hami melons;

[0009] Step S5: Evaluation of high-quality varieties of Hami melons.

[0010] Further, in step S1, the collection of Hami melon data is to collect historical data related to Hami melons, and the historical data related to Hami melons includes Hami melon variety types, planting environment data, Hami melon growth data, pest and disease data, market data, and evaluation grades.

[0011] Further, in step S2, the preprocessing of Hami melon data is to perform data cleaning, data standardization, and dataset construction on the collected historical data related to Hami melons to obtain a training dataset and a test dataset.

[0012] Further, in step S3, the construction of the evaluation model for high-quality varieties of Hami melons is to complete the construction of the evaluation model for high-quality varieties of Hami melons based on the ensemble learning method, which specifically includes the following steps:

[0013] Step S31: Initial data weight; Set the initial weight of each data in the training dataset to ; where N is the number of data in the training dataset;

[0014] Step S32: Training of weak evaluator units; According to the current data weight, use a support vector machine to train the training dataset to obtain a weak evaluator m r ; where m r is the weak evaluator obtained from the r-th training, and r is the training times index;

[0015] Step S33: Calculation of error rate; Calculate the error rate of the weak evaluator m r on the training dataset ;

[0016] Step S34: Update the weight of the weak evaluator; Adjust the weight of the weak evaluator according to the error rate;

[0017] Step S35: Design of rate change regulation factor; Calculate the rate change regulation factor based on the average error rate and the weight of the weak evaluator; The formula used is as follows:

[0018] ;

[0019] ;

[0020] wherein, and are the average error rates of the first r times and the first r - 1 times of training respectively, is the error rate of the weak evaluator obtained in the a - th training on the training data set, r and a are training times indices, β r+1 and β r are the rate change regulation factors at the (r + 1)-th and r - th trainings respectively, and are the error rates of the weak evaluators obtained at the (r + 1)-th and r - th trainings on the training data set respectively;

[0021] Step S36: Update the data weights; introduce a calibration coefficient to update the data weights; the formula used is as follows:

[0022] ;

[0023] wherein, [[ID= thirty-five]] and are the weights of the data x i at the (r + 1)-th and r - th trainings respectively, x i is the i-th data in the training data set, i is the data index, is the smoothing term, is the classification label of the weak evaluator m a for the data x i , m a is the weak evaluator obtained in the a - th training, ρ r is the calibration coefficient at the r - th training, R is the maximum number of trainings, W r is the weight of the weak evaluator m r , y i is the true label of the data x i , is the indicator function; when , is 1, otherwise is 0;

[0024] Step S37: Update the coefficients; preset that the value range of the initial calibration coefficient ρ0 is [1, 1.5], the value range of the initial peak threshold coefficient is [2.5, 3], the value range of the initial valley threshold coefficient is [1, 1.2] and the proportion threshold S wantThe value range of is [0.1, 0.2]; if the current training times r is less than the maximum training times R, calculate the proportion of misclassified data at the r-th training

[0025] Step S371: If , then go to Step S38 to stop the training of the weak evaluator;

[0026] Step S372: If , then take the average of the peak threshold coefficient and the valley threshold coefficient as the calibration coefficient at the (r + 1)-th training , the valley threshold coefficient is equal to , the peak threshold coefficient remains unchanged, and go to Step S32 to perform the next training of the weak evaluator; where and are the peak threshold coefficient and the valley threshold coefficient at the r-th training respectively, and are the peak threshold coefficient and the valley threshold coefficient at the (r + 1)-th training respectively;

[0027] Step S373: If , then take the average of the peak threshold coefficient and the valley threshold coefficient as the calibration coefficient at the (r + 1)-th training , the peak threshold coefficient is equal to , the valley threshold coefficient remains unchanged, and go to Step S32 to perform the next training of the weak evaluator;

[0028] Step S38: Integration; based on the weights of the weak evaluators, perform weighted integration on the trained weak evaluators to obtain a honeydew melon high-quality variety evaluation model.

[0029] Furthermore, in Step S4, the parameter optimization of the honeydew melon high-quality variety evaluation model specifically includes the following steps:

[0030] Step S41: Initial search position; for the initial calibration coefficient ρ0, the initial peak threshold coefficient , the initial valley threshold coefficient and the proportion threshold S want in the honeydew melon high-quality variety evaluation model, establish a parameter search space, randomly initialize H search positions within the parameter search space, use each search position to represent a set of model parameters, and take the cross-entropy loss of the honeydew melon high-quality variety evaluation model established based on the model parameters for the test data set as the fitness value corresponding to the search position;

[0031] Step S42: Design the guiding value; including the following steps:

[0032] Step S421: Calculate the tendency value; calculate the tendency value of each search position according to the fitness value; the formula used is as follows:

[0033] ;

[0034] In the formula, is the h-th search position at the t-th iteration, is 's tendency value, is the global optimal position at the t-th iteration, and are respectively and 's fitness values, h is the search position index, and t is the iteration number index;

[0035] Step S422: Calculate the fluctuation parameter; generate a random number within the range of (0, 1) for each search position , if , then the fluctuation parameter ; otherwise, the fluctuation parameter ; where is 's corresponding random number, is 's corresponding fluctuation parameter;

[0036] Step S423: Calculate the guiding value; calculate the guiding value corresponding to each search position in each dimension according to the fluctuation parameter and the global optimal position; the formula used is as follows:

[0037] ;

[0038] In the formula, and are respectively and 's values in the d-th dimension, is 's corresponding guiding value, d is the dimension index, is a random number generated within the range of (0, 1), and are respectively the upper and lower limits of the parameter search space in the d-th dimension;

[0039] Step S43: Update the search position; update each dimension of each search position according to the guiding value and the fitness value; the formula used is as follows:

[0040] ;

[0041] wherein, is the value of the h-th search position at the (t + 1)-th iteration in the d-th dimension, D is the total dimension of the parameter search space, is a random search position at the t-th iteration, and are respectively and 's fitness values, is the L2 norm; is to generate a D-dimensional column vector, where each element is a random number within the range of (0, 1); is the sign function;

[0042] Step S44: Determine the optimal model parameters; preset the fitness threshold. When there is a search position whose fitness value is less than the fitness threshold, the model parameters represented by this search position are used as the optimal model parameters, and a high-quality cantaloupe variety evaluation model is established based on the optimal model parameters; otherwise, if the maximum number of iterations is reached, return to Step S41 to re-initialize the search position; otherwise, increment the iteration count by 1 and return to Step S42 to continue the search.

[0043] Furthermore, in Step S5, the high-quality cantaloupe variety evaluation is to collect real-time cantaloupe-related data. After data cleaning and data standardization of the collected real-time cantaloupe-related data, it is input into the high-quality cantaloupe variety evaluation model established based on the optimal model parameters for evaluation, obtaining the evaluation level and generating an evaluation report.

[0044] The high-quality cantaloupe variety evaluation system provided by the present invention includes a cantaloupe data collection module, a cantaloupe data preprocessing module, a module for constructing a high-quality cantaloupe variety evaluation model, a module for optimizing the parameters of the high-quality cantaloupe variety evaluation model, and a high-quality cantaloupe variety evaluation module;

[0045] The cantaloupe data collection module collects historical cantaloupe-related data and sends the data to the cantaloupe data preprocessing module;

[0046] The cantaloupe data preprocessing module performs data cleaning, data standardization, and dataset construction processing on the collected historical cantaloupe-related data, and sends the data to the module for constructing a high-quality cantaloupe variety evaluation model;

[0047] The module for constructing a high-quality cantaloupe variety evaluation model, based on the ensemble learning method, calculates the rate change regulation factor according to the average error rate and the weight calculation rate of the weak evaluator, introduces the calibration coefficient to update the data weights, and updates the calibration coefficient, the peak threshold coefficient, and the valley threshold coefficient, and performs weighted integration on the weak evaluators to obtain a high-quality cantaloupe variety evaluation model, and sends the data to the module for optimizing the parameters of the high-quality cantaloupe variety evaluation model;

[0048] The optimization module for the parameters of the high-quality cantaloupe variety evaluation model calculates the tendency value and fluctuation parameters of each search position according to the fitness value, obtains the guiding value corresponding to each search position in each dimension, updates each dimension of each search position respectively according to the guiding value and the fitness value, finds the optimal model parameters, and sends the data to the high-quality cantaloupe variety evaluation module;

[0049] The high-quality cantaloupe variety evaluation module collects real-time cantaloupe-related data, performs data cleaning and data standardization processing, and then inputs it into the high-quality cantaloupe variety evaluation model established based on the optimal model parameters for evaluation, obtains the evaluation grade and generates an evaluation report.

[0050] The beneficial effects achieved by the present invention using the above solution are as follows:

[0051] (1) Aiming at the problem that in the existing high-quality cantaloupe variety evaluation methods, the quality of cantaloupe varieties is affected by multiple dimensions and there are complex non-linear relationships between factors, and the interaction between cantaloupe variety characteristics cannot be effectively processed, resulting in inaccurate and unreliable evaluation results. This solution is based on the ensemble learning method, trains weak evaluators through support vector machines, adjusts the weights of weak evaluators according to the error rate, calculates the rate change regulation factor according to the average error rate and the weights of weak evaluators, introduces a calibration coefficient to update the data weights, and updates the calibration coefficient, peak threshold coefficient and valley threshold coefficient, and performs weighted integration on the weak evaluators to obtain a high-quality cantaloupe variety evaluation model, which can more accurately and reliably evaluate the quality of cantaloupe varieties, effectively process the non-linear relationship between cantaloupe variety characteristics, and effectively avoid overfitting and underfitting, thereby improving the accuracy, stability and reliability of high-quality cantaloupe variety evaluation.

[0052] (2) Aiming at the problem that in the existing high-quality cantaloupe variety evaluation methods, the variety evaluation of cantaloupe needs to determine the final variety quality according to multiple characteristics, and a single parameter configuration is difficult to adapt to all characteristics, resulting in the inability to accurately distinguish the quality of different cantaloupe varieties. This solution uses the search position to represent the model parameters, calculates the tendency value and fluctuation parameters of each search position according to the fitness value, obtains the guiding value corresponding to each search position in each dimension, updates each dimension of each search position respectively according to the guiding value and the fitness value, finds the optimal model parameters, makes the evaluation process more efficient and accurate, enables the established high-quality cantaloupe variety evaluation model to have better classification accuracy and generalization ability, can adapt to different variety characteristics, and thus can more accurately distinguish and evaluate the quality of cantaloupe varieties. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic flow chart of the high-quality cantaloupe variety evaluation method based on big data provided by the present invention;

[0054] Figure 2 Schematic diagram of the cantaloupe high-quality variety evaluation system based on big data provided by the present invention;

[0055] Figure 3 Flow chart of step S3;

[0056] Figure 4 Flow chart of step S4.

[0057] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. Detailed implementation manners

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation to the present invention.

[0060] Embodiment 1, refer to Figure 1 , the cantaloupe high-quality variety evaluation method provided by the present invention, the method includes the following steps:

[0061] Step S1: Cantaloupe data collection; collect historical cantaloupe-related data;

[0062] Step S2: Cantaloupe data preprocessing; perform data cleaning, data standardization and dataset construction processing on the collected historical cantaloupe-related data;

[0063] Step S3: Construct a cantaloupe high-quality variety evaluation model; based on the ensemble learning method, calculate the rate change adjustment factor according to the average error rate and the weight calculation rate of the weak evaluator, introduce a calibration coefficient to update the data weight, and update the calibration coefficient, peak threshold coefficient and valley threshold coefficient, and perform weighted integration on the weak evaluators to obtain a cantaloupe high-quality variety evaluation model;

[0064] Step S4: Optimization of the parameters of the high-quality cantaloupe variety evaluation model; calculate the trend value and fluctuation parameter of each search position according to the fitness value, obtain the guiding value corresponding to each search position in each dimension, update each dimension of each search position respectively according to the guiding value and the fitness value, and find the optimal model parameters;

[0065] Step S5: Evaluation of high-quality cantaloupe varieties; collect real-time cantaloupe-related data, perform data cleaning and data standardization processing, and then input it into the high-quality cantaloupe variety evaluation model established based on the optimal model parameters for evaluation, obtain the evaluation level and generate an evaluation report.

[0066] Example 2, refer to Figure 1 , based on the above example, in step S1, the cantaloupe data collection is to collect historical cantaloupe-related data;

[0067] The historical cantaloupe-related data includes cantaloupe variety types, planting environment data, cantaloupe growth data, pest and disease data, market data, and evaluation levels;

[0068] The planting environment data includes meteorological data and soil data. The meteorological data includes temperature, humidity, light duration, light intensity, and precipitation. The soil data includes soil fertility, soil type, and soil pH value;

[0069] The cantaloupe growth data includes growth cycle data, plant morphology data, and fruit growth data. The plant morphology data includes the height of the plant, the thickness of the stem, the number of leaves, the size of the leaves, and the color of the leaves. The fruit growth data includes fruit weight, sugar content, and color;

[0070] The pest and disease data includes pest and disease types, pest and disease occurrence frequencies, and pest and disease damage degrees;

[0071] The market data includes price and sales volume;

[0072] The evaluation levels include excellent, good, average, and poor, and the evaluation levels are used as data labels.

[0073] Example 3, refer to Figure 1 , based on the above example, in step S2, the cantaloupe data preprocessing is to perform data cleaning, data standardization, and dataset construction processing on the collected historical cantaloupe-related data to obtain a training dataset and a test dataset;

[0074] The data cleaning includes missing value processing and outlier processing. The missing value processing is to fill the missing data using the K-nearest neighbor algorithm, and the outlier processing is to identify and delete outliers using the box plot method;

[0075] The data standardization includes numerical data standardization and categorical data encoding. The numerical data standardization maps all numerical data to a unified range using the maximum-minimum scaling method, and the categorical data encoding converts categorical data into numerical data using One-Hot encoding. The categorical data includes the types of honeydew melon varieties, the soil types in the planting environment data, the colors of the leaves in the honeydew melon growth data, the types of pests and diseases in the pest and disease data, and the evaluation grades. The numerical data is all the data in the historical honeydew melon-related data except the categorical data.

[0076] The construction of the dataset is to construct a training dataset and a test dataset based on the historical honeydew melon-related data after data cleaning and data standardization.

[0077] Example 4, refer to Figure 1 and Figure 3 , based on the above example, in step S3, the evaluation of the quality of honeydew melon varieties not only involves influencing factors in multiple dimensions, but also includes complex non-linear relationships between these factors. Based on the ensemble learning method, the construction of the high-quality honeydew melon variety evaluation model is completed, accurately considering the interaction of these multiple factors to ensure that the final evaluation model can accurately and reliably identify and evaluate the quality of honeydew melon varieties. The construction of the high-quality honeydew melon variety evaluation model specifically includes the following:

[0078] Step S31: Initial data weight; Set the initial weight of each data in the training dataset to ; where N is the number of data in the training dataset. Different honeydew melon data have different impacts on the evaluation, but initially, the key data cannot be effectively identified. Therefore, equal initial weights are set to avoid over-relying on partial data initially and ensure the stability and reliability of the evaluation.

[0079] Step S32: Train the weak evaluator unit; According to the current data weight, use the support vector machine to train the training dataset to obtain the weak evaluator m r ; where m r is the weak evaluator obtained in the r-th training, and r is the training times index. The evaluation of honeydew melon varieties involves multiple factors with complex non-linear relationships. The support vector machine is used to train the weak evaluator. Each training is based on the weight of the current data, and the learning of data features is gradually strengthened through repeated training to effectively handle the non-linear relationships between honeydew melon variety features.

[0080] Step S33: Calculate the error rate; Calculate the error rate of the weak evaluator m r on the training dataset ; Judging the performance of data during the training process through the error rate can quantify the deficiencies in each training, adjust the training strategy in a timely manner, and avoid the negative impact of error accumulation on the model. The formula used is as follows:

[0081] ;

[0082] In the formula, is the error rate of the weak evaluator m r on the training dataset, x i is the i-th data in the training dataset, and i is the data index. is the weight of the data x i at the r-th training, y i is the data x i 's true label. is the classification label of the weak evaluator m r for the data x i . is the indicator function; when , is 1, otherwise is 0.

[0083] Step S34: Update the weight of the weak evaluator; adjust the weight of the weak evaluator according to the error rate; by dynamically adjusting the weight of the weak evaluator, the influence of the well-performing evaluator is strengthened, and the evaluation accuracy of the final model is improved. The formula used is as follows:

[0084] ;

[0085] In the formula, W r is the weight of the weak evaluator m r .

[0086] Step S35: Design the rate change regulation factor; calculate the rate change regulation factor according to the average error rate and the weight of the weak evaluator; the error fluctuation of the evaluator will be affected by the complexity of the melon data. Over-adjustment will lead to overfitting, while insufficient adjustment will not be able to effectively improve the performance. By regulating the error change trend, it is possible to better avoid overfitting or underfitting of the model and improve the stability and accuracy of the evaluation results. The formula used is as follows:

[0087] ;

[0088] ;

[0089] In the formula, and are the average error rates of the first r and the first r - 1 trainings respectively, is the error rate of the weak evaluator obtained in the a-th training on the training dataset, and r and a are the training times indices, βr+1 and β r are the rate change regulation factors during the (r + 1)-th and r-th trainings respectively, and are the error rates of the weak evaluators obtained during the (r + 1)-th and r-th trainings on the training dataset respectively;

[0090] Step S36: Update data weights; introduce a calibration coefficient to update the data weights; some cantaloupe data will be misclassified during the initial training due to the complexity of their features, resulting in the need to adjust the data weights to help the model better identify these cantaloupe data. Update the data weights based on the calibration coefficient, effectively focusing the model's attention on the data that is difficult to classify, thereby improving the overall evaluation effect; the formula used is as follows:

[0091] ;

[0092] In the formula, and are the weights of data x i during the (r + 1)-th and r-th trainings respectively, x i is the i-th data in the training dataset, i is the data index, is the smoothing term, is the weight of the weak evaluator m a for data x i , m a is the weak evaluator obtained during the a-th training, ρ r is the calibration coefficient during the r-th training, R is the maximum number of trainings, W r is the weight of the weak evaluator m r , y i is the true label of data x i , is the indicator function; when , is 1, otherwise is 0;

[0093] Step S37: Update coefficients; preset the value range of the initial calibration coefficient ρ0 to be [1, 1.5], the value range of the initial peak threshold coefficient to be [2.5, 3], the value range of the initial valley threshold coefficient to be [1, 1.2] and the proportion threshold S want to be [0.1, 0.2]; if the current number of trainings r is less than the maximum number of trainings R, then calculate the proportion of misclassified data during the r-th training , update the calibration coefficient, peak threshold coefficient, and valley threshold coefficient based on the data ratio; otherwise, go to step S38 to stop the training of the weak evaluator; ensuring a more refined and flexible training process through the updated coefficients, thereby avoiding overfitting or underfitting and improving the adaptability of the model to different cantaloupe data; including the following steps:

[0094] Step S371: If , then go to step S38 to stop the training of the weak evaluator;

[0095] Step S372: If , then take the average of the peak threshold coefficient and the valley threshold coefficient as the calibration coefficient for the (r + 1)-th training , the valley threshold coefficient is equal to , the peak threshold coefficient remains unchanged, , and go to step S32 to perform the next training of the weak evaluator; where and are the peak threshold coefficient and the valley threshold coefficient for the r-th training respectively, and are the peak threshold coefficient and the valley threshold coefficient for the (r + 1)-th training respectively;

[0096] Step S373: If , then take the average of the peak threshold coefficient and the valley threshold coefficient as the calibration coefficient for the (r + 1)-th training , the peak threshold coefficient is equal to , the valley threshold coefficient remains unchanged, , and go to step S32 to perform the next training of the weak evaluator;

[0097] Step S38: Integration; based on the weights of the weak evaluators, perform weighted integration on the trained weak evaluators to obtain a cantaloupe high-quality variety evaluation model ; where R final is the final number of training times; by weighted integrating multiple weak evaluators, better synthesize the advantages of each weak evaluator, improve the final evaluation accuracy, and be able to more accurately predict the high-quality varieties of cantaloupe.

[0098] By performing the above operations, in view of the problem in the existing evaluation method for high-quality cantaloupe varieties that the quality of cantaloupe varieties is affected by multiple dimensions and there are complex non-linear relationships between factors, and the interaction between cantaloupe variety characteristics cannot be effectively processed, resulting in inaccurate and unreliable evaluation results, this solution is based on the ensemble learning method. Weak estimators are trained by support vector machines, the weights of the weak estimators are adjusted according to the error rate, the rate change regulation factor is calculated based on the average error rate and the weights of the weak estimators, the data weights are updated by introducing the calibration coefficient, and the calibration coefficient, peak threshold coefficient and valley threshold coefficient are updated. The weak estimators are weighted and integrated to obtain an evaluation model for high-quality cantaloupe varieties, which can more accurately and reliably evaluate the quality of cantaloupe varieties, effectively process the non-linear relationship between cantaloupe variety characteristics, and effectively avoid overfitting and underfitting, thereby improving the accuracy, stability and reliability of the evaluation of high-quality cantaloupe varieties.

[0099] Example Five, refer to Figure 1 and Figure 4 , based on the above example, in step S4, the variety evaluation of cantaloupe needs to determine the final variety quality according to multiple characteristics, and the relationship between these characteristics is complex and non-linear. By performing parameter optimization in the cantaloupe evaluation model, the parameter configuration that can optimize the evaluation result is found, so that the cantaloupe evaluation has better classification accuracy and generalization ability; the parameter optimization of the high-quality cantaloupe variety evaluation model specifically includes the following steps:

[0100] Step S41: Initial search position; for the initial calibration coefficient ρ0, initial peak threshold coefficient , initial valley threshold coefficient and the ratio threshold S want in the high-quality cantaloupe variety evaluation model, a parameter search space is established. H search positions are randomly initialized within the parameter search space, and each search position represents a set of model parameters. The cross-entropy loss of the high-quality cantaloupe variety evaluation model established based on the model parameters for the test data set is used as the fitness value corresponding to the search position; the high-quality evaluation of cantaloupe varieties involves multiple complex characteristics, and these characteristics vary greatly between different varieties. A single parameter configuration is difficult to adapt to all characteristics. By randomly initializing multiple search positions within the parameter space, corresponding to different parameter combinations, it is ensured that different characteristics and evaluation indicators of cantaloupe varieties can be widely explored without being limited to a single parameter configuration;

[0101] Step S42: Design the guiding value; including the following steps:

[0102] Step S421: Calculate the tendency value; calculate the tendency value of each search position according to the fitness value; the characteristic differences between different varieties need to be adapted by adjusting the tendency value. By guiding the search towards a better variety characteristic evaluation model, overfitting of the model to low-quality varieties can be reduced. The formula used is as follows:

[0103] ;

[0104] In the formula, is the h-th search position at the t-th iteration, is 's tendency value, is the global optimal position at the t-th iteration, and are respectively and 's fitness values, h is the search position index, and t is the iteration number index;

[0105] Step S422: Calculate the fluctuation parameter; generate a random number in the range (0, 1) for each search position. If , then the fluctuation parameter ; otherwise, the fluctuation parameter ; where is the random number corresponding to , and is the fluctuation parameter corresponding to ; The design of the fluctuation parameter allows the model to adjust the search stability according to the different characteristics of the cantaloupe varieties, ensuring that the cantaloupe variety evaluation model can still maintain the exploration ability when the characteristic complexity is high and can be stably optimized when the characteristics are simple;

[0106] Step S423: Calculate the guiding value; calculate the guiding value corresponding to each search position in each dimension according to the fluctuation parameter and the global optimal position; the guiding value enables the model to adapt to the evaluation requirements of different varieties of cantaloupe, helps optimize the search process, enables the evaluation model to adaptively adjust parameters among different varieties, and better captures the key characteristics affecting the judgment of high-quality varieties. The formula used is as follows:

[0107] ;

[0108] In the formula, and are respectively and 's values in the d-th dimension, is 's corresponding guiding value, d is the dimension index, is a random number generated in the range (0, 1), and They are the upper and lower limits of the parameter search space in the d-th dimension respectively;

[0109] Step S43: Update the search positions; update each dimension of each search position according to the guiding value and the fitness value; this update method can continuously optimize the parameters, enabling the final evaluation model to better adapt to the characteristics of each variety and improving the evaluation accuracy and reliability; the formula used is as follows:

[0110] ;

[0111] In the formula, is the value of the h-th search position in the d-th dimension at the (t + 1)-th iteration, D is the total dimension of the parameter search space, is a random search position at the t-th iteration, is 's fitness value, is the L2 norm; is to generate a D-dimensional column vector, where each element is a random number within the range of (0, 1); is the sign function. When , , when , , when , ;

[0112] Step S44: Determine the optimal model parameters; preset the fitness threshold. When the fitness value of a search position is less than the fitness threshold, the model parameters represented by this search position are used as the optimal model parameters, and a high-quality cantaloupe variety evaluation model is established based on the optimal model parameters; otherwise, if the maximum number of iterations is reached, return to Step S41 to re-initialize the search positions; otherwise, increment the number of iterations by 1 and return to Step S42 to continue the search.

[0113] By performing the above operations, for the problem in the existing high-quality cantaloupe variety evaluation method that the variety evaluation of cantaloupe needs to determine the final variety quality based on multiple features, and a single parameter configuration is difficult to adapt to all features, resulting in the inability to accurately distinguish the quality of different cantaloupe varieties, this solution uses the search positions to represent the model parameters, calculates the tendency value and fluctuation parameters of each search position according to the fitness value, obtains the guiding value corresponding to each search position in each dimension, updates each dimension of each search position according to the guiding value and the fitness value, finds the optimal model parameters, making the evaluation process more efficient and accurate, enabling the established high-quality cantaloupe variety evaluation model to have better classification accuracy and generalization ability, and being able to adapt to different variety characteristics, so as to more accurately distinguish and evaluate the quality of cantaloupe varieties.

[0114] Example 6. Refer to Figure 1 , this example is based on the above example. In step S5, for the evaluation of high-quality cantaloupe varieties, relevant real-time cantaloupe data is collected. After data cleaning and data standardization of the collected real-time cantaloupe relevant data, it is input into the high-quality cantaloupe variety evaluation model established based on the optimal model parameters for evaluation, obtaining the evaluation level and generating an evaluation report; the real-time cantaloupe relevant data includes cantaloupe variety types, planting environment data, cantaloupe growth data, pest and disease data, and market data.

[0115] Example 7. Refer to Figure 2 , this example is based on the above example. The high-quality cantaloupe variety evaluation system based on big data provided by the present invention includes a cantaloupe data collection module, a cantaloupe data preprocessing module, a module for constructing a high-quality cantaloupe variety evaluation model, a module for optimizing the parameters of the high-quality cantaloupe variety evaluation model, and a high-quality cantaloupe variety evaluation module;

[0116] The cantaloupe data collection module collects historical cantaloupe relevant data and sends the data to the cantaloupe data preprocessing module;

[0117] The cantaloupe data preprocessing module performs data cleaning, data standardization, and dataset construction processing on the collected historical cantaloupe relevant data, and sends the data to the module for constructing a high-quality cantaloupe variety evaluation model;

[0118] The module for constructing a high-quality cantaloupe variety evaluation model is based on the ensemble learning method, calculates the rate change adjustment factor according to the average error rate and the weight calculation rate of the weak evaluator, introduces a calibration coefficient to update the data weights, and updates the calibration coefficient, peak threshold coefficient, and valley threshold coefficient, and performs weighted integration on the weak evaluators to obtain a high-quality cantaloupe variety evaluation model, and sends the data to the module for optimizing the parameters of the high-quality cantaloupe variety evaluation model;

[0119] The module for optimizing the parameters of the high-quality cantaloupe variety evaluation model calculates the tendency value and fluctuation parameter of each search position according to the fitness value, obtains the guiding value corresponding to each search position in each dimension, updates each dimension of each search position respectively according to the guiding value and the fitness value, finds the optimal model parameters, and sends the data to the high-quality cantaloupe variety evaluation module;

[0120] The high-quality cantaloupe variety evaluation module collects real-time cantaloupe relevant data, performs data cleaning and data standardization processing, and then inputs it into the high-quality cantaloupe variety evaluation model established based on the optimal model parameters for evaluation, obtaining the evaluation level and generating an evaluation report.

[0121] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0122] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention.

[0123] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural modes and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. A method for evaluating high-quality varieties of Hami melons based on big data, characterized in that: The method includes the following steps: Step S1: Hami melon data collection; collect historical Hami melon-related data; Step S2: Hami melon data preprocessing; Step S3: Construct a high-quality variety evaluation model for Hami melons; Step S4: Optimize the parameters of the high-quality variety evaluation model for Hami melons; Step S5: Evaluate the high-quality varieties of Hami melons; collect real-time Hami melon-related data, perform data cleaning and data standardization, and then input it into the high-quality variety evaluation model of Hami melons established based on the optimal model parameters for evaluation; In step S3, it includes step S35: Design a rate variation regulation factor; calculate the rate variation regulation factor according to the average error rate and the weight of the weak evaluator; the formula used is as follows: ; ; Wherein, and are the average error rates of the first r times and the first r - 1 times of training respectively, is the error rate of the weak estimator obtained in the a-th training on the training data set. r and a are training times indices, and β r+1 and β r are the rate change regulation factors at the (r + 1)-th and r-th trainings respectively, and [[ID= 2. The method for evaluating high-quality cantaloupe varieties based on big data according to claim 1, wherein: In step S3, the construction of the high-quality variety evaluation model for Hami melons is to complete the construction of the high-quality variety evaluation model for Hami melons based on the ensemble learning method, including the following steps: Step S31: Initial data weight; Set the initial weight of each data in the training dataset to ; where N is the number of data in the training dataset; Step S32: Train the weak evaluator unit; according to the current data weights, use a support vector machine to train the training data set to obtain the weak evaluator m r ; where m r is the weak evaluator obtained from the r-th training; Step S33: Calculate the error rate; Calculate the weak evaluator m based on the data weights r Error rate on the training dataset ; Step S34: Update the weight of the weak evaluator; adjust the weight of the weak evaluator according to the error rate; Step S35: Design a rate variation regulation factor; Step S36: Update the data weight; introduce a calibration coefficient to update the data weight; the formula used is as follows: ; wherein, and are the weights of data x during the (r + 1)-th and r-th training respectively, i x i is the i-th data in the training dataset, i is the data index, is the smoothing term, is the classification label of the weak evaluator m a for data x i m a is the weak evaluator obtained from the a-th training, ρ r is the calibration coefficient during the r-th training, R is the maximum number of trainings, W r is the weight of the weak evaluator m r y i is the true label of data x i ; is the indicator function; when holds, is 1, otherwise is 0; Step S37: Update coefficients; the preset value range of the initial criterion coefficient ρ0 is [1, 1.5], the initial peak threshold coefficient is in the range of [2.5, 3], the initial valley threshold coefficient is in the range of [1, 1.2], and the ratio threshold S want is in the range of [0.1, 0.2]; if the current training times r is less than the maximum training times R, calculate the proportion of misclassified data at the r-th training , and update the criterion coefficient, peak threshold coefficient and valley threshold coefficient based on the data proportion; otherwise, go to step S38 to stop the training of the weak evaluator; Step S38: Ensemble; based on the weight of the weak evaluator, perform weighted ensemble on the trained weak evaluators to obtain the high-quality variety evaluation model of Hami melons.

3. The method for evaluating high-quality cantaloupe varieties based on big data according to claim 2, characterized in that: In step S37, the update coefficient specifically includes the following steps: Step S371: If , then go to step S38 and stop the training of the weak evaluator; Step S372: If , then use the average of the peak threshold coefficient and the valley threshold coefficient as the calibration coefficient for the (r + 1)-th training , and the valley threshold coefficient is equal to . The peak threshold coefficient remains unchanged, and go to Step S32 to perform the training of the next weak evaluator; where and are the peak threshold coefficient and the valley threshold coefficient for the (r + 1)-th training, respectively; Step S373: If , then use the average of the peak threshold coefficient and the valley threshold coefficient as the calibration coefficient for the (r + 1)-th training . The peak threshold coefficient is equal to . The valley threshold coefficient remains unchanged, and go to step S32 to perform the training of the next weak evaluator.

4. The method for evaluating high-quality cantaloupe varieties based on big data according to claim 1, characterized in that: In step S4, the optimization of the parameters of the high-quality variety evaluation model for Hami melons specifically includes the following steps: Step S41: Initial search positions; for the initial calibration coefficient ρ0, the initial peak threshold coefficient , the initial valley threshold coefficient and the proportional threshold S want in the melon high-quality variety evaluation model, a parameter search space is established. H search positions are randomly initialized within the parameter search space. Each search position represents a set of model parameters. The cross-entropy loss of the melon high-quality variety evaluation model established based on the model parameters for the test dataset is used as the fitness value corresponding to the search position; Step S42: Design a guiding value; Step S43: Update the search position; update each dimension of each search position according to the guiding value and the fitness value; the formula used is as follows: ; Wherein, is the value of the h-th search position at the (t + 1)-th iteration in the d-th dimension, D is the total dimension of the parameter search space, is the h-th search position at the t-th iteration, is a random search position at the t-th iteration, and are respectively and 's fitness values, is the L2 norm; is to generate a D-dimensional column vector, where each element is a random number within the range of (0, 1); is the sign function, h is the search position index, t is the iteration number index, d is the dimension index, is 's guiding value corresponding to the value in the d-th dimension; Step S44: Determine the optimal model parameters; preset a fitness threshold. When the fitness value of a search position is less than the fitness threshold, the model parameters represented by this search position are used as the optimal model parameters, and a high-quality variety evaluation model of Hami melons is established based on the optimal model parameters; otherwise, if the maximum number of iterations is reached, return to step S41 to re-initialize the search position; otherwise, increment the number of iterations by 1 and return to step S42 to continue the search.

5. The method for evaluating high-quality cantaloupe varieties based on big data according to claim 4, wherein: In step S42, the design of the guiding value specifically includes the following steps: Step S421: Calculate the tendency value; calculate the tendency value of each search position according to the fitness value; the formula used is as follows: ; Wherein, is the h-th search position at the t-th iteration, is 's tendency value, is the global optimal position at the t-th iteration, and are respectively and 's fitness values; Step S422: Calculate the fluctuation parameter; generate a random number within the range of (0, 1) for each search position , if , then the fluctuation parameter ; otherwise, the fluctuation parameter ; where is the corresponding random number,[[]] is the corresponding fluctuation parameter; Step S423: Calculate the guiding value; calculate the guiding value corresponding to each dimension of each search position according to the fluctuation parameter and the global optimal position; the formula used is as follows: ; Wherein, and are respectively and the values in the d-th dimension, is the corresponding guiding value, d is the dimension index, is to generate a random number within the range of (0, 1), and are respectively the upper and lower limits of the parameter search space in the d-th dimension.

6. The method for evaluating high-quality cantaloupe varieties based on big data according to claim 1, wherein: In step S1, the Hami melon data collection is to collect historical Hami melon-related data, and the historical Hami melon-related data includes Hami melon variety types, planting environment data, Hami melon growth data, pest and disease data, market data, and evaluation grades.

7. The method for evaluating high-quality cantaloupe varieties based on big data according to claim 1, characterized in that: In step S5, the evaluation of the high-quality varieties of Hami melons is to collect real-time Hami melon-related data. After data cleaning and data standardization of the collected real-time Hami melon-related data, it is input into the high-quality variety evaluation model of Hami melons established based on the optimal model parameters for evaluation to obtain an evaluation grade and generate an evaluation report.

8. The method for evaluating high-quality varieties of Hami melons based on big data according to claim 1, wherein: In step S2, the cantaloupe data preprocessing is to perform data cleaning, data standardization, and dataset construction on the collected historical cantaloupe-related data to obtain a training dataset and a test dataset.

9. A high-quality cantaloupe variety evaluation system based on big data, which is used to implement the high-quality cantaloupe variety evaluation method based on big data according to any one of claims 1-8, and is characterized in that: It includes a cantaloupe data acquisition module, a cantaloupe data preprocessing module, a cantaloupe high-quality variety evaluation model construction module, a cantaloupe high-quality variety evaluation model parameter optimization module, and a cantaloupe high-quality variety evaluation module; The cantaloupe data acquisition module collects historical cantaloupe-related data and sends the data to the cantaloupe data preprocessing module; The cantaloupe data preprocessing module performs data cleaning, data standardization, and dataset construction on the collected historical cantaloupe-related data, and sends the data to the cantaloupe high-quality variety evaluation model construction module; The cantaloupe high-quality variety evaluation model construction module, based on the ensemble learning method, calculates the rate change adjustment factor according to the average error rate and the weight calculation rate of the weak evaluator, introduces a calibration coefficient to update the data weights, and updates the calibration coefficient, peak threshold coefficient, and valley threshold coefficient, and performs weighted integration on the weak evaluators to obtain a cantaloupe high-quality variety evaluation model, and sends the data to the cantaloupe high-quality variety evaluation model parameter optimization module; The cantaloupe high-quality variety evaluation model parameter optimization module calculates the tendency value and fluctuation parameter of each search position according to the fitness value, obtains the guiding value corresponding to each search position in each dimension, updates each dimension of each search position according to the guiding value and the fitness value, finds the optimal model parameters, and sends the data to the cantaloupe high-quality variety evaluation module; The cantaloupe high-quality variety evaluation module collects real-time cantaloupe-related data, performs data cleaning and data standardization processing, and then inputs it into the cantaloupe high-quality variety evaluation model established based on the optimal model parameters for evaluation, obtains the evaluation level and generates an evaluation report.

Citation Information

Cited By

  • Watermelon resistance breeding optimization method and system based on cultivation data identification traceability

    CN121503787A