A Method for Generating an Evaluation Set of Large Models in the Field of Natural Resources
By establishing a data distribution model and annotation model, combining kernel density estimation and weighted average fusion algorithm, the problem of inaccurate data annotation quality distribution estimation in the natural resources field is solved, and the labeling quality and model performance of the evaluation set are improved.
Patent Information
- Application Number
- CN202510233228.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art is difficult to accurately reflect the true annotation quality distribution of data in the natural resource field, resulting in the impact of the quality of the evaluation set and model performance.
By establishing a data distribution model and annotation model, and combining statistical inference methods, the annotation quality distribution of the data set is obtained. Then, the labeled mass distribution is approximated using kernel density estimation to determine the center and bandwidth of the kernel function. Finally, the weighted average fusion algorithm is used to optimize the unbiased estimate of the kernel function to obtain the fusion estimate of the annotated quality of the data set.
The labeling quality of the large model evaluation set in the natural resources field is improved, and the non-uniform distribution and non-standard labeling noise characteristics are accurately depicted, which enhances the estimation accuracy and reliability of the labeling quality distribution.
Smart Images

Figure CN119721511B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural resources, and more specifically, to a method for generating a large model evaluation set in the field of natural resources. Background Art
[0002] In the field of natural resources, with the development of big data and artificial intelligence technologies, constructing a high-quality large model evaluation set has become a key link in promoting model performance improvement. Existing data annotation methods usually rely on simple annotation models, assuming that the data distribution is uniform and the annotation noise is standard. However, the data in the field of natural resources is complex, the data distribution is often non-uniform, and the annotation noise may also exhibit non-standard characteristics. This situation makes it difficult for existing methods to accurately reflect the true annotation quality distribution of the data when dealing with data in the field of natural resources, thereby affecting the quality of the evaluation set and the performance of the model.
[0003] In terms of technical principles, traditional annotation quality assessment methods mainly rely on a single annotation model, lacking in-depth analysis and modeling of data distribution characteristics. As a non-parametric estimation method, kernel density estimation has certain advantages in data distribution modeling. However, in the application in the field of natural resources, the corresponding relationship between the number of data sources and the number of kernel functions has not been fully considered, nor has it been considered how to further improve the fusion estimation accuracy of annotation quality through optimization algorithms.
[0004] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: existing annotation methods cannot effectively handle the non-uniform distribution and non-standard annotation noise of data in the field of natural resources, resulting in inaccurate estimation of the annotation quality distribution; at the same time, there is a lack of an effective fusion optimization mechanism to comprehensively integrate the estimation results of multiple kernel functions and further improve the reliability of annotation quality. These problems limit the generation of high-quality evaluation sets, thereby affecting the performance and application effects of large models in the field of natural resources. Summary of the Invention
[0005] The present invention provides a method for generating a large model evaluation set in the field of natural resources.
[0006] In the first aspect of the present invention, there is provided a method for generating a large model evaluation set in the field of natural resources, including:
[0007] S1. Obtaining the annotation quality distribution of a data set through statistical inference based on a data distribution model and an annotation model, where the data distribution model includes a non-uniform data distribution, and the annotation model includes non-standard annotation noise;
[0008] S2. Set the number of kernel functions for kernel density estimation to the number of data sources in the natural resources field, and perform an expansion approximation on the labeled quality distribution through kernel density estimation to determine the center and bandwidth of each kernel function respectively;
[0009] S3. Take the center of each of the kernel functions as an unbiased estimate value, and perform data fusion optimization through a weighted average fusion algorithm to obtain a fusion estimate value of the labeled quality of the data set. Select data based on the labeled quality to form a data set to form an evaluation set.
[0010] Further, in step S1, obtaining the labeled quality distribution of the data set through statistical inference based on the data distribution model and the labeling model includes:
[0011] S11. Establish a data distribution model and a labeling model;
[0012] S12. Obtain the labeled quality distribution based on the data distribution model and the labeling model according to statistical inference;
[0013] S13. Perform a weighted approximation on the non-uniform data distribution and the non-standard labeling noise through kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard labeling noise respectively;
[0014] S14. Obtain an approximate representation of the labeled quality distribution based on the approximate representations of each non-uniform distribution and non-standard labeling noise and the relationship between each non-uniform distribution and non-standard labeling noise and the labeled quality distribution.
[0015] Further, in step S12, obtaining the labeled quality distribution based on the data distribution model and the labeling model according to statistical inference includes:
[0016] Predict the prior distribution of the current sample based on the labeled quality distribution of the previous sample and the data distribution model;
[0017] Update the prior distribution of the current sample based on the labeling model to obtain the labeled quality distribution of the current sample. The calculation method is: divide the product of the labeled value distribution of sample k based on data features through the labeling model and the prior distribution of the labeled values based on the previous sample by the prior distribution of the labeled values based on the previous sample.
[0018] Further, in step S13, performing a weighted approximation on the non-uniform data distribution and the non-standard labeling noise respectively through kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard labeling noise respectively includes:
[0019] Construct a kernel density estimation model with m kernel functions;
[0020] Perform a weighted sum of the kernel functions of the kernel density estimation model to obtain an approximation of each of the non-uniform data distribution and the non-standard annotation noise.
[0021] Further, in step S14, based on the approximate representations of the non-uniform distributions and non-standard annotation noises and the relationships between the non-uniform distributions, non-standard annotation noises, and the annotation quality distribution, to obtain an approximate representation of the annotation quality distribution, including:
[0022] Approximate the annotation quality distribution of the previous sample through kernel density estimation;
[0023] Substitute the approximate representation of the non-uniform data distribution into the data distribution model for approximation of the data distribution model;
[0024] Substitute the approximate representation of the non-standard annotation noise into the annotation model to obtain an approximation of the annotation value distribution obtained through the annotation model;
[0025] Substitute the approximation of the data distribution model and the annotation quality distribution of the previous sample k - 1 into the prior distribution to obtain an approximation of the prior distribution;
[0026] Substitute the approximation of the prior distribution and the approximation of the annotation value distribution into the posterior distribution to obtain an approximate representation of the annotation quality distribution.
[0027] Further, taking the centers of the respective kernel functions as unbiased estimates, and performing data fusion optimization through a weighted average fusion algorithm to obtain a fused estimate value of the annotation quality of the data set, including:
[0028] Set the fused estimate value as the sum of the products of the fusion weights corresponding to each unbiased estimate value and the unbiased estimate value, and satisfy that the sum of the fusion weights is 1;
[0029] Determine the optimization objective function as minimizing the mean square error of the fused estimate value, and obtain the initial mean square error based on the covariance matrix or cross-covariance matrix between the initial fusion weights and the unbiased estimate values;
[0030] Calculate the gradient of the objective function with respect to each of the fusion weights;
[0031] Update the fusion weights based on the gradient;
[0032] Calculate the new mean square error based on the updated fusion weights. If the new mean square error is less than the mean square error before this iteration, update the mean square error to the mean square error obtained in this iteration; otherwise, retain the mean square error and fusion weights before the update;
[0033] Iteratively update the fusion weights until a preset condition is satisfied to obtain the final fusion weights;
[0034] An integrated estimated value of the dataset annotation quality is obtained based on the final integrated weight and the corresponding unbiased estimated value.
[0035] Further, the objective function is equivalently determined and expressed as:
[0036]
[0037] where represents the objective function, represents the integrated weight of the i-th unbiased estimated value, represents the variance of the i-th unbiased estimated value, represents the covariance between the i-th unbiased estimated value and the j-th unbiased estimated value.
[0038] In a second aspect of the present invention, a device for generating an evaluation set of a large model in the field of natural resources is provided, including:
[0039] An inference module, configured to obtain the annotation quality distribution of the dataset through statistical inference based on a data distribution model and an annotation model, where the data distribution model includes a non-uniform data distribution and the annotation model includes non-standard annotation noise;
[0040] An approximation module, configured to set the number of kernel functions of kernel density estimation to the number of data sources in the field of natural resources, and perform expansion approximation on the annotation quality distribution through kernel density estimation to determine the center and bandwidth of each kernel function;
[0041] An integration module, configured to use the centers of the respective kernel functions as unbiased estimated values, and perform data integration optimization through a weighted average integration algorithm to obtain an integrated estimated value of the dataset annotation quality.
[0042] In a third aspect of the present invention, an electronic device is provided, where the electronic device includes: at least one processor, a memory, and an input-output unit; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the method according to any one of the first aspect.
[0043] In a fourth aspect of the present invention, a computer-readable storage medium is provided, which includes instructions that, when running on a computer, cause the computer to execute the method according to any one of the first aspect.
[0044] The above embodiments of the present invention have at least the following beneficial effects: The present invention can effectively improve the annotation quality of the evaluation set of large models in the field of natural resources. By establishing a data distribution model and an annotation model, and combining statistical inference methods, it is possible to accurately characterize the non-uniform distribution of data and the characteristics of non-standard annotation noise, thereby providing a more reliable basis for estimating the annotation quality distribution. In addition, by using kernel density estimation to expand and approximate the annotation quality distribution and matching the number of kernel functions with the number of data sources in the field of natural resources, the accuracy and flexibility of the annotation quality distribution estimation can be further improved. At the same time, based on the weighted average fusion algorithm to optimize the unbiased estimated values of each kernel function, the advantages of multiple estimation results can be integrated, the error can be reduced, and finally a high-quality dataset annotation quality fusion estimated value can be obtained.
[0045] The present invention can also provide higher-quality data support for the evaluation of large models in the field of natural resources. By optimizing the process of estimating the annotation quality distribution, it is possible to effectively reduce the errors caused by uneven data distribution and annotation noise, thereby improving the accuracy and reliability of the evaluation set. This is of great significance for improving the application effect and performance of large models in the field of natural resources, and can better meet the model evaluation needs in the complex data environment of the natural resources field, promoting the development and application of related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of example and not limitation, wherein:
[0047] Figure 1 is a schematic flowchart of a method for generating an evaluation set of a large model in the field of natural resources provided by an embodiment of the present invention;
[0048] Figure 2 is a schematic structural diagram of a device for generating an evaluation set of a large model in the field of natural resources provided by an embodiment of the present invention;
[0049] Figure 3 schematically shows a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to convey the scope of the present invention fully to those skilled in the art.
[0051] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0052] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0053] The following refers to Figure 1 , Figure 1 which is a schematic flowchart of a method for generating a large model evaluation set in the field of natural resources provided for an embodiment of the present invention. As Figure 1 shown, a method 100 for generating a large model evaluation set in the field of natural resources includes:
[0054] S1. Obtain the annotation quality distribution of the data set through statistical inference based on the data distribution model and the annotation model, where the data distribution model includes a non-uniform data distribution, and the annotation model includes non-standard annotation noise;
[0055] S2. Set the number of kernel functions of kernel density estimation to the number of data sources in the field of natural resources, and perform expansion approximation on the annotation quality distribution through kernel density estimation to determine the center and bandwidth of each kernel function;
[0056] S3. Use the centers of each of the kernel functions as unbiased estimated values, and perform data fusion optimization through a weighted average fusion algorithm to obtain a fusion estimated value of the annotation quality of the data set.
[0057] It should be noted that the present invention relates to a method for generating a large model evaluation set in the field of natural resources, aiming to improve the accuracy and reliability of the evaluation set by optimizing the data annotation quality distribution. The method includes three main steps: First, based on the data distribution model and the annotation model, obtain the annotation quality distribution of the data set through statistical inference. The data distribution model is used to describe the non-uniform distribution characteristics of the data, while the annotation model is used to describe the non-standard noise in the data annotation process. Through statistical inference, the annotation quality distribution of the data set can be accurately estimated.
[0058] Specifically, the establishment of the data distribution model and the annotation model is a key step in this method. The data distribution model can include various non-uniform data distribution forms, such as normal distribution, Poisson distribution, etc., while the annotation model can include different types of annotation noises, such as Gaussian noise, uniform noise, etc. Kernel density estimation is used to perform an unfolding approximation of the annotation quality distribution, that is, the unfolding of the approximate representation of the annotation quality distribution. The number of kernel functions is set to the number of data sources in the natural resources field to ensure the accuracy and flexibility of the estimation. The kernel function in kernel density estimation can be selected as Gaussian kernel, triangular kernel, etc., and the bandwidth parameter is adjusted according to the data characteristics.
[0059] More specifically, the data fusion optimization is achieved through a weighted average fusion algorithm. The centers of the respective kernel functions serve as unbiased estimators, and data fusion optimization is carried out through the weighted average fusion algorithm to obtain a fused estimated value of the annotation quality of the data set. The optimization objective is to minimize the mean square error of the fused estimated value. By iteratively updating the fusion weights, the optimal fusion weights are finally obtained. Preferably, the gradient descent method can be used to calculate the gradient of the objective function with respect to each fusion weight, and the fusion weights are updated according to the gradient to ensure the convergence and stability of the fusion process.
[0060] In some embodiments, in step S1, obtaining the annotation quality distribution of the data set through statistical inference based on the data distribution model and the annotation model includes:
[0061] S11. Establish the data distribution model and the annotation model, expressed as:
[0062]
[0063] where, represents the data sample number, represents the sample at the data feature, represents the sample data distribution model, represents a non-uniform data distribution, represents the sample annotation value, represents the sample annotation model, represents non-standard annotation noise;
[0064] S12. Obtain the annotation quality distribution based on the data distribution model and the annotation model according to statistical inference;
[0065] S13. Perform a weighted approximation of the non-uniform data distribution and the non-standard annotation noise through kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard annotation noise respectively;
[0066] S14. Obtain an approximate representation of the annotation quality distribution based on the approximate representations of each non-uniform distribution and non-standard annotation noise, as well as the relationships between each non-uniform distribution and non-standard annotation noise and the annotation quality distribution.
[0067] It should be noted that based on the establishment of the data distribution model and the annotation model, the process of obtaining the annotation quality distribution through statistical inference further refines the model construction and inference methods. The data distribution model is expressed as , where represents the data characteristics at sample , represents the data distribution model of sample , represents the non-uniform data distribution. The annotation model is expressed as , where represents the annotation value of sample , represents the annotation model of sample , represents the non-standard annotation noise. By performing weighted approximation on the non-uniform data distribution and the non-standard annotation noise through kernel density estimation, their respective approximate representations can be obtained, thus providing a basis for the estimation of the annotation quality distribution.
[0068] Specifically, the parameter settings of the data distribution model and the annotation model are the key to implementing this method. In the data distribution model, can be a linear or non-linear function used to describe the dynamic change process of the data. For example, a polynomial function or a neural network can be used to approximate complex data distributions. The non-uniform data distribution can be a random variable with a specific probability distribution, such as a Gaussian distribution or a Laplace distribution, and its parameters (such as the mean and variance) can be adjusted according to the actual data characteristics. In the annotation model, can be a simple linear mapping or a complex non-linear mapping used to describe the relationship between data characteristics and annotation values. The non-standard annotation noise can be Gaussian noise, uniform noise, or other types of noise, and its parameters (such as the noise intensity) can also be set according to the actual errors in the annotation process. The number of kernel functions in the kernel density estimation matches the number of data sources in the natural resources field. The kernel function can be selected from Gaussian kernels, uniform kernels, triangular kernels, etc., and the bandwidth parameter is optimized and selected according to the distribution characteristics of the data.
[0069] Preferably, in order to further improve the accuracy of the annotation quality distribution estimation, the operation steps can be refined in the following ways: First, for the selection of the kernel function in the kernel density estimation, the Gaussian kernel function can be adopted because it has good smoothness and locality and can better adapt to the characteristics of non-uniform data distribution. Second, the bandwidth parameter can be optimally selected by the cross-validation method to ensure the accuracy and stability of the kernel density estimation. In addition, for the approximate representation of non-uniform data distribution and non-standard annotation noise, the adaptive weighting method can be adopted to dynamically adjust the weights according to the local characteristics of the data, thereby improving the accuracy of the approximation. Finally, the approximation of the annotation quality distribution can be further optimized by the Bayesian inference method, combining the prior distribution and the observed data to obtain a more accurate posterior distribution estimation.
[0070] In some embodiments, in step S12, obtaining the annotation quality distribution based on statistical inference according to the data distribution model and the annotation model includes:
[0071] Predicting the prior distribution of the current sample based on the annotation quality distribution of the previous sample and the data distribution model, expressed as:
[0072]
[0073] where, represents the prior distribution based on the annotation value of the previous sample, represents the data feature distribution of sample k obtained through the data distribution model, represents the annotation quality distribution of the previous sample k-1;
[0074] Updating the prior distribution of the current sample based on the annotation model to obtain the annotation quality distribution of the current sample, expressed as:
[0075]
[0076] where, represents the annotation value distribution of sample k based on the data features through the annotation model, represents the annotation quality distribution of the current sample k.
[0077] It should be noted that in the statistical inference process of the annotation quality distribution of the present invention, by combining the data distribution model and the annotation model, the calculation methods of the prior distribution and the posterior distribution are further refined. Specifically, based on the annotation quality distribution of the previous sample and the data distribution model, the prior distribution of the current sample is predicted, expressed as . Where, represents the prior distribution based on the annotation value of the previous sample, Represents the data feature distribution of sample k obtained through the data distribution model. Represents the previous sample of the annotation quality distribution. The prior distribution is updated through the annotation model to obtain the annotation quality distribution of the current sample . This process utilizes the idea of Bayesian inference, combines prior knowledge and observed data, and dynamically updates the estimation of the annotation quality distribution.
[0078] Specifically, the calculation of the prior distribution and the posterior distribution is the core link in the estimation of the annotation quality distribution. In the calculation of the prior distribution, it is possible to use the data distribution model and the non-uniform data distribution to describe the data feature distribution of the sample . For example, if the data distribution model is a linear dynamic system, it can be a linear function, and it can be Gaussian noise. In the update of the posterior distribution, the annotation model describes the relationship between the data features and the annotation values, and the non-standard annotation noise can be Gaussian noise or other types of noise. Through Bayes' formula , combining the prior distribution and the observed data, the estimation of the annotation quality distribution can be dynamically updated. This process not only considers the dynamic changes of the data but also combines the noise characteristics in the annotation process, thereby improving the accuracy of the estimation of the annotation quality distribution.
[0079] Preferably, in order to further optimize the estimation process of the annotation quality distribution, the calculations of the prior distribution and the posterior distribution can be refined. For example, in the calculation of the prior distribution, time series analysis methods such as Kalman filtering or particle filtering can be introduced to better handle the dynamic changes of the data. For the non-standard annotation noise , an adaptive noise modeling method can be adopted to dynamically adjust the noise parameters according to the statistical characteristics of the annotation data.
[0080] Furthermore, in the update of the posterior distribution, the Monte Carlo method can be used for numerical calculation to improve the calculation efficiency and accuracy. In an alternative solution, for the parameter estimation of the data distribution model and the annotation model, maximum likelihood estimation or Bayesian estimation methods can be adopted and optimized in combination with the actual data, thereby further improving the accuracy and reliability of the estimation of the annotation quality distribution.
[0081] In some embodiments, in step S13, the non-uniform data distribution and the non-standard annotation noise are respectively weighted and approximated through kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard annotation noise respectively, including:
[0082] Construct a kernel density estimation model with m kernel functions, expressed as:
[0083]
[0084] where, represents the kernel density estimation model, represents the kernel function, represents the center of the i-th kernel function, represents the bandwidth of the i-th kernel function;
[0085] Perform a weighted sum of the kernel functions of the kernel density estimation model to obtain approximations of the non-uniform data distribution and the non-standard annotation noise respectively;
[0086] The approximation of the non-uniform data distribution is expressed as:
[0087]
[0088] where, is the approximation of the non-uniform data distribution, is used for approximating the number of kernel functions, is the weight corresponding to the i-th kernel function in the approximation of the non-uniform data distribution;
[0089] The approximation of the non-standard annotation noise is expressed as:
[0090]
[0091] where, is the approximation of the non-standard annotation noise, is used for approximating the number of kernel functions, represents the weight corresponding to the i-th kernel function in the approximation of the non-standard annotation noise.
[0092] It should be noted that in the process of approximating the annotation quality distribution in the present invention, the non-uniform data distribution and the non-standard annotation noise are weighted and approximated through kernel density estimation, further refining the approximation method of the annotation quality distribution. Kernel density estimation is a non-parametric estimation method that approximates complex data distributions through the weighted sum of multiple kernel functions. In the present invention, the kernel density estimation model is expressed as , where represents the kernel density estimation model, represents the kernel function, represents the center of the i-th kernel function, represents the The bandwidth of a kernel function. Through kernel density estimation, approximate representations of non-uniform data distributions and non-standard annotation noises can be obtained respectively, providing a more accurate basis for estimating the annotation quality distribution.
[0093] Specifically, the selection of the kernel function and parameter setting in kernel density estimation are the key points for approximating the annotation quality distribution. The kernel function can be selected from Gaussian kernel, uniform kernel, triangular kernel or other kernels, and these kernel functions have different smoothing characteristics and locality. For example, the Gaussian kernel function is suitable for approximating most data distributions due to its good smoothing property and locality. The center and bandwidth of the kernel function are the key parameters of kernel density estimation. The center can be determined by local feature points of the data, and the bandwidth determines the smoothing degree of the kernel function, which is usually adjusted according to the distribution characteristics of the data and the noise level. For the approximate representation of non-uniform data distributions and the approximate representation of non-standard annotation noises and weights of the kernel function can be used to achieve this. The number of kernel functions can be set according to the number and complexity of data sources, and the weights can be determined by optimization methods (such as minimizing the mean square error).
[0094] Preferably, in order to further optimize the kernel density estimation process, the selection of the kernel function and parameter setting can be refined. For example, in the selection of the kernel function, a mixed kernel function can be selected in combination with the distribution characteristics of the data, such as the combination of Gaussian kernel and triangular kernel, to better adapt to complex data distributions. For the bandwidth parameter , an adaptive bandwidth method can be adopted to dynamically adjust the bandwidth according to the local data density, thereby improving the accuracy of kernel density estimation. In addition, the weights of the kernel function can be adjusted by cross-validation or Bayesian optimization methods to ensure the stability and accuracy of kernel density estimation. In an alternative solution, a regularization term, such as L1 or L2 regularization, can be introduced to prevent overfitting and further improve the reliability of approximating the annotation quality distribution.
[0095] In some embodiments, in step S14, based on the approximate representations of each non-uniform distribution and non-standard annotation noise and the relationship between each non-uniform distribution and non-standard annotation noise and the annotation quality distribution to obtain an approximate representation of the annotation quality distribution, including:
[0096] Approximating the annotation quality distribution of the previous sample through kernel density estimation, expressed as:
[0097]
[0098] Among them, represents the annotation quality distribution of the previous sample k - 1, is the weight corresponding to the i-th kernel function in the approximate representation of the annotation quality distribution of the previous sample;
[0099] Substitute the approximate representation of the non-uniform data distribution into the data distribution model for approximation of the data distribution model, expressed as:
[0100]
[0101] Substitute the approximate representation of the non-standard annotation noise into the annotation model to obtain an approximation of the annotation value distribution obtained through the annotation model, expressed as:
[0102]
[0103] Substitute the approximation of the data distribution model and the annotation quality distribution of the previous sample k - 1 into the prior distribution to obtain an approximation of the prior distribution, expressed as:
[0104]
[0105] Substitute the approximation of the prior distribution and the approximation of the annotation value distribution into the posterior distribution to obtain an approximate representation of the annotation quality distribution, expressed as:
[0106]
[0107] It should be noted that in the process of approximate representation of the annotation quality distribution in the present invention, the application of kernel density estimation is further refined. By applying kernel density estimation to the annotation quality distribution of the previous sample, as well as the approximations of the data distribution model and the annotation model, an approximate representation of the annotation quality distribution is obtained. This process utilizes the flexibility and non-parametric characteristics of kernel density estimation, and can better adapt to the complexity of non-uniform data distribution and non-standard annotation noise. Specifically, kernel density estimation approximates the annotation quality distribution of the previous sample, combines the approximation results of the data distribution model and the annotation model, and finally obtains an approximate representation of the annotation quality distribution of the current sample, providing a basis for subsequent data fusion optimization.
[0108] Specifically, the application of kernel density estimation in the approximation of the annotation quality distribution involves multiple key steps. First, for the annotation quality distribution of the previous sample , approximate it through kernel density estimation, expressed as , where is the weight, is the center of the kernel function, is the bandwidth. Secondly, the approximation of the data distribution model can be represented by substituting the approximation result of the non-uniform data distribution into the model as . Similarly, the approximation of the annotation value distribution of the annotation model can be represented by substituting the approximation result of the non-standard annotation noise as . Combining these approximation results with the update formulas of the prior distribution and the posterior distribution, an approximate representation of the annotation quality distribution is finally obtained. In terms of parameter settings, the number of kernel functions can be selected according to the complexity of the data source, and the weights can be determined by an optimization method (such as minimizing the mean square error), and the bandwidth can be adjusted by cross-validation or heuristic methods.
[0109] Preferably, in order to further optimize the approximation process of the annotation quality distribution, the parameter selection and calculation method of kernel density estimation can be refined. For example, in the selection of the kernel function, a Gaussian kernel function combined with a local adaptive bandwidth method can be adopted to better adapt to the local changes of the data distribution. For the weights of the kernel functions , they can be dynamically adjusted by Bayesian optimization methods to ensure the accuracy and stability of kernel density estimation.
[0110] Furthermore, a regularization term can be introduced to prevent overfitting, especially when dealing with high-dimensional data. In an alternative solution, for the approximate representation of the annotation quality distribution, a hierarchical kernel density estimation method can be adopted, dividing the data into multiple levels for modeling, thereby improving the fitting ability for complex data distributions. This method can further improve the accuracy of the annotation quality distribution estimation while maintaining the computational efficiency.
[0111] In some embodiments, taking the centers of the respective kernel functions as unbiased estimates and performing data fusion optimization through a weighted average fusion algorithm to obtain a fused estimate of the dataset annotation quality includes:
[0112] Setting the fused estimate as the sum of the products of the fusion weights corresponding to each unbiased estimate and the unbiased estimate, and satisfying that the sum of the fusion weights is 1;
[0113] Determining the optimization objective function as minimizing the mean square error of the fused estimate, and obtaining the initial mean square error based on the covariance matrix or cross-covariance matrix between the initial fusion weights and the unbiased estimates;
[0114] Calculating the gradient of the objective function with respect to each of the fusion weights;
[0115] Updating the fusion weights based on the gradient;
[0116] Calculate the new mean square error based on the updated fusion weights. If the new mean square error is less than the mean square error before this iteration, update the mean square error to the mean square error obtained in this iteration; otherwise, retain the mean square error and fusion weights before the update.
[0117] Iteratively update the fusion weights until a preset condition is met to obtain the final fusion weights.
[0118] Obtain the fused estimate value of the final dataset annotation quality based on the final fusion weights and the corresponding unbiased estimate values.
[0119] It should be noted that in the process of fused estimation of the annotation quality distribution in the present invention, the unbiased estimate values obtained by kernel density estimation are optimized through a weighted average fusion algorithm to obtain the fused estimate value of the dataset annotation quality. The core of this process lies in setting the fused estimate value as the optimization result of the objective function, where the sum of the fusion weights is 1, and the objective function is to minimize the mean square error of the fused estimate value. By iteratively updating the fusion weights, the optimal fused estimate value is finally obtained, thereby improving the reliability and accuracy of the annotation quality.
[0120] Specifically, the implementation of the weighted average fusion algorithm involves multiple key steps and parameter settings. First, the calculation formula for the fused estimate value is , where is the fusion weight of the th unbiased estimate value, and is the th unbiased estimate value. The objective function is defined as minimizing the mean square error of the fused estimate value, that is, , where is the variance of the th unbiased estimate value, and is the covariance between the th and the th unbiased estimate values. The initial fusion weights can be set by uniform distribution or based on prior knowledge, and then the weights are updated by calculating the gradient of the objective function with respect to each fusion weight. In each iteration, if the new mean square error is less than the current value, update the fusion weights and the mean square error; otherwise, retain the current value. This process is optimized through iteration to finally obtain the optimal fusion weights and fused estimate values.
[0121] Preferably, to further optimize the fusion estimation process, refinement can be carried out in the definition of the objective function and the optimization method. For example, the objective function can introduce a regularization term, such as L2 regularization, to prevent overfitting and improve the stability of the fusion estimation. In terms of the optimization method, in addition to the gradient descent method, more efficient optimization algorithms, such as the Newton method or the conjugate gradient method, can be adopted to accelerate the convergence speed and improve the computational efficiency. In addition, for the initial setting of the fusion weights, it can be adaptively initialized based on the variance and covariance matrix of the unbiased estimates, rather than a simple uniform distribution. In an alternative solution, a dynamic weight adjustment mechanism can be considered to dynamically adjust the weights according to the error feedback of each iteration, so as to better adapt to the differences in the annotation quality of different data sources. This method can further improve the robustness and adaptability of the annotation quality distribution estimation while maintaining the accuracy of the fusion estimation.
[0122] In some embodiments, the objective function is equivalently determined and expressed as:
[0123]
[0124] Wherein, represents the objective function, represents the fusion weight of the th unbiased estimate, represents the variance of the th unbiased estimate, represents the covariance between the th unbiased estimate and the th unbiased estimate.
[0125] It should be noted that in the process of fusion estimation of the annotation quality distribution of the present invention, the equivalent form of the objective function is further clarified to optimize the calculation of the fusion weights. The objective function is expressed as , wherein is the fusion weight of the th unbiased estimate, is the variance of the th unbiased estimate, is the covariance between the th and the th unbiased estimates. This form decomposes the mean square error of the fusion estimate into a combination of variance and covariance, providing a clearer mathematical basis for optimizing the fusion weights, thereby further improving the accuracy and reliability of the annotation quality distribution estimation.
[0126] Specifically, the equivalent form of the objective function involves the calculation of fusion weights, variances, and covariances. The fusion weight is a parameter that needs to be adjusted in the optimization process, and its sum is 1, that is, . The variance The uncertainty representing each unbiased estimate can be obtained by performing statistical analysis on the outputs of each kernel function. The covariance describes the correlation between different unbiased estimates and reflects the degree of information sharing between them. During the optimization process, the initial fusion weights can be allocated based on the variances of the unbiased estimates. For example, estimates with smaller variances can be assigned larger weights. The optimization objective is to minimize the objective function , and the weights are updated by calculating the gradient of the objective function with respect to the fusion weights until the preset convergence conditions are met, such as the mean squared error no longer decreasing significantly or the number of iterations reaching the upper limit. This process ensures that the fused estimate has the minimum uncertainty in a statistical sense by dynamically adjusting the fusion weights.
[0127] Preferably, to further optimize the fusion estimation process, the optimization strategy and parameter settings of the objective function can be refined. For example, constraint conditions can be introduced, such as restricting the range of the fusion weights (e.g., 0 ≤ ≤ 1), to avoid some weights being too large or too small, thereby improving the stability of the fusion result. In terms of the choice of optimization algorithm, in addition to the gradient descent method, more advanced optimization algorithms such as genetic algorithms or particle swarm optimization algorithms can also be adopted. These algorithms can better handle complex non-linear optimization problems, especially in the case of multiple local optimal solutions.
[0128] Furthermore, for the calculation of variances and covariances, a dynamic adjustment mechanism can be introduced to update these parameters based on the error feedback of each iteration, so as to better adapt to the changes in the data. In an alternative solution, a multi-objective optimization framework can be considered to simultaneously minimize the mean squared error and maximize the diversity of the fused estimates, in order to further improve the performance of the annotation quality distribution estimation.
[0129] The above embodiments of the present invention have the following beneficial effects: The present invention can effectively improve the annotation quality of the large model evaluation set in the field of natural resources through statistical inference and kernel density estimation methods. Specifically, based on the data distribution model and the annotation model, it can accurately capture the non-uniform distribution of the data and the characteristics of non-standard annotation noise, thereby providing a reliable basis for the estimation of the annotation quality distribution. By expanding and approximating the annotation quality distribution through kernel density estimation and reasonably setting the number of kernel functions, the accuracy and flexibility of the annotation quality distribution estimation can be improved, ensuring the high quality of the evaluation set.
[0130] In addition, the present invention can optimize the unbiased estimated values of each kernel function through a weighted average fusion algorithm, integrate the advantages of multiple estimation results, reduce errors, and obtain a high-quality fused estimated value of the dataset annotation quality. This method can not only improve the accuracy and reliability of the annotation quality, but also provide higher-quality data support for the evaluation of large models in the field of natural resources, meet the model evaluation requirements in complex data environments, and promote the development and application of related technologies in the field of natural resources.
[0131] As Figure 2 shown, a large model evaluation set generation device 200 in the field of natural resources according to some embodiments, the device 200 includes:
[0132] An inference module 201, configured to obtain the annotation quality distribution of a dataset through statistical inference based on a data distribution model and an annotation model, where the data distribution model includes a non-uniform data distribution, and the annotation model includes non-standard annotation noise;
[0133] An approximation module 202, configured to set the number of kernel functions of kernel density estimation to the number of data sources in the field of natural resources, and perform expansion approximation on the annotation quality distribution through kernel density estimation to determine the center and bandwidth of each kernel function;
[0134] A fusion module 203, configured to use the centers of the respective kernel functions as unbiased estimated values, and perform data fusion optimization through a weighted average fusion algorithm to obtain a fused estimated value of the dataset annotation quality.
[0135] It can be understood that the various modules described in the large model evaluation set generation device 200 in the field of natural resources correspond to the respective steps in the large model evaluation set generation method described in the reference Figure 1 description. Thus, the operations, features, and beneficial effects described above for the large model evaluation set generation method in the field of natural resources also apply to the large model evaluation set generation device 200 in the field of natural resources and the modules included therein, and will not be elaborated herein.
[0136] Next, referring to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present invention. The electronic device in some embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The terminal device shown is only an example and should not impose any limitations on the functions and usage ranges of the embodiments of the present invention.
[0137] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0138] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included. Figure 3 Each block shown in
[0139] Furthermore, the storage medium of the embodiments of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions may be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. And the foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.
[0140] The above description is only some preferred embodiments of the present invention and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present invention.
Claims
1. A method for generating a large model evaluation set in the field of natural resources, characterized in that: include: S1. Obtaining the annotation quality distribution of the data set by statistical inference based on the data distribution model and the annotation model, wherein the data distribution model includes non-uniform data distribution, and the annotation model includes non-standard annotation noise; S2. Setting the number of kernel functions of kernel density estimation to the number of data sources in the field of natural resources, and performing an expansion approximation on the annotation quality distribution through kernel density estimation to determine the center and bandwidth of each kernel function; S3, taking the center of each kernel function as an unbiased estimate, performing data fusion optimization through a weighted average fusion algorithm to obtain a fusion estimate of the annotation quality of the data set, and selecting data based on the annotation quality to form a data set to form an evaluation set; In step S1, obtaining the annotation quality distribution of the data set through statistical inference based on the data distribution model and the annotation model includes: S11. Establish data distribution model and annotation model; S12, obtaining the annotation quality distribution according to statistical inference based on the data distribution model and the annotation model; S13, performing weighted approximation on the non-uniform data distribution and the non-standard annotation noise by kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard annotation noise respectively; S14, obtaining an approximate representation of the annotation quality distribution based on the approximate representation of each non-uniform distribution and non-standard annotation noise and the relationship between each non-uniform distribution and non-standard annotation noise and the annotation quality distribution; In step S12, the annotation quality distribution is obtained according to statistical inference based on the data distribution model and the annotation model; including: Based on the annotation quality distribution and data distribution model of the previous sample, the prior distribution of the current sample is predicted; Based on the labeling model, the prior distribution of the current sample is updated to obtain the labeling quality distribution of the current sample, and the calculation method is: the product of the labeling value distribution of sample k based on the data feature through the labeling model and the prior distribution based on the labeling value of the previous sample is divided by the prior distribution based on the labeling value of the previous sample; In step S13, weighted approximation is performed on the non-uniform data distribution and the non-standard annotation noise by kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard annotation noise, respectively, including: Construct a kernel density estimation model with m kernel functions; expressed as: in, represents the kernel density estimation model, represents the kernel function, represents the center of the i-th kernel function, represents the bandwidth of the i-th kernel function; Performing weighted summation on the kernel functions of the kernel density estimation model to obtain respective approximations of the non-uniform data distribution and the non-standard annotation noise; The approximate representation of the non-uniform data distribution is: in, is an approximate representation of the non-uniform data distribution, For approximation The number of kernel functions, is the weight corresponding to the i-th kernel function in the approximate representation of non-uniform data distribution; The approximate expression of the non-standard annotation noise is: in, is an approximate representation of non-standard annotation noise, For approximation The number of kernel functions, Represents the weight corresponding to the i-th kernel function in the approximate representation of non-standard annotation noise.
2. The method for generating a large model evaluation set in the field of natural resources according to claim 1, characterized in that: In step S14, based on the approximate representations of the non-uniform distributions and the non-standard annotation noises and the relationship between the non-uniform distributions and the non-standard annotation noises and the annotation quality distribution, an approximate representation of the annotation quality distribution is obtained, including: Approximating the labeled quality distribution of the previous sample by kernel density estimation; Substituting an approximate representation of the non-uniform data distribution into the data distribution model to approximate the data distribution model; Substituting an approximate representation of the non-standard annotation noise into the annotation model to obtain an approximation of the distribution of annotation values obtained by the annotation model; Substituting the approximation of the data distribution model and the labeled quality distribution of the previous sample k - 1 into the prior distribution to obtain an approximation of the prior distribution; The approximation of the prior distribution and the approximation of the label value distribution are substituted into the posterior distribution to obtain an approximate representation of the label quality distribution.
3. The method for generating a large model evaluation set in the field of natural resources according to claim 2, characterized in that: The center of each kernel function is used as an unbiased estimate, and data fusion optimization is performed through a weighted average fusion algorithm to obtain a fusion estimate of the annotation quality of the data set, including: The fused estimated value is set to be the sum of the product of the fusion weight corresponding to each unbiased estimated value and the unbiased estimated value, and the sum of the fusion weights is 1; Determining the optimization objective function to minimize the mean square error of the fused estimate, and obtaining the initial mean square error based on the covariance matrix or cross-covariance matrix between the initial fusion weights and each unbiased estimate; Calculating the gradient of the objective function with respect to each of the fusion weights; Updating the fusion weight based on the gradient; Calculate a new mean square error based on the updated fusion weights. If the new mean square error is smaller than the mean square error before this iteration, update the mean square error to the mean square error obtained in this iteration. Otherwise, retain the mean square error and fusion weights before the update. Iteratively updating the fusion weight until a preset condition is met to obtain a final fusion weight; A final fusion estimate of the data set annotation quality is obtained based on the final fusion weight and the corresponding unbiased estimate.
4. The method for generating a large model evaluation set in the field of natural resources according to claim 3, characterized in that: The objective function is equivalently determined as: in, represents the objective function, represents the fusion weight of the i-th unbiased estimate, represents the variance of the ith unbiased estimate, represents the covariance between the ith unbiased estimator and the jth unbiased estimator.
5. A device for generating a large model evaluation set in the field of natural resources, characterized in that: include: An inference module, configured to obtain a label quality distribution of a data set by statistical inference based on a data distribution model and a labeling model, wherein the data distribution model includes a non-uniform data distribution and the labeling model includes non-standard labeling noise; An approximation module is used to set the number of kernel functions of kernel density estimation to the number of data sources in the field of natural resources, and to approximate the annotation quality distribution through kernel density estimation to determine the center and bandwidth of each kernel function; A fusion module is used to use the centers of the kernel functions as unbiased estimates, and perform data fusion optimization through a weighted average fusion algorithm to obtain a fusion estimate of the annotation quality of the data set; Obtaining the annotation quality distribution of the data set through statistical inference based on the data distribution model and the annotation model includes: S11. Establish data distribution model and annotation model; S12, obtaining the annotation quality distribution according to statistical inference based on the data distribution model and the annotation model; S13, performing weighted approximation on the non-uniform data distribution and the non-standard annotation noise by kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard annotation noise respectively; S14, obtaining an approximate representation of the annotation quality distribution based on the approximate representation of each non-uniform distribution and non-standard annotation noise and the relationship between each non-uniform distribution and non-standard annotation noise and the annotation quality distribution; In step S12, the annotation quality distribution is obtained according to statistical inference based on the data distribution model and the annotation model; including: Based on the annotation quality distribution and data distribution model of the previous sample, the prior distribution of the current sample is predicted; Based on the labeling model, the prior distribution of the current sample is updated to obtain the labeling quality distribution of the current sample, and the calculation method is: the product of the labeling value distribution of sample k based on the data feature through the labeling model and the prior distribution based on the labeling value of the previous sample is divided by the prior distribution based on the labeling value of the previous sample; In step S13, weighted approximation is performed on the non-uniform data distribution and the non-standard annotation noise by kernel density estimation to obtain approximate representations of the non-uniform data distribution and the non-standard annotation noise, respectively, including: Construct a kernel density estimation model with m kernel functions; expressed as: in, represents the kernel density estimation model, represents the kernel function, represents the center of the i-th kernel function, represents the bandwidth of the i-th kernel function; Performing weighted summation on the kernel functions of the kernel density estimation model to obtain respective approximations of the non-uniform data distribution and the non-standard annotation noise; The approximate representation of the non-uniform data distribution is: in, is an approximate representation of the non-uniform data distribution, For approximation The number of kernel functions, is the weight corresponding to the i-th kernel function in the approximate representation of non-uniform data distribution; The approximate expression of the non-standard annotation noise is: in, is an approximate representation of non-standard annotation noise, For approximation The number of kernel functions, Represents the weight corresponding to the i-th kernel function in the approximate representation of non-standard annotation noise.
6. An electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, it implements the method for generating a large model evaluation set in the field of natural resources as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method for generating a large model evaluation set in the field of natural resources as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Multi-source photoelectric detection data fusion method, device, equipment and medium
CN118885979A