Method, device and storage medium for improving repeatability of spectral prediction results

By eliminating data using the median and range thresholds and constructing a normal distribution, the problem of poor repeatability in spectral prediction results is solved, improving the reliability and efficiency of spectral prediction, reducing equipment costs, and making it suitable for solid sample analysis in the near-infrared industry.

CN119494044BActive Publication Date: 2026-04-24INTELLIGENT ANALYSIS SERVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTELLIGENT ANALYSIS SERVICE CO LTD
Filing Date
2024-10-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency, high cost, long time, and low reliability in improving the repeatability of spectral prediction results, especially with poor repeatability of spectral models under multiple measurements.

Method used

By acquiring the sampling data of the sample to be tested, data is removed using the median and range thresholds, M normal distributions are constructed, and the sampling results are determined based on the median of each normal distribution, thus achieving spectral prediction.

Benefits of technology

It improves the repeatability and reliability of spectral prediction results, reduces the need for high-precision spectral acquisition equipment, lowers R&D costs, and is applicable to near-infrared industry analysis of solid samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119494044B_ABST
    Figure CN119494044B_ABST
Patent Text Reader

Abstract

The application relates to a method and device for improving the repeatability of spectral prediction results and a storage medium, and belongs to the technical field of spectral measurement. The method comprises the following steps: obtaining sampling data of a to-be-measured sample, each sampling data comprising N sub-sample data; for each sampling data, performing data rejection on the N sub-sample data based on the median of the N sub-sample data in the sampling data and a pre-set range threshold, to obtain processed sub-sample data; constructing M normal distributions based on the processed sub-sample data based on the median; determining a sampling result of the sampling data based on the average value of the median of each normal distribution, and the sampling result is used for spectral prediction. Since the sub-sample data in any distribution form is converted into a connection composed of M normal distributions, the repeatability of the spectral prediction result can be improved, meanwhile, a spectral acquisition device with higher repeatability accuracy is not required, and the research and development cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a method, apparatus, and storage medium for improving the repeatability of spectral prediction results, belonging to the field of spectral measurement technology. Background Technology

[0002] Spectral prediction repeatability refers to the degree of consistency among spectral prediction results obtained when multiple spectral measurements are performed on the same sample under the same experimental conditions and predictions are made using a spectral model. In other words, spectral prediction repeatability measures the stability and reliability of spectral prediction results. If the spectral prediction results of the spectral model have good repeatability, then the performance of spectral analysis is better; therefore, it is necessary to improve the repeatability of spectral prediction results.

[0003] Traditional methods for improving the repeatability of spectral prediction results include the following:

[0004] 1. Increase the number of spectral acquisitions to improve the signal-to-noise ratio; however, this method will result in excessively long detection time, and the temperature effect will be very obvious after the sample has been exposed to light for a long time, which will affect the accuracy of sample measurement.

[0005] 2. Increase the number of sample collection subsamples and average multiple results; however, this will also result in excessively long testing time, and the influence of temperature effects will be very large, leading to greater differences in test results between multiple tests.

[0006] 3. Improve the spectral repeatability of the spectral acquisition equipment (i.e., improve equipment performance); however, this method has high research and development costs, low efficiency, and long time cycle, and cannot produce significant results in the short term.

[0007] 4. The spectral prediction results are filtered out using the median upper and lower limits; however, if the median deviates, even with the addition of a filtering mechanism, the repeatability of the spectral prediction results will deteriorate.

[0008] 5. Use Mahalanobis distance to eliminate spectra; although this method can effectively eliminate spectra that deviate from the spectral prediction results, the prerequisite is that the quality of the spectrum is good (i.e., high signal-to-noise ratio and good repeatability), and it is also necessary to improve the performance of the equipment to solve the problem of poor repeatability.

[0009] 6. Segmented modeling strategy, and in quantitative analysis, segmented modeling of different contents of the same substance; although it can improve the repeatability within the content range of the analyte, segmented modeling is complex to use, and if wavelength drift occurs in the equipment when the sample is located at the inflection point of different contents, the accuracy of quantitative analysis of the sample will deteriorate. Summary of the Invention

[0010] This application provides a method, apparatus, and storage medium for improving the repeatability of spectral prediction results, which can solve the problem of poor repeatability and low user reliability in multiple measurements; at the same time, it eliminates the need for spectral acquisition equipment with higher repeatability accuracy, thereby reducing research and development costs. This application provides the following technical solution:

[0011] Firstly, a method for improving the repeatability of spectral prediction results is provided, the method comprising:

[0012] Acquire sampling data of the sample to be tested, each sampling data includes N sub-sample data; where N is an integer greater than 1;

[0013] For each sampled data, the N subsamples are removed based on the median of the N subsamples and a pre-set range threshold to obtain the processed subsamples.

[0014] Based on the median, M normal distributions are constructed for the processed sample data; M is an odd number greater than or equal to 3.

[0015] The sampling result of the sampled data is determined based on the average of the median of each normal distribution, and the sampling result is used for spectral prediction.

[0016] Optionally, constructing M normal distributions for the processed sample data based on the median includes:

[0017] Based on the median and a pre-set first distribution range threshold, a first data range consisting of data located at both ends of the median is determined; wherein, the first distribution range threshold is less than the range threshold;

[0018] Within the first data range, a median normal distribution is constructed based on the smaller data volume of the subsamples at both ends of the median.

[0019] For each sub-data range that is within the processed sub-sample data and outside the first data range, determine the sub-median of the sub-sample data within the sub-data range;

[0020] A second data range is determined based on the sub-median and the second distribution range threshold corresponding to the sub-data range; the second distribution threshold is less than the range threshold, and the sum of the second distribution range threshold corresponding to each sub-data range and the first distribution range threshold is the range threshold;

[0021] Within the second data range, a bi-directional normal distribution is constructed based on the end with the smaller amount of data at both ends of the sub-median.

[0022] Optionally, the method further includes:

[0023] Based on the distribution of the processed sample data, the first distribution range threshold and the second distribution range threshold are set.

[0024] Optionally, setting the first distribution range threshold and the second distribution range threshold based on the distribution of the processed sample data includes:

[0025] If the distribution of the processed sample data is a normal distribution, then the difference between the first distribution range threshold and the range threshold is set to be less than a preset threshold, and the second distribution range threshold is set to be less than the difference.

[0026] If the distribution of the processed sample data is M-shaped, then the second distribution range threshold is set to be greater than the first distribution range threshold.

[0027] Optionally, for each sampled data, the process of removing data from the N subsamples based on the median of the N subsamples and a pre-set range threshold to obtain processed subsample data includes:

[0028] Determine the median of the N sample data;

[0029] The data filtering range is obtained by using the sum of the median and the range threshold as the upper limit and the difference between the median and the range threshold as the lower limit.

[0030] Data outside the data filtering range is removed to obtain the processed sample data.

[0031] Optionally, the range threshold is set based on the repeatability of sample subsample results within the modeling set of the spectral model.

[0032] Optionally, the value of M is 3.

[0033] Secondly, an apparatus for improving the repeatability of spectral prediction results is provided, the apparatus comprising:

[0034] The data acquisition module is used to acquire sampling data of the sample to be tested. Each sampling data includes N sub-sample data, where N is an integer greater than 1.

[0035] The data filtering module is used to remove data from N subsamples for each sampled data based on the median of N subsamples and a preset range threshold, so as to obtain processed subsample data.

[0036] The data processing module is used to construct M normal distributions for the processed sample data based on the median; where M is an odd number greater than or equal to 3.

[0037] The result determination module is used to determine the sampling result of the sampled data based on the average of the median of each normal distribution, and the sampling result is used for spectral prediction.

[0038] Thirdly, an apparatus for improving the repeatability of spectral prediction results is provided, the apparatus comprising a processor and a memory; the memory stores a program, which is loaded and executed by the processor to implement the method for improving the repeatability of spectral prediction results as described in the first aspect.

[0039] Fourthly, a computer-readable storage medium is provided, wherein a program is stored therein, the program being loaded and executed by the processor to implement the method for improving the repeatability of spectral prediction results as described in the first aspect.

[0040] The beneficial effects of this application include: acquiring sampling data of the sample to be tested, each sampling data comprising N sub-sample data; for each sampling data, removing data from the N sub-sample data based on the median and a pre-set range threshold to obtain processed sub-sample data; constructing M normal distributions for the processed sub-sample data based on the median; determining the sampling result of the sampling data based on the average of the medians of each normal distribution, and using the sampling result for spectral prediction; solving the problem of poor repeatability and low user confidence in multiple measurement results; and improving the repeatability of spectral prediction results by converting sub-sample data of any distribution form into a combination of M normal distributions, while eliminating the need for spectral acquisition equipment with higher repeatability accuracy, thus reducing R&D costs.

[0041] Furthermore, in this application, the repeatability range of different test samples can be adjusted by setting different range thresholds, making it applicable to solid samples in the near-infrared industry analysis market. Moreover, it eliminates the need to consider the degree of mixing between samples; even with multiple mixtures of the same substance, the repeatability remains controllable. The system integration cost is low, and the efficiency is higher.

[0042] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0043] Figure 1 This is a flowchart of a method for improving the repeatability of spectral prediction results provided in one embodiment of this application;

[0044] Figure 2 This is a flowchart of a method for improving the repeatability of spectral prediction results provided in another embodiment of this application;

[0045] Figure 3 This is a block diagram of an apparatus for improving the repeatability of spectral prediction results according to an embodiment of this application;

[0046] Figure 4 This is a block diagram of an apparatus for improving the repeatability of spectral prediction results provided in one embodiment of this application. Detailed Implementation

[0047] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.

[0048] In quantitative analysis used to measure solid particles, data is typically acquired through multiple sampling sessions. This involves averaging the data from multiple samplings to arrive at the final measurement. However, if the collected spectra show anomalies, the model's predicted values ​​may also become abnormal, causing data deviation. Poor repeatability from multiple samplings leads to lower data reliability and negatively impacts user experience.

[0049] Based on this, this application provides a data processing method to reduce costs and improve the repeatability of multiple data collection results, thereby enhancing users' confidence in the measurement data.

[0050] Figure 1 This is a flowchart of a method for improving the repeatability of spectral prediction results according to an embodiment of this application. The method includes at least the following steps:

[0051] Step 101: Obtain the sampling data of the sample to be tested. Each sampling data includes N subsample data.

[0052] Where N is an integer greater than 1.

[0053] Typically, a single data collection of a sample is obtained through multiple subsample collections, meaning that each data collection includes N subsamples.

[0054] For example, if the value of N is 100, then the sampling data obtained each time includes 100 subsample results corresponding to 100 spectra collected.

[0055] Step 102: For each sampled data, the N subsamples are removed based on the median of the N subsamples and a pre-set range threshold to obtain the processed subsamples.

[0056] In one example, for each sampled data, the N subsamples are filtered out based on the median of the N subsamples and a pre-set range threshold to obtain processed subsamples. This process includes: determining the median of the N subsamples; using the sum of the median and the range threshold as the upper limit and the difference between the median and the range threshold as the lower limit to obtain the data filtering range; and removing data outside the data filtering range to obtain the processed subsamples.

[0057] In other words, the data filtering range is median ± m, where m is the range threshold. Data samples within the filtering range are retained, while those outside the range are removed, resulting in processed sample data. Therefore, the range of the processed sample data is 2|m|, which is twice the absolute value of m.

[0058] In one example, the range threshold is based on the repeatability setting of sample sub-results within the modeling set of the spectral model. That is, measurements of all samples are collected from the modeling set of the spectral model, obtained based on multiple measurements under the same or similar conditions; for each sample, the range of its measurements is calculated, i.e., the difference between the maximum and minimum values. By analyzing these range values, the repeatability of the sample results is evaluated. If the ranges of most samples are small, the results have high repeatability; conversely, if the ranges are large, it may indicate poor repeatability. Based on the repeatability analysis, a range threshold can be set. This range threshold can include most sample results with high repeatability and exclude those with abnormally large ranges. The set range threshold is used to remove sample results whose ranges exceed the threshold. Doing so helps improve the accuracy and reliability of the model.

[0059] Step 103: Based on the median, construct M normal distributions for the processed sample data.

[0060] Where M is an odd number greater than or equal to 3.

[0061] Because the distribution of sample data collected multiple times can vary, such as normal distribution, random distribution, M-type distribution, and trend distribution, different sample data distributions can cause most sample data to deviate from the median. This leads to a deviation of the overall data from the median after averaging, resulting in poor repeatability of the final average result. In this embodiment, instead of directly averaging the N sample data, M normal distributions are constructed from the processed sample data. That is, sample data of any distribution form is transformed into a combination of M normal distributions. This solves the problem of different distributions of sample data collected multiple times, thereby improving the repeatability of the final average result.

[0062] In one example, based on the median, M normal distributions are constructed for the processed sample data, including steps 1031-1035:

[0063] Step 1031: Based on the median and a pre-set first distribution range threshold, determine the first data range consisting of data points located at both ends of the median. The first distribution range threshold is less than the range threshold.

[0064] Assuming a first distribution range threshold n is given, the first data range can be determined based on the median, that is, the data range consisting of the median ± n.

[0065] Step 1032: Within the first data range, construct a median normal distribution based on the smaller data volume of the two subsamples at either end of the median.

[0066] In this step, a standard normal distribution is constructed based on the side with fewer sample results within the range of the median. Excess data within the first data range is removed, thus completing the construction of the first normal distribution. At this point, the influence of excess deviation data on the average result can be eliminated.

[0067] Specifically, a normal distribution is constructed by taking the smaller sample size at both ends of the median as the benchmark. This involves taking the side with lower probability density as the benchmark and removing data with higher probability density from the other side, so that the constructed normal distribution is a standard normal distribution that is symmetrical from left to right.

[0068] Step 1033: For each sub-data range that is within the processed sub-sample data and outside the first data range, determine the sub-median of the sub-sample data within the sub-data range.

[0069] Taking a value of M of 3 as an example, each sub-data range that is located within the processed sub-sample data and outside the first data range includes: sub-data ranges located on both sides outside the first data range.

[0070] When the value of M is an odd number greater than 3, each sub-data range located within the processed subsample data and outside the first data range can be obtained by dividing the data ranges on both sides of the first data range equally according to (M-1) / 2, or by dividing according to a preset data interval. This embodiment does not limit the way the sub-data range is divided.

[0071] Step 1034: Determine the second data range based on the sub-median and the second distribution range threshold corresponding to the sub-data range; wherein the second distribution threshold is less than the range threshold, and the sum of the second distribution range threshold and the first distribution range threshold corresponding to each sub-data range is the range threshold.

[0072] Taking a value of M of 3 as an example, in this case, in addition to the first data range, there are two sub-data ranges. Assuming that the second distribution range thresholds k1 and k2 corresponding to the two sub-data ranges are given, then one second data range includes the data range obtained by sub-median ± k1 of one sub-data range; the other second data range includes the data range obtained by sub-median ± k2 of the other sub-data range. At this time, |m|>|n|+|k1|+|k2|.

[0073] When the value of M is an odd number greater than 3, different sub-data ranges correspond to a second distribution range threshold. The second distribution range thresholds corresponding to different sub-data ranges may be the same or different. This embodiment does not limit the value of the second distribution range threshold.

[0074] Step 1035: Within the second data range, construct a bi-directional normal distribution based on the end with the smaller amount of data at both ends of the sub-median.

[0075] For each second data range, the side with the lower probability density at both ends of the sub-median is used as the standard, and the data with the higher probability density on the other side are removed, so that the constructed normal distribution is a standard normal distribution that is symmetrical from left to right, thus obtaining the bi-directional normal distribution corresponding to the second data range.

[0076] In this embodiment, before step 1031, the method further includes: setting a first distribution range threshold and a second distribution range threshold based on the distribution of the processed sample data.

[0077] In one example, if the distribution of the processed sample data is normal, the difference between the first distribution range threshold and the range threshold is set to be less than a preset threshold, and the second distribution range threshold is set to be less than the difference.

[0078] In another example, if the processed sample data follows an M-shaped distribution, then the second distribution range threshold is set to be greater than the first distribution range threshold. Here, the "M" in the M-shaped distribution does not represent a numerical value; rather, it refers to a data distribution plot that resembles the letter "M." This distribution characteristic indicates that the data has two distinct peaks and a trough in the middle, forming a special form of bimodal distribution.

[0079] Step 104: Determine the sampling result of the sampled data based on the average of the median of each normal distribution. The sampling result is used for spectral prediction.

[0080] In this way, by selecting different distribution range thresholds, the range threshold for repeatability of multiple sampling results can be set. That is, within the range of |m|, if the values ​​of |k1| and |k2| are greater than |n| and the larger they are, the larger the repeatability range of multiple sampling results will be. If |k1| and |k2| are extremely small or infinitesimal, the repeatability range of multiple sampling results can be 0. This achieves controllable repeatability range.

[0081] In summary, the method for improving the repeatability of spectral prediction results provided in this embodiment acquires sampling data of the sample to be tested, with each sampling data including N sub-sample data. For each sampling data, the N sub-sample data are removed based on the median and a pre-set range threshold, resulting in processed sub-sample data. Based on the median, M normal distributions are constructed for the processed sub-sample data. The sampling result is determined based on the average of the medians of each normal distribution, and the sampling result is used for spectral prediction. This method can solve the problem of poor repeatability and low user confidence in multiple measurement results. Since sub-sample data of any distribution form is converted into a combination of M normal distributions, the repeatability of spectral prediction results can be improved. At the same time, it eliminates the need for spectral acquisition equipment with higher repeatability accuracy, thus reducing research and development costs.

[0082] Furthermore, in this application, the repeatability range of different test samples can be adjusted by setting different range thresholds, making it applicable to solid samples in the near-infrared industry analysis market. Moreover, it eliminates the need to consider the degree of mixing between samples; even with multiple mixtures of the same substance, the repeatability remains controllable. The system integration cost is low, and the efficiency is higher.

[0083] To better understand the method for improving the repeatability of spectral prediction results provided in this application, the method is illustrated below using M=3 as an example. (Refer to...) Figure 2 The method includes the following steps:

[0084] Step 21: Obtain the sampling data of the sample to be tested. Each sampling data includes N subsample data.

[0085] Step 22: Calculate the median of N samples, set the range threshold to m, and remove the sample data outside the range of ±m around the median to obtain the processed sample data.

[0086] Step 23: Construct an intermediate normal distribution based on the median value of the processed subsample data within the range of ±n, and remove subsample data that deviate significantly from the current range of ±n.

[0087] Step 24: Using the sample data outside the ±n range and within the ±m range, set k1 and k2 on the left and right respectively to construct two normal distributions to obtain two-end normal distributions, so as to remove the sample data with large deviations on the left and right sides between the ±n range and the ±m range.

[0088] Step 25: Take the average of the medians of the three normal distributions to obtain a sampling result.

[0089] Figure 3 This is a block diagram of an apparatus for improving the repeatability of spectral prediction results according to an embodiment of this application. The apparatus includes at least the following modules: a data acquisition module 310, a data filtering module 320, a data processing module 330, and a result determination module 340.

[0090] The data acquisition module 310 is used to acquire sampling data of the sample to be tested, and each sampling data includes N sub-sample data; where N is an integer greater than 1.

[0091] The data filtering module 320 is used to remove data from the N subsamples of each sampled data based on the median of the N subsamples and a preset range threshold, so as to obtain the processed subsamples.

[0092] Data processing module 330 is used to construct M normal distributions for the processed sample data based on the median; where M is an odd number greater than or equal to 3.

[0093] The result determination module 340 is used to determine the sampling result of the sampled data based on the average of the median of each normal distribution, and the sampling result is used for spectral prediction.

[0094] For relevant details, please refer to the above method implementation examples.

[0095] It should be noted that the apparatus for improving the repeatability of spectral prediction results provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus for improving the repeatability of spectral prediction results can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus for improving the repeatability of spectral prediction results provided in the above embodiments and the method embodiments for improving the repeatability of spectral prediction results belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0096] Figure 4 This is a block diagram of an apparatus for improving the repeatability of spectral prediction results according to an embodiment of this application. The apparatus includes at least a processor 401 and a memory 402.

[0097] Processor 401 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0098] Memory 402 may include one or more computer-readable storage media, which may be non-transitory. Memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 402 is used to store at least one instruction, which is executed by processor 401 to implement the method for improving the repeatability of spectral prediction results provided in the method embodiments of this application.

[0099] In some embodiments, the apparatus for improving the repeatability of spectral prediction results may also include: a peripheral device interface and at least one peripheral device. The processor 401, memory 402, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, the peripheral device includes, but is not limited to: radio frequency circuitry, a touch display screen, audio circuitry, and a power supply.

[0100] Of course, the apparatus for improving the repeatability of spectral prediction results may also include fewer or more components, and this embodiment is not limited thereto.

[0101] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the method for improving the repeatability of spectral prediction results in the above-described method embodiments.

[0102] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the method for improving the repeatability of spectral prediction results in the above-described method embodiments.

[0103] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0104] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for improving the repeatability of spectral prediction results, characterized in that, The method includes: Acquire sampling data of the sample to be tested, each sampling data includes N sub-sample data; where N is an integer greater than 1; For each sampled data, the N subsamples are removed based on the median of the N subsamples and a pre-set range threshold to obtain the processed subsamples. Based on the median, M normal distributions are constructed for the processed sample data; M is an odd number greater than or equal to 3. The sampling result of the sampled data is determined based on the average of the median of each normal distribution, and the sampling result is used for spectral prediction; wherein, constructing M normal distributions for the processed sample data based on the median includes: Based on the median and a pre-set first distribution range threshold, a first data range consisting of data located at both ends of the median is determined; wherein, the first distribution range threshold is less than the range threshold; Within the first data range, a median normal distribution is constructed based on the smaller data volume of the subsamples at both ends of the median. For each sub-data range that is within the processed sub-sample data and outside the first data range, determine the sub-median of the sub-sample data within the sub-data range; A second data range is determined based on the sub-median and the second distribution range threshold corresponding to the sub-data range; the second distribution range threshold is less than the range threshold, and the sum of the second distribution range threshold corresponding to each sub-data range and the first distribution range threshold is the range threshold; Within the second data range, a bi-directional normal distribution is constructed based on the end with the smaller amount of data at both ends of the sub-median.

2. The method according to claim 1, characterized in that, The method further includes: Based on the distribution of the processed sample data, the first distribution range threshold and the second distribution range threshold are set.

3. The method according to claim 2, characterized in that, The step of setting the first distribution range threshold and the second distribution range threshold based on the distribution of the processed sample data includes: If the distribution of the processed sample data is a normal distribution, then the difference between the first distribution range threshold and the range threshold is set to be less than a preset threshold, and the second distribution range threshold is set to be less than the difference. If the distribution of the processed sample data is M-shaped, then the second distribution range threshold is set to be greater than the first distribution range threshold.

4. The method according to any one of claims 1 to 3, characterized in that, For each sampled data point, the N subsamples are truncated based on the median of the N subsamples and a pre-set range threshold to obtain processed subsamples, including: Determine the median of the N sample data; The data filtering range is obtained by using the sum of the median and the range threshold as the upper limit and the difference between the median and the range threshold as the lower limit. Data outside the data filtering range is removed to obtain the processed sample data.

5. The method according to claim 4, characterized in that, The range threshold is set based on the repeatability of sample subsample results within the modeling set of the spectral model.

6. The method according to any one of claims 1 to 3, characterized in that, The value of M is 3.

7. An apparatus for improving the repeatability of spectral prediction results, characterized in that, The device includes: The data acquisition module is used to acquire sampling data of the sample to be tested. Each sampling data includes N sub-sample data, where N is an integer greater than 1. The data filtering module is used to remove data from N subsamples for each sampled data based on the median of N subsamples and a preset range threshold, so as to obtain processed subsample data. The data processing module is used to construct M normal distributions for the processed sample data based on the median; where M is an odd number greater than or equal to 3. The result determination module is used to determine the sampling result of the sampled data based on the average of the median of each normal distribution, and the sampling result is used for spectral prediction. The step of constructing M normal distributions for the processed sample data based on the median includes: Based on the median and a pre-set first distribution range threshold, a first data range consisting of data located at both ends of the median is determined; wherein, the first distribution range threshold is less than the range threshold; Within the first data range, a median normal distribution is constructed based on the smaller data volume of the subsamples at both ends of the median. For each sub-data range that is within the processed sub-sample data and outside the first data range, determine the sub-median of the sub-sample data within the sub-data range; A second data range is determined based on the sub-median and the second distribution range threshold corresponding to the sub-data range; the second distribution range threshold is less than the range threshold, and the sum of the second distribution range threshold corresponding to each sub-data range and the first distribution range threshold is the range threshold; Within the second data range, a bi-directional normal distribution is constructed based on the end with the smaller amount of data at both ends of the sub-median.

8. An apparatus for improving the repeatability of spectral prediction results, characterized in that, The device includes a processor and a memory; the memory stores a program that is loaded and executed by the processor to implement the method for improving the repeatability of spectral prediction results as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a program that is loaded and executed by a processor to implement the method for improving the repeatability of spectral prediction results as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Element probabilistic prediction model training method and element probabilistic prediction method

    CN114965441A