A method for improving representativeness and reliability of wind measurement data

By screening the wind measurement data multiple times and using interpolation model processing, the problem of poor wind measurement data quality in complex mountain wind farms is solved, the representativeness and reliability of the data are improved, and the completeness and accuracy of the data are ensured.

CN119669213BActive Publication Date: 2025-05-16BEIJING XIACHU TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510192490.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-16
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the quality of wind measurement data, especially in complex mountain wind farms. Due to freezing and other factors, wind measurement equipment often fails or stops testing, resulting in a large amount of unreasonable or missed data in the data, affecting the representativeness and reliability of the data.

Method used

By filtering the wind measurement data in all scene areas multiple times, unrepresentative data are eliminated, and the screened data is interpolated using a pre-trained interpolation model to improve the integrity and reliability of the data, and finally verify the interpolated data to ensure the integrity of the data during the transmission process.

Benefits of technology

The representativeness and reliability of wind measurement data are improved. Through multiple screening and interpolation processing, the calculation amount is reduced and the data is ensured, which solves the problems of missing data and unreasonable data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669213B_ABST
    Figure CN119669213B_ABST
Patent Text Reader

Abstract

The present invention provides a method for improving the representativeness and reliability of wind measurement data, and relates to the technical field of wind measurement data processing, wherein the method comprises: performing a first screening of wind measurement data of all scene areas according to historical wind measurement data, and obtaining areas where data in the same scene meets the representativeness requirements; calculating the average difference of wind measurement data of each station in the screened area, and performing a second screening of the station wind measurement data according to the average difference, and obtaining representative data that meets the standards in each area; inputting the representative data into a pre-trained interpolation model, outputting interpolated data, and verifying the interpolated data with the data to be stored at the receiving end, so that the data to be stored is complete interpolated data. This solution can make the wind measurement data more representative and reliable, and ensure the accuracy of wind resource assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wind measurement data processing, and in particular to a method for improving the representativeness and reliability of wind measurement data. Background Art

[0002] Wind measurement data mainly refers to measurement data related to wind speed and wind direction. These data play an important role in many fields such as meteorology, environmental monitoring, aviation and navigation. At the same time, these data can truly and comprehensively reflect the wind resource conditions in the area where the wind tower is located, providing a reliable basis for the design, construction and operation of wind farms.

[0003] Existing technologies usually cannot guarantee the quality of wind measurement data collected from wind towers, especially in wind farms located in complex mountainous areas. Affected by freezing and other factors, wind measurement equipment often fails or stops measuring, resulting in a large amount of unreasonable or missing data in the wind measurement data.

[0004] Based on this, there is an urgent need for a method to improve the representativeness and reliability of wind measurement data to solve the above technical problems. Summary of the invention

[0005] The embodiment of the present invention provides a method for improving the representativeness and reliability of wind measurement data, which can effectively improve the representativeness and reliability of wind measurement data.

[0006] In a first aspect, an embodiment of the present invention provides a method for improving the representativeness and reliability of wind measurement data, comprising:

[0007] Perform a first screening of the wind measurement data of all scene areas based on historical wind measurement data to obtain areas where the data in the same scene meets the representative requirements;

[0008] Calculate the average difference of wind measurement data of each station in the screened area, and perform a second screening on the station wind measurement data based on the average difference to obtain representative data that meets the standards in each area;

[0009] The representative data is input into a pre-trained interpolation model, interpolation data is output, and the interpolation data is verified with the data to be stored at the receiving end, so that the data to be stored is complete interpolation data.

[0010] In a second aspect, an embodiment of the present invention further provides a device for improving the representativeness and reliability of wind measurement data, including:

[0011] A first screening module is used to perform a first screening of the wind measurement data of all scene areas according to the historical wind measurement data to obtain areas where the data in the same scene meets the representativeness requirement;

[0012] The second screening module is used to calculate the average difference of the wind measurement data of each station in the screened area, and perform a second screening on the wind measurement data of the station according to the average difference to obtain representative data that meets the standards in each area;

[0013] The processing module is used to input the representative data into a pre-trained interpolation model, output interpolation data, and verify the interpolation data with the data to be stored at the receiving end, so that the data to be stored is complete interpolation data.

[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.

[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, enables the computer to execute the method described in any embodiment of this specification.

[0016] The embodiment of the present invention provides a method for improving the representativeness and reliability of wind measurement data. First, the wind measurement data of all scene areas are screened to obtain the area corresponding to each scene that meets the representativeness requirements; then the difference between the wind measurement data of the site in the screened area and the wind measurement data of the site to be evaluated is calculated, and the representativeness of the data of each site is measured according to the difference, and the representative data in the same area is screened for the second time; then the data screened for the second time is interpolated with a trained neural network to improve the integrity and reliability of the data; finally, the interpolated data is verified to avoid loss or tampering during data transmission. The method improves the representativeness of the wind measurement data while reducing the amount of calculation by screening the wind measurement data multiple times, and ensures the reliability and integrity of the data by interpolating the screened data and verifying the transmitted data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 is a flow chart of a method for improving representativeness and reliability of wind measurement data provided by an embodiment of the present invention;

[0019] Figure 2 is a hardware architecture diagram of an electronic device provided by an embodiment of the present invention;

[0020] Figure 3 It is a structural diagram of a device for improving the representativeness and reliability of wind measurement data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0022] As mentioned above, the prior art usually causes a large amount of unreasonable or missing data in the wind measurement data due to failures or suspension of wind measurement equipment, making it difficult to ensure the quality of the wind measurement data collected from the wind tower.

[0023] Based on this, the concept of the present invention is to first screen all wind measurement data multiple times, eliminate unrepresentative data, and then interpolate and transmit and verify the remaining data to ensure the reliability of the data.

[0024] The specific implementation of the above concept is described below.

[0025] Please refer to Figure 1 , an embodiment of the present invention provides a method for improving the representativeness and reliability of wind measurement data, the method comprising:

[0026] Step 100, performing a first screening of the wind measurement data of all scene areas according to the historical wind measurement data, and obtaining areas whose data meet the representativeness requirement under the same scene;

[0027] Step 102, calculating the average difference of wind measurement data of each station in the screened area, and performing a second screening on the wind measurement data of the stations according to the average difference to obtain representative data that meets the standards in each area;

[0028] Step 104, input the representative data into a pre-trained interpolation model, output interpolation data, and verify the interpolation data with the data to be stored at the receiving end, so that the data to be stored is complete interpolation data.

[0029] In an embodiment of the present invention, the wind measurement data of all scene areas are first screened to obtain the area corresponding to each scene that meets the representativeness requirements; then the difference between the wind measurement data of the site in the screened area and the wind measurement data of the site to be evaluated is calculated, and the representativeness of the data of each site is measured according to the difference, and the representative data in the same area is screened for the second time; then the trained neural network is used to interpolate the data screened for the second time to improve the integrity and reliability of the data; finally, the interpolated data is verified to avoid loss or tampering during data transmission. This method improves the representativeness of the wind measurement data while reducing the amount of calculation by screening the wind measurement data multiple times, and ensures the reliability and integrity of the data by interpolating the screened data and verifying the transmitted data.

[0030] Described below Figure 1 How the various steps are performed.

[0031] First, with respect to step 100, the wind measurement data of all scene areas are first screened according to the historical wind measurement data to obtain areas where the data in the same scene meets the representativeness requirement.

[0032] Wind farm wind towers are usually widely distributed. For vast plains, the data from wind towers are usually not affected by scene changes. However, for complex scenes such as high-altitude areas, the wind data at different altitudes in the same area will also vary greatly. Taking a mountain range as an example, the same mountain range can be regarded as the same large area, and the bottom, middle and top of the mountain belong to different scenes. The scene at the bottom of the mountain is divided into multiple sub-areas, and each sub-area is affected by multiple factors, resulting in the wind data not being representative of the scene. It can be seen that for complex terrain environments, directly screening data from all wind measurement data for resource assessment is not only computationally intensive, but also difficult to reflect the differences between different sub-areas.

[0033] Therefore, this embodiment adopts a multi-layer screening method to filter out data from representative areas under the same scene according to the different divided areas. For example, multiple areas with the same bottom of the mountain scene are screened, and the wind measurement data of the area that best represents the mountain scene is selected for subsequent processing. This can greatly reduce the workload of calculation and improve the efficiency of resource assessment.

[0034] It is understandable that the above process only takes the mountain scene as an example. For other complex terrain environments, the actual area division results shall prevail, and no further elaboration will be given here.

[0035] In an embodiment of the present invention, the first screening includes the following processes: correcting the wind direction data and wind speed data of all areas in the same scene; calculating the difference between the historical wind measurement data and the corrected wind measurement data in the same area, and establishing a distribution histogram with the difference as the horizontal axis and the number of times the difference is reached as the vertical axis; performing a hypothesis test on the corrected wind measurement data in the current area according to the distribution histogram, and based on a preset significance level, screening and determining that the area with the least significant data difference in the test results is the area that meets the representativeness requirements.

[0036] Specifically, the wind measurement data of all areas of the same scene are first corrected. For example, if the scene at the bottom of the mountain is divided into six areas, the wind direction data of all areas are corrected for magnetic declination, and the wind speed data are corrected for tower shadow, etc., to obtain the corrected data.

[0037] The first screening process takes wind speed data as an example, and combines the historical wind speed data of these areas to determine the most representative wind speed data for each area in the past. These historical data can be monthly wind speed data or annual wind speed data. Considering that directly judging representativeness through wind speed data is easily affected by various factors, this embodiment uses wind power density to characterize the wind energy resources of each area.

[0038] That is to say, the difference between the standard wind power density determined by the historical wind speed data and the wind power density to be selected determined by the corrected wind speed data in the same area is calculated, and a distribution histogram is established with the difference as the horizontal axis and the number of times the difference is reached as the vertical axis.

[0039] The t-test in the hypothesis test can infer whether the difference between two numbers is statistically significant. In this embodiment, the hypothesis test is used to determine the difference between the standard wind power density and the wind power density to be selected. That is to say, the distribution histogram established above shows that the difference in wind power density is normally distributed. Under ideal conditions, the wind power density to be selected should be the same as the standard wind power density, that is, the difference is 0. Therefore, by judging the difference between the mean of the actual difference and the ideal difference of 0, it is determined whether the wind measurement data in the area is similar to the historical standard data. If they are similar, it can be considered that the data in the area is representative, so that subsequent processing can be performed.

[0040] Specifically, let the null hypothesis in the hypothesis test be H 0 means the mean of the difference is equal to the ideal difference; alternative hypothesis H 1 means that the mean difference is not equal to the ideal difference, then the wind power density to be selected is available data, and it is assumed that is from a normal distribution The difference sample is calculated by the following formula to calculate the test statistic h of the sample:

[0041]

[0042]

[0043]

[0044] In the formula, is the sample mean; S 2 is the sample variance; S is the sample standard deviation; For the i Sample individuals; μ is the assumed population mean; m is the sample size.

[0045] Through degrees of freedom m –1 and significance level α Find the test statistic distribution table to get the standard test statistic. If the standard test statistic is less than the calculated test statistic, reject the null hypothesis; otherwise, do not reject the null hypothesis. α The meaning is: the probability or risk of being rejected when the null hypothesis is correct. Usually the significance level is 0.05 or 0.01, but because the actual application situation is different, the significance level should be selected according to the actual situation.

[0046] All areas corresponding to each scene are screened by the above method, and areas that meet the representative requirements in each scene are selected.

[0047] It is worth noting that the remaining data collected by the observation sites, such as temperature, humidity, etc., are not the key data considered in the embodiment of the present invention, and the screening process for such data is relatively mature. Therefore, the screening of such data can use the method provided in this embodiment, or other methods can be used, which will not be elaborated here.

[0048] Then, for step 102, the average difference of the wind measurement data of each station in the screened area is calculated, and the wind measurement data of the stations are secondly screened according to the average difference to obtain representative data that meets the standards in each area.

[0049] In the embodiment of the present invention, the second screening process includes the following steps:

[0050] Calculate the first difference between the wind measurement data of the site to be evaluated in the area and the wind measurement data of all remaining sites in the same area, and determine the average difference of the site to be evaluated based on the first difference :

[0051]

[0052] in, is the wind measurement data of the site to be evaluated at time t; is the wind measurement data of the kth station at the tth time; T is the total duration; K is the number of all stations in the same area;

[0053] The quotient of the square of the first error and the average difference and the square of the first error in an ideal state is calculated to obtain a representative measurement value P:

[0054]

[0055] in, is the first error, ; is the wind measurement data of the site to be evaluated at time t under ideal conditions; is the wind measurement data of the kth station at the tth time under ideal conditions; when it is an ideal state =0;

[0056] The site wind measurement data whose representative measurement value meets the preset standard is determined as the representative data.

[0057] Specifically, after obtaining the representative area of ​​each scene, it is necessary to evaluate and screen each observation site in the area. First, select a site as the site to be evaluated, then calculate the first difference between the remaining sites in the area and the site to be evaluated, and calculate the defined average difference according to the above formula. When the average difference is closer to 0, it means that the wind measurement data of the site to be evaluated is closer to the mean of the entire area, and the representativeness of the observation site is better.

[0058] Furthermore, in order to more intuitively compare the representativeness of each site, we assume that the first difference follows a normal distribution with a mean of 0, and then take the sum of the squares of the first difference and the average difference, divided by the sum of the squares of the first difference under the ideal state, that is, the case where the average difference is 0, as shown in the above formula, to calculate the representativeness measurement value P used to characterize the representativeness of the site to be evaluated in the area.

[0059] Finally, the value is used for comparison and screening to determine representative data that meets the preset standards.

[0060] With respect to step 104, the representative data is input into a pre-trained interpolation model, interpolation data is output, and the interpolation data is verified with the data to be stored at the receiving end, so that the data to be stored is complete interpolation data.

[0061] Since some observation sites may have a lot of missing data due to environmental factors, in order to ensure the accuracy of subsequent resource assessment based on representative data, it is necessary to interpolate the data to obtain complete data for the observation site over a long period of time.

[0062] Taking into account that the missing mode of some data is not random, but has a continuous missing feature, and multiple data are missing at one time, this embodiment proposes a data interpolation method based on random forest, including the following steps: determining the missing status of the representative data; wherein, when the representative data is missing at only one altitude layer, the representative data is interpolated on the same tower, and when the representative data is missing at multiple altitude layers, the representative data is interpolated on different towers; according to the missing status, reference data without missing data is determined, and the reference data and the representative data to be interpolated are input into a trained interpolation model to output the interpolated data.

[0063] Specifically, random forest is an integrated learning method, a classification or regression model composed of multiple decision trees. It uses the collective wisdom of multiple decision trees to improve the overall prediction accuracy by integrating the predictions of each decision tree; the basic idea of ​​random forest is to first use the self-help resampling technology to randomly extract multiple samples with replacement from the original training sample set to generate a new training sample set, and the capacity of each sample is the same as the number of samples in the original training set; then, multiple decision trees are constructed based on the self-help sample set to form a random forest; finally, based on the input sample to be classified / regressed, the random forest votes on the output results of each decision tree to determine the final prediction result.

[0064] In this embodiment, the established initial random forest model needs to be trained before interpolating the representative data, and the interpolation model is obtained by training in the following manner: the original data set without missing data is divided into a training set and a test set, and a plurality of sub-training sets are randomly selected from the training set; a classification decision tree is established for each of the sub-training sets, and the classification decision tree is split and pruned according to a preset loss function to obtain a plurality of decision models; all the decision models are screened and evaluated using the test set, and the decision model with the most accurate classification result is determined to be the interpolation model.

[0065] Specifically, the optimal interpolation method of the missing data is determined by analyzing the missing data, so that the complete data without missing data is selected from the corresponding historical data as the original data set for model training. Assuming that the original data set without missing data is M, when training the model, 90% of the data set M is selected as the training set data, and the remaining 10% is used as the test set. Then, multiple sub-training sets are extracted from the training data set, and a classification regression decision tree model is established for each sub-training set. The model includes three basic processes: splitting, pruning, and selection. In the splitting stage, the decision tree divides the input and prediction features by recursively processing the left and right child nodes, so that it can handle both continuous and discrete data. In the pruning stage, according to the complexity of the pruning cost, starting from the initial maximum decision tree, the child node with the smallest contribution of training data entropy to the overall performance is selected as the next pruning object until only the root node is left. In this way, the decision tree will generate a series of nested pruning trees. Finally, the prediction performance of each pruned tree is evaluated using an independent test set, and the K-fold cross-validation method can be used to select the optimal decision tree.

[0066] Furthermore, in the splitting and pruning phase of the training process, it is necessary to prevent overfitting by reducing the complexity of the decision tree. However, excessive simplicity of the decision tree can easily lead to large errors in the prediction results. Therefore, it is necessary to define a loss function as a constraint condition for the training process:

[0067]

[0068] In the formula, is the prediction error; For the I The number of leaf nodes in a tree; A parameter that balances the degree of fit to the training data and the complexity of the tree; For the I The complexity of the tree. When increasing from 0, we get I One of the trees The tree sequence, for each fixed parameter , there is an optimal tree sequence that minimizes the prediction error.

[0069] By building a decision tree for each sub-training set, and using the above formula as the loss function for splitting and pruning, the number of r Multiple decision models , through the following formula r The decision number model is used for voting screening and evaluation, and the interpolation model with the most accurate classification results can be obtained. :

[0070]

[0071] In the formula, is the adaptability function; Y is the attribute classification variable output by the model.

[0072] Finally, the missing data to be interpolated are input into the interpolation model to obtain complete interpolation data.

[0073] Since it is necessary to perform resource assessment based on the interpolated wind measurement data and store the interpolated data for subsequent use, it is necessary to further verify the interpolated data and the data transmitted to the user end to ensure that there is no data loss during the data transmission process, so as to make the results obtained by the evaluation based on the data more reliable.

[0074] Considering that if data loss occurs during data transmission, the system will automatically fill in a default value to replace the lost data, which makes it difficult to directly determine whether the data is lost when receiving data. Therefore, this embodiment proposes a method for converting data into an image with visual characteristics, so that subsequent processing processes can more quickly determine whether the data is lost.

[0075] In an embodiment of the present invention, the verification process includes the following steps: extracting first data information for characterizing the interpolated data and second data information for characterizing the data to be stored; wherein the data information includes: data size, data type, data value range, data default value and update time; performing data conversion on both the first data information and the second data information to obtain first image attribute information and second image attribute information; wherein the image attribute information includes image size, image format, image resolution, image noise and image saturation; based on the first image attribute information and the second image attribute information, verifying the interpolated data and the data to be stored.

[0076] Further, the first image attribute information and the second image attribute information are converted respectively to obtain a first image and a second image; it is determined whether there is a difference between the first image and the second image; if so, it is determined that the data to be stored is incomplete data and the interpolated data is retransmitted, otherwise it is determined that the data to be stored is complete data.

[0077] Specifically, a transmission time period is first specified, the wind measurement data transmitted within the time period are arranged to form a multi-order matrix, and the matrix is ​​converted into a corresponding image according to the data information. In this embodiment, the correspondence between the data information and the image attributes is as follows: the image size is used to characterize the data size, that is, the number of rows and columns of the matrix; the image format is used to characterize the data type; the image resolution is used to characterize the data value range; the image noise is used to characterize the data default value; and the image saturation is used to characterize the update time.

[0078] Generally speaking, since the amount of data default values ​​is an important factor in determining whether data is lost during data transmission, the image noise, which is the most obvious image information, is used to characterize the data default values ​​of the data to be stored after transmission. For example, the more default values ​​are supplemented after the data to be stored is lost, the more image noise is obtained by conversion. By comparing the first image converted from the interpolated data and the second image converted from the data to be stored, it can be easily found that there are more image noise in the image corresponding to the data to be stored, so as to determine that the interpolated data is different from the data to be stored, that is, the data is lost during the transmission process; further, the image resolution is used to characterize the value range of the interpolated data, for example, the smaller the size of the wind measurement data, the smaller the resolution of the corresponding converted image pixels; the image saturation is used to characterize the update time of the data, for example, the later the update time of the text file, the lower the corresponding converted image resolution and image saturation. The values ​​of the image size, image resolution and image saturation are not specifically limited here, as long as they can reflect the changes with the text size and time.

[0079] In this embodiment, by comparing the first image and the second image, it can be intuitively felt whether there is a difference between the data before and after transmission, so as to assist the detection of the data to be stored. When no difference is found between the first image and the second image after comparison, it is considered that the received data to be stored is not lost. When a difference is found between the first image and the second image after comparison, it is considered that the wind measurement data is lost during the data transmission process. At this time, the corresponding data to be stored needs to be deleted, and the wind measurement data in the corresponding time period needs to be retransmitted, and then the above operation is repeated for the newly transmitted data until there is no difference between the data before and after transmission.

[0080] like Figure 2 , Figure 3 As shown, the embodiment of the present invention provides a device for improving the representativeness and reliability of wind measurement data. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, Figure 2 As shown, it is a hardware architecture diagram of an electronic device where a device for improving the representativeness and reliability of wind measurement data provided by an embodiment of the present invention is located. Figure 2 In addition to the processor, memory, network interface, and non-volatile memory shown, the electronic device in the embodiment may also include other hardware, such as a forwarding chip responsible for processing messages, etc. Taking software implementation as an example, Figure 3As shown, as a device in a logical sense, the CPU of the electronic device in which it is located reads the corresponding computer program in the non-volatile memory into the internal memory and runs it. This embodiment provides a device for improving the representativeness and reliability of wind measurement data, including:

[0081] The first screening module 300 is used to perform a first screening of the wind measurement data of all scene areas according to the historical wind measurement data, and obtain the areas where the data in the same scene meets the representativeness requirement;

[0082] The second screening module 302 is used to calculate the average difference of the wind measurement data of each station in the screened area, and perform a second screening on the wind measurement data of the station according to the average difference to obtain representative data that meets the standards in each area;

[0083] The processing module 304 is used to input the representative data into a pre-trained interpolation model, output interpolation data, and verify the interpolation data with the data to be stored at the receiving end, so that the data to be stored is complete interpolation data.

[0084] In an embodiment of the present invention, when the first screening module 300 performs a first screening of the wind measurement data of all scene areas according to the historical wind measurement data to obtain the areas where the data in the same scene meets the representative requirements, it is specifically used to perform the following operations: correct the wind direction data and wind speed data of all areas in the same scene; calculate the difference between the historical wind measurement data and the corrected wind measurement data in the same area, and establish a distribution histogram with the difference as the horizontal axis and the number of times the difference is reached as the vertical axis; perform a hypothesis test on the corrected wind measurement data in the current area according to the distribution histogram, and according to a preset significance level, screen and determine that the area with the least significant data difference in the test results is the area that meets the representative requirements.

[0085] In the embodiment of the present invention, the second screening module 302 is specifically used to perform the following operations when the average difference of the wind measurement data of each station in the area obtained by the calculation screening is performed, and the wind measurement data of the station is secondly screened according to the average difference to obtain representative data that meets the standards in each area: calculate the first difference between the wind measurement data of the station to be evaluated in the area and the wind measurement data of all the remaining stations in the same area, and determine the average difference of the station to be evaluated according to the first difference :

[0086]

[0087] in, is the wind measurement data of the site to be evaluated at time t; is the wind measurement data of the kth station at the tth time; T is the total duration; K is the number of all stations in the same area;

[0088] The quotient of the square of the first error and the average difference and the square of the first error in an ideal state is calculated to obtain a representative measurement value P:

[0089]

[0090] in, is the first error, ; is the wind measurement data of the site to be evaluated at time t under ideal conditions; is the wind measurement data of the kth station at the tth time under ideal conditions; when it is an ideal state =0;

[0091] The site wind measurement data whose representative measurement value meets the preset standard is determined as the representative data.

[0092] In an embodiment of the present invention, when the processing module 304 inputs the representative data into a pre-trained interpolation model and outputs the interpolated data, it is specifically used to perform the following operations: determine the missing status of the representative data; when the representative data is missing at only one altitude layer, perform same-tower interpolation on the representative data; when the representative data is missing at multiple altitude layers, perform different-tower interpolation on the representative data; determine reference data without missing data based on the missing status, input the reference data and the representative data to be interpolated into the trained interpolation model, and output the interpolated data.

[0093] In an embodiment of the present invention, the interpolation model is trained in the following manner: an original data set without missing data is divided into a training set and a test set, and a plurality of sub-training sets are randomly selected from the training set; a classification decision tree is established for each of the sub-training sets, and the classification decision tree is split and pruned according to a preset loss function to obtain a plurality of decision models; all the decision models are screened and evaluated using the test set to determine that the decision model with the most accurate classification result is the interpolation model.

[0094] In the embodiment of the present invention, when the processing module 304 verifies the interpolated data and the data to be stored at the receiving end so that the data to be stored is complete interpolated data, it is specifically used to perform the following operations: extracting first data information used to characterize the interpolated data and second data information used to characterize the data to be stored; wherein the data information includes: data size, data type, data value range, data default value and update time; performing data conversion on both the first data information and the second data information to obtain first image attribute information and second image attribute information; wherein the image attribute information includes image size, image format, image resolution, image noise and image saturation; based on the first image attribute information and the second image attribute information, verifying the interpolated data and the data to be stored.

[0095] In an embodiment of the present invention, when the processing module 304 verifies the interpolated data and the data to be stored based on the first image attribute information and the second image attribute information, it is specifically used to perform the following operations: convert the first image attribute information and the second image attribute information respectively to obtain a first image and a second image; determine whether there is a difference between the first image and the second image; if so, determine that the data to be stored is incomplete data and retransmit the interpolated data; otherwise, determine that the data to be stored is complete data.

[0096] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on a device for improving the representativeness and reliability of wind measurement data. In other embodiments of the present invention, a device for improving the representativeness and reliability of wind measurement data may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0097] The information interaction, execution process and other contents between the modules in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention. For the specific contents, please refer to the description in the embodiment of the method of the present invention, and no further description is given here.

[0098] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, a method for improving the representativeness and reliability of wind measurement data in any embodiment of the present invention is implemented.

[0099] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor executes a method for improving the representativeness and reliability of wind measurement data in any embodiment of the present invention.

[0100] Specifically, a system or device equipped with a storage medium can be provided, on which software program code that implements the functions of any of the above-mentioned embodiments is stored, and a computer (or CPU or MPU) of the system or device can be enabled to read and execute the program code stored in the storage medium.

[0101] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0102] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0103] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0104] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or to a memory provided in an expansion module connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or expansion module is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0105] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical factors in the process, method, article or device including the elements.

[0106] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various media that can store program codes.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for improving the representativeness and reliability of wind measurement data, characterized in that: include: Perform a first screening of the wind measurement data of all scene areas based on historical wind measurement data to obtain areas where the data in the same scene meets the representative requirements; Calculate the average difference of wind measurement data of each station in the screened area, and perform a second screening on the station wind measurement data based on the average difference to obtain representative data that meets the standards in each area; Inputting the representative data into a pre-trained interpolation model, outputting interpolated data, and verifying the interpolated data with the data to be stored at the receiving end, so that the data to be stored is complete interpolated data; The step of inputting the representative data into a pre-trained interpolation model and outputting interpolation data comprises: Determine the missing status of the representative data; when the representative data is missing at only one altitude layer, interpolate the representative data on the same tower; when the representative data is missing at multiple altitude layers, interpolate the representative data on different towers; Determining reference data without data missing according to the missing data situation, and inputting the reference data and representative data to be interpolated into a trained interpolation model, and outputting the interpolated data; The interpolation model is trained in the following way: An original data set without missing data is divided into a training set and a test set, and a plurality of sub-training sets are randomly selected from the training set; A classification decision tree is established for each of the sub-training sets, and the classification decision tree is split and pruned according to a preset loss function to obtain multiple decision models; The loss function is: In the formula, is the prediction error; For the I The number of leaf nodes in a tree; A parameter that balances the degree of fit to the training data and the complexity of the tree; For the I The complexity of the tree; Using the test set to screen and evaluate all the decision models, and determining that the decision model with the most accurate classification result is the interpolation model; The interpolation model is: In the formula, is the adaptability function; Y is the attribute classification variable output by the model; r is the total number of decision models; is the i-th decision model among r decision models; The verifying the interpolated data and the data to be stored at the receiving end includes: The wind measurement data transmitted within the target transmission time period are arranged to form a multi-order matrix, and the matrix is ​​converted into a corresponding image according to the data information, where the corresponding relationship between the data information and the image attributes is: the image size is used to characterize the data size, that is, the number of rows and columns of the matrix; the image format is used to characterize the data type; the image resolution is used to characterize the data value range; the image noise is used to characterize the data default value; the image saturation is used to characterize the update time; Extracting first data information for characterizing the interpolated data and second data information for characterizing the data to be stored; wherein the data information includes: data size, data type, data value range, data default value and update time; Performing data conversion on the first data information and the second data information to obtain first image attribute information and second image attribute information; wherein the image attribute information includes image size, image format, image resolution, image noise, and image saturation; Based on the first image attribute information and the second image attribute information, verifying the interpolated data and the data to be stored; Converting the first image attribute information and the second image attribute information respectively to obtain a first image and a second image; Determine whether there is a difference between the first image and the second image; if so, determine that the data to be stored is incomplete data and retransmit the interpolated data; otherwise, determine that the data to be stored is complete data.

2. The method according to claim 1, characterized in that The wind measurement data of all scene areas are first screened according to the historical wind measurement data to obtain areas where the data in the same scene meets the representative requirements, including: The wind direction data and wind speed data of all areas in the same scene are corrected; Calculate the difference between the historical wind measurement data and the corrected wind measurement data in the same area, and establish a distribution histogram with the difference as the horizontal axis and the number of times the difference is reached as the vertical axis; A hypothesis test is performed on the corrected wind measurement data in the current area according to the distribution histogram, and according to a preset significance level, the area with the least significant data difference in the test result is screened and determined as the area that meets the representativeness requirement.

3. The method according to claim 1, characterized in that The average difference of wind measurement data of each station in the region obtained by the calculation and screening is then subjected to a second screening of the wind measurement data of the station according to the average difference to obtain representative data that meets the standards in each region, including: Calculate the first difference between the wind measurement data of the site to be evaluated in the area and the wind measurement data of all remaining sites in the same area, and determine the average difference of the site to be evaluated based on the first difference : in, is the wind measurement data of the site to be evaluated at time t; is the wind measurement data of the kth station at the tth time; T is the total duration; K t is the number of sites in the same area at time t; The quotient of the square of the first error and the average difference and the square of the first error in the ideal state is calculated to obtain a representative measurement value P: in, is the first error, ; is the wind measurement data of the site to be evaluated at time t under ideal conditions; is the wind measurement data of the kth station at the tth time under ideal conditions; when it is an ideal state =0; The site wind measurement data whose representative measurement value meets the preset standard is determined as the representative data.

4. A device for improving the representativeness and reliability of wind measurement data, characterized in that: include: A first screening module is used to perform a first screening of the wind measurement data of all scene areas according to the historical wind measurement data to obtain areas where the data in the same scene meets the representativeness requirement; The second screening module is used to calculate the average difference of the wind measurement data of each station in the screened area, and perform a second screening on the wind measurement data of the station according to the average difference to obtain representative data that meets the standards in each area; The processing module is used to input the representative data into a pre-trained interpolation model, output interpolation data, and verify the interpolation data with the data to be stored at the receiving end, so that the data to be stored is complete interpolation data.

5. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 3 is implemented.

6. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Anemometer tower data processing method and device

    CN110427357A

  • Complex terrain wind power plant extreme gale wind speed prediction method based on CFD

    CN118228621A