Method and system for determining temperature representative value
Through the combination of adaptive dynamic time regularization and mean European distance, combined with agglomeration hierarchical clustering and generalized extreme value distribution model, the problem of inaccurate temperature effect representative values and low efficiency of abnormal detection of time-course temperature data in the existing technology is solved, accurate temperature seasonal division and efficient abnormal detection are achieved, and the accuracy of the representative values of temperature effect are improved.
Patent Information
- Application Number
- CN202510724456.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
It is difficult for the prior art to accurately obtain representative values of temperature action, and traditional methods are inefficient in detecting abnormal time-range temperature data, making it difficult to effectively divide typical temperature seasons.
The adaptive dynamic time regularization algorithm (Adaptive DTW) and mean Euclidean distance were used to calculate the differential distance matrix of temperature change characteristics, combined with the agglomeration hierarchical clustering analysis to divide the seasonal patterns of temperature changes, typical summer and winter temperature months were determined, and the generalized extreme value distribution model and transcendence probability were used to calculate the representative value of temperature action.
Accurately divide typical temperature seasons and efficiently detect abnormal temperature data, improve the accuracy of the representative value of temperature action, and ensure the safety and durability of the structure under different temperature conditions.
Smart Images

Figure CN120234571A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of statistics, and in particular to a method and system for determining a temperature representative value. Background Art
[0002] In the time-history analysis methods of temperature data, most of them still use the more subjective seasonal month division for the division of typical seasons. However, there are significant differences in temperature characteristics and representative seasons in different regions, and the existing seasonal division methods often lack effectiveness and rationality. In addition, in terms of anomaly detection of time-history temperature data, traditional methods are still mainly based on probability elimination modes such as extreme value elimination and box plots. For large-scale time-history temperature data with implicit trends, these methods are difficult to efficiently extract potential features from the data. In addition, most of the research on representative values of temperature effects on steel-concrete composite structures is still based on the fitting of the extreme value distribution model of the recurrence period. There is a lack of reasonable division of typical temperature seasons and reliable data outlier detection. Studying representative values of temperature effects helps to accurately predict and evaluate the stress conditions of steel-concrete composite structures under different temperature conditions, thereby ensuring the safety and durability of the structure.
[0003] In summary, how to accurately obtain the representative value of temperature effect is an important issue that needs to be solved urgently. Summary of the invention
[0004] The embodiments of the present invention provide a method and system for determining a temperature representative value, which can solve the problem of inaccurate temperature action representative values obtained in the prior art.
[0005] An embodiment of the present invention provides a method for determining a temperature representative value, comprising the following steps: The abnormal temperature data of the steel-concrete composite structure within one year is obtained, and the difference distance matrix of the temperature change characteristics of different months is obtained by using the adaptive dynamic time warping algorithm Adaptive DTW and the mean Euclidean distance. The difference distance matrix of Adaptive DTW and the difference distance matrix of mean Euclidean distance are weighted to obtain a comprehensive distance matrix. Agglomerative hierarchical cluster analysis was performed on the comprehensive distance matrix to classify the months with similar temperature change characteristics within a year into the same category, obtain the seasonal pattern of temperature change, and determine the typical summer temperature month and the typical winter temperature month; Temperature measurement points on the steel-concrete composite structure are selected, and the daily extreme temperature differences of the measurement points are obtained in typical summer temperature months and typical winter temperature months. The generalized extreme value distribution model is used to obtain the probability distribution of the daily extreme temperature differences. According to the probability distribution, the exceedance probability is used to obtain the representative values of temperature effects with different recurrence periods.
[0006] Further, the specific steps for obtaining the abnormal temperature data of the steel-concrete composite structure within one year include: identifying the data with sharp changes in the annual temperature data time series of the steel-concrete composite structure using the gradient detection method; wherein, the annual temperature data time series is used to represent the temperature change data at different time points within one year; screening out the temperature values that deviate abnormally from the normal range from the annual temperature data time series using the box plot detection method; identifying the outliers from the annual temperature data time series using the Isolation Forest algorithm; and obtaining the abnormal temperature data within one year according to the identification and screening results.
[0007] Further, for performing agglomerative hierarchical clustering analysis on the comprehensive distance matrix and dividing the months with similar temperature change characteristics within one year into the same category to obtain the seasonal pattern of temperature change, the specific steps include: Performing fitting replacement on the abnormal temperature data within one year using the adaptive interpolation algorithm, and performing smoothing processing on the replacement result using Gaussian filtering; Classifying the smoothing processing result by month, using Dynamic Time Warping (DTW) to pair the 12 months pairwise, and obtaining the similarity between the temperature sequences of the 12 months pair by pair to generate a symmetric DTW distance matrix DTW _ dist : ; Among them, n represents the largest month 12 of the year; Obtaining the mean value of the data for each month, and obtaining the Euclidean distance matrix based on the mean value ; Normalizing the DTW distance matrix and the Euclidean distance matrix, and weighting the normalized result: ; Among them, represents the weight; For the distance matrix Performing clustering using the agglomerative hierarchical clustering method, dividing the 12 months into k clusters, and corresponding each cluster to a seasonal pattern and taking this seasonal pattern as a temperature change pattern.
[0008] Further, for obtaining the probability distribution of the daily extreme temperature difference using the Generalized Extreme Value Distribution model, the specific steps include: Taking the daily temperature difference extreme value between two detection points as a random variable x , and performing probability fitting using the Generalized Extreme Value Distribution , the formula is: ; Among them, is the location parameter, controlling the center of the data; is the scale parameter, which controls the width or variability of the data; is the shape parameter, which determines the tail behavior of the distribution; Cumulative distribution function , and the formula is: ; Determine the distribution parameters by maximum likelihood estimation, and determine the probability distribution of the daily extreme temperature difference according to the distribution parameters.
[0009] Furthermore, the use of the exceedance probability to obtain the representative values of temperature actions for different return periods specifically includes the following steps: According to the probability distribution of the daily extreme temperature difference for different return periods, based on a 50-year return period, use the exceedance probability to obtain the corresponding representative values of temperature actions. The formula is: ; ; ; ; ; Among them, X is the extreme value of the daily temperature difference, is the typical number of days in summer during the 50-year return period; is the typical number of days in winter during the 50-year return period; is the random variable of the summer daily temperature difference; is the random variable of the winter daily temperature difference; is the exceedance probability with a return period of 50 years. During the 100-year design reference period, for extreme temperature events with a return period of 50 years, the expected number of times exceeding the temperature threshold is 2 times, = 0.02; is the shape parameter in summer; is the shape parameter in winter; is the scale parameter in summer; is the scale parameter in winter; is the location parameter in summer; is the location parameter in winter; Inverse solve the extreme value of the temperature difference , that is, the representative value of the temperature action.
[0010] The embodiment of the present invention provides a system for determining the representative value of temperature, including: A data acquisition module, which is used to acquire the abnormal temperature data of a steel-concrete composite structure within one year, and respectively obtain the difference distance matrices of the temperature change characteristics of different months by using the Adaptive DTW (Adaptive Dynamic Time Warping) algorithm and the mean Euclidean distance, and perform weighted processing on the difference distance matrix of Adaptive DTW and the difference distance matrix of the mean Euclidean distance to obtain a comprehensive distance matrix; A typical seasonal temperature month acquisition module, which is used to perform agglomerative hierarchical clustering analysis on the comprehensive distance matrix, divide the months with similar temperature change characteristics within one year into the same category, obtain the seasonal pattern of temperature change, and determine the typical summer temperature months and the typical winter temperature months; A temperature action representative value acquisition module, which is used to select the temperature measurement points on the steel-concrete composite structure, obtain the daily extreme temperature difference of the measurement points in the typical summer temperature months and the typical winter temperature months, and use the generalized extreme value distribution model for the daily extreme temperature difference to obtain the probability distribution of the daily extreme temperature difference; according to the probability distribution, use the exceedance probability to obtain the temperature action representative values of different return periods.
[0011] The embodiment of the present invention provides a method and a system for determining a temperature representative value. Compared with the prior art, its beneficial effects are as follows: Use the Adaptive DTW (Adaptive Dynamic Time Warping) algorithm and the mean Euclidean distance to respectively obtain the difference distance matrices of the temperature change characteristics of different months, and perform agglomerative hierarchical clustering analysis on the temperature change characteristics, divide the months with similar temperature change characteristics into the same category, obtain the seasonal pattern of temperature change, determine the typical summer temperature months and the typical winter temperature months, reasonably divide the typical temperature seasons, and finally obtain the temperature action representative value according to the reasonable division result, achieving the purpose of accurately obtaining the temperature action representative value. Description of the Drawings
[0012] Figure 1 It is a calculation flow chart of a method for determining a temperature representative value provided by an embodiment of the present invention; Figure 2 It is an isolation forest algorithm flow chart of a method for determining a temperature representative value provided by an embodiment of the present invention; Figure 3 It is an adaptive interpolation algorithm flow chart of an interpolation fitting and Gaussian filtering processing flow chart of a method for determining a temperature representative value provided by an embodiment of the present invention; Figure 4 It is a local Gaussian smoothing algorithm processing flow chart of an interpolation fitting and Gaussian filtering processing flow chart of a method for determining a temperature representative value provided by an embodiment of the present invention; Figure 5 It is an agglomerative hierarchical clustering algorithm flow chart of Adaptive DTW and Euclidean distance of a method for determining a temperature representative value provided by an embodiment of the present invention; Figure 6 A dendrogram of hierarchical clustering for the classification demonstration diagram of randomly generated data of a method for determining a temperature representative value provided by an embodiment of the present invention; Figure 7 The hierarchical clustering results of adaptive DTW and Euclidean distance for the classification demonstration diagram of randomly generated data of a method for determining a temperature representative value provided by an embodiment of the present invention. Detailed implementation manners
[0013] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific implementation manners of the present invention will be given in conjunction with the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0014] Refer to Figure 1 , an embodiment of the present invention provides a method for determining a temperature representative value, including the following steps: Step 1: Obtain the abnormal temperature data of the steel-concrete composite structure within one year, use the Adaptive Dynamic Time Warping (Adaptive DTW) algorithm and the mean Euclidean distance to obtain the difference distance matrices of the temperature change characteristics of different months respectively, and perform weighted processing on the difference distance matrix of Adaptive DTW and the difference distance matrix of the mean Euclidean distance to obtain a comprehensive distance matrix.
[0015] Step 2: Perform agglomerative hierarchical clustering analysis on the comprehensive distance matrix, divide the months with similar temperature change characteristics within one year into the same category to obtain the seasonal pattern of temperature change, and determine the typical summer temperature months and typical winter temperature months. Select the temperature measurement points on the steel-concrete composite structure, obtain the daily extreme temperature difference of the measurement points in the typical summer temperature months and typical winter temperature months, and use the generalized extreme value distribution model for the daily extreme temperature difference to obtain the probability distribution of the daily extreme temperature difference; according to the probability distribution, use the exceedance probability to obtain the temperature action representative values of different return periods.
[0016] In view of the deficiencies in the prior art, the present invention proposes an improved method, which uses Adaptive DTW and the mean Euclidean distance + agglomerative hierarchical clustering algorithm to classify the months with similar temperature difference change characteristics, so as to determine the typical summer and typical winter temperature months in the annual temperature time history change. This method can effectively improve the data analysis efficiency, has good applicability in the analysis of unequal long time series data, and reduces the influence of regional temperature changes on the analysis results.
[0017] In terms of detecting outliers in temperature data, the present invention combines multiple methods such as gradient detection, box plot detection, and isolation forest algorithm to comprehensively detect abnormal temperature data. By analyzing the time-course temperature change trend, an adaptive interpolation algorithm is used to fit and replace the outliers, and finally, a local Gaussian smoothing algorithm is used to eliminate the influence of short-term fluctuations to improve the reliability and accuracy of the data.
[0018] In addition, based on the generalized extreme value distribution model (GEV), the present invention conducts a probability distribution analysis on the measured daily extreme temperature, and thus proposes a calculation method for the representative value of temperature action based on the recurrence period. This method can more accurately reflect the distribution characteristics of temperature extremes, and further provide a reliable basis for the calculation of the representative value of temperature action.
[0019] To achieve the above invention objectives, the technical solutions adopted by the present invention include the following steps: S1: Batch read temperature data.
[0020] In this step, through the matlab program, the time-course of temperature data from different sensors or temperature measurement points is read. Through standardization processing, the data format is ensured to be consistent to prepare for subsequent processing.
[0021] S2: Detect and evaluate abnormal temperatures based on gradient detection and isolation forest algorithm.
[0022] In this step, outliers in temperature data are detected through gradient detection, box plot detection, and isolation forest algorithm. The gradient detection method is used to identify the sharply changing parts of temperature data, the box plot detection is used to screen out temperature values that deviate abnormally from the normal range, and the isolation forest algorithm is used to efficiently detect outliers in high-dimensional data. By combining multiple detection methods, the accuracy and robustness of outlier detection are improved. After detection, the proportion of each type of abnormal data is evaluated to analyze the data stability and reliability of the sensor measurement points. As Figure 2 shown.
[0023] S3: Interpolate and fit abnormal temperature data and perform Gaussian filtering processing.
[0024] The detected abnormal data is fitted and replaced by an adaptive interpolation algorithm. The adaptive interpolation algorithm uses a regional marking method to identify long and short NaN segments (data anomalies) and process them differently. Small missing segments (less than 10 data points) are completed by cubic interpolation, while long missing segments are deleted to avoid the accumulation of interpolation errors and ensure interpolation accuracy. Cubic interpolation can calculate reasonable temperature values based on the local trend of the data to avoid the interference of data fluctuations on the overall analysis results. Subsequently, Gaussian filtering is used to smooth the data. The local Gaussian smoothing algorithm only smoothes short-term fluctuating data to avoid information loss caused by over-smoothing, ensuring that the long-term trend information of the time series is not affected, thereby improving the stability and reliability of the data. This step helps to ensure the accuracy of the final temperature data so that it conforms to the actual temperature change law. Figure 3 The figure shows the flow chart of the adaptive interpolation algorithm. Figure 4 This is the flowchart of the local Gaussian smoothing algorithm.
[0025] S4: Typical temperature seasons are divided based on Adaptive DTW+mean Euclidean distance+hierarchical clustering algorithm.
[0026] Adaptive DTW+mean Euclidean distance+agglomerative hierarchical clustering algorithm is used to perform cluster analysis on annual temperature data. Adaptive dynamic time warping Adaptive DTW and mean Euclidean distance are used to obtain the difference distance matrix of temperature change characteristics in different months. The weight parameter is combined with Adaptive DTW and Euclidean distance to calculate the comprehensive distance matrix. Agglomerative hierarchical clustering analysis is performed based on the comprehensive distance matrix to divide the months with similar temperature difference change characteristics into the same category. The hierarchical clustering algorithm obtains multiple temperature change patterns by continuously merging similar data points or dividing different clusters. By analyzing various clusters, the typical summer temperature month and the typical winter temperature month are determined, avoiding the subjectivity in traditional methods and achieving accurate clustering of unequal time series. Adaptive DTW calculates the nonlinear alignment distance between time series, effectively solving the problem of time axis misalignment, and significantly reducing the computational complexity of traditional DTW, while Euclidean distance provides computational efficiency. After normalization, distance weighted fusion is performed based on adjustable weight parameters to obtain optimized clustering effects and reduce the impact that may be caused by regional temperature changes. Figure 5 shown.
[0027] S5: Probability distribution analysis of daily extreme temperature difference.
[0028] By statistically analyzing the daily extreme temperature difference data, fitting its probability distribution, and determining the distribution type of the extreme values. The Generalized Extreme Value (GEV) distribution model is used to fit the temperature extreme data, study the variation law of the daily temperature difference under different return periods, and calculate the occurrence probability of the extreme temperature difference. This analysis helps predict the occurrence frequency of different extreme temperature events and provides data support for the calculation of the representative values of subsequent temperature effects.
[0029] S6: Calculation of the representative value of temperature effect based on the return period and the exceedance probability.
[0030] Based on the probability distribution results in S5 and combined with the Generalized Extreme Value distribution model, the representative values of temperature effects under different return periods are calculated. These representative values reflect the extreme temperature events that may occur within different return periods and provide a basis for evaluating the impact of temperature on structures. By calculating the exceedance probability, the occurrence risk of extreme temperature events can be further quantified, and thus provide a scientific basis for predicting the impact of temperature changes on structural health.
[0031] During the long-term on-site measurement of structural temperature, the detection equipment may be interfered by factors such as construction and environmental changes, resulting in abnormal fluctuations in temperature data. To accurately identify these abnormal data, the present invention adopts the gradient detection and Isolation Forest algorithm in step S2, and these two methods can effectively detect and eliminate abnormal data.
[0032] Gradient detection identifies abnormal fluctuations by calculating the change rate (i.e., local slope) between adjacent data points. The basic idea of this method is that if the change amplitude between two adjacent data points is too large, it may be due to abnormal data or sudden environmental changes. By calculating the difference of data points, abnormal fluctuations can be captured and marked for processing.
[0033] Specifically, the steps of gradient detection are as follows: Calculate the change amplitude between each pair of adjacent data points, that is, the difference value.
[0034] If the difference value exceeds the preset threshold, it is considered that the data corresponding to this point has an abnormal change.
[0035] Replace the data values of these abnormal points with NaN for subsequent processing.
[0036] In gradient detection, the local slope (i.e., the change rate of data) can be expressed by calculating the difference of adjacent data points: (1).
[0037] Where, and are data points in the time series, is the time step, that is, the time difference between adjacent data points.
[0038] Then, for each local slope, determine whether it exceeds the set slope threshold.
[0039] (2).
[0040] The beneficial effects of the above solution are as follows: This method can quickly and effectively identify drastic changes in temperature data, thereby timely marking and eliminating these abnormal data to ensure data accuracy. The gradient detection method can quickly identify abnormal data with overly drastic temperature changes, and is easy to operate, suitable for preliminary anomaly detection in real-time monitoring.
[0041] Isolation Forest is an unsupervised machine learning algorithm widely used in anomaly detection in high-dimensional data. This algorithm makes judgments based on the "isolation" between data points. Abnormal data points are usually easier to isolate than normal data points. By constructing multiple random trees, the isolation tree randomly selects features and splitting values to divide the data into two subsets. For normal data points, they are usually difficult to isolate, while abnormal data points, due to their significantly different properties, are easily isolated during the tree construction process.
[0042] In the implementation of the present invention, the Isolation Forest scores the "isolation degree" of each data point through multiple isolation trees, and points with higher scores are considered outliers.
[0043] The beneficial effects of the above solution are as follows: This algorithm is efficient and robust, and is particularly suitable for anomaly detection of large-scale data.
[0044] Data preset: First, extract temperature data from the data table, and then set the parameters of the Isolation Forest, including the number of trees and the subsample size of each tree.
[0045] Construct isolation trees: For each tree, the algorithm randomly extracts subsamples from the dataset, selects a feature for splitting until the depth of the tree reaches the preset maximum depth or the number of split data points is less than the specified sample size.
[0046] Calculate the path length of each data point: Calculate the path length of each data point through each tree. The path length represents the number of steps required to reach the leaf node from the root node. For outliers, due to their properties being significantly different from most data, they are usually easier to isolate, so their path lengths are shorter.
[0047] (3).
[0048] Where: is the data point Path length in the tree ; is the number of subsamples of the tree.
[0049] Calculate the anomaly score: The anomaly score for each data point is obtained by integrating its path lengths in all isolation trees. Data points with lower scores are considered anomalies, while those with higher scores are normal points.
[0050] (4).
[0051] Among them, is a constant related to the number of samples: (5).
[0052] Among them, is the number of subsamples in the dataset; is the Euler's constant.
[0053] Determine the anomalies: By setting a threshold (e.g., 95th percentile), mark the data points with higher anomaly scores as outliers.
[0054] if s ( x ) > threshold (6), mark the point as an outlier.
[0055] Replace the outliers: Replace the detected outliers with NaN (missing values) for subsequent processing. Finally, update the data table and return the processed temperature data.
[0056] In data analysis, directly removing data after anomaly detection may have a negative impact on time series data with an inherent trend. Especially when the outliers are just short-term fluctuations or measurement errors, it may destroy the potential information in the data. In step S3, the adaptive interpolation algorithm can calculate reasonable temperature values based on the local trend of the data, avoiding interference from data fluctuations on the overall analysis results. By using the adaptive interpolation algorithm, the trend of the data can be effectively restored without disturbing the overall change pattern of the data. This interpolation method calculates reasonable values based on the local trend of the data, thus avoiding the information loss caused by directly removing outliers.
[0057] The beneficial effects of the above solution are as follows: The adaptive interpolation algorithm distinguishes short-term and long-term missing data through the region marking method and adopts different processing strategies. For short-term missing data (less than 10 consecutive missing points), Hermite interpolation is used to complete the data to maintain the continuity of the data; while for long-term missing data (10 or more consecutive missing points), NaN is deleted to prevent the accumulation of interpolation errors from affecting the authenticity of the data. In addition, by de-duplicating and sorting the time data, the stability of the interpolation calculation is ensured, and the adaptability of the interpolation method is improved. Hermite interpolation is a method of interpolation based on given data points and local derivative information. Different from traditional linear interpolation or Lagrange interpolation methods, Hermite interpolation not only considers the positions of the data points but also utilizes the derivative information of the data points, making the interpolation curve smoother, especially suitable for processing time series data containing noise or outliers.
[0058] (7).
[0059] Where: is the cubic Hermite interpolation function. is the normalized parameter, and are the abscissas of the known data points. and are the function values corresponding to the data points. and are the derivative (tangent slope) values corresponding to the data points.
[0060] To further improve the reliability and stability of the data after interpolation fitting, Gaussian filtering is used to optimize the temperature data. Gaussian filtering is a technique widely used in signal processing and data smoothing. It smooths the data through weighted averaging and can effectively remove the noise in the data. The local Gaussian smoothing algorithm adopts a dynamic window calculation method and adaptively adjusts the smoothing parameter according to the change of the short-term time interval. For the short-term fluctuation problem, the algorithm calculates the median of the local time interval and constructs the smoothing weight based on the Gaussian kernel function. The weight value decreases as the distance between the data point and the center point increases. Its core idea is to calculate the weight through a Gaussian function and then smooth the temperature data. The mathematical expression of the Gaussian function is: (8).
[0061] Where: is the smoothed data at time point ; is the original time series data point; is the Gaussian weight, defined as: (9).
[0062] Wherein: is the time point smoothed data at; is an adaptive window, dynamically calculated according to the median of the short - term time interval: , where represents the set of time intervals within the window before and after the current time point, ensuring that the smoothing process targets short - term fluctuations.
[0063] Since the smoothing calculation may lack sufficient neighborhood data at the data endpoints, the present invention introduces an endpoint mean constraint: (10).
[0064] Wherein, is the number of points within the endpoint region, used to prevent abnormal deviations in the endpoint data.
[0065] The beneficial effects of the above - mentioned solution are as follows: In the processing of temperature data, the local Gaussian smoothing algorithm adopts a dynamic window calculation method, adaptively adjusts the smoothing parameters according to the changes in the short - term time interval. For the short - term fluctuation problem, the algorithm calculates the median of the local time interval and constructs the smoothing weight based on the Gaussian kernel function. Since the Gaussian weight decays exponentially with the increase of the time distance, the points far from the center have less influence on the current time point, thus keeping the long - term trend unchanged and only eliminating short - term abnormal fluctuations. In addition, the algorithm introduces a mean constraint mechanism at the data boundary to ensure the smoothness of the endpoint data while maintaining the overall data trend unchanged. The local Gaussian smoothing algorithm is particularly suitable for processing data containing random noise, and it can effectively retain the main trend of the data and eliminate unnecessary details.
[0066] In S4, the Adaptive DTW + mean Euclidean distance + hierarchical clustering algorithm is used to divide the typical temperature seasons and classify the measured temperature data by month. Using the Adaptive Dynamic Time Warping (Adaptive DTW) algorithm, the similarity between the temperature sequences of 12 months is calculated pairwise to generate a symmetric DTW distance matrix D: .
[0067] The DTW distance matrix is a multi - dimensional symmetric matrix, representing the similarity relationship between months.
[0068] At the same time, calculate the mean of the data for each month and calculate the Euclidean distance based on the mean , after normalizing the DTW distance matrix and the Euclidean distance matrix, weighted calculation is performed, and the weighted formula is as follows: .
[0069] The agglomerative hierarchical clustering method is adopted for the weighted comprehensive distance matrix. After clustering is completed, the 12 months are divided into k clusters, and each cluster corresponds to a similar seasonal pattern. Among them, represents the weight, which is generally set to 0.3 to 0.5. Although DTW can handle data with different sequence lengths, if the monthly data is less, excessive stretching will affect the accuracy of the extracted distance feature matrix. The more missing data in the month, the lower the weight.
[0070] The beneficial effects of the above solution are as follows: Since DTW can handle time series of different lengths and calculate similarity through alignment, it is very suitable for capturing the changing trends in seasonal time series. At the same time, the adaptive algorithm significantly reduces the computational complexity of the DTW distance matrix. The Euclidean distance provides computational efficiency. After normalization, distance weighting fusion is performed based on adjustable weight parameters to obtain an optimized clustering effect.
[0071] In S5, according to the clustering analysis results, different seasonal patterns are analyzed respectively. Taking the extreme values of the daily temperature difference at two measuring points as random variables, probability fitting is performed using the generalized extreme value distribution, and its probability density formula is as follows: (13).
[0072] Its cumulative distribution function formula is as follows: (14).
[0073] Among them, : location parameter, controlling the center of the data; : scale parameter, controlling the width or variability of the data; : shape parameter, determining the tail behavior of the distribution: >0: having a heavy tail (Frechet distribution); =0: Gumbel distribution, applicable to the extreme values of exponential distribution; <0: having a bounded tail (Weibull distribution).
[0074] The distribution parameters are determined using maximum likelihood estimation, thereby determining the distribution form of the daily extreme temperature difference.
[0075] In S6, according to the fitted probability distribution form in S5, based on the 50-year return period and the exceedance probability, the corresponding representative value of the temperature action is proposed, and the calculation formula is as follows: (15).
[0076] (16).
[0077] (17).
[0078] (18).
[0079] (19).
[0080] Wherein: X is the extreme value of daily temperature difference, is the typical number of days in summer during the 50-year return period; is the typical number of days in winter during the 50-year return period; is the random variable of daily temperature difference in summer; is the random variable of daily temperature difference in winter; is the exceedance probability with a return period of 50 years. During the 100-year design reference period, for extreme temperature events with a return period of 50 years, the expected number of times exceeding the temperature threshold is 2 times, = 0.02; is the shape parameter in summer; is the shape parameter in winter; is the scale parameter in summer; is the scale parameter in winter; is the location parameter in summer; is the location parameter in winter. By back-solving the extreme temperature difference that is, the representative value of temperature action.
[0081] The specific effects of the present invention include: Due to the discontinuity of the time series and the variability of the time interval caused by on-site conditions in temperature monitoring, the adaptive interpolation algorithm uses the region marking method to identify long and short NaN segments (data anomalies) and process them differently. Small segment missing (less than 10 data points) is interpolated and completed, while long segment missing is deleted to avoid the accumulation of interpolation errors and ensure interpolation accuracy. The local Gaussian smoothing algorithm uses the dynamic window size calculation method to adjust the Gaussian weights according to the change of the short-term time interval. Calculate the median of the local time interval to avoid information loss caused by over-smoothing. Only smooth the short-term fluctuating data to ensure that the long-term trend information of the time series is not affected. The Adaptive DTW and the mean Euclidean distance obtain the difference distance matrix of the temperature change characteristics in different months, use the weight parameter to combine the Adaptive DTW and the Euclidean distance, calculate the comprehensive distance matrix, perform hierarchical clustering analysis based on the comprehensive distance matrix, divide the months with similar temperature change characteristics into the same category, obtain multiple temperature change patterns, determine the typical summer temperature months and typical winter temperature months through multiple temperature change patterns, reasonably divide the typical temperature seasons, and finally obtain the representative value of temperature action according to the reasonable division result, achieving the purpose of accurately obtaining the representative value of temperature action.
[0082] An embodiment of the present invention provides a system for determining a representative temperature value, including: A data acquisition module, configured to acquire abnormal temperature data of a steel-concrete composite structure within one year, obtain a differential distance matrix of temperature change characteristics for different months by using the Adaptive Dynamic Time Warping (Adaptive DTW) algorithm and the mean Euclidean distance respectively, and perform weighted processing on the differential distance matrix of Adaptive DTW and the differential distance matrix of the mean Euclidean distance to obtain a comprehensive distance matrix.
[0083] A typical seasonal temperature month acquisition module, configured to perform agglomerative hierarchical clustering analysis on the comprehensive distance matrix, divide months with similar temperature change characteristics within one year into the same category, obtain the seasonal pattern of temperature change, and determine the typical summer temperature months and the typical winter temperature months.
[0084] A representative temperature action value acquisition module, configured to select temperature measurement points on the steel-concrete composite structure, obtain the daily extreme temperature difference of the measurement points in the typical summer temperature months and the typical winter temperature months, and use the generalized extreme value distribution model for the daily extreme temperature difference to obtain the probability distribution of the daily extreme temperature difference; according to the probability distribution, obtain the representative temperature action values for different return periods by using the exceedance probability.
[0085] A specific embodiment is as follows: This embodiment discloses a method for determining a representative temperature value, and the specific steps are as follows: S1. First, extract the time-history temperature data, and use machine learning algorithms to clean and fit the data.
[0086] S2. Then, according to the results of data cleaning, calculate the temperature difference data with months as the basic unit, obtain the differential distance matrix of temperature change characteristics for different months by using Adaptive DTW and the mean Euclidean distance, combine the normalized Adaptive DTW distance matrix and the Euclidean distance matrix with weight parameters, calculate the comprehensive distance matrix by weighting, and perform agglomerative hierarchical clustering analysis according to the comprehensive distance matrix to divide the typical seasons.
[0087] S3. Secondly, analyze the data of the divided typical seasons respectively, use the daily extreme value of the temperature difference between two measurement points as a random variable, perform probability distribution fitting on it by using the generalized extreme value distribution, and use the maximum likelihood estimation to estimate the distribution parameters.
[0088] S4. Finally, based on the fitted probability distribution results, and based on the return period and the exceedance probability, propose the representative temperature action values under the typical seasons.
[0089] Finally, a classification demonstration diagram of a method for determining a representative temperature value, such as Figure 6 shown as a dendrogram of hierarchical clustering, Figure 7It is the hierarchical clustering result based on adaptive DTW + Euclidean distance.
[0090] The method of the present invention is a calculation method for the representative value of the combined structural temperature action based on machine learning algorithms and the generalized extreme value distribution, and is applicable to engineering structures with long-term measured temperature data. According to the isolation forest algorithm, in the form of unsupervised and labeled, it can more efficiently analyze the outliers deviating from and the potential trends in the measured data, and improve the effectiveness and reliability of outlier detection. At the same time, by using the Adaptive DTW + mean Euclidean distance + hierarchical clustering algorithm, the problem of inaccurate season division caused by seasonal differences in different regions is solved. By extracting the characteristic values of data in different months, the classification problem of the typical seasons of the temperature time history data can be calculated more efficiently, and it is applicable to the problem of typical season division when analyzing seasonal differences in different regions.
[0091] The above embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A method for determining a representative temperature value, characterized in that, It includes the following steps: Obtain the abnormal temperature data of the steel-concrete composite structure within one year. Use the Adaptive Dynamic Time Warping (AdaptiveDTW) algorithm and the mean Euclidean distance to obtain the difference distance matrices of the temperature change characteristics for different months respectively. Perform weighted processing on the difference distance matrix of Adaptive DTW and the difference distance matrix of the mean Euclidean distance to obtain the comprehensive distance matrix; Perform agglomerative hierarchical clustering analysis on the comprehensive distance matrix, divide the months with similar temperature change characteristics within one year into the same category to obtain the seasonal pattern of temperature change, and determine the typical summer temperature months and typical winter temperature months; Select the temperature measurement points on the steel-concrete composite structure, obtain the daily extreme temperature differences of the measurement points in the typical summer temperature months and typical winter temperature months, and use the generalized extreme value distribution model for the daily extreme temperature differences to obtain the probability distribution of the daily extreme temperature differences; According to the probability distribution, use the exceedance probability to obtain the representative values of temperature actions for different return periods.
2. The method for determining a representative temperature value according to claim 1, wherein The specific steps for obtaining the abnormal temperature data of the steel-concrete composite structure within one year include: Use the gradient detection method to identify the data with sharp changes in the annual temperature data time series of the steel-concrete composite structure; among them, the annual temperature data time series is used to represent the temperature change data at different time points within one year; Use the box plot detection method to screen out the temperature values that deviate abnormally from the normal range from the annual temperature data time series; Use the Isolation Forest algorithm to identify the outliers from the annual temperature data time series; According to the identification and screening results, obtain the abnormal temperature data within one year.
3. The method for determining a temperature representative value according to claim 2, characterized in that The specific steps for performing agglomerative hierarchical clustering analysis on the comprehensive distance matrix, dividing the months with similar temperature change characteristics within one year into the same category to obtain the seasonal pattern of temperature change include: Perform fitting replacement on the abnormal temperature data within one year using the adaptive interpolation algorithm, and perform smoothing processing on the replacement result using Gaussian filtering; Classify the smoothing results by month. Use dynamic time warping (DTW) to pair up the 12 months two by two, and obtain the similarity between the temperature sequences of the 12 months pair by pair to generate a symmetric DTW distance matrix DTW _ dist : ; Among them, n represents the maximum month 12 of a year; Obtain the mean value of the data for each month, and obtain the Euclidean distance matrix based on the mean value ; Normalize the DTW distance matrix and the Euclidean distance matrix, and perform weighting on the normalized results: ; Among them, represents the weight; For the distance matrix Use the agglomerative hierarchical clustering method for clustering, divide the 12 months into k clusters, each cluster corresponds to a seasonal pattern and use this seasonal pattern as a temperature change pattern.
4. The method for determining a temperature representative value according to claim 1, wherein The specific steps for using the generalized extreme value distribution model for the daily extreme temperature differences to obtain the probability distribution of the daily extreme temperature differences include: Taking the extreme value of the daily temperature difference between two detection points as a random variable x , probability fitting is performed using the generalized extreme value distribution , and the formula is: ; Among them, is the location parameter, controlling the center of the data; is the scale parameter, controlling the width or variability of the data; is the shape parameter, determining the tail behavior of the distribution; Cumulative distribution function , the formula is: ; Use maximum likelihood estimation to determine the distribution parameters, and determine the probability distribution of the daily extreme temperature differences according to the distribution parameters.
5. The method for determining a representative temperature value according to claim 4, wherein The specific steps for using the exceedance probability to obtain the representative values of temperature actions for different return periods include: According to the probability distribution of the daily extreme temperature differences for different return periods, based on a 50-year return period, use the exceedance probability to obtain the corresponding representative values of temperature actions. The formula is: ; ; ; ; ; Among them, X is the extreme value of daily temperature difference, is the typical number of days in summer during the 50-year return period; is the typical number of days in winter during the 50-year return period; is the random variable of daily temperature difference in summer; is the random variable of daily temperature difference in winter; is the exceedance probability with a return period of 50 years. During the 100-year design reference period, for the extreme temperature event with a return period of 50 years, the expected number of times exceeding the temperature threshold is 2 times, = 0.02; is the shape parameter in summer; is the shape parameter in winter; is the scale parameter in summer; is the scale parameter in winter; is the location parameter in summer; is the location parameter in winter; Solve for the extreme value of temperature difference , which is the representative value of temperature effect.
6. A system for determining a representative temperature value, characterized in that, It includes: A data acquisition module, which is used to obtain the abnormal temperature data of the steel-concrete composite structure within one year. Use the Adaptive Dynamic Time Warping (Adaptive DTW) algorithm and the mean Euclidean distance to obtain the difference distance matrices of the temperature change characteristics for different months respectively. Perform weighted processing on the difference distance matrix of Adaptive DTW and the difference distance matrix of the mean Euclidean distance to obtain the comprehensive distance matrix; A typical seasonal temperature month acquisition module is used to perform agglomerative hierarchical clustering analysis on the comprehensive distance matrix, divide the months with similar temperature change characteristics within a year into the same category, obtain the seasonal pattern of temperature change, and determine the typical summer temperature months and typical winter temperature months; A temperature action representative value acquisition module is used to select temperature measurement points on the steel-concrete composite structure, obtain the daily extreme temperature difference of the measurement points in the typical summer temperature months and typical winter temperature months, and use the generalized extreme value distribution model for the daily extreme temperature difference to obtain the probability distribution of the daily extreme temperature difference; According to the probability distribution, the representative values of temperature actions with different return periods are obtained using the exceedance probability.
Citation Information
Patent Citations
Highway steel box girder bridge temperature gradient mode evaluation method
CN107391823A
Machine tool temperature sensitive point selection and thermal error modeling method
CN118605384A
Temperature and humidity trend analysis and early warning method and system for high-voltage cable
CN119382321A
Vertical temperature gradient fatigue load spectrum of long-life railway steel box girder bridge and construction method
CN119598826A
Casting mold breakout prediction method based on feature vectors and hierarchical clustering
WO2020119156A1
Cited By
Transformer temperature signal time sequence analysis method
CN115876339A