A photovoltaic support pile foundation uplift bearing capacity detection method based on reinforcement learning

CN122346722BActive Publication Date: 2026-09-18POWERCHINA BEIJING ENG CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610307363.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-22
Publication Date
2026-09-18
Estimated Expiration
2046-04-22

AI Technical Summary

Technical Problem

传统静载荷检测方法依赖大型加载设备,检测过程耗时长、成本高,不适用于大规模批量桩基的快速评估,且在高风沙、高盐碱或软土环境下布设困难;基于工程经验的承载力预测方法受限于地质非线性特征与结构参数差异,泛化性差,预测精度波动大,易导致评估结果偏保守或失准;现有数据驱动模型多采用固定结构的深度神经网络,难以充分表达桩基受力响应过程中的动态趋势、非线性扰动与跨模态耦合信息,导致状态建模能力不足;此外,传统承载力分类方式采用静态阈值划分,忽略了桩基结构演化过程中承载力分布的连续性与风险过渡特征,难以实现精准等级判断与动态风险分级

Benefits of technology

本发明通过引入改进型TimeMixer网络与改进型SAC模型,实现了对光伏支架桩基抗拔检测数据的智能化建模与承载力预测,有效解决了现有技术中依赖静载荷试验设备检测效率低、经验公式精度不足、桩基状态建模能力弱等问题。首先,本发明通过采集多源抗拔检测数据并进行预处理,构建抗拔检测特征张量,为状态建模提供高质量输入基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122346722B_ABST
    Figure CN122346722B_ABST
Patent Text Reader

Abstract

The application discloses a photovoltaic support pile foundation uplift bearing capacity detection method based on reinforcement learning, which comprises the following steps: step one, collecting the uplift detection data of the target photovoltaic support pile foundation; step two, data preprocessing is performed on the uplift detection data; step three, a pile foundation trend state modeling is performed through an improved TimeMixer network; step four, an improved SAC model is constructed and initialized in a reinforcement learning environment; step five, strategy training is performed based on the improved SAC model, and a predicted bearing capacity is output; step six, threshold division is performed on the predicted bearing capacity, and a safety level label is obtained; and step seven, the uplift bearing capacity detection result of the target photovoltaic support pile foundation is generated. Through the improved TimeMixer network and the improved SAC model, the precision and efficiency of the photovoltaic support pile foundation uplift bearing capacity detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent detection and engineering structure safety assessment technology, and in particular to a method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning. Background Technology

[0002] With the increasing scale of photovoltaic power plant construction and the continuous expansion of engineering scenarios in complex geological environments, the demand for the assessment and quality control of the pull-out bearing capacity of photovoltaic support foundation structures has significantly increased. Current projects mainly use static load tests or empirical formula estimation methods to test and determine the pull-out performance of photovoltaic support pile foundations. However, in practical applications, the following problems commonly exist: Traditional static load testing methods rely on large loading equipment, which is time-consuming and costly, making them unsuitable for rapid assessment of large-scale batch pile foundations. Furthermore, they are difficult to deploy in environments with high winds and sand, high salinity, or soft soil. Bearing capacity prediction methods based on engineering experience are limited by geological nonlinear characteristics and differences in structural parameters, resulting in poor generalization, large fluctuations in prediction accuracy, and a tendency for conservative or inaccurate assessment results. Existing data-driven models often employ deep neural networks with fixed structures, which struggle to fully express the dynamic trends, nonlinear disturbances, and cross-modal coupling information during the pile foundation's stress response, leading to insufficient state modeling capabilities. In addition, traditional bearing capacity classification methods use static threshold divisions, ignoring the continuity of bearing capacity distribution and risk transition characteristics during the pile foundation's structural evolution, making it difficult to achieve accurate level judgment and dynamic risk grading.

[0003] Therefore, how to provide a reinforcement learning-based method for detecting the pull-out bearing capacity of photovoltaic support pile foundations is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] One objective of this invention is to propose a method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning. This invention combines an improved TimeMixer network and an improved SAC model, and describes in detail how to extract the pile foundation detection state vector through the improved TimeMixer network, generate the predicted bearing capacity and determine the safety level through the improved SAC model. It has the advantages of high prediction accuracy, high detection efficiency and strong adaptability.

[0005] A method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to an embodiment of the present invention includes the following steps: Step 1: Collect pull-out test data of the target photovoltaic support pile foundation; Step 2: Perform data preprocessing on the pull-out detection data to generate the pull-out detection feature tensor; Step 3: Input the pull-out detection feature tensor into the improved TimeMixer network, perform pile foundation trend state modeling, and generate pile foundation detection state vector. The improved TimeMixer network includes a time difference modeling module, a joint channel guidance module, a time channel hierarchical modeling module, and a state generation module. Step 4: Based on the pile foundation detection state vector, construct a reinforcement learning environment and initialize an improved SAC model, which includes a dual-scale policy network, a dual Q-value network, and a target Q-value network; Step 5: Execute strategy training based on the improved SAC model, and input the pile foundation detection state vector of the target photovoltaic support pile foundation into the trained improved SAC model to output the predicted bearing capacity; Step 6: Apply threshold classification to the predicted bearing capacity to obtain safety level labels; Step 7: Based on the predicted bearing capacity and safety level label, generate the pull-out bearing capacity test results of the target photovoltaic support pile foundation.

[0006] Optionally, step one specifically includes: The pull-out test data includes the loading force sequence, pile top displacement sequence, pile structure parameters, and geological environment parameters; The structural parameters of the pile include pile diameter, pile length, and depth of penetration into the soil; The geological environmental parameters include the main soil layer type, natural moisture content, penetration resistance coefficient, and relative density at the location of the target photovoltaic support pile foundation.

[0007] Optionally, step two specifically includes: The loading force sequence and pile top displacement sequence are time-aligned according to the set time step order. The Z-score standardization method is used to filter and remove abnormal data in the loading force sequence and pile top displacement sequence. The missing data in the loading force sequence and pile top displacement sequence are filled by linear interpolation method to obtain the standard loading force sequence and standard pile top displacement sequence. The structural parameters of the pile are subjected to minimum-maximum normalization to obtain the structural parameter feature vector. The main soil layer category in the geological environment parameters is transformed into a soil layer category vector through a trainable embedding mapping matrix. The natural water content, penetration resistance coefficient and relative density are subjected to minimum-maximum normalization and then arranged into a geological environment parameter vector in sequence. The standard loading force sequence and the standard pile top displacement sequence are subjected to fast Fourier transform to obtain the complex spectrum of the loading force in the frequency domain and the complex spectrum of the pile top displacement in the frequency domain. Based on the complex spectrum of the loading force in the frequency domain and the complex spectrum of the pile top displacement in the frequency domain, the dominant frequency of the loading force, the maximum amplitude frequency of the loading force, the total power of the loading force, the dominant frequency of the pile top displacement, the maximum amplitude frequency of the pile top displacement, and the total power of the pile top displacement are extracted and spliced ​​together to form a frequency domain response feature sequence. The standard load force sequence, standard pile top displacement sequence, structural parameter feature vector, geological environment parameter vector, and frequency domain response feature sequence are dimensionally aligned and feature-stitched to obtain the standard pull-out test feature tensor.

[0008] Optionally, step three specifically includes: In the temporal difference modeling module, the first-order temporal difference is calculated for all time steps and feature dimensions of the pull-out detection feature tensor to obtain the first-order difference tensor. The mean difference of the pull-out detection feature tensor with a sliding window size of 2 and the mean difference of the pull-out detection feature tensor with a sliding window size of 3 are calculated respectively to obtain the second-order difference tensor and the third-order difference tensor. Based on the time step of the third-order difference tensor, the first-order difference tensor and the second-order difference tensor are truncated to obtain aligned first-order difference tensors and aligned second-order difference tensors. The first-order difference tensor, the second-order difference tensor, and the third-order difference tensor are concatenated along the feature dimension to obtain the trend-enhancing feature tensor. In the joint channel guidance module, structural parameter feature vectors and geological environment parameter vectors are obtained and concatenated along the feature dimension to obtain a joint guidance vector; The joint guiding vector is linearly mapped to the channel mapping bias vector through a trainable channel mapping matrix, and then normalized using the Sigmoid activation function to obtain the channel weight vector. The channel weight vector and the trend enhancement feature tensor are multiplied element-wise along the feature dimension to obtain the channel guidance feature tensor. The time channel hierarchical modeling module includes an L-layer TimeMixer hybrid module, where L is a positive integer greater than or equal to 4. The time channel hierarchical modeling module adopts a hierarchical sparse strategy to generate channel-enhanced feature tensors. In the state generation module, the channel enhancement feature tensor is averaged in the time dimension to obtain the pile foundation detection state vector.

[0009] Optionally, the time channel hierarchical modeling module employs a hierarchical sparse strategy to generate channel-enhanced feature tensors, specifically including: In the first and second layers of the TimeMixer module, the channel feature vectors of each channel of the channel-guided feature tensor are sparsely modulated to obtain a sparse activation feature tensor. Then, the sparse activation feature tensor is subjected to time mixing and channel mixing operations to output a channel-mixed feature tensor. Specifically: The logits vector is transformed into a linear mapping, and the logits vector has a dimension of 2, which respectively represent the score of the current channel and the score of the current channel. The logits vector is transformed into a channel gating probability vector using the Gumbel Softmax activation function. The channel gating probability vector has a dimension of 2, representing the probability of retaining the current channel and the probability of discarding the current channel, respectively. The probability of retaining the current channel is extracted as the sparse activation weight of the current channel. The sparse activation weights are used to sparsely modulate the channel feature vectors to obtain sparse feature vectors, and the sparse feature vectors of all channels are combined into a sparse activation feature tensor. The sparse activation feature tensor is subjected to layer normalization and then input into a multilayer perceptron in the time dimension for temporal hybrid modeling to obtain a temporal modeling feature tensor. The temporal modeling feature tensor is then residually connected with the sparse activation feature tensor to obtain a temporal hybrid feature tensor. The temporal fusion feature tensor is subjected to layer normalization and then input into a multilayer perceptron in the channel dimension for channel fusion modeling to obtain the channel modeling feature tensor. The channel modeling feature tensor is then residually connected with the temporal fusion feature tensor to obtain the channel fusion feature tensor. In the third layer and the subsequent TimeMixer mixing module, sparse modulation is canceled, and the channel mixing feature tensor is subjected to time mixing and channel mixing operations to output the channel enhanced feature tensor.

[0010] Optionally, step four specifically includes: The pile foundation detection state vectors are arranged into the state space of the reinforcement learning environment according to the set sample number order. The action space of the reinforcement learning environment is set up, the action space includes several actions, and each action corresponds to a set of load-bearing capacity prediction parameters; A reward function is constructed, which is a weighted combination of a carrying capacity prediction error term, an action change penalty term, and a strategy distribution entropy term. The load-bearing capacity prediction error term is the negative absolute difference between the actual load-bearing capacity and the predicted load-bearing capacity corresponding to the action; the action change penalty term is the negative Euclidean distance between the action difference between the current sample number and the previous sample number; the policy distribution entropy term is the negative information entropy of the fused policy distribution. In a reinforcement learning environment, initialize the improved SAC model; The dual-scale policy network includes a fast policy network and a slow policy network. The fast policy network consists of three fully connected layers, each followed by a ReLU activation function. The slow policy network adopts a time-series modeling structure based on a sliding window, performs one-dimensional convolution on the pile foundation detection state vector within a set sliding window interval, and extracts cross-sample state evolution features through a gated recurrent unit. The dual Q-value network includes a first Q-value network and a second Q-value network. Both the first Q-value network and the second Q-value network are composed of a four-layer perceptron structure. The first and second layers are fully connected layers, the third layer is a ReLU activation layer, and the fourth layer is a linear regression layer. The target Q-value network includes a first target Q-value network and a second target Q-value network, and the target Q-value network has the same structure as the dual Q-value network.

[0011] Optionally, step five specifically includes: During the policy training process, the pile foundation detection state vector of the current sample number is input into the dual-scale policy network. The short-term policy distribution is generated through the fast policy network, and the long-term policy distribution is generated through the slow policy network. The short-term policy distribution and the long-term policy distribution are weighted and fused to generate the fused policy distribution. Based on the fusion strategy distribution, the pile foundation detection state vector of the current sample number is sampled for action, and the pile foundation detection state vector and the sampled action are concatenated to form a state-action pair. The state-action pair is input into a dual Q-value network. The first Q-value is calculated and generated through the first Q-value network, and the second Q-value is calculated and generated through the second Q-value network. The minimum value between the first Q-value and the second Q-value is selected as the estimated Q-value. The sampling action of the current sample number is used to calculate the predicted bearing capacity through a single MLP structure. The actual bearing capacity, the predicted bearing capacity, the sampling action of the current sample number and the previous sample number, and the fusion strategy distribution are substituted into the reward function to calculate the reward value of the current sample number. The state action pair of the next sample number is input into the target Q-value network. The first target Q-value is generated by the first target Q-value network, and the second target Q-value is generated by the second target Q-value network. The minimum value between the first target Q-value and the second target Q-value is selected as the candidate target Q-value. The target Q value is obtained by weighting the reward value of the current sample number, the candidate target Q value, and the negative log probability of the fusion strategy distribution of the next sample number. The mean square error between the estimated Q value and the target Q value is calculated to obtain the Critic loss function, and the gradient backpropagation update of the dual Q value network is performed by minimizing the Critic loss function. Calculate the expected difference between the negative log probability of the fusion policy distribution for the current sample number and the estimated Q value to obtain the policy loss function, and perform gradient update on the dual-scale policy network by minimizing the policy loss function; The target Q-value network is updated using a moving average method; When the number of training rounds reaches the set maximum number of rounds, stop policy training and obtain the improved SAC model that has been trained. The pile foundation detection state vector of the target photovoltaic support pile foundation is input into the trained improved SAC model, and the predicted bearing capacity is output.

[0012] Optionally, step six specifically includes: Set bearing capacity thresholds T1 and T2, wherein bearing capacity threshold T1 is greater than bearing capacity threshold T2; If the predicted bearing capacity is greater than or equal to the bearing capacity threshold T1, the safety level label is safe. If the predicted bearing capacity is greater than or equal to the bearing capacity threshold T2 and less than the bearing capacity threshold T1, then the safety level label is a warning. If the predicted bearing capacity is less than the bearing capacity threshold T2, the safety level label is dangerous.

[0013] Optionally, the tensile bearing capacity test results include pile foundation number, predicted bearing capacity, safety level label, and test time.

[0014] The beneficial effects of this invention are: This invention, by introducing an improved TimeMixer network and an improved SAC model, achieves intelligent modeling and bearing capacity prediction of pull-out test data for photovoltaic support pile foundations. This effectively solves the problems of low testing efficiency due to reliance on static load testing equipment, insufficient accuracy of empirical formulas, and weak pile foundation condition modeling capabilities in existing technologies. Firstly, this invention collects and preprocesses multi-source pull-out test data to construct a pull-out test feature tensor, providing a high-quality input foundation for condition modeling.

[0015] Secondly, this invention proposes an improved TimeMixer network, which includes a time difference modeling module, a joint channel guidance module, a time channel hierarchical modeling module, and a state generation module. This improved TimeMixer network can fully extract multi-order time difference trend information from the pile foundation pull-out detection features and integrate pile structural parameters and geological environment parameters for channel weight adjustment, effectively enhancing feature representation capabilities and adaptability to engineering scenarios. By adjusting the channel activation intensity through a hierarchical sparsity strategy, the improved TimeMixer network's response sensitivity to key variables is improved, avoiding modeling interference caused by channel redundancy.

[0016] Furthermore, in the reinforcement learning strategy modeling stage, this invention constructs an improved SAC model combining a dual-scale policy network, a dual Q-value network, and a target Q-value network. By constructing a reward function that includes a bearing capacity prediction error term, an action change penalty term, and a policy distribution entropy term, dynamic optimization and continuous output of the pile foundation detection strategy are achieved. During training, short-term and long-term strategies are integrated to enhance the improved SAC model's adaptability to different pile foundation structures and geological conditions. Through action sampling and Q-value estimation, policy bias is effectively reduced, ensuring the reliability of the bearing capacity prediction results. Finally, thresholding of the predicted bearing capacity and outputting safety level labels provide a rapid and accurate basis for pile foundation bearing capacity detection in photovoltaic power plant construction and operation.

[0017] In summary, this invention can achieve high-precision prediction and grade determination of pile foundation bearing capacity under complex working conditions, and has the advantages of strong data adaptability, good modeling effect, high inference efficiency and low deployment cost. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of a reinforcement learning-based method for detecting the pull-out bearing capacity of photovoltaic support pile foundations proposed in this invention; Figure 2 This is a flowchart of the improved TimeMixer network structure in the reinforcement learning-based photovoltaic support pile foundation pull-out bearing capacity detection method proposed in this invention; Figure 3 This is a flowchart of the improved SAC model modeling structure in the reinforcement learning-based photovoltaic support pile foundation pull-out bearing capacity detection method proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figures 1-3 A reinforcement learning-based method for detecting the pull-out bearing capacity of photovoltaic support pile foundations includes the following steps: Step 1: Collect pull-out test data of the target photovoltaic support pile foundation; Step 2: Perform data preprocessing on the pull-out detection data to generate the pull-out detection feature tensor; Step 3: Input the pull-out detection feature tensor into the improved TimeMixer network, perform pile foundation trend state modeling, and generate pile foundation detection state vector. The improved TimeMixer network includes a time difference modeling module, a joint channel guidance module, a time channel hierarchical modeling module, and a state generation module. Step 4: Based on the pile foundation detection state vector, construct a reinforcement learning environment and initialize an improved SAC model, which includes a dual-scale policy network, a dual Q-value network, and a target Q-value network; Step 5: Execute strategy training based on the improved SAC model, and input the pile foundation detection state vector of the target photovoltaic support pile foundation into the trained improved SAC model to output the predicted bearing capacity; Step 6: Apply threshold classification to the predicted bearing capacity to obtain safety level labels; Step 7: Based on the predicted bearing capacity and safety level label, generate the pull-out bearing capacity test results of the target photovoltaic support pile foundation.

[0021] In this embodiment, step one specifically includes: The pull-out test data includes the loading force sequence, pile top displacement sequence, pile structure parameters, and geological environment parameters; The structural parameters of the pile include pile diameter, pile length, and depth of penetration into the soil; The geological environmental parameters include the main soil layer type, natural moisture content, penetration resistance coefficient, and relative density at the location of the target photovoltaic support pile foundation.

[0022] In this embodiment, step two specifically includes: The loading force sequence and pile top displacement sequence are time-aligned according to the set time step order. The Z-score standardization method is used to filter and remove abnormal data in the loading force sequence and pile top displacement sequence. The missing data in the loading force sequence and pile top displacement sequence are filled by linear interpolation method to obtain the standard loading force sequence and standard pile top displacement sequence. The structural parameters of the pile are subjected to minimum-maximum normalization to obtain the structural parameter feature vector. The main soil layer category in the geological environment parameters is transformed into a soil layer category vector through a trainable embedding mapping matrix. The natural water content, penetration resistance coefficient and relative density are subjected to minimum-maximum normalization and then arranged into a geological environment parameter vector in sequence. In this invention, the generation process of the soil layer category vector is as follows: A set of primary soil layer categories is defined, which includes several primary soil layer categories. Each primary soil layer category is encoded as a unique hot vector using a numbering method. For example, if the set of primary soil layer categories includes {clay, silty clay, sand, gravelly soil, fill, and silty soil}, then the unique hot vector for "clay" is [1,0,0,0,0,0], and the unique hot vector for "silty clay" is [0,1,0,0,0,0]. A trainable embedding mapping matrix is ​​defined, and the unique hot vector is multiplied by the trainable embedding mapping matrix to obtain the soil layer category vector.

[0023] The standard loading force sequence and the standard pile top displacement sequence are subjected to fast Fourier transform to obtain the complex spectrum of the loading force in the frequency domain and the complex spectrum of the pile top displacement in the frequency domain. Based on the complex spectrum of the loading force in the frequency domain and the complex spectrum of the pile top displacement in the frequency domain, the dominant frequency of the loading force, the maximum amplitude frequency of the loading force, the total power of the loading force, the dominant frequency of the pile top displacement, the maximum amplitude frequency of the pile top displacement, and the total power of the pile top displacement are extracted and spliced ​​together to form a frequency domain response feature sequence. In this invention, the time-domain signal is converted into a frequency-domain signal using a Fast Fourier Transform (FFT). This allows for the extraction of the dominant frequency, maximum amplitude frequency, and total power of the loading force and pile top displacement, reflecting the dominant vibration components, energy distribution, and stiffness characteristics of the pile-soil relationship. These frequency-domain features are then concatenated into a frequency-domain response feature sequence, enhancing the improved TimeMixer network's ability to identify pile foundation bearing behavior and deformation patterns. This also strengthens its sensitivity to frequency-domain variation trends under different geological conditions, thereby improving overall prediction accuracy and generalization ability.

[0024] The standard load force sequence, standard pile top displacement sequence, structural parameter feature vector, geological environment parameter vector, and frequency domain response feature sequence are dimensionally aligned and feature-stitched to obtain the standard pull-out test feature tensor. The standard pull-out test feature tensor structure is time step × feature dimension, where the feature dimension includes the feature dimensions of all preprocessed pull-out test data, specifically 14 + d, where d represents the dimension of the soil layer category vector.

[0025] In this embodiment, step three specifically includes: In the temporal difference modeling module, the first-order temporal difference is calculated for all time steps and feature dimensions of the pull-out detection feature tensor to obtain the first-order difference tensor. The mean difference of the pull-out detection feature tensor with a sliding window size of 2 and the mean difference of the pull-out detection feature tensor with a sliding window size of 3 are calculated respectively to obtain the second-order difference tensor and the third-order difference tensor. Based on the time step of the third-order difference tensor, the first-order difference tensor and the second-order difference tensor are truncated to obtain aligned first-order difference tensors and aligned second-order difference tensors. The first-order difference tensor, the second-order difference tensor, and the third-order difference tensor are concatenated along the feature dimension to obtain the trend-enhancing feature tensor. In this invention, the time-difference modeling module effectively enhances the modeling capability of trend changes and dynamic disturbances in pile foundation pull-out test data by introducing a multi-order difference feature modeling mechanism. The first-order difference tensor characterizes the local rate of change of the original features, while the second-order and third-order difference tensors further extract short- to medium-term trend evolution features through moving average differencing, exhibiting stronger smoothness and robustness. Time-step alignment ensures the consistency of the multi-order difference tensors in the time dimension, which are then spliced ​​and fused in the feature dimension to form a trend-enhanced feature tensor containing multi-scale change information, significantly improving the sensitivity and stability of the improved TimeMixer network to changes in pile foundation mechanical response.

[0026] In the joint channel guidance module, structural parameter feature vectors and geological environment parameter vectors are obtained and concatenated along the feature dimension to obtain a joint guidance vector; The joint guiding vector is linearly mapped to the channel mapping bias vector through a trainable channel mapping matrix, and then normalized using the Sigmoid activation function to obtain the channel weight vector. The channel weight vector and the trend enhancement feature tensor are multiplied element-wise along the feature dimension to obtain the channel guidance feature tensor. In this invention, structural parameter feature vectors and geological environment parameter vectors are concatenated to form a joint guiding vector reflecting the physical properties of the pile foundation and the conditions of the construction site. A normalized channel weight vector is then generated through a trainable channel mapping matrix, enabling the improved TimeMixer network to selectively adjust the importance of each feature channel based on different pile types and geological conditions. This design method effectively enhances the improved TimeMixer network's ability to identify and generalize key bearing capacity variation characteristics in complex geological environments, thereby improving the accuracy and environmental adaptability of bearing capacity prediction results.

[0027] The time channel hierarchical modeling module includes an L-layer TimeMixer hybrid module, where L is a positive integer greater than or equal to 4. The time channel hierarchical modeling module adopts a hierarchical sparse strategy to generate channel-enhanced feature tensors. In the state generation module, the channel enhancement feature tensor is averaged in the time dimension to obtain the pile foundation detection state vector.

[0028] In this embodiment, the time channel hierarchical modeling module adopts a hierarchical sparse strategy to generate channel-enhanced feature tensors, specifically including: In the first and second layers of the TimeMixer module, the channel feature vectors of each channel of the channel-guided feature tensor are sparsely modulated to obtain a sparse activation feature tensor. Then, the sparse activation feature tensor is subjected to time mixing and channel mixing operations to output a channel-mixed feature tensor. Specifically: The logits vector is transformed into a linear mapping, and the logits vector has a dimension of 2, which respectively represent the score of the current channel and the score of the current channel. The logits vector is transformed into a channel gating probability vector using the Gumbel Softmax activation function. The channel gating probability vector has a dimension of 2, representing the probability of retaining the current channel and the probability of discarding the current channel, respectively. The probability of retaining the current channel is extracted as the sparse activation weight of the current channel. For example, suppose the current channel, after linear mapping, yields a logits vector of [2.1, 0.3], representing the score for retaining or discarding the current channel. After Gumbel Softmax activation, a channel gating probability vector of [0.91, 0.09] is obtained, where 0.91 is the sparse activation weight of this channel, indicating that the improved TimeMixer network has a high degree of confidence in retaining the current channel. The sparse activation weight is then multiplied with the current channel features in a weighted manner to adjust the sparsity of the feature output and filter information.

[0029] The sparse activation weights are used to sparsely modulate the channel feature vectors to obtain sparse feature vectors, and the sparse feature vectors of all channels are combined into a sparse activation feature tensor. The sparse activation feature tensor is subjected to layer normalization and then input into a multilayer perceptron in the time dimension for temporal hybrid modeling to obtain a temporal modeling feature tensor. The temporal modeling feature tensor is then residually connected with the sparse activation feature tensor to obtain a temporal hybrid feature tensor. The temporal fusion feature tensor is subjected to layer normalization and then input into a multilayer perceptron in the channel dimension for channel fusion modeling to obtain the channel modeling feature tensor. The channel modeling feature tensor is then residually connected with the temporal fusion feature tensor to obtain the channel fusion feature tensor. In the third layer and the subsequent TimeMixer mixing module, sparse modulation is canceled, and the channel mixing feature tensor is subjected to time mixing and channel mixing operations to output the channel enhanced feature tensor.

[0030] In this invention, the hierarchical sparse strategy significantly enhances the expressive efficiency and channel structure adaptability of feature modeling. By introducing a sparse modulation mechanism in the first two TimeMixer mixing modules, the retention probability of each channel is dynamically calculated based on the Gumbel Softmax activation function, enabling adaptive selection of the importance of different channels. Only channel information with discriminative power for the current task is retained, thereby effectively suppressing redundant or invalid feature interference and reducing the risk of overfitting caused by feature redundancy. Simultaneously, the sparse activation feature tensor is deeply modeled using a multilayer perceptron in both the temporal and channel dimensions, and residual connections are fused, further improving the improved TimeMixer network's ability to model key temporal dynamics and semantic interactions between channels. From the third layer onwards, sparse modulation is eliminated, focusing on deep fusion of all channel features, which is beneficial for capturing global semantic and trend information. In summary, the hierarchical sparse strategy balances the flexibility of channel selection with the hierarchical adaptability of modeling depth, improving the noise resistance and generalization performance of the improved TimeMixer network.

[0031] In this embodiment, step four specifically includes: The pile foundation detection state vectors are arranged into the state space of the reinforcement learning environment according to the set sample number order. The action space of the reinforcement learning environment is set up, the action space includes several actions, and each action corresponds to a set of load-bearing capacity prediction parameters; A reward function is constructed, which is a weighted combination of a carrying capacity prediction error term, an action change penalty term, and a strategy distribution entropy term. The load-bearing capacity prediction error term is the negative absolute difference between the actual load-bearing capacity and the predicted load-bearing capacity corresponding to the action; the action change penalty term is the negative Euclidean distance between the action difference between the current sample number and the previous sample number; the policy distribution entropy term is the negative information entropy of the fused policy distribution. In a reinforcement learning environment, initialize the improved SAC model; The dual-scale policy network includes a fast policy network and a slow policy network. The fast policy network consists of three fully connected layers, each followed by a ReLU activation function. The slow policy network adopts a time-series modeling structure based on a sliding window, performs one-dimensional convolution on the pile foundation detection state vector within a set sliding window interval, and extracts cross-sample state evolution features through a gated recurrent unit. In this invention, a dual-scale policy network is designed to balance the instantaneous response characteristics of pile foundation detection status with the evolutionary trend across time windows, thereby improving the decision-making accuracy and stability of the policy network at different time scales. Specifically, the fast policy network, through a three-layer fully connected structure, rapidly responds to current state information, making it suitable for capturing local changes and sudden behaviors. The slow policy network introduces a sliding window mechanism and a gated loop structure to perform convolution and temporal modeling on historical state sequences, extracting long-term dependent dynamic evolutionary features and enhancing the robustness and predictive foresight of the policy. Through dual-scale collaborative modeling, short-term mutations and long-term trends can be integrated, achieving better control and adjustment of the generated strategy under complex pile foundation detection environments.

[0032] The dual Q-value network includes a first Q-value network and a second Q-value network. Both the first Q-value network and the second Q-value network are composed of a four-layer perceptron structure. The first and second layers are fully connected layers, the third layer is a ReLU activation layer, and the fourth layer is a linear regression layer. The target Q-value network includes a first target Q-value network and a second target Q-value network, and the target Q-value network has the same structure as the dual Q-value network.

[0033] In this embodiment, step five specifically includes: During the policy training process, the pile foundation detection state vector of the current sample number is input into the dual-scale policy network. The short-term policy distribution is generated through the fast policy network, and the long-term policy distribution is generated through the slow policy network. The short-term policy distribution and the long-term policy distribution are weighted and fused to generate the fused policy distribution. Based on the fusion strategy distribution, the pile foundation detection state vector of the current sample number is sampled for action, and the pile foundation detection state vector and the sampled action are concatenated to form a state-action pair. The state-action pair is input into a dual Q-value network. The first Q-value is calculated and generated through the first Q-value network, and the second Q-value is calculated and generated through the second Q-value network. The minimum value between the first Q-value and the second Q-value is selected as the estimated Q-value. The sampling action of the current sample number is used to calculate the predicted bearing capacity through a single MLP structure. The actual bearing capacity, the predicted bearing capacity, the sampling action of the current sample number and the previous sample number, and the fusion strategy distribution are substituted into the reward function to calculate the reward value of the current sample number. The state action pair of the next sample number is input into the target Q-value network. The first target Q-value is generated by the first target Q-value network, and the second target Q-value is generated by the second target Q-value network. The minimum value between the first target Q-value and the second target Q-value is selected as the candidate target Q-value. The target Q value is obtained by weighting the reward value of the current sample number, the candidate target Q value, and the negative log probability of the fusion strategy distribution of the next sample number. The mean square error between the estimated Q value and the target Q value is calculated to obtain the Critic loss function, and the gradient backpropagation update of the dual Q value network is performed by minimizing the Critic loss function. Calculate the expected difference between the negative log probability of the fusion policy distribution for the current sample number and the estimated Q value to obtain the policy loss function, and perform gradient update on the dual-scale policy network by minimizing the policy loss function; The target Q-value network is updated using a moving average method; When the number of training rounds reaches the set maximum number of rounds, stop policy training and obtain the improved SAC model that has been trained. The pile foundation detection state vector of the target photovoltaic support pile foundation is input into the trained improved SAC model, and the predicted bearing capacity is output.

[0034] In this embodiment, step six specifically includes: Set bearing capacity thresholds T1 and T2, wherein bearing capacity threshold T1 is greater than bearing capacity threshold T2; If the predicted bearing capacity is greater than or equal to the bearing capacity threshold T1, the safety level label is safe. If the predicted bearing capacity is greater than or equal to the bearing capacity threshold T2 and less than the bearing capacity threshold T1, then the safety level label is a warning. If the predicted bearing capacity is less than the bearing capacity threshold T2, the safety level label is dangerous.

[0035] In this embodiment, the tensile bearing capacity test results include pile foundation number, predicted bearing capacity, safety level label, and test time.

[0036] Example 1 To verify the feasibility of this invention in practice, the method was applied to the pile foundation uplift bearing capacity testing of a 200MW photovoltaic power station project in a plateau region. This region has a complex geological structure and soft soil, with a large number of hidden substandard pile foundations that, while not significantly damaged, have insufficient actual bearing capacity. Traditional uplift tests rely on manual judgment and static load measurement, which is not only inefficient but also highly subjective and prone to large errors in determining test results, making it difficult to meet the rapid testing requirements of high-intensity and high-density construction environments.

[0037] In the implementation scenario, an embedded acquisition device is first used to perform pull-out force tests on each target photovoltaic support pile foundation to obtain raw detection signals, including information such as stress changes, displacement changes, and vibration response over time. The raw detection signals are preprocessed using an edge computing module to generate a pull-out detection feature tensor, which is then fed into an inference engine containing an improved TimeMixer network to identify the stress response trend and perturbation characteristics throughout the loading process, extracting the pile foundation detection state vector. This pile foundation detection state vector is then input into an improved SAC model trained and optimized using field samples, outputting a predicted bearing capacity. Based on different bearing capacity threshold ranges set by standard engineering specifications, the model automatically determines the safety level label and generates a test report, which is synchronized to the construction quality control platform, providing construction personnel with immediate safety feedback and handling suggestions.

[0038] To further verify the performance advantages of the method of this invention, a comparative experiment was conducted with the following three detection methods: Comparison Scheme A is the traditional manual static load method, Comparison Scheme B is a deep learning prediction model based on LSTM, and Comparison Scheme C is a regression prediction model based on XGBoost. The comparative experiments were conducted under the same test conditions, selecting 100 representative pile foundation samples. Prediction and judgment were performed on each of the four methods. Evaluation indicators included mean prediction error (MAE), maximum relative error (MRE), grade determination accuracy, single pile detection time, false alarm rate, and false negative rate. The statistical results are shown in Table 1.

[0039] Table 1. Comparison of performance of different methods in pile foundation tensile bearing capacity testing

[0040] As shown in Table 1, the method of this invention exhibits significant advantages in predicting tensile bearing capacity and assessing safety levels. The average prediction error of the method of this invention is 4.3 kN, lower than that of comparative scheme A (9.6 kN), comparative scheme B (6.9 kN), and comparative scheme C (5.6 kN), indicating that the method of this invention effectively improves the accuracy of bearing capacity prediction. The method of this invention controls the maximum relative error to 7.8%, which is lower than that of the comparative schemes (18.4%, 14.1%, and 11.7%), indicating stronger error control capabilities and the ability to meet the on-site requirements for high-precision testing. The safety level determination accuracy of the method of this invention reaches 96.5%, significantly higher than that of the comparative schemes, demonstrating that the invention possesses higher reliability and accuracy in pile foundation bearing capacity assessment.

[0041] Furthermore, the method of this invention demonstrates outstanding detection efficiency, with a single pile detection time of only 2.6 seconds, significantly lower than the 156.8 seconds per pile of comparative scheme A, and superior to the 7.4 seconds and 6.1 seconds per pile of comparative schemes B and C, respectively, thus significantly improving the detection efficiency at photovoltaic construction sites. Regarding anomaly identification capabilities, the method of this invention has a false alarm rate of only 1.2% and a false negative rate of 1.5%, exhibiting higher stability and scenario adaptability compared to the comparative schemes. This indicates that the method of this invention can effectively avoid misjudgment and false negative detection of dangerous pile foundations, thereby enhancing the engineering safety assurance capability.

[0042] The method of this invention not only outperforms the comparative schemes in terms of prediction accuracy and precision, but also achieves breakthrough improvements in several key indicators such as detection speed, anomaly control and safety level assessment, providing more intelligent, efficient and safe detection support for large-scale photovoltaic support pile foundation construction.

[0043] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Collect pull-out test data of the target photovoltaic support pile foundation; the pull-out test data includes the loading force sequence, pile top displacement sequence, pile structural parameters and geological environment parameters; Step 2: Perform data preprocessing on the pull-out detection data to generate the pull-out detection feature tensor; Step 3: Input the pull-out detection feature tensor into the improved TimeMixer network, perform pile foundation trend state modeling, and generate pile foundation detection state vector. The improved TimeMixer network includes a time difference modeling module, a joint channel guidance module, a time channel hierarchical modeling module, and a state generation module. In the temporal difference modeling module, the first-order temporal difference is calculated for all time steps and feature dimensions of the pull-out detection feature tensor to obtain the first-order difference tensor. The mean difference of the pull-out detection feature tensor with a sliding window size of 2 and the mean difference of the pull-out detection feature tensor with a sliding window size of 3 are calculated respectively to obtain the second-order difference tensor and the third-order difference tensor. Based on the time step of the third-order difference tensor, the first-order difference tensor and the second-order difference tensor are truncated to obtain aligned first-order difference tensors and aligned second-order difference tensors. The first-order difference tensor, the second-order difference tensor, and the third-order difference tensor are concatenated along the feature dimension to obtain the trend-enhancing feature tensor. In the joint channel guidance module, structural parameter feature vectors and geological environment parameter vectors are obtained and concatenated along the feature dimension to obtain a joint guidance vector; The joint guiding vector is linearly mapped to the channel mapping bias vector through a trainable channel mapping matrix, and then normalized using the Sigmoid activation function to obtain the channel weight vector. The channel weight vector and the trend enhancement feature tensor are multiplied element-wise along the feature dimension to obtain the channel guidance feature tensor. The time channel hierarchical modeling module includes an L-layer TimeMixer hybrid module, where L is a positive integer greater than or equal to 4. The time channel hierarchical modeling module adopts a hierarchical sparse strategy to generate channel-enhanced feature tensors. In the state generation module, the channel enhancement feature tensor is averaged in the time dimension to obtain the pile foundation detection state vector; Step 4: Based on the pile foundation detection state vector, construct a reinforcement learning environment and initialize an improved SAC model, which includes a dual-scale policy network, a dual Q-value network, and a target Q-value network; The dual Q-value network includes a first Q-value network and a second Q-value network. Both the first Q-value network and the second Q-value network are composed of a four-layer perceptron structure. The first and second layers are fully connected layers, the third layer is a ReLU activation layer, and the fourth layer is a linear regression layer. The target Q-value network includes a first target Q-value network and a second target Q-value network, and the target Q-value network has the same structure as the dual Q-value network. Step 5: Execute strategy training based on the improved SAC model, and input the pile foundation detection state vector of the target photovoltaic support pile foundation into the trained improved SAC model to output the predicted bearing capacity; During the policy training process, the pile foundation detection state vector of the current sample number is input into the dual-scale policy network. The short-term policy distribution is generated through the fast policy network, and the long-term policy distribution is generated through the slow policy network. The short-term policy distribution and the long-term policy distribution are weighted and fused to generate the fused policy distribution. Based on the fusion strategy distribution, the pile foundation detection state vector of the current sample number is sampled for action, and the pile foundation detection state vector and the sampled action are concatenated to form a state-action pair. The state-action pair is input into a dual Q-value network. The first Q-value is calculated and generated through the first Q-value network, and the second Q-value is calculated and generated through the second Q-value network. The minimum value between the first Q-value and the second Q-value is selected as the estimated Q-value. The state action pair of the next sample number is input into the target Q-value network. The first target Q-value is generated by the first target Q-value network, and the second target Q-value is generated by the second target Q-value network. The minimum value between the first target Q-value and the second target Q-value is selected as the candidate target Q-value. Step 6: Apply threshold classification to the predicted bearing capacity to obtain safety level labels; Step 7: Based on the predicted bearing capacity and safety level label, generate the pull-out bearing capacity test results of the target photovoltaic support pile foundation.

2. The method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to claim 1, characterized in that, The structural parameters of the pile include pile diameter, pile length, and depth of penetration into the soil; The geological environmental parameters include the main soil layer type, natural moisture content, penetration resistance coefficient, and relative density at the location of the target photovoltaic support pile foundation.

3. The method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to claim 1, characterized in that, Step two specifically includes: The loading force sequence and pile top displacement sequence are time-aligned according to the set time step order. The Z-score standardization method is used to filter and remove abnormal data in the loading force sequence and pile top displacement sequence. The missing data in the loading force sequence and pile top displacement sequence are filled by linear interpolation method to obtain the standard loading force sequence and standard pile top displacement sequence. The structural parameters of the pile are subjected to minimum-maximum normalization to obtain the structural parameter feature vector. The main soil layer category in the geological environment parameters is transformed into a soil layer category vector through a trainable embedding mapping matrix. The natural water content, penetration resistance coefficient and relative density are subjected to minimum-maximum normalization and then arranged into a geological environment parameter vector in sequence. The standard loading force sequence and the standard pile top displacement sequence are subjected to fast Fourier transform to obtain the complex spectrum of the loading force in the frequency domain and the complex spectrum of the pile top displacement in the frequency domain. Based on the complex spectrum of the loading force in the frequency domain and the complex spectrum of the pile top displacement in the frequency domain, the dominant frequency of the loading force, the maximum amplitude frequency of the loading force, the total power of the loading force, the dominant frequency of the pile top displacement, the maximum amplitude frequency of the pile top displacement, and the total power of the pile top displacement are extracted and spliced ​​together to form a frequency domain response feature sequence. The standard load force sequence, standard pile top displacement sequence, structural parameter feature vector, geological environment parameter vector, and frequency domain response feature sequence are dimensionally aligned and feature-stitched to obtain the standard pull-out test feature tensor.

4. The method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to claim 1, characterized in that, The time channel hierarchical modeling module employs a hierarchical sparse strategy to generate channel-enhanced feature tensors, specifically including: In the first and second layers of the TimeMixer module, the channel feature vectors of each channel of the channel-guided feature tensor are sparsely modulated to obtain a sparse activation feature tensor. Then, the sparse activation feature tensor is subjected to time mixing and channel mixing operations to output a channel-mixed feature tensor. Specifically: The logits vector is transformed into a linear mapping, and the logits vector has a dimension of 2, which respectively represent the score of the current channel and the score of the current channel. The logits vector is transformed into a channel gating probability vector using the Gumbel Softmax activation function. The channel gating probability vector has a dimension of 2, representing the probability of retaining the current channel and the probability of discarding the current channel, respectively. The probability of retaining the current channel is extracted as the sparse activation weight of the current channel. The sparse activation weights are used to sparsely modulate the channel feature vectors to obtain sparse feature vectors, and the sparse feature vectors of all channels are combined into a sparse activation feature tensor. The sparse activation feature tensor is subjected to layer normalization and then input into a multilayer perceptron in the time dimension for temporal hybrid modeling to obtain a temporal modeling feature tensor. The temporal modeling feature tensor is then residually connected with the sparse activation feature tensor to obtain a temporal hybrid feature tensor. The temporal fusion feature tensor is subjected to layer normalization and then input into a multilayer perceptron in the channel dimension for channel fusion modeling to obtain the channel modeling feature tensor. The channel modeling feature tensor is then residually connected with the temporal fusion feature tensor to obtain the channel fusion feature tensor. In the third layer and the subsequent TimeMixer mixing module, sparse modulation is canceled, and the channel mixing feature tensor is subjected to time mixing and channel mixing operations to output the channel enhanced feature tensor.

5. The method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to claim 1, characterized in that, Step four specifically includes: The pile foundation detection state vectors are arranged into the state space of the reinforcement learning environment according to the set sample number order. The action space of the reinforcement learning environment is set up, the action space includes several actions, and each action corresponds to a set of load-bearing capacity prediction parameters; A reward function is constructed, which is a weighted combination of a carrying capacity prediction error term, an action change penalty term, and a strategy distribution entropy term. The load-bearing capacity prediction error term is the negative absolute difference between the actual load-bearing capacity and the predicted load-bearing capacity corresponding to the action; the action change penalty term is the negative Euclidean distance between the action difference between the current sample number and the previous sample number; the policy distribution entropy term is the negative information entropy of the fused policy distribution. In a reinforcement learning environment, initialize the improved SAC model; The dual-scale policy network includes a fast policy network and a slow policy network. The fast policy network consists of three fully connected layers, each followed by a ReLU activation function. The slow policy network adopts a time-series modeling structure based on a sliding window, performs one-dimensional convolution on the pile foundation detection state vector within a set sliding window interval, and extracts cross-sample state evolution features through a gated recurrent unit.

6. The method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to claim 1, characterized in that, Step five specifically includes: The sampling action of the current sample number is used to calculate the predicted bearing capacity through a single MLP structure. The actual bearing capacity, the predicted bearing capacity, the sampling action of the current sample number and the previous sample number, and the fusion strategy distribution are substituted into the reward function to calculate the reward value of the current sample number. The target Q value is obtained by weighting the reward value of the current sample number, the candidate target Q value, and the negative log probability of the fusion strategy distribution of the next sample number. Calculate the mean square error between the estimated Q value and the target Q value to obtain the Critic loss function, and update the dual Q-value network by gradient backpropagation by minimizing the Critic loss function. Calculate the expected difference between the negative log probability of the fusion policy distribution for the current sample number and the estimated Q value to obtain the policy loss function, and perform gradient update on the dual-scale policy network by minimizing the policy loss function; The target Q-value network is updated using a moving average method; When the number of training rounds reaches the set maximum number of rounds, stop policy training and obtain the improved SAC model that has been trained. The pile foundation detection state vector of the target photovoltaic support pile foundation is input into the trained improved SAC model, and the predicted bearing capacity is output.

7. The method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to claim 1, characterized in that, Step six specifically includes: Set bearing capacity thresholds T1 and T2, wherein bearing capacity threshold T1 is greater than bearing capacity threshold T2; If the predicted bearing capacity is greater than or equal to the bearing capacity threshold T1, the safety level label is safe. If the predicted bearing capacity is greater than or equal to the bearing capacity threshold T2 and less than the bearing capacity threshold T1, then the safety level label is a warning. If the predicted bearing capacity is less than the bearing capacity threshold T2, the safety level label is dangerous.

8. The method for detecting the pull-out bearing capacity of photovoltaic support pile foundations based on reinforcement learning according to claim 1, characterized in that, The results of the pull-out bearing capacity test include the pile number, predicted bearing capacity, safety level label, and test time.

Citation Information

Patent Citations

  • Method and device for evaluating distributed photovoltaic bearing capacity of power distribution network in real time

    CN119578702A

  • Simulation apparatus of wire electric discharge machine having function of determining welding positions of core using machine learning

    US20170151618A1