A data processing method, device, and method for constructing a digital twin model of a crane
The improved PCA method for crane data processing integrates diverse data formats and frequencies, enhancing data quality and representation for accurate digital twin models, supporting real-time monitoring and predictive maintenance.
Patent Information
- Application Number
- CN202510443929.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional methods are difficult to effectively integrate and process multi-source heterogeneous data of cranes, resulting in compatibility issues and waste of resources during data storage and analysis, and traditional PCA methods cannot effectively handle nonlinear relationships in crane data.
The improved PCA method is adopted to select kernel functions to feature map the integrated data, implicitly project data into high-dimensional feature space, calculate the covariance matrix and perform feature decomposition, select feature vectors with a large contribution rate for dimensionality reduction, and monitor the data distribution changes in real time for kernel function and parameter optimization.
It realizes the deep fusion of multi-source heterogeneous data, retains key information, improves the quality and representativeness of data after dimensionality reduction, is suitable for the construction of complex crane digital twin models, and optimizes model parameters through a closed-loop feedback mechanism to ensure the accuracy and adaptability of the model.
Smart Images

Figure CN119961664B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of cranes, and particularly to a data processing method and a method for constructing a crane digital twin model. Background Art
[0002] As an indispensable heavy machinery in industrial production, the real-time monitoring of the operating status of cranes is crucial for ensuring production safety and improving operation efficiency. With the wide application of Internet of Things technology, the operating data of cranes can be collected in real time by sensors and transmitted to the cloud for analysis and processing. However, due to the large number of sensors, the collected data often has high-dimensional characteristics, which not only increases the difficulty of data storage and processing, but also may lead to inaccurate analysis results. Therefore, researching effective data dimensionality reduction methods has important practical significance for the analysis of crane Internet of Things data.
[0003] During the construction process of a crane digital twin system, various types of data are involved, including geometric data, physical property data, operating status data, etc. These data are respectively sourced from different sensors, systems or databases, and have different data formats, precisions and update frequencies. Traditional data processing methods are difficult to efficiently integrate them, resulting in compatibility problems and resource waste during data transmission, storage and analysis. For example, geometric data may be stored in the CAD model format, while operating status data is collected in real time by sensors and exists in the form of time series data. Unifying these data for digital twin model construction faces huge challenges. Summary of the Invention
[0004] On the one hand, this application provides a data processing method, which has the advantage of being able to effectively process the non-linear relationships in crane data, perform dimensionality reduction operations in a high-dimensional feature space, better retain the key information of the data, improve the quality and representativeness of the data after dimensionality reduction, and make the dimensionality reduction result more suitable for the construction of complex crane digital twin models.
[0005] The technical solution is as follows: A method for processing crane digital twin data based on improved PCA, including the following steps:
[0006] S1 Data collection and preprocessing: Collect the operating status data of the crane and perform preprocessing;
[0007] S2 Data integration: Correlate the collected different types of data according to the predefined crane data model and semantic rules;
[0008] S3 Data dimensionality reduction: Perform dimensionality reduction operations on the integrated data using improved PCA, including the following steps:
[0009] Select a kernel function to perform feature mapping on the integrated data, and map the original data into the high-dimensional feature space F;
[0010] Calculate the covariance matrix C in the high-dimensional feature space F, and perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues and eigenvectors ;
[0011] Calculate the contribution rate of the eigenvalues, select the first K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P;
[0012] Project the original data onto the principal components to achieve dimensionality reduction.
[0013] Furthermore, in step S1, the operating status data of the crane is collected by deploying sensors on the crane and connecting the design database and operation management system of the crane. The sensors include displacement sensors, pressure sensors, and temperature sensors.
[0014] Furthermore, in step S1, the preprocessing of the crane operating status data includes:
[0015] Check and clean the format of the operating status data, and remove outliers and error data;
[0016] For geometric data, extract relevant model information from the design database, and perform simplification and feature extraction.
[0017] Furthermore, in step S2, using the timestamp as the key index, match the operating status data at the same moment with the corresponding geometric data and physical performance data; calibrate the spatial position information of the data to ensure the spatial consistency of different data.
[0018] Furthermore, in step S3, select the kernel function according to the following steps:
[0019] S301: First, calculate the complexity index of the data. The complexity index is the non-linear index, distribution entropy, and eigenvalue decay rate;
[0020] (1) Calculate the non-linear index of the data:
[0021] ;
[0022] : Original data points;
[0023] : The projection matrix of the first k principal components extracted by PCA in the local neighborhood of the point ;
[0024] : Euclidean norm;
[0025] (2) Calculate the distribution entropy of the data:
[0026] ;
[0027] ;
[0028] : Gaussian kernel function, ;
[0029] h: Bandwidth parameter, determined by the Silverman criterion: , being the standard deviation of the data;
[0030] (3) Calculate the eigenvalue decay rate of the data:
[0031] Normalize the eigenvalues of the original data covariance matrix, and calculate the number of principal components required for the first 95% cumulative contribution rate. The calculation formula is:
[0032] ;
[0033] Perform eigenvalue decomposition on the original data covariance matrix, and normalize the eigenvalues ;
[0034] Calculate the cumulative contribution rate of the first k principal components: ;
[0035] S302: Then, according to the non - linear index, distribution entropy, and eigenvalue decay rate of the data, select the kernel function in the following way:
[0036] If the non - linear index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, select the Gaussian kernel function;
[0037] If the eigenvalue decay rate of the data is less than the preset threshold γ, select the polynomial kernel function, and determine the optimal order d through cross - validation;
[0038] If neither of the above conditions is met, adopt the mixed kernel function, and the mixed kernel function is a weighted combination of the Gaussian kernel and the linear kernel;
[0039] The decision rule formula for kernel function selection is:
[0040] Gaussian kernel RBF: ;
[0041] Polynomial kernel: ;
[0042] Mixed kernel: ;
[0043] , is the weight.
[0044] Furthermore, (1) optimize the Gaussian kernel bandwidth σ in the Gaussian kernel function:
[0045] The objective function is to maximize the between-class variance ratio:
[0046] ;
[0047] represents the trace of the matrix;
[0048] represents the sum of between-class differences, represents the sum of within-class differences;
[0049] : between-class scatter matrix:
[0050] ;
[0051] where c is the number of classes, is the number of samples in the i-th class, is the mean vector of the i-th class, and μ is the mean vector of all data;
[0052] : within-class scatter matrix:
[0053] ;
[0054] where C i represents the sample set of the i-th class;
[0055] Use the Bayesian optimization method for optimization;
[0056] (2) Adjust the order d of the polynomial kernel function:
[0057] Calculate the mutual information between dimensions :
[0058] ;
[0059] If there exists , and δ is the threshold, then .
[0060] Furthermore, in step S3, monitor the change of data distribution in real time: recalculate the data complexity index every interval of time T. If the index exceeds the preset range, trigger the re-optimization of the kernel function type and parameters; update the projection matrix P online.
[0061] Furthermore, in step S3, for the original data x i , the selected kernel function , the covariance matrix C is:
[0062]
[0063] where m is the number of data samples;
[0064] The contribution rate of the eigenvalue is:
[0065]
[0066] where n is the total number of eigenvalues;
[0067] Combine the selected K eigenvectors into a projection matrix:
[0068]
[0069] where K is the number of eigenvector groups;
[0070] The data after dimensionality reduction is:
[0071] .
[0072] On the other hand, the present application provides a crane digital twin data processing device based on improved PCA, including:
[0073] A data acquisition module for collecting the operating state data of the crane and performing preprocessing;
[0074] A data integration module for associating different types of collected data according to a predefined crane data model and semantic rules;
[0075] A data dimensionality reduction module for performing dimensionality reduction operations on the integrated data using improved PCA, including the following steps:
[0076] Select a kernel function to perform feature mapping on the integrated data, and map the original data to a high-dimensional feature space F;
[0077] Calculate the covariance matrix C in the high-dimensional feature space F, perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues and eigenvectors ;
[0078] Calculate the contribution rate of the eigenvalues, select the top K eigenvectors with larger contribution rates as the principal components, and construct a projection matrix P;
[0079] Project the original data onto the principal components to achieve dimensionality reduction.
[0080] In another aspect, the present application provides a data processing terminal, including a memory and a processor. The memory stores a computer program. When the processor calls and executes the computer program, the method for processing crane digital twin data based on improved PCA as described above is implemented.
[0081] In another aspect, the present application provides a computer-readable medium. The computer-readable medium stores a computer program. When the computer program is called and executed by a computer, the method for processing crane digital twin data based on improved PCA as described above is implemented.
[0082] In another aspect, the present application provides a method for constructing a crane digital twin model. The crane digital twin model is constructed by using the data obtained by the method for processing crane digital twin data based on improved PCA as described above; the constructed crane digital twin model is optimized and calibrated by using the actual crane operation data, and the model parameters are adjusted by using a machine learning algorithm.
[0083] In summary, the beneficial effects of the present application are as follows:
[0084] 1. The method for processing crane digital twin data based on improved PCA provided by the present application realizes the deep fusion of multi-source heterogeneous data such as geometry, physics, and operation. This fusion method not only considers the data format and content, but also fully utilizes the spatio-temporal attributes and semantic information of the data to ensure the integrity and accuracy of the data during the integration process, providing comprehensive data support for the digital twin model.
[0085] 2. The method for processing crane digital twin data based on improved PCA provided by the present application introduces a kernel function in PCA, breaking through the linear limitation of traditional PCA and being able to effectively process the non-linear relationships in crane data. The dimensionality reduction operation is carried out in the high-dimensional feature space, better retaining the key information of the data, improving the quality and representativeness of the data after dimensionality reduction, and making the dimensionality reduction result more suitable for the construction of complex crane digital twin models.
[0086] 3. The method for constructing a crane digital twin model provided by the present application constructs a closed-loop feedback mechanism between the digital twin model and the actual crane operation data, and can continuously optimize the model parameters and structure according to the actual operation situation. Through continuous learning and improvement, the digital twin model can more accurately reflect the real state of the crane, predict potential faults in advance, and provide strong guarantee for the safe operation and maintenance of the crane. Description of the Drawings
[0087] Figure 1 Steps Schematic of a Method for Processing Crane Digital Twin Data Based on Improved PCA in the Present Application Detailed Embodiments
[0088] The specific implementation manners of the present application will be described in detail below in conjunction with the accompanying drawings.
[0089] A specific embodiment of the present application provides a method for processing digital twin data of a crane based on improved PCA, as Figure 1 follows:
[0090] S1 Data acquisition and preprocessing: Acquire the operating state data of the crane and perform preprocessing;
[0091] In step S1, the operating state data of the crane is acquired by deploying sensors on the crane and connecting the design database and operation management system of the crane. The sensors include displacement sensors, pressure sensors, and temperature sensors. The sensors collect the operating state data in real time and transmit it to the data acquisition module at a certain frequency.
[0092] In step S1, the preprocessing of the operating state data of the crane includes:
[0093] Perform format check and cleaning on the operating state data to remove outliers and error data;
[0094] For geometric data, extract relevant model information from the design database, simplify and extract features, and convert it into a format suitable for subsequent processing. For example, simplify the 3D CAD model into key geometric dimension and shape feature vectors.
[0095] S2 Data integration: Correlate the collected different types of data according to the predefined crane data model and semantic rules;
[0096] In step S2, using the timestamp as the key index, match the operating state data at the same moment with the corresponding geometric data and physical property data; for example, when the crane lifts a heavy object at a certain moment, integrate the lifting weight (operating state data) at that moment with the geometric structure data of the crane boom and the physical property data of the material (such as strength, stiffness). At the same time, calibrate the spatial position information of the data to ensure the spatial consistency of different data.
[0097] S3 Data dimensionality reduction: Perform dimensionality reduction operations on the integrated data using improved PCA, including the following steps:
[0098] Select a kernel function to perform feature mapping on the integrated data, and map the original data to the high-dimensional feature space F.
[0099] In the processing of digital twin data of the crane, the selection of the kernel function needs to comprehensively consider data characteristics, computational complexity, and model construction requirements. In this embodiment, the kernel function is a Gaussian kernel function or a polynomial kernel function.
[0100] If there are complex non - linear relationships in crane data and the distribution is relatively complex, the Gaussian kernel function is a good choice. During the operation of the crane, the relationship between its physical performance data (such as stress and strain) and operating state data (such as speed and acceleration) may be non - linear, and this relationship is difficult to describe with a simple mathematical model. The Gaussian kernel function can map the data into a high - dimensional feature space, effectively handle this complex non - linear relationship, making the features of the data more obvious in the high - dimensional space, facilitating subsequent dimensionality reduction operations and model construction. In addition, the Gaussian kernel function has strong adaptability to data, can handle noise and outliers in the data, and improve the stability of data processing.
[0101] When there are polynomial relationships between some dimensions in crane data, the polynomial kernel function is more appropriate. There may be a certain polynomial - form association between the geometric dimension data and the operating state data of the crane. The polynomial kernel function can capture this relationship by adjusting the order, and can, to a certain extent, explore the deep - seated connections between data. Moreover, the polynomial kernel function is relatively simple in calculation, with a low computational complexity. In the case of a large amount of data, it can improve the efficiency of data processing.
[0102] Specifically, in step S3, the kernel function is selected according to the following steps:
[0103] S301: First, calculate the complexity index of the data. The complexity index is the non - linearity index, distribution entropy, and eigenvalue decay rate;
[0104] (1) Calculate the non - linearity index (NonlinearityIndex, NLI) of the data:
[0105] ;
[0106] : The original data points;
[0107] : At point The projection matrix of the first k principal components extracted by PCA on the local neighborhood (such as the 10% nearest - neighbor samples) of;
[0108] : Euclidean norm;
[0109] By calculating the local linear reconstruction error of the original data under low - dimensional projection, the non - linear degree of the data is evaluated. The larger the error, the stronger the non - linearity.
[0110] (2) Calculate the distribution entropy (DistributionEntropy, H) of the data:
[0111] ;
[0112] ;
[0113] : Gaussian kernel function, ;
[0114] h: Bandwidth parameter, determined by Silverman's criterion: , is the standard deviation of the data;
[0115] Use kernel density estimation to calculate the entropy value of the data distribution. The higher the entropy value, the more complex the data distribution.
[0116] (3) Calculate the data eigenvalue decay rate:
[0117] Normalize the eigenvalues of the original data covariance matrix, and calculate the number of principal components required for the first 95% cumulative contribution rate. The calculation formula is:
[0118] ;
[0119] Perform eigen decomposition on the original data covariance matrix, and normalize the eigenvalues ;
[0120] Calculate the cumulative contribution rate of the first k principal components: ;
[0121] Normalize the eigenvalues of the original data covariance matrix, and calculate the number of principal components required for the first 95% cumulative contribution rate. The fewer the number, the stronger the linear separability of the data. That is, the smaller the EDR, the closer the data is to linear separability.
[0122] S302: Then, according to the non - linear index, distribution entropy and eigenvalue decay rate of the data, select the kernel function in the following way:
[0123] If the non - linear index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, select the Gaussian kernel function;
[0124] If the eigenvalue decay rate of the data is less than the preset threshold γ, select the polynomial kernel function, and determine the optimal order d through cross - validation;
[0125] If none of the above conditions are met, adopt the mixed kernel function, which is a weighted combination of the Gaussian kernel and the linear kernel;
[0126] The decision rule formula for kernel function selection is:
[0127] Gaussian kernel RBF: ;
[0128] Polynomial kernel: ;
[0129] Mixed kernel: ;
[0130] , is the weight, which is dynamically allocated based on the information retention rate of the validation set (such as the classification accuracy of the data after dimensionality reduction). Typically, the threshold = 0.7, = 2.0, = 0.3 (determined based on experimental statistics). The polynomial order d is selected to be the optimal value through cross-validation (such as ).
[0131] Optimize the kernel function parameters. For the Gaussian kernel, optimize the bandwidth parameter σ by maximizing the ratio of between-class variance (the class separability of the data after dimensionality reduction), and use the Bayesian optimization algorithm to search for the optimal value. For the polynomial kernel, adaptively adjust the order d based on the feature interaction intensity (calculate the correlation between dimensions through mutual information), ensuring that the high-order terms are only used for strongly correlated dimensions to avoid overfitting.
[0132] (1) Optimize the Gaussian kernel bandwidth σ in the Gaussian kernel function:
[0133] The objective function is to maximize the ratio of between-class variance:
[0134] ;
[0135] represents the trace of the matrix, that is, the sum of the diagonal elements of the matrix;
[0136] represents the sum of between-class differences, represents the sum of within-class differences; by maximizing the ratio of between-class differences to within-class differences, the data after dimensionality reduction is more separable between classes, while reducing the dispersion within classes. This metric is commonly used in supervised dimensionality reduction tasks to improve the performance of classification models.
[0137] : Between-class scatter matrix:
[0138] ;
[0139] where c is the number of classes, is the number of samples in the i-th class, is the mean vector of the i-th class, and μ is the mean vector of all data;
[0140] : Within-class scatter matrix:
[0141] ;
[0142] Among them, C i represents the sample set of the i-th class;
[0143] The Bayesian optimization method is used for optimization, and iterative search .
[0144] It should be noted that in the supervised dimensionality reduction scenario of Gaussian kernel PCA (or kernel method): The data is implicitly mapped to the high-dimensional feature space F through the Gaussian kernel . The between-class and within-class scatter matrices S B and S w are calculated in the high-dimensional feature space F, and the structure of F is determined by the kernel parameter σ. If σ is too small, the data may be too scattered in the high-dimensional space (the within-class difference increases). If σ is too large, the data may be over-smoothed (the between-class difference decreases).
[0145] The trace Tr(S B ) and Tr(S w ) will change with the change of σ. Therefore, it is necessary to optimize σ to maximize the between-class separability.
[0146] Adjust σ through Bayesian optimization, and its core steps include:
[0147] 1. Initialization: Select an initial σ and calculate the kernel matrix K(σ).
[0148] 2. Scatter matrix calculation: Based on K(σ) and class labels, indirectly obtain Tr(S B (σ)) and Tr(S w (σ)).
[0149] 3. Objective function evaluation: Calculate Tr(S B ) / Tr(S w ) as the optimization objective.
[0150] 4. Iterative update: Adjust σ through Bayesian optimization, and repeat steps 2 and 3 until convergence.
[0151] The between-class and within-class scatter matrices S B and S w are functions of σ because their calculations are based on the high-dimensional space after kernel mapping, and the kernel mapping is controlled by σ.
[0152] Application of the kernel trick: In practice, there is no need to explicitly calculate , but indirectly express the trace of the scatter matrix through the kernel matrix.
[0153] Optimization significance: Adjust σ to maximize the between-class difference and minimize the within-class difference, thereby improving the classification performance after dimensionality reduction.
[0154] (2) Adjust the order d of the polynomial kernel function:
[0155] Determine the feature interaction strength: Calculate the mutual information between dimensions :
[0156] ;
[0157] Adaptive rule: If there exists , where δ is a threshold, then . δ is the threshold for determining whether the mutual information is "significant", based on the mutual information strength between dimensions. During specific implementation, experimental tuning is required, and the specific value of δ needs to be determined through cross-validation or the effect of the actual task.
[0158] Dynamic feedback adjustment: During the operation of the digital twin model, monitor the change of data distribution in real time: Recalculate the data complexity index every time interval T. If the index exceeds the preset range, trigger the re-optimization of the kernel function type and parameters; Update the projection matrix P through online learning to ensure that the dimensionality reduction model continuously adapts to the dynamic changes of the crane operation state.
[0159] (1) Real-time monitoring and trigger condition:
[0160] Recalculate NLI, H, and EDR every time interval T (such as 1 hour). If any index changes by more than 10% (such as ), trigger re-optimization.
[0161] (2) Online learning to update the projection matrix:
[0162] Update the projection matrix P using incremental PCA (IPCA):
[0163] .
[0164] Typically, taking the port gantry crane as an example, the initial data analysis is as follows:
[0165] NLI = 0.85, H = 2.3, EDR = 0.25.
[0166] Trigger the selection of the Gaussian kernel, and Bayesian optimization gives σ = 1.2.
[0167] The classification accuracy after dimensionality reduction is increased by 12% (from 78% to 90%).
[0168] Dynamic adjustment during operation: After 3 months, due to structural wear, H drops to 1.8, triggering a hybrid kernel ( = 0.6, = 0.4).
[0169] After updating the projection matrix, the prediction error remains within 3%. Through the above method, this application clarifies the mathematical logic of kernel function selection and parameter optimization, and combines the dynamic feedback mechanism to significantly improve the technical effect.
[0170] The above kernel function determination method has the following effects:
[0171] 1. Data-driven kernel function decision-making mechanism;
[0172] Bind the data complexity quantification indicators (nonlinear index, distribution entropy, eigenvalue decay rate) to the kernel function selection, break through the limitations of traditional empirical selection or fixed kernel functions, and significantly improve the self-adaptability of the dimensionality reduction process.
[0173] 2. Decoupling of parameter optimization and dimensionality correlation;
[0174] For polynomial kernels, propose an order adjustment method based on the feature interaction intensity to avoid dimensional redundancy caused by a globally unified order, reduce the computational overhead, and improve the model generalization ability while doing so.
[0175] 3. Online dynamic adjustment ability;
[0176] Through a real-time monitoring and feedback mechanism, the kernel function and parameters can be automatically adjusted according to changes in the crane operating environment (such as sudden load changes, structural fatigue) to ensure the long-term accuracy of the digital twin model.
[0177] Calculate the covariance matrix C in the high-dimensional feature space F. For the original data x i , the selected kernel function , the covariance matrix C is:
[0178]
[0179] where m is the number of data samples.
[0180] Perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues and eigenvectors. Calculate the contribution rate of the eigenvalues. The contribution rate of the eigenvalues is:
[0181]
[0182] where n is the total number of eigenvalues.
[0183] Select the first K eigenvectors with larger contribution rates as the principal components, construct a projection matrix P, and combine the selected K eigenvectors into a projection matrix:
[0184]
[0185] where K is the number of eigenvector groups.
[0186] Project the original data onto the principal components to achieve dimensionality reduction. The dimensionality-reduced data is :
[0187] .
[0188] This embodiment further illustrates the present solution with an application example of data processing for large port cranes.
[0189] Overview of crane application scenarios:
[0190] Select a large port gantry crane equipped with a variety of sensors, including a high-precision displacement sensor for measuring the lifting height of the spreader, an amplitude sensor for monitoring the horizontal movement amplitude of the spreader, a speed sensor for detecting the lifting and running speeds, and an acceleration sensor for capturing the acceleration changes during the movement process. These sensors collect the operation status data in real time. Strain gauges are installed at key parts of the crane structure to obtain physical property data such as structural stress and strain. At the same time, environmental monitoring equipment collects data such as temperature, humidity, wind speed, and wind direction of the surrounding environment. In addition, the control system of the crane records control-related information such as operation instructions and operation modes, and detailed geometric dimension data such as boom length and bridge structure dimensions are stored in its design database.
[0191] Data acquisition and preliminary processing:
[0192] Acquisition and processing of operation status data: The displacement sensor collects the lifting height data every 0.1 second. During a typical loading and unloading operation process, 6000 data points are collected within 10 minutes, and the data range is between 0 - 30 meters. After preliminary screening, obvious abnormal mutation values (such as data exceeding the normal lifting height range by ±5 meters) are removed. The speed sensor and the acceleration sensor collect data synchronously. The speed data range is between 0 - 2m / s, and the acceleration data is between -2 - 2m / s². Similarly, outlier rejection and data smoothing processing are performed, and the moving average method is used to calculate the average of adjacent 5 data points to reduce noise interference.
[0193] Acquisition and processing of physical property data: The stress and strain data collected by the strain gauges are updated every 1 second. 600 data points are obtained within 10 minutes. The stress data range is between 0 - 500MPa, and the strain data is normalized so that its value is between 0 - 1 for unified analysis with other data.
[0194] Acquisition and processing of environmental data: The environmental monitoring equipment collects temperature (10 - 40°C), humidity (30% - 80%), wind speed (0 - 15m / s), and wind direction (0 - 360°) data every minute, and a total of 10 groups of data are collected within 10 minutes. For the wind speed and wind direction data, they are converted into horizontal and vertical components according to trigonometric relationships for subsequent combination with the force analysis of the crane.
[0195] Geometry and Control Data Extraction and Conversion: Extract key geometric dimensions such as the boom length of 50 meters and the bridge width of 20 meters from the design database, and convert them into dimensionless relative ratio data, such as the ratio of the boom length to the bridge width is 2.5. Classify and code the operation instructions and operation mode data of the control system. For example, different lifting operation instructions are divided into 5 categories, and the operation modes are divided into 3 types, which are represented by the numbers 1-5 and 1-3 respectively.
[0196] Data Integration Process:
[0197] Integration Based on Timestamp: Integrate the operation status data, physical performance data, environmental data, and corresponding control data collected within the same minute at 1-minute time intervals. For example, at the 5th minute, data such as the lifting height, speed, acceleration, stress and strain, temperature, humidity, wind speed and direction components, as well as the current operation instructions and operation modes at that moment are combined into a data record.
[0198] Association Based on Spatial Location: According to the structural layout of the crane, associate the physical performance data collected by strain gauges with the geometric data of the corresponding structural components. For example, the data collected by the strain gauges at the root of the boom is associated with the geometric dimensions and material properties of the boom, establishing a spatial correspondence relationship between the data for analyzing the interaction between structural forces and geometric structures.
[0199] Data Dimensionality Reduction Operations:
[0200] (1) Kernel Function Selection and Feature Mapping: Select the polynomial kernel function (where d = 3) to perform feature mapping on the integrated data. Map the original data vector containing multi-dimensional data such as lifting height, speed, acceleration, stress and strain, environmental data, geometric ratio, and control coding to the high-dimensional feature space F.
[0201] (2) Covariance Matrix Calculation and Eigenvalue Decomposition: Calculate the covariance matrix C in the high-dimensional feature space. Assuming there are n = 60 integrated data samples within 10 minutes, the calculation formula is . Perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues and eigenvectors .
[0202] (3) Principal Component Selection and Dimensionality Reduction Calculation: Calculate the contribution rate of the eigenvalues , select the first k eigenvectors whose cumulative contribution rate reaches 90% as the principal components to construct the projection matrix P. Finally, project the original data onto the principal components to achieve dimensionality reduction. The low-dimensional data can be used for the subsequent construction and analysis of the crane digital twin model, effectively reducing the dimensionality and complexity of the data while retaining key information.
[0203] Through the above example applications, the actual operation process and effects of the patented method in complex crane data processing are demonstrated, providing strong data support for the intelligent management and digital twin applications of cranes.
[0204] Another specific embodiment of the present application provides a crane digital twin data processing device based on improved PCA, including:
[0205] A data acquisition module for collecting the operating status data of the crane and performing preprocessing;
[0206] A data integration module for correlating different types of collected data according to a predefined crane data model and semantic rules;
[0207] A data dimensionality reduction module for performing dimensionality reduction operations on the integrated data using improved PCA, including the following steps:
[0208] Select a kernel function to perform feature mapping on the integrated data, and map the original data into a high-dimensional feature space F;
[0209] Calculate the covariance matrix C in the high-dimensional feature space F, and perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues and eigenvectors ;
[0210] Calculate the contribution rate of the eigenvalues, select the first K eigenvectors with larger contribution rates as the principal components, and construct a projection matrix P;
[0211] Project the original data onto the principal components to achieve dimensionality reduction.
[0212] Another specific embodiment of the present application provides a data processing terminal, including a memory and a processor. When the processor calls and executes the computer program stored in the memory, the above-mentioned crane digital twin data processing method based on improved PCA is implemented.
[0213] Another specific embodiment of the present application provides a computer-readable medium storing a computer program, which, when called and executed by a computer, implements the above-mentioned crane digital twin data processing method based on improved PCA.
[0214] Another specific embodiment of the present application provides a method for constructing a crane digital twin model. The crane digital twin model is constructed using the data obtained by the above-mentioned crane digital twin data processing method based on improved PCA; the constructed crane digital twin model is optimized and calibrated using the actual crane operation data, and the model parameters are adjusted using machine learning algorithms.
[0215] Using the data after dimensionality reduction as input, a digital twin model of the crane is constructed by adopting appropriate modeling methods (such as physics-based modeling, machine learning modeling, etc.). For example, using the neural network algorithm in machine learning, with the dimensionality-reduced operating state data and geometric and physical feature data as the input layer and the performance indicators of the crane (such as lifting capacity, stability, etc.) as the output layer, a prediction model is constructed. By comparing the model output with the actual crane operation data, error metrics (such as mean square error, absolute error, etc.) are calculated. According to the error feedback, optimization methods such as the backpropagation algorithm are used to adjust the model parameters, continuously optimize the digital twin model, and improve its accuracy and reliability.
[0216] Through the above system and method, the effective integration and dimensionality reduction processing of different types of data in the crane digital twin technology are realized, providing a solid data foundation and technical support for the intelligent monitoring, fault prediction, and optimized operation of the crane. In practical applications, the system parameters and algorithms can be appropriately adjusted and optimized according to the models, working environments, and application requirements of different cranes to achieve the best processing effect.
[0217] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A data processing method, characterized in that, It includes the following steps: S1 Data acquisition and preprocessing: Acquire the operating status data of the crane and perform preprocessing; S2 Data integration: According to the predefined crane data model and semantic rules, correlate different types of acquired data; Using the timestamp as the key index, match the operating status data at the same moment with the corresponding geometric data and physical performance data; Calibrate the spatial position information of the data to ensure the spatial consistency of different data; S3 Data dimensionality reduction: Perform dimensionality reduction on the integrated data using improved PCA, including the following steps: Select a kernel function to perform feature mapping on the integrated data, and map the original data to the high-dimensional feature space F; Calculate the covariance matrix C in the high-dimensional feature space F, and perform eigen-decomposition on the covariance matrix C to obtain the eigenvalues and eigenvectors ; Calculate the contribution rate of the eigenvalues, select the top K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P; Project the original data onto the principal components to achieve dimensionality reduction; Select the kernel function according to the following steps: First, calculate the complexity index of the data. The complexity index is the nonlinear index, distribution entropy, and eigenvalue decay rate; Then, according to the nonlinear index, distribution entropy, and eigenvalue decay rate of the data, select the kernel function in the following way: If the nonlinear index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, select the Gaussian kernel function; If the eigenvalue decay rate of the data is less than the preset threshold γ, select the polynomial kernel function and determine the optimal order d through cross-validation; If none of the above conditions are met, use the mixed kernel function, which is a weighted combination of the Gaussian kernel and the linear kernel.
2. The data processing method according to claim 1, wherein In step S1, the operating status data of the crane is acquired by deploying sensors on the crane and connecting the design database and operation management system of the crane. The sensors include displacement sensors, pressure sensors, and temperature sensors.
3. The data processing method according to claim 2, wherein In step S1, the preprocessing of the crane operating status data includes: Perform format check and cleaning on the operating status data, and remove outliers and error data; For the geometric data, extract relevant model information from the design database and perform simplification and feature extraction.
4. The data processing method according to claim 1, wherein In step S3, select the kernel function according to the following steps: S301: First, calculate the complexity index of the data. The complexity index is the nonlinear index, distribution entropy, and eigenvalue decay rate; (1) Calculate the nonlinear index of the data: ; : Original data point; : The projection matrix of the first k principal components extracted by PCA on the local neighborhood of the point ; : Euclidean norm; (2) Calculate the distribution entropy of the data: ; ; : Gaussian kernel function, ; h: Bandwidth parameter, determined by the Silverman criterion: , is the standard deviation of the data; (3) Calculate the eigenvalue decay rate of the data: Normalize the eigenvalues of the original data covariance matrix, and calculate the number of principal components required for the first 95% cumulative contribution rate. The calculation formula is: ; Perform eigen-decomposition on the covariance matrix of the original data and normalize the eigenvalues ; Calculate the cumulative contribution rate of the first k principal components: ; S302: Then, according to the nonlinear index, distribution entropy, and eigenvalue decay rate of the data, select the kernel function in the following way: If the nonlinear index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, select the Gaussian kernel function; If the eigenvalue decay rate of the data is less than the preset threshold γ, select the polynomial kernel function and determine the optimal order d through cross-validation; If none of the above conditions are met, use the mixed kernel function, which is a weighted combination of the Gaussian kernel and the linear kernel; The decision rule formula for kernel function selection is: Gaussian kernel RBF: ; Polynomial kernel: ; Hybrid core: ; , is the weight.
5. The data processing method according to claim 4, wherein (1) Optimize the Gaussian kernel bandwidth σ in the Gaussian kernel function: The objective function is to maximize the between-class variance ratio: ; represents the trace of a matrix; represents the sum of the differences between classes, represents the sum of the differences within classes; : Between-class scatter matrix: ; where c is the number of classes, is the number of samples in the i-th class, is the mean vector of the i-th class, and μ is the mean vector of all data; : Within-class scatter matrix: ; Among them, C i represents the sample set of the i-th category; Use the Bayesian optimization method for optimization; (2) Adjust the order d of the polynomial kernel function: Calculate the mutual information between dimensions : ; If there exists , and δ is a threshold value, then .
6. The data processing method according to claim 5, wherein In step S3, monitor the change of data distribution in real time: recalculate the data complexity index every time interval T. If the index exceeds the preset range, trigger the re-optimization of the kernel function type and parameters; Update the projection matrix P online.
7. The data processing method according to claim 1, characterized in that In step S3, for the original data x i , the selected kernel function , and the covariance matrix C is: ; Where m is the number of data samples; The contribution rate of the eigenvalue is: ; Where n is the total number of eigenvalues; Combine the selected K eigenvectors into a projection matrix: ; Where K is the number of eigenvector groups; The data after dimensionality reduction is as follows: 。 8. A data processing device, characterized in that, Include: A data acquisition module for acquiring the operating status data of the crane and performing preprocessing; A data integration module for correlating different types of acquired data according to a predefined crane data model and semantic rules; Using the timestamp as the key index, match the operating status data at the same moment with the corresponding geometric data and physical performance data; calibrate the spatial position information of the data to ensure the spatial consistency of different data; A data dimensionality reduction module for performing dimensionality reduction operations on the integrated data using improved PCA, including the following steps: Select a kernel function to perform feature mapping on the integrated data, and map the original data to the high-dimensional feature space F; Calculate the covariance matrix C in the high-dimensional feature space F, and perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues and eigenvectors ; Calculate the contribution rate of the eigenvalues, select the top K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P; Project the original data onto the principal components to achieve dimensionality reduction; Select a kernel function according to the following steps: First, calculate the data complexity index, and the complexity index is the nonlinearity index, distribution entropy, and eigenvalue decay rate; Then, according to the nonlinearity index, distribution entropy, and eigenvalue decay rate of the data, select a kernel function in the following manner: If the nonlinearity index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, select the Gaussian kernel function; If the eigenvalue decay rate of the data is less than the preset threshold γ, select the polynomial kernel function and determine the optimal order d through cross-validation; If none of the above conditions are met, use a mixed kernel function, and the mixed kernel function is a weighted combination of the Gaussian kernel and the linear kernel.
9. A method for constructing a digital twin model of a crane, characterized in that, Construct a crane digital twin model using the data obtained by the data processing method described in any one of claims 1-7; optimize and calibrate the constructed crane digital twin model using the actual crane operation data, and adjust the model parameters using machine learning algorithms.
Citation Information
Patent Citations
Construction method of digital twinborn body of portal crane
CN114644292A
A digital twin and a method for a heavy-duty vehicle
EP4250168A1