Data processing method and device and crane digital twinborn model construction method

Through the improved PCA method, the kernel function is selected for feature mapping and dimensionality reduction, which solves the problems of high dimensionality and data integration of crane data, realizes efficient data integration and dimensionality reduction, and provides more suitable data construction for digital twin models.

CN119961664AActive Publication Date: 2025-05-09NANJING SPECIAL EQUIP SAFETY SUPERVISION & INSPECTION INST

Patent Information

Application Number
CN202510443929.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Crane data has high-dimensional characteristics, which leads to increased difficulty in data storage and processing, inaccurate analysis results, and traditional data processing methods are difficult to efficiently integrate multiple types of data, resulting in compatibility issues and waste of resources.

Method used

The improved principal component analysis (PCA) method is adopted to feature map the data by selecting appropriate kernel functions, and the dimensionality reduction operation is carried out in the high-dimensional feature space, retaining the key information of the data, and improving the data quality and representativeness after dimensionality reduction.

Benefits of technology

The deep fusion of multi-source heterogeneous data is achieved to ensure the integrity and accuracy of the data integration process, and provide comprehensive data support for the digital twin model. The dimensionality reduction results are more suitable for the construction of complex crane digital twin models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961664A_ABST
    Figure CN119961664A_ABST
Patent Text Reader

Abstract

The invention discloses a crane digital twin data processing method based on improved PCA (Principal Component Analysis). According to the method, main characteristics of Internet of Things data are analyzed, characteristic values and characteristic vectors are calculated by adopting a covariance matrix method so as to judge a mechanical state influence principal component factor, and then crane Internet of Things data dimension reduction calculation based on improved PCA is carried out according to the contribution rate of the crane state influence principal component factor, so that the mechanical state influence principal component factor is judged. And crane state data dimension reduction is realized. And instance verification is carried out to prove that the data information loss is small after dimension reduction by adopting the PCA method, the data after dimension reduction shows better performance in classification and prediction tasks, and the training speed of the crane data large model and the accuracy degree of operation fault early warning are improved. According to the method, effective dimension reduction of the crane state data is realized, and the problems that data acquired by real-time monitoring of the operation state of the Internet of Things in the research process of crane intelligent supervision, digital twinning and the like is complex, high in dimension, inaccurate in storage, processing and analysis and the like are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of crane technology, and in particular to a data processing method and a method for constructing a crane digital twin model. Background Art

[0002] As an indispensable heavy machinery in industrial production, real-time monitoring of the operation status of cranes is crucial to ensure production safety and improve work efficiency. With the widespread application of Internet of Things technology, the operation data of cranes can be collected in real time through sensors and transmitted to the cloud for analysis and processing. However, due to the large number of sensors, the collected data often has high-dimensional characteristics, which not only increases the difficulty of data storage and processing, but also may lead to inaccurate analysis results. Therefore, studying effective data dimensionality reduction methods is of great practical significance for the analysis of crane Internet of Things data.

[0003] In the process of building a crane digital twin system, various types of data are involved, including geometric data, physical performance data, operating status data, etc. These data come from different sensors, systems or databases, with different data formats, precision and update frequencies. Traditional data processing methods make it difficult to integrate them efficiently, resulting in compatibility issues and resource waste during data transmission, storage and analysis. For example, geometric data may be stored in the CAD model format, while operating status data is collected by sensors in real time and exists in the form of time series data. It is a huge challenge to unify these data for the construction of digital twin models. Summary of the invention

[0004] On the one hand, the present application provides a data processing method, which has the advantage of being able to effectively process nonlinear relationships in crane data, perform dimensionality reduction operations in high-dimensional feature space, better retain key information of the data, improve the quality and representativeness of the data after dimensionality reduction, and make the dimensionality reduction results more suitable for the construction of complex crane digital twin models.

[0005] The technical solution is as follows: A crane digital twin data processing method based on improved PCA includes the following steps: S1 Data collection and preprocessing: collect the crane's operating status data and preprocess it; S2 Data Integration: Correlate different types of collected data based on predefined crane data models and semantic rules; S3 data dimensionality reduction: Improved PCA is used to reduce the dimensionality of the integrated data, including the following steps: Select a kernel function to perform feature mapping on the integrated data, and map the original data into a high-dimensional feature space F; Calculate the covariance matrix C in the high-dimensional feature space F, perform eigendecomposition on the covariance matrix C, and obtain the eigenvalues and the eigenvector ; Calculate the contribution rate of the eigenvalue, select the first K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P; Projecting the original data onto the principal components achieves dimensionality reduction.

[0006] Furthermore, in step S1, the operation status data of the crane is collected by deploying sensors on the crane and connecting the design database and operation management system of the crane, wherein the sensors include displacement sensors, pressure sensors and temperature sensors.

[0007] Furthermore, in step S1, the preprocessing of the crane operation status data includes: Check and clean the format of the running status data to remove abnormal values ​​and erroneous data; For geometric data, relevant model information is extracted from the design database for simplification and feature extraction.

[0008] Furthermore, in step S2, the operating status data at the same time is matched with the corresponding geometric data and physical performance data using the timestamp as the key index; the spatial position information of the data is calibrated to ensure the spatial consistency of different data.

[0009] Furthermore, in step S3, a kernel function is selected according to the following steps: S301: First, calculate the complexity index of the data, which is the nonlinear index, distribution entropy and eigenvalue decay rate; (1) Calculate the nonlinear index of the data: ; : original data points; : At point The projection matrix of the first k principal components extracted by PCA on the local neighborhood of ; : Euclidean norm; (2) Calculate the distribution entropy of data: ; ; : Gaussian kernel function, ; h: bandwidth parameter, determined by the Silverman criterion: , is the data standard deviation; (3) Calculate the data eigenvalue decay rate: Normalize the eigenvalues ​​of the covariance matrix of the original data and calculate the number of principal components required for the first 95% cumulative contribution rate. The calculation formula is: ; Perform eigendecomposition on the original data covariance matrix and normalize the eigenvalues ; Calculate the cumulative contribution rate of the first k principal components: ; S302: Then, according to the nonlinear index, distribution entropy and eigenvalue decay rate of the data, a kernel function is selected in the following manner: If the nonlinear index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, the Gaussian kernel function is selected; If the eigenvalue decay rate of the data is less than the preset threshold γ, a polynomial kernel function is selected, and the optimal order d is determined through cross-validation; If the above conditions are not met, a hybrid kernel function is used, which is a weighted combination of a Gaussian kernel and a linear kernel; The decision rule formula for kernel function selection is: Gaussian Kernel RBF: ; Polynomial kernel: ; Hybrid Core: ; , is the weight.

[0010] Furthermore, (1) the Gaussian kernel bandwidth σ in the Gaussian kernel function is optimized: The objective function is to maximize the inter-class variance ratio: ; represents the trace of a matrix; represents the sum of the differences between classes, represents the sum of intra-class differences; : Inter-class scatter matrix: ; Where c is the number of categories, is the number of samples in the i-th category, is the mean vector of the i-th class, μ is the mean vector of all data; : Intra-class scatter matrix: ; Among them, Ci represents the sample set of the i-th category; Bayesian optimization method is used for optimization; (2) Adjust the order d of the polynomial kernel function: Compute mutual information between dimensions : ; If exists , δ is the threshold, then .

[0011] Furthermore, in step S3, the data distribution changes are monitored in real time: the data complexity index is recalculated at each interval T, and if the index exceeds the preset range, the kernel function type and parameter re-optimization is triggered; and the projection matrix P is updated by online learning.

[0012] Further, in step S3, for the original data x i , the selected kernel function , the covariance matrix C is:

[0013] Where m is the number of data samples; The contribution rate of the eigenvalue is:

[0014] Where n is the total number of eigenvalues; Combine the selected K eigenvectors into a projection matrix:

[0015] Where K is the number of feature vector groups; Data after dimensionality reduction for: .

[0016] On the other hand, the present application provides a crane digital twin data processing device based on improved PCA, comprising: Data acquisition module, used to collect the operating status data of the crane and perform pre-processing; Data integration module, used to associate different types of collected data according to predefined crane data models and semantic rules; The data dimension reduction module is used to perform dimension reduction operation on the integrated data using improved PCA, including the following steps: Select a kernel function to perform feature mapping on the integrated data, and map the original data into a high-dimensional feature space F; Calculate the covariance matrix C in the high-dimensional feature space F, perform eigendecomposition on the covariance matrix C, and obtain the eigenvalues and the eigenvector ; Calculate the contribution rate of the eigenvalue, select the first K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P; Projecting the original data onto the principal components achieves dimensionality reduction.

[0017] On the other hand, the present application provides a data processing terminal, including a memory and a processor, wherein the memory stores a computer program, and when the processor calls and executes the computer program, the crane digital twin data processing method based on improved PCA as described above is implemented.

[0018] On the other hand, the present application provides a computer-readable medium, which stores a computer program. When the computer program is called and executed by a computer, it implements the crane digital twin data processing method based on improved PCA as described above.

[0019] On the other hand, the present application provides a method for constructing a digital twin model of a crane, which constructs a digital twin model of a crane using the data obtained by the crane digital twin data processing method based on the improved PCA as described above; optimizes and calibrates the constructed digital twin model of the crane using actual crane operation data, and adjusts the model parameters using a machine learning algorithm.

[0020] In summary, the beneficial effects of this application are: 1. The crane digital twin data processing method based on improved PCA provided in this application realizes the deep fusion of multi-source heterogeneous data such as geometry, physics, and operation. This fusion method not only considers the format and content of the data, but also makes full use of the spatiotemporal attributes and semantic information of the data to ensure the integrity and accuracy of the data during the integration process, and provide comprehensive data support for the digital twin model.

[0021] 2. The crane digital twin data processing method based on improved PCA provided in this application introduces kernel function in PCA, breaks through the linear limitation of traditional PCA, and can effectively process nonlinear relationships in crane data. Dimensionality reduction operation is performed in high-dimensional feature space, which better retains the key information of the data, improves the quality and representativeness of the data after dimensionality reduction, and makes the dimensionality reduction result more suitable for the construction of complex crane digital twin models.

[0022] 3. The crane digital twin model construction method provided in this application constructs a closed-loop feedback mechanism between the digital twin model and the actual crane operation data, which can continuously optimize the model parameters and structure according to the actual operation conditions. Through continuous learning and improvement, the digital twin model can more accurately reflect the real state of the crane, predict potential failures in advance, and provide strong guarantees for the safe operation and maintenance of the crane. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Schematic diagram of the steps of a crane digital twin data processing method based on improved PCA in this application DETAILED DESCRIPTION

[0024] The specific implementation of the present application is described in detail below with reference to the accompanying drawings.

[0025] A specific embodiment of the present application provides a crane digital twin data processing method based on improved PCA, such as Figure 1 , including the following steps: S1 Data collection and preprocessing: collect the crane's operating status data and preprocess it; In step S1, the operation status data of the crane is collected by deploying sensors on the crane and connecting the design database and operation management system of the crane. The sensors include displacement sensors, pressure sensors and temperature sensors. The sensors collect the operation status data in real time and transmit it to the data acquisition module at a certain frequency.

[0026] In step S1, the preprocessing of the crane operation status data includes: Check and clean the format of the running status data to remove abnormal values ​​and erroneous data; For geometric data, relevant model information is extracted from the design database, simplified and feature extracted, and converted into a format suitable for subsequent processing. For example, a 3D CAD model is simplified into key geometric dimensions and shape feature vectors.

[0027] S2 Data Integration: Correlate different types of collected data based on predefined crane data models and semantic rules; In step S2, the timestamp is used as the key index to match the operating status data at the same time with the corresponding geometric data and physical performance data; for example, when a crane lifts a heavy object at a certain moment, the lifting weight (operating status data) at that moment is integrated with the geometric structure data of the crane boom and the physical performance data of the material (such as strength and stiffness). At the same time, the spatial position information of the data is calibrated to ensure the consistency of different data in space.

[0028] S3 data dimensionality reduction: Improved PCA is used to reduce the dimensionality of the integrated data, including the following steps: Select the kernel function to perform feature mapping on the integrated data and map the original data into the high-dimensional feature space F.

[0029] In the processing of crane digital twin data, the selection of kernel function needs to comprehensively consider data characteristics, computational complexity and model construction requirements. In this embodiment, the kernel function is a Gaussian kernel function or a polynomial kernel function.

[0030] If there are complex nonlinear relationships in the crane data and the distribution is relatively complex, the Gaussian kernel function is a good choice. During the operation of the crane, the relationship between its physical performance data (such as stress and strain) and operating status data (such as speed and acceleration) may be nonlinear, and this relationship is difficult to describe with a simple mathematical model. The Gaussian kernel function can map the data to a high-dimensional feature space and effectively handle this complex nonlinear relationship, making the characteristics of the data in the high-dimensional space more obvious, which is convenient for subsequent dimensionality reduction operations and model construction. In addition, the Gaussian kernel function has strong adaptability to data, can handle noise and outliers in the data, and improve the stability of data processing.

[0031] When there is a polynomial relationship between certain dimensions in the crane data, the polynomial kernel function is more appropriate. There may be a polynomial relationship between the crane's geometric dimension data and its operating status data. The polynomial kernel function can capture this relationship by adjusting the order, and can to some extent explore the deep connection between the data. Moreover, the polynomial kernel function is relatively simple in calculation and has low computational complexity. When the amount of data is large, it can improve the efficiency of data processing.

[0032] Specifically, in step S3, the kernel function is selected according to the following steps: S301: First, calculate the complexity index of the data, which is the nonlinear index, distribution entropy and eigenvalue decay rate; (1) Calculate the nonlinearity index (NLI) of the data: ; : original data points; : At point The projection matrix of the first k principal components extracted by PCA on the local neighborhood of (such as the 10% samples of the nearest neighbors); : Euclidean norm; The nonlinearity of the data is evaluated by calculating the local linear reconstruction error of the original data under the low-dimensional projection. The larger the error, the stronger the nonlinearity.

[0033] (2) Calculate the distribution entropy (DistributionEntropy, H) of the data: ; ; : Gaussian kernel function, ; h: bandwidth parameter, determined by the Silverman criterion: , is the data standard deviation; Kernel density estimation is used to calculate the entropy of data distribution. The higher the entropy value, the more complex the data distribution.

[0034] (3) Calculate the data eigenvalue decay rate: Normalize the eigenvalues ​​of the covariance matrix of the original data and calculate the number of principal components required for the first 95% cumulative contribution rate. The calculation formula is: ; Perform eigendecomposition on the original data covariance matrix and normalize the eigenvalues ; Calculate the cumulative contribution rate of the first k principal components: ; The eigenvalues ​​of the original data covariance matrix are normalized, and the number of principal components required for the first 95% cumulative contribution rate is calculated. The smaller the number, the stronger the linear separability of the data. That is, the smaller the EDR, the closer the data is to linear separability.

[0035] S302: Then, according to the nonlinear index, distribution entropy and eigenvalue decay rate of the data, a kernel function is selected in the following manner: If the nonlinear index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, the Gaussian kernel function is selected; If the eigenvalue decay rate of the data is less than the preset threshold γ, a polynomial kernel function is selected, and the optimal order d is determined through cross-validation; If the above conditions are not met, a hybrid kernel function is used, which is a weighted combination of a Gaussian kernel and a linear kernel; The decision rule formula for kernel function selection is: Gaussian Kernel RBF: ; Polynomial kernel: ; Hybrid Core: ; , is the weight. It is dynamically allocated by the information retention rate of the validation set (such as the classification accuracy of the data after dimensionality reduction). Typically, the threshold =0.7, =2.0, = 0.3 (determined based on experimental statistics). The polynomial order d is selected by cross-validation to obtain the optimal value (e.g. ).

[0036] Optimize kernel function parameters. For Gaussian kernel, optimize bandwidth parameter σ by maximizing inter-class variance ratio (class separability of data after dimensionality reduction), and use Bayesian optimization algorithm to search for the optimal value. For polynomial kernel, adaptively adjust order d based on feature interaction strength (calculating correlation between dimensions through mutual information), ensuring that high-order terms are only used for strongly correlated dimensions to avoid overfitting.

[0037] (1) Optimize the Gaussian kernel bandwidth σ in the Gaussian kernel function: The objective function is to maximize the inter-class variance ratio: ; Represents the trace of the matrix, that is, the sum of the diagonal elements of the matrix; represents the sum of the differences between classes, Represents the sum of intra-class differences; by maximizing the ratio of inter-class differences to intra-class differences, the reduced-dimensional data is made more separable between categories while reducing the degree of dispersion within the categories. This indicator is often used in supervised dimensionality reduction tasks to improve the performance of classification models.

[0038] : Inter-class scatter matrix: ; Where c is the number of categories, is the number of samples in the i-th category, is the mean vector of the i-th class, μ is the mean vector of all data; : Intra-class scatter matrix: ; Among them, C i represents the sample set of the i-th category; Use Bayesian optimization method for optimization and iterative search .

[0039] It should be noted that in the supervised dimensionality reduction scenario of Gaussian kernel PCA (or kernel method): the data is processed by Gaussian kernel Implicitly mapped to a high-dimensional feature space F. The inter-class and intra-class divergence matrix S B and S w It is calculated in a high-dimensional feature space F, and the structure of F is determined by the kernel parameter σ. If σ is too small, the data may be too dispersed in the high-dimensional space (increase the intra-class difference). If σ is too large, the data may be over-smoothed (reduce the inter-class difference).

[0040] The trace of the divergence matrix Tr(S B ) and Tr(S w) will change with the change of σ, so it is necessary to maximize the inter-class separability by optimizing σ.

[0041] The core steps of adjusting σ through Bayesian optimization include: 1. Initialization: Select an initial σ and calculate the kernel matrix K(σ).

[0042] 2. Scatter matrix calculation: Based on K(σ) and category labels, Tr(S B (σ)) and Tr(S w (σ)).

[0043] 3. Objective function evaluation: Calculate Tr(S B ) / Tr(S w ), as the optimization target.

[0044] 4. Iterative update: Adjust σ through Bayesian optimization and repeat steps 2 and 3 until convergence.

[0045] The inter-class and intra-class scatter matrix S B and S w are functions of σ, because their calculations are based on the high-dimensional space after kernel mapping, which is controlled by σ.

[0046] Application of kernel tricks: no explicit calculation required , but rather indirectly expresses the trace of the divergence matrix through the kernel matrix.

[0047] Optimization significance: Adjust σ to maximize the difference between classes and minimize the difference within classes, thereby improving the classification performance after dimensionality reduction.

[0048] (2) Adjust the order d of the polynomial kernel function: Determining the strength of feature interactions: Calculating mutual information between dimensions : ; Adaptive rules: If present , δ is the threshold, then δ is the threshold for judging whether the mutual information is "significant", which is based on the mutual information strength between dimensions. Experimental tuning is required for specific implementation, and the specific value of δ needs to be determined through cross-validation or actual task results.

[0049] Dynamic feedback adjustment: During the operation of the digital twin model, the data distribution changes are monitored in real time: the data complexity index is recalculated at every interval T. If the index exceeds the preset range, the kernel function type and parameters are re-optimized; the projection matrix P is updated through online learning to ensure that the dimensionality reduction model continues to adapt to the dynamic changes in the crane's operating status.

[0050] (1) Real-time monitoring and triggering conditions: NLI, H, and EDR are recalculated every time interval T (e.g., 1 hour). If any indicator changes by more than 10% (e.g., ), triggering re-optimization.

[0051] (2) Online learning updates the projection matrix: Incremental PCA (IPCA) updates the projection matrix P: .

[0052] Typically, taking a port gantry crane as an example, the initial data analysis is as follows: NLI=0.85, H=2.3, EDR=0.25.

[0053] Trigger Gaussian kernel selection and Bayesian optimization to obtain σ=1.2.

[0054] After dimensionality reduction, the classification accuracy increased by 12% (from 78% to 90%).

[0055] Dynamic adjustment during operation: After 3 months, due to structural wear, H dropped to 1.8, triggering the hybrid core ( =0.6, =0.4).

[0056] After updating the projection matrix, the prediction error remains within 3%. Through the above method, this application clarifies the mathematical logic of kernel function selection and parameter optimization, and combines the dynamic feedback mechanism to significantly improve the technical effect.

[0057] The above kernel function determination method has the following effects: 1. Data-driven kernel function decision-making mechanism; Binding the quantitative indicators of data complexity (nonlinear index, distribution entropy, eigenvalue decay rate) with the selection of kernel functions breaks through the limitations of traditional empirical selection or fixed kernel functions and significantly improves the adaptability of the dimensionality reduction process.

[0058] 2. Decoupling parameter optimization from dimension correlation; For the polynomial kernel, an order adjustment method based on feature interaction strength is proposed to avoid dimensional redundancy caused by global unified order, reduce computational overhead and improve model generalization ability.

[0059] 3. Online dynamic adjustment capability; Through real-time monitoring and feedback mechanism, the kernel function and parameters can be automatically adjusted as the crane operating environment changes (such as load mutation, structural fatigue), ensuring the long-term accuracy of the digital twin model.

[0060] Calculate the covariance matrix C in the high-dimensional feature space F, for the original data x i , the selected kernel function , the covariance matrix C is:

[0061] Where m is the number of data samples.

[0062] Perform eigendecomposition on the covariance matrix C to obtain eigenvalues ​​and eigenvectors. Calculate the contribution rate of the eigenvalue, which is:

[0063] Where n is the total number of eigenvalues.

[0064] Select the first K eigenvectors with larger contribution rates as principal components, construct the projection matrix P, and combine the selected K eigenvectors into a projection matrix:

[0065] Where K is the number of feature vector groups.

[0066] The original data is projected onto the principal component to achieve dimensionality reduction. The data after dimensionality reduction is : .

[0067] This embodiment further illustrates the solution by taking the application example of data processing of large port cranes.

[0068] Overview of crane application scenarios: A large gantry crane at a port was selected and equipped with a variety of sensors, including high-precision displacement sensors for measuring the lifting height of the spreader, amplitude sensors for monitoring the horizontal movement amplitude of the spreader, speed sensors for detecting the lifting and running speeds, and acceleration sensors for capturing the acceleration changes during the movement. These sensors collect operating status data in real time. Strain gauges are installed at key parts of the crane structure to obtain physical performance data such as structural stress and strain, while environmental monitoring equipment is used to collect data such as the temperature, humidity, wind speed and direction of the surrounding environment. In addition, the crane's control system records control-related information such as operating instructions and operating modes, and its design database stores detailed geometric dimension data, such as boom length and bridge structure dimensions.

[0069] Data collection and preliminary processing: Data collection and processing of operation status: The displacement sensor collects lifting height data every 0.1 seconds. In a typical loading and unloading operation, 6,000 data points are collected within 10 minutes, and the data range is between 0-30 meters. After preliminary screening, obviously abnormal mutation values ​​(such as data exceeding the normal lifting height range of ±5 meters) are removed. The speed sensor and acceleration sensor collect data synchronously, with the speed data range of 0-2m / s and the acceleration data between -2-2m / s². Abnormal value removal and data smoothing are also performed, and the moving average method is used to average the 5 adjacent data points to reduce noise interference.

[0070] Physical property data collection and processing: The stress-strain data collected by the strain gauge is updated every 1 second, and 600 data points are obtained within 10 minutes. The stress data range is 0-500MPa. The strain data is normalized so that its value is between 0-1, so as to be analyzed uniformly with other data.

[0071] Environmental data collection and processing: Environmental monitoring equipment collects temperature (10-40°C), humidity (30% -80%), wind speed (0 -15m / s) and wind direction (0-360°) data once a minute, and collects 10 sets of data in 10 minutes. For wind speed and direction data, it is converted into horizontal and vertical components according to the trigonometric function relationship so that it can be combined with the crane force analysis later.

[0072] Extraction and conversion of geometric and control data: Extract key geometric dimensions such as the crane boom length of 50 meters and the bridge width of 20 meters from the design database and convert them into dimensionless relative proportional data, such as the ratio of boom length to bridge width is 2.5. The control system's operating instructions and operating mode data are classified and coded, such as different lifting operation instructions are divided into 5 categories and operating modes are divided into 3 types, represented by numbers 1-5 and 1-3 respectively.

[0073] Data integration process: Integration based on timestamp: Integrate the operating status data, physical performance data, environmental data and corresponding control data collected within the same minute at a time interval of 1 minute. For example, at the 5th minute, the lifting height, speed, acceleration, stress strain, temperature, humidity, wind speed and direction components, and current operation instructions and operation mode data at that moment are combined into one data record.

[0074] Association based on spatial position: According to the structural layout of the crane, the physical performance data collected by the strain gauges are associated with the corresponding geometric data of the structural components. For example, the data collected by the strain gauges at the root of the boom are associated with the geometric dimensions and material properties of the boom, and the spatial correspondence between the data is established to analyze the interaction between the structural force and the geometric structure.

[0075] Data dimensionality reduction operation: (1) Kernel function selection and feature mapping: Selecting polynomial kernel function (Here d = 3) Perform feature mapping on the integrated data. Map the original data vector containing multi-dimensional data such as lifting height, speed, acceleration, stress strain, environmental data, geometric proportion and control coding to the high-dimensional feature space F.

[0076] (2) Covariance matrix calculation and feature decomposition: Calculate the covariance matrix C in the high-dimensional feature space. Assuming there are n = 60 data samples integrated within 10 minutes, the calculation formula is: . Perform eigendecomposition on the covariance matrix C and obtain the eigenvalue and the eigenvector .

[0077] (3) Principal component selection and dimensionality reduction calculation: Calculate the contribution rate of eigenvalues , select the first k eigenvectors whose contribution rate reaches 90% as the principal components to construct the projection matrix P. Finally, project the original data onto the principal components to achieve dimensionality reduction. It can be used for subsequent crane digital twin model construction and analysis, effectively reducing the dimension and complexity of the data while retaining key information.

[0078] Through the above example applications, the actual operation process and effect of this patented method in complex crane data processing are demonstrated, providing strong data support for the intelligent management and digital twin applications of cranes.

[0079] Another specific embodiment of the present application provides a crane digital twin data processing device based on improved PCA, comprising: Data acquisition module, used to collect the operating status data of the crane and perform pre-processing; Data integration module, used to associate different types of collected data according to predefined crane data models and semantic rules; The data dimension reduction module is used to perform dimension reduction operation on the integrated data using improved PCA, including the following steps: Select a kernel function to perform feature mapping on the integrated data, and map the original data into a high-dimensional feature space F; Calculate the covariance matrix C in the high-dimensional feature space F, perform eigendecomposition on the covariance matrix C, and obtain the eigenvalues and the eigenvector ; Calculate the contribution rate of the eigenvalue, select the first K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P; Projecting the original data onto the principal components achieves dimensionality reduction.

[0080] Another specific embodiment of the present application provides a data processing terminal, including a memory and a processor, wherein the memory stores a computer program, and when the processor calls and executes the computer program, the crane digital twin data processing method based on improved PCA as described above is implemented.

[0081] Another specific embodiment of the present application provides a computer-readable medium, which stores a computer program. When the computer program is called and executed by a computer, it implements the crane digital twin data processing method based on improved PCA as described above.

[0082] Another specific embodiment of the present application provides a method for constructing a digital twin model of a crane, which constructs a digital twin model of a crane using data obtained by the crane digital twin data processing method based on the improved PCA as described above; optimizes and calibrates the constructed digital twin model of the crane using actual crane operation data, and adjusts the model parameters using a machine learning algorithm.

[0083] Using the reduced-dimensional data as input, a suitable modeling method (such as physics-based modeling, machine learning modeling, etc.) is used to build a crane digital twin model. For example, a neural network algorithm in machine learning is used to build a prediction model with the reduced-dimensional operating status data and geometric physical feature data as the input layer and the crane's performance indicators (such as lifting capacity, stability, etc.) as the output layer. By comparing the model output with the actual crane operation data, the error indicators (such as mean square error, absolute error, etc.) are calculated. Based on the error feedback, the model parameters are adjusted using optimization methods such as the back propagation algorithm to continuously optimize the digital twin model and improve its accuracy and reliability.

[0084] Through the above system and method, the effective integration and dimensionality reduction of different types of data in the crane digital twin technology are realized, providing a solid data foundation and technical support for the intelligent monitoring, fault prediction and optimized operation of the crane. In practical applications, the system parameters and algorithms can be appropriately adjusted and optimized according to the models, working environments and application requirements of different cranes to achieve the best processing effect.

[0085] The above is only a preferred implementation of the present application. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present application, and these all fall within the scope of protection of the present application.

Claims

1. A data processing method, characterized in that: The following steps are involved: S1 Data collection and preprocessing: collect the crane's operating status data and preprocess it; S2 Data Integration: Correlate different types of collected data based on predefined crane data models and semantic rules; S3 data dimensionality reduction: Improved PCA is used to reduce the dimensionality of the integrated data, including the following steps: Select a kernel function to perform feature mapping on the integrated data, and map the original data into a high-dimensional feature space F; Calculate the covariance matrix C in the high-dimensional feature space F, perform eigendecomposition on the covariance matrix C, and obtain the eigenvalues and the eigenvector ; Calculate the contribution rate of the eigenvalue, select the first K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P; Projecting the original data onto the principal components achieves dimensionality reduction.

2. The data processing method according to claim 1, characterized in that: In step S1, the operation status data of the crane is collected by deploying sensors on the crane and connecting the design database and operation management system of the crane, wherein the sensors include displacement sensors, pressure sensors and temperature sensors.

3. The data processing method according to claim 2, characterized in that: In step S1, the preprocessing of the crane operation status data includes: Check and clean the format of the running status data to remove abnormal values ​​and erroneous data; For geometric data, relevant model information is extracted from the design database for simplification and feature extraction.

4. The data processing method according to claim 1, characterized in that: In step S2, the operating status data at the same time is matched with the corresponding geometric data and physical performance data using the timestamp as the key index; the spatial position information of the data is calibrated to ensure the spatial consistency of different data.

5. The data processing method according to claim 1, characterized in that: In step S3, a kernel function is selected according to the following steps: S301: First, calculate the complexity index of the data, which is the nonlinear index, distribution entropy and eigenvalue decay rate; (1) Calculate the nonlinear index of the data: ; : original data points; : At point The projection matrix of the first k principal components extracted by PCA on the local neighborhood of ; : Euclidean norm; (2) Calculate the distribution entropy of data: ; ; : Gaussian kernel function, ; h: bandwidth parameter, determined by the Silverman criterion: , is the data standard deviation; (3) Calculate the data eigenvalue decay rate: Normalize the eigenvalues ​​of the covariance matrix of the original data and calculate the number of principal components required for the first 95% cumulative contribution rate. The calculation formula is: ; Perform eigendecomposition on the original data covariance matrix and normalize the eigenvalues ; Calculate the cumulative contribution rate of the first k principal components: ; S302: Then, according to the nonlinear index, distribution entropy and eigenvalue decay rate of the data, a kernel function is selected in the following manner: If the nonlinear index of the data is greater than the preset threshold α and the distribution entropy is greater than the preset threshold β, the Gaussian kernel function is selected; If the eigenvalue decay rate of the data is less than the preset threshold γ, a polynomial kernel function is selected, and the optimal order d is determined through cross-validation; If the above conditions are not met, a hybrid kernel function is used, which is a weighted combination of a Gaussian kernel and a linear kernel; The decision rule formula for kernel function selection is: Gaussian Kernel RBF: ; Polynomial kernel: ; Hybrid Core: ; , is the weight.

6. The data processing method according to claim 5, characterized in that: (1) Optimize the Gaussian kernel bandwidth σ in the Gaussian kernel function: The objective function is to maximize the inter-class variance ratio: ; represents the trace of a matrix; represents the sum of the differences between classes, represents the sum of intra-class differences; : Inter-class scatter matrix: ; Where c is the number of categories, is the number of samples in the i-th category, is the mean vector of the i-th class, μ is the mean vector of all data; : Intra-class scatter matrix: ; Among them, C i represents the sample set of the i-th category; Bayesian optimization method is used for optimization; (2) Adjust the order d of the polynomial kernel function: Compute mutual information between dimensions : ; If exists , δ is the threshold, then .

7. The data processing method according to claim 6, characterized in that: In step S3, the data distribution changes are monitored in real time: the data complexity index is recalculated at each interval T, and if the index exceeds the preset range, the kernel function type and parameters are re-optimized; Online learning updates the projection matrix P.

8. The data processing method according to claim 1, characterized in that: In step S3, for the original data x i , the kernel function selected , the covariance matrix C is: ; Where m is the number of data samples; The contribution rate of the eigenvalue is: ; Where n is the total number of eigenvalues; Combine the selected K eigenvectors into a projection matrix: ; Where K is the number of feature vector groups; Data after dimensionality reduction for: 。 9. A data processing device, characterized in that: include: Data acquisition module, used to collect the operating status data of the crane and perform pre-processing; Data integration module, used to associate different types of collected data according to predefined crane data models and semantic rules; The data dimension reduction module is used to perform dimension reduction operation on the integrated data using improved PCA, including the following steps: Select a kernel function to perform feature mapping on the integrated data, and map the original data into a high-dimensional feature space F; Calculate the covariance matrix C in the high-dimensional feature space F, perform eigendecomposition on the covariance matrix C, and obtain the eigenvalues and the eigenvector ; Calculate the contribution rate of the eigenvalue, select the first K eigenvectors with larger contribution rates as the principal components, and construct the projection matrix P; Projecting the original data onto the principal components achieves dimensionality reduction.

10. A method for constructing a digital twin model of a crane, characterized in that: Construct a crane digital twin model using the data obtained by the data processing method described in any one of claims 1 to 8; optimize and calibrate the constructed crane digital twin model using actual crane operation data, and adjust the model parameters using a machine learning algorithm.

Citation Information

Patent Citations

  • Fishery administration monitoring system infrared target identification method based on PCNN and PCANet, and storage medium

    CN110738166A

  • Wind driven generator fault diagnosis based on long-term and short-term memory model recurrent neural network

    CN111241748A

  • Rolling bearing abnormity detection method based on mixed kernel function convex hull approximation

    CN111307459A

  • Digital twin-driven mechanical arm modeling, control and monitoring integrated system

    CN111496781A

  • Construction method of digital twinborn body of portal crane

    CN114644292A

Cited By

  • Intelligent inspection method and system for hoisting machinery based on multi-modal large model

    CN122133529A

  • Intelligent Inspection Method and System for Lifting Machinery Based on Multimodal Large Model

    CN122133529B