Load weighing model training method, load weighing method and related equipment
By constructing and training an initial model with rational functions as the basis function, the problem of inaccurate load weighing in the prior art is solved, and accurate weighing of the working mechanical load is achieved.
Patent Information
- Application Number
- CN202510491537.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing weighing algorithm based on moment balance has inaccurate problems in load weighing, mainly due to the complex modeling, many considerations and the difficulty in eliminating the influence of complex nonlinear factors, such as centroid shift, oil temperature changes, vibration inertia, etc.
A training method for load weighing models is adopted. By obtaining historical working conditions information and corresponding load labels, the initial model is constructed with a rational function as the basis function, and the model is obtained by training the model until the preset training stop condition is met.
This method can accurately realize the load weighing of the working machine. Through continuous learning and exploration of training data, it approximates the complex nonlinear relationship between the load and the working condition information, improving the accuracy of the weighing.
Smart Images

Figure CN120030914A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of engineering machinery, and in particular to a training method for a load weighing model, a load weighing method and related equipment. Background Art
[0002] Automatic load weighing technology is an important research direction in the field of engineering machinery, which aims to achieve real-time and accurate monitoring of load weight during operation. In recent years, with the development of sensor technology and data processing algorithms, dynamic weighing technology has been widely used.
[0003] At present, the load weighing system consists of two parts: a data acquisition device and a data processing system. The data acquisition device consists of a variety of state sensors and is installed in a specific location according to needs. Most data processing systems used for material mass calculation establish a dynamic balance model of torque and then estimate the material mass based on dynamic data. However, in the existing weighing algorithm based on torque balance, the modeling is complex and there are many factors to consider. It is often necessary to add a lot of compensation, which is time-consuming and labor-intensive, and it is difficult to eliminate the influence of complex nonlinear factors, such as center of mass offset, oil temperature changes, vibration inertia, etc., which leads to inaccurate load weighing. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a load weighing model training method, a load weighing method and related equipment to solve the problem of inaccurate load weighing caused by the weighing algorithm in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a method for training a load weighing model, the method comprising: Acquire a training data set, the training data set includes a plurality of training samples, each training sample includes historical working condition information of the working machine and a load label of the working machine corresponding to the historical working condition information, wherein the historical working condition information includes at least one of angle information of a working device of the working machine, pressure information of a hydraulic system of the working machine, and speed-related information of the working device of the working machine; Construct an initial model, the basis functions of which are rational functions; Taking historical working condition information as input and load labels as output, the initial model is trained until the preset training stop condition is met, and the load weighing model corresponding to the operating machinery is obtained. The load weighing model is used to load weigh the operating machinery.
[0006] In the embodiment of the present application, constructing an initial model includes: A rational function is used as a basis function in a neural network, and an absolute value operation and an offset operation are performed on the denominator of the rational function to obtain a modified rational function, so as to obtain an initial model based on the modified rational function.
[0007] In the embodiment of the present application, the initial model is:
[0008] in, is the initial model, is the scaling factor, is the first coefficient, is the second coefficient, m and n are the degrees of the polynomial in the initial model, and x is the input quantity.
[0009] In the embodiment of the present application, historical working condition information is used as input and load labels are used as output to train the initial model until a preset training stop condition is met, thereby obtaining a load weighing model corresponding to the operating machine, including: For each training sample, perform the following steps: Input the historical operating condition information into the initial model to obtain the load prediction result corresponding to the historical operating condition information; Determine the loss function value of the initial model based on the load prediction results and load labels; When the loss function value does not meet the preset training stop condition, the initial model is updated to obtain an updated initial model, and the historical operating condition information is returned to the initial model to obtain the load prediction result corresponding to the historical operating condition information, until the preset training stop condition is met to obtain the load weighing model.
[0010] In an embodiment of the present application, when the loss function value does not meet the preset training stop condition, the initial model is updated to obtain an updated initial model, including: When the loss function value does not meet the training stop condition, the model parameters of the initial model are adjusted to obtain an optimized initial model; In the optimized initial model, when the absolute value of the weight of the neuron corresponding to the hidden layer is less than the preset value, the neuron is pruned to obtain a pruned initial model; The pruned initial model is accurately trained to obtain an updated initial model.
[0011] In the embodiment of the present application, the pruned initial model is accurately trained to obtain an updated initial model, including: The updated initial model is functionally adjusted using a preset objective function to obtain an adjusted initial model, wherein the objective function is used to standardize the expression form of the initial model; The adjusted initial model is trained using the training data set to obtain an updated initial model.
[0012] In an embodiment of the present application, when the loss function value does not meet the preset training stop condition, the model parameters of the initial model are adjusted to obtain an updated initial model, including: When the loss function value does not meet the preset training stop condition, the gradient of the model parameters is determined using the pre-acquired gradient expression; According to the gradient, the model parameters of the initial model are adjusted.
[0013] In the embodiment of the present application, the loss function value of the initial model is determined according to the load prediction result and the load label, including: Add a regularization term to the preset loss function to obtain the target loss function; The loss function value of the initial model is determined using the target loss function, load prediction results, and load labels.
[0014] In a second aspect, an embodiment of the present application provides a load weighing method, the method comprising: Obtain the current working condition information of the operating machinery; Input the current working condition information into the load weighing model to determine the load of the operating machinery; The load weighing model is a load weighing model trained according to the first aspect.
[0015] In a third aspect, an embodiment of the present application provides a training device for a load weighing model, the device comprising: A first acquisition module is used to acquire a training data set, the training data set includes a plurality of training samples, each training sample includes historical working condition information of the working machine and a load label of the working machine corresponding to the historical working condition information, wherein the historical working condition information includes at least one of angle information of a working device of the working machine, pressure information of a hydraulic system of the working machine, and speed-related information of the working device of the working machine; A building module is used to build an initial model, and the basis functions of the initial model are rational functions; The training module is used to train the initial model with historical working condition information as input and load labels as output until the preset training stop conditions are met, so as to obtain the load weighing model corresponding to the operating machinery. The load weighing model is used to load weigh the operating machinery.
[0016] In a fourth aspect, an embodiment of the present application provides a load weighing device, the device comprising: The second acquisition module is used to acquire the current working condition information of the operating machine; A determination module, used for inputting the current working condition information into a load weighing model to determine the load of the working machine; The load weighing model is a load weighing model trained according to the first aspect.
[0017] In a fifth aspect, an embodiment of the present application provides a processor configured to execute the training method of the load weighing model according to the first aspect and / or the load weighing method according to the second aspect.
[0018] In a sixth aspect, an embodiment of the present application provides a machine-readable storage medium having instructions stored thereon, the instructions being used to enable a machine to execute a training method for a load weighing model according to the first aspect, and / or a load weighing method according to the second aspect.
[0019] The technical solution provided in the embodiment of the present application, first, is based on constructing an initial model. For the complex nonlinear relationships existing in the load weighing problem of the operating machinery, the model can explore the appropriate function combination by continuously learning the training data, so as to approximately represent these relationships. Secondly, this initial model is composed of rational functions. Since rational functions have powerful expressive power, they can excellently fit the nonlinear relationship between the load of the operating machinery and the historical working condition information, which means that the model theoretically has the ability to approximate any complex nonlinear function. For this reason, the model can well fit the nonlinear relationship between the load and the working condition information. After further training, the load weighing model is finally obtained. Using this load weighing model, the load weighing of the operating machinery can be accurately achieved.
[0020] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the accompanying drawings: Figure 1 A schematic diagram of a flow chart of a method for training a load weighing model according to an embodiment of the present application is schematically shown; Figure 2 A schematic diagram of a network structure of an initial model according to an embodiment of the present application is shown; Figure 3 A schematic diagram of a network structure of an initial model after pruning according to an embodiment of the present application is shown; Figure 4 A schematic diagram of a flow chart of a load weighing method according to an embodiment of the present application is schematically shown; Figure 5 A schematic diagram of the structure of a training device for a load weighing model provided in an embodiment of the present application is shown; Figure 6 A schematic diagram of the structure of a load weighing device provided according to an embodiment of the present application is shown; Figure 7 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0023] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations. In the embodiments of this application, some existing solutions in the industry such as certain software, components, and models may be mentioned, which should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.
[0024] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0025] It should be noted that the training method of the load weighing model and the load weighing method provided in the subsequent embodiments of the present application can be applied to working machinery, including but not limited to excavators, loaders and bulldozers. In order to clearly explain the technical solution, the training method of the load weighing model and the load weighing method applied to the excavator are used as an example to explain the embodiments.
[0026] Figure 1 The flowchart of a method for training a load weighing model according to an embodiment of the present application is schematically shown. Figure 1 As shown, an embodiment of the present application provides a method for training a load weighing model, the method comprising: Step 101: obtaining a training data set, the training data set including a plurality of training samples, each training sample including historical working condition information of a working machine and a load label of the working machine corresponding to the historical working condition information, wherein the historical working condition information includes at least one of angle information of a working device of the working machine, pressure information of a hydraulic system of the working machine, and speed-related information of the working device of the working machine; In an embodiment of the present application, the training data set includes multiple training samples. Each training sample consists of historical working condition information of an operating machine and a corresponding load label of the operating machine. The working condition information can be collected by installing a sensor in the operating machine. The operating machine can be a mechanical device such as an excavator or a crane.
[0027] In one example, the working device of a working machine will be at different angles under different working conditions, and the angles can be obtained by installing angle sensors of the boom, dipper arm, and bucket on the working machine. Changes in these angles will directly affect the stress of the load and the overall stability of the working machine. The hydraulic system of the working machine is a key part of driving the working device. When the load changes, the pressure of the hydraulic system will change accordingly in order to maintain the normal operation of the working device. The pressure information can be obtained from pressure sensors installed in the boom cylinder, dipper arm cylinder, and bucket cylinder. The speed-related information of the working device is also related to the load. Speed-related information can be obtained from sensors such as acceleration sensors of the boom, dipper arm, and bucket.
[0028] The information of each state sensor under different loads is collected and normalized to obtain the historical working condition information. The load information is then added as a label to the historical working condition information to construct a training data set.
[0029] Step 102: construct an initial model, where the basis functions of the initial model are rational functions.
[0030] The Kolmogorov-Arnold theorem shows that any multivariate continuous function can be accurately approximated by a composite of single-variable continuous functions. Based on this theorem, the initial model constructed in this application uses rational functions as basis functions. A rational function is a function form obtained by dividing two polynomial functions, and has a strong function expression ability. It can flexibly fit various complex curves and functional relationships through different combinations of coefficients and degrees. In the problem of load weighing of operating machinery, there is a complex nonlinear relationship between the load and the historical working condition information, and this characteristic of rational functions enables it to adapt well to the modeling needs of such complex relationships. For example, by adjusting the first coefficient, the second coefficient and the degree in the rational function, the function can better fit the nonlinear mapping relationship between the working device angle, the hydraulic system pressure, the working device speed and other information and the load.
[0031] Step 103: Take the historical working condition information as input and the load label as output, train the initial model until the preset training stop condition is met, and obtain the load weighing model corresponding to the operating machinery. The load weighing model is used to load weigh the operating machinery.
[0032] The initial model is trained by taking the historical working condition information obtained as the input of the initial model and the corresponding load label as the output. During the training process, the model will calculate according to the input historical working condition information to obtain the load prediction result. Then, by comparing the load prediction result and the load label, the loss function value of the initial model is determined. The loss function is used to measure the degree of difference between the model prediction result and the true value. If the loss function value does not meet the preset training stop condition, it means that there is still a large deviation between the model prediction result and the true value, and the model parameters of the initial model need to be adjusted. The purpose of adjusting the model parameters is to enable the model to output a load prediction result that is closer to the true value in subsequent calculations. After continuously adjusting the model parameters and recalculating the loss function value, when the loss function value corresponding to the updated initial model meets the preset training stop condition, it is considered that the model has achieved a good prediction performance, and the model obtained at this time is the load weighing model corresponding to the operating machinery. The load weighing model can be used to accurately weigh the load of the operating machinery during the actual operation process. The preset training stop condition can be that the loss function value converges to a minimum value, or the loss function value no longer converges, or the number of updates of the loss function value reaches a preset number of times, etc.
[0033] In the training method provided in the embodiment of the present application, first, an initial model is constructed based on the Kolmogorov-Arnold theorem. For the complex nonlinear relationships existing in the load weighing problem of the operating machinery, the model can explore suitable function combinations by continuously learning training data, so as to approximately represent these relationships. Secondly, this initial model is composed of rational functions. Since rational functions have powerful expressive power, they can excellently fit the nonlinear relationship between the load of the operating machinery and the historical working condition information, which means that the model theoretically has the ability to approximate any complex nonlinear function. For this reason, the model can well fit the nonlinear relationship between the load and the working condition information. After further training, a load weighing model is finally obtained. Using this load weighing model, the load weighing of the operating machinery can be accurately achieved.
[0034] In one embodiment of the present application, constructing an initial model includes: Rational functions are used as basis functions in KAN (Kolmogorov - Arnold Network, a neural network based on the Kolmogorov-Arnold theorem), and the denominator of the rational function is operated with an absolute value and an offset to obtain a modified rational function, so as to obtain an initial model based on the modified rational function.
[0035] In this embodiment, the Kolmogorov-Arnold theorem provides a theoretical basis for constructing the initial model, and the original function in the KAN network is fitted with the B-spline (Basis Spline, B-spline) as the basis function. In this embodiment, a rational function can be used instead of the B-spline as a new basis function to construct the initial model.
[0036] Common rational function forms can be expressed as: , in, is the initial model, is the scaling factor, is the first coefficient, is the second coefficient, m and n are the degrees of the polynomial, and x is the input quantity, where the input quantity corresponds to the historical working condition information of the operating machinery, such as the angle of the working device, the pressure of the hydraulic system, or the speed of the working device.
[0037] Furthermore, in order to solve the problem of the original rational function, its denominator is processed as follows. First, the denominator is operated on by absolute value, that is, The main purpose of this step is to prevent the denominator from being negative under certain input data, because a negative denominator may cause the sign of the function value to change, resulting in abnormal output of the model. For example, when the polynomial in the denominator is negative under certain operating conditions, taking the absolute value can ensure that the denominator is always positive, making the range of rational functions more stable and predictable.
[0038] Secondly, an offset operation is performed on the denominator after taking the absolute value. Assuming the offset is ε (ε is a constant greater than zero, usually set according to the characteristics of the actual working condition data and the experience of model training), the corrected denominator is . Optionally, ε in this embodiment can be 1. The purpose of adding an offset operation is to further prevent the denominator from approaching zero. In the actual working conditions of the operating machinery, some extreme situations may occur, causing the denominator to approach zero, which will cause the value of the rational function to increase sharply, affecting the training and prediction effects of the model. By adding an offset, even when the denominator is close to zero, the value of the denominator is always greater than ε, thereby effectively avoiding abnormal fluctuations in the function value and enhancing the stability of the model.
[0039] The initial model obtained through the above process can be:
[0040] in, is the initial model, is the scaling factor, is the first coefficient, is the second coefficient, m and n are the degrees of the polynomial in the initial model, and x is the input quantity.
[0041] In this embodiment, by performing absolute value operations and adding offset operations on the denominator of the rational function, the problem of abnormal function values caused by zero or negative denominators is effectively avoided. In the training process of the load weighing model of the operating machinery, even in the face of complex and changeable working condition data, the initial model can maintain a stable output, and the model training will not be interrupted or erroneous results will not occur due to special circumstances of the denominator. This greatly improves the reliability of model training, reduces the number of times retraining is required due to model instability, and saves time and computing resources.
[0042] In one embodiment of the present application, historical working condition information is used as input and load labels are used as output to train the initial model until a preset training stop condition is met, thereby obtaining a load weighing model corresponding to the operating machine, including: For each training sample, perform the following steps: Input the historical operating condition information into the initial model to obtain the load prediction result corresponding to the historical operating condition information; Determine the loss function value of the initial model based on the load prediction results and load labels; When the loss function value does not meet the preset training stop condition, the initial model is updated to obtain an updated initial model, and the historical operating condition information is returned to the initial model to obtain the load prediction result corresponding to the historical operating condition information, until the preset training stop condition is met to obtain the load weighing model.
[0043] In this embodiment, the historical working condition information in the training sample is input into the initial model. Based on its own structure and parameter settings, the initial model processes and calculates the input historical working condition information, and finally outputs the load prediction result corresponding to the historical working condition information. For example, after the initial model receives the historical working condition information such as the boom angle, boom cylinder pressure, and boom acceleration of the excavator working device, it outputs a predicted value of the current load weight after internal function calculation and parameter weighting.
[0044] According to the load prediction results and load labels, the loss function value of the initial model is determined. The loss function is a key indicator to measure the degree of difference between the model prediction results and the true value. In this embodiment, the loss function used can be the mean square error (MSE) function, and its calculation formula is: , Where MSE is the loss function value, N is the number of training samples, is the load label of the kth sample (i.e. the true load value), It is the load prediction result of the model for the k-th sample. By calculating the loss function value, the deviation between the current prediction result of the model and the true value can be clearly understood.
[0045] When the loss function value does not meet the preset training stop condition, it is necessary to adjust the model parameters of the initial model to obtain an updated initial model. The preset training stop condition can be that the loss function value converges to a minimum value, for example, the loss function value is less than a preset threshold, such as 0.01; it can also be that after a certain number of iterative trainings, the model performance no longer improves significantly. The process of adjusting the model parameters is a process of iterative optimization. By continuously changing the parameter values in the model, the model can output a load prediction result closer to the true value in subsequent calculations. After each adjustment, the historical working condition information needs to be input into the updated initial model again, and the loss function value is calculated again to determine whether it meets the preset training stop condition. When the loss function value corresponding to the updated initial model meets the preset training stop condition, the training process ends, and the model obtained at this time is the load weighing model corresponding to the construction machinery.
[0046] In this embodiment, by processing and iteratively optimizing each training sample, the model can continuously learn the complex relationship between the historical working condition information and the load, and gradually reduce the error between the prediction result and the true load label. After multiple iterative trainings, the model can accurately capture the change law of the load under various working conditions, thereby improving the accuracy of load weighing.
[0047] In an embodiment of the present application, when the loss function value does not meet the preset training stop condition, updating the initial model to obtain an updated initial model includes: When the loss function value does not meet the training stop condition, adjusting the model parameters of the initial model to obtain an optimized initial model; In the optimized initial model, when the absolute value of the weight of the neuron corresponding to the hidden layer is less than the preset value, pruning the neuron to obtain a pruned initial model; Performing precise training on the pruned initial model to obtain an updated initial model.
[0048] In this embodiment, during the training process, when the loss function value does not meet the training stop condition, the Adam optimizer is used to adjust the model parameters of the initial model. Specifically, the Adam optimizer dynamically calculates the learning rate of each parameter based on the current gradient information and the previously calculated gradient first-order moment and second-order moment estimates. The model parameters can converge toward the optimal value faster by adaptively adjusting the gradient. For example, for a parameter w in the model, the Adam optimizer calculates a suitable update amount Δw based on its corresponding gradient gw and the previous gradient accumulation information, thereby updating the value of the parameter w. After adjustment by the Adam optimizer, the optimized initial model is obtained.
[0049] The above-mentioned gradient information can be calculated according to the following formula: ; in, is the gradient, for , for , for The derivative of for The derivative of .
[0050] From this we can see that The square of the denominator of the gradient To ensure that the gradient does not change too drastically, the derivative of the absolute value operation ( is mathematically smooth, so the gradient calculation remains stable, making Avoid gradient explosion numerically and converge more easily.
[0051] As an example, during the training process, the training data can be divided into multiple batches for training. In the present invention, the batch size is set to 250, the step size is set to 50, and the learning rate is set to 0.001. The performance analysis is performed with a step size of 50, that is, the performance of the model is evaluated every 50 training steps (which can be the number of iterations or the number of batches). The learning rate is set to 0.001. The learning rate determines the step size of the model at each parameter update. If the learning rate is too large, the model may skip the optimal solution during the training process, resulting in failure to converge; if the learning rate is too small, the training speed of the model will be very slow, requiring a lot of training time and computing resources.
[0052] It should be noted that if Figure 2 As shown, Figure 2is a network structure diagram of the initial model. The initial model in this embodiment includes at least 4 layers of network structure, wherein the input layer is related to the number of input quantities of the working condition information and can be set to 12. The hidden layer includes at least 2 layers of network structure for feature extraction. The number of outputs of the output layer can be 1, i.e., load.
[0053] In the optimized initial model, when the absolute value of the weight of the neuron corresponding to the hidden layer is less than the preset value, the neuron is pruned. Figure 3 As shown, Figure 3 The network structure diagram of the pruned initial model. The preset value is usually set according to the scale and complexity of the model and the actual training experience, for example, it can be set to 0.01. The purpose of pruning neurons is to reduce the redundant structure of the model and remove neurons that contribute less to the model performance, thereby simplifying the model structure, reducing the consumption of computing resources, and avoiding overfitting. For example, in a neural network model with multiple hidden layers, if the absolute value of the weight of a hidden layer neuron is always small, it means that the neuron has little effect on the output result during the calculation process of the model. By pruning the neuron, the model can be made more concise and efficient. After pruning, the pruned initial model is obtained.
[0054] In one embodiment, the pruned initial model is accurately trained to obtain an updated initial model, including: Using a preset objective function to perform function adjustment on the updated initial model to obtain an adjusted initial model; The adjusted initial model is trained using the training data set to obtain an updated initial model.
[0055] In this embodiment, the historical working condition data of the operating machinery is first analyzed to determine the distribution characteristics and laws of the data. For example, through the analysis of a large amount of crane operation data, it is found that there is a nonlinear trigonometric function relationship between the boom angle and the load, and there is a power function relationship between the hydraulic system pressure and the load. Based on these analysis results, a suitable objective function is selected, such as a custom function that includes a combination of trigonometric functions and power functions:
[0056] in, For custom functions, They represent the historical working condition information such as working device angle, hydraulic system pressure, working device speed, etc. ... and b 1 ...b nare learnable parameters. During the training process, the historical operating condition information and corresponding load labels in the training data set are used to continuously adjust these parameters through optimization algorithms (such as gradient descent algorithms) so that the function f gradually approaches the objective function.
[0057] After completing the function adjustment, enter the precision training stage. During the precision training process, reduce the learning rate of the optimization algorithm, for example, adjust the learning rate from 0.001 to 0.0001, so as to perform more precise parameter fine-tuning. At the same time, increase the number of training iterations to fully train the model. In each iteration, use the training data to calculate the gradient of the loss function to the remaining parameters of the model (such as the weights of the neuron connections in the neural network), and update the parameter values according to the gradient information. During the training process, the validation data set is regularly used to evaluate the performance indicators of the model, such as mean square error, mean absolute error, etc. When the performance indicators of the model on the validation data set meet the preset accuracy requirements, stop training. At this point, the parameter values learned by the model determine the mapping relationship between the quality and each state information. For example, when the mean square error of the model is stably less than 0.01 on the validation set, it is considered that the model has obtained an accurate mapping relationship and can be used for load weighing of operating machinery.
[0058] In this embodiment, the updated initial model is adjusted by using the preset objective function to obtain the adjusted initial model, which provides the model with an expression capability that is more in line with the physical laws of the operating machinery, enabling the model to more accurately fit the complex relationship between the load and historical working condition information. Based on this, precise training further fine-tunes the model parameters, eliminates minor deviations in the model prediction, and improves the prediction accuracy of the model.
[0059] In one embodiment of the present application, when the loss function value does not meet the preset training stop condition, the model parameters of the initial model are adjusted to obtain an updated initial model, including: When the loss function value does not meet the preset training stop condition, the gradient of the model parameters is determined using the pre-acquired gradient expression; According to the gradient, the model parameters of the initial model are adjusted.
[0060] In this embodiment, when the loss function value does not meet the preset training stop condition, the gradient expression pre-acquired in the above embodiment is used to determine the gradient of the model parameters. As an example, during the training process, the current model parameter values and the input historical operating condition information are substituted into the pre-acquired gradient expression to calculate the gradient values of the model parameters w and b at this time. Assume that the calculated gradient of the weight w is gw, and the gradient of the bias b is gb. Then, according to the calculated gradient, adjust the parameters of the initial model. The common parameter adjustment method is based on the idea of gradient descent, that is, the new parameter value is equal to the current parameter value minus the product of the learning rate η and the gradient value. For the weight w, the adjusted new value w new = w-η×gw; for bias b, the adjusted new value b new = b-η×gb. The learning rate η is a pre-set hyperparameter that determines the step size of each parameter adjustment. By continuously adjusting the model parameters according to the gradient, the model can gradually optimize in the direction of reducing the loss function value until the preset training stop condition is met.
[0061] In this embodiment, training is performed by pre-acquiring the gradient expression. During the model training process, there is no need to approximate the gradient through complex numerical calculations, which greatly reduces the amount of calculation, saves computing resources and time, and improves training efficiency.
[0062] In an embodiment of the present application, before determining the loss function value of the initial model according to the load prediction result and the load label, the method further includes: Add a regularization term to the preset loss function to obtain the target loss function; The loss function value of the initial model is determined using the target loss function, load prediction results, and load labels.
[0063] In this embodiment, the regularization term includes an L1 regularization term and an L2 regularization term. The regularization term is added to the preset loss function to obtain a target loss function.
[0064] As an example, adding the L1 regularization term to the loss function, the target loss function L can be obtained as: ; In this embodiment, is the regularization parameter, is the absolute value of the weight of the initial model, n is the number of weights, and for other specific parameter descriptions, please refer to the above embodiment, which will not be repeated in this embodiment. By adding a regularization term to the loss function, the model parameters are constrained, which avoids excessive growth and unreasonable distribution of the model parameters, thereby effectively suppressing the overfitting phenomenon. The generalization ability of the model is improved.
[0065] In one embodiment of the present application, the method further includes: taking historical working condition information as input and load labels as output, training the initial model until a preset training stop condition is met, and obtaining a load weighing model corresponding to the working machine. Convert the mapping relationship in the load weighing model into a mathematical expression; Write mathematical expressions into the embedded controller of the working machine to enable the working machine to weigh the load.
[0066] In this embodiment, the trained load weighing model reflects the mapping relationship between the historical working condition information of the working machine (such as working device angle, hydraulic system pressure, working device speed, etc.) and the load. In this application, it is necessary to convert this complex mapping relationship into a clear mathematical expression.
[0067] After obtaining the mathematical expression, write it into the embedded controller of the working machine. The embedded controller is the core part of the working machine control system, responsible for collecting sensor data (i.e. historical working condition information) and processing it in real time. First, it is necessary to select the appropriate programming language and development tools according to the hardware and software environment of the embedded controller. Then, write the mathematical expression into the corresponding code module. In the code, define the input variables (corresponding to the historical working condition information) and the output variables (corresponding to the load), and program and implement it according to the operation logic of the mathematical expression. During the operation of the working machine, the embedded controller collects information such as the working device angle, hydraulic system pressure and working device speed in real time, calls the above code module, calculates the load weight according to the mathematical expression, and uses the result for the control and display functions of the working machine.
[0068] The model mapping relationship is converted into a mathematical expression and written into the embedded controller, which avoids the complicated model loading and running process, greatly improves the speed of load weighing, and realizes the application of the model.
[0069] Figure 4 The following schematically shows a flow chart of a load weighing method according to an embodiment of the present application. Figure 4 As shown, in an embodiment of the present application, the method includes: Step 401, obtaining current working condition information of the operating machine; Step 402, input the current working condition information into the load weighing model to determine the load of the operating machine.
[0070] In this embodiment, the load weighing model is a load weighing model trained according to the above embodiment. During the operation of the operating machine, various sensors are used to collect the current working condition information of the operating machine in real time. These sensors include but are not limited to angle sensors, pressure sensors and speed sensors.
[0071] The sorted current working condition information is used as input and transmitted to the load weighing model. The load weighing model has been trained with a large amount of historical working condition data and corresponding load labels, and has learned the complex mapping relationship between historical working condition information and load. When the current working condition information is input, the model performs a series of operations and processing on the input data based on its internal structure and trained parameters. Finally, the load prediction value of the operating machinery is output. This load prediction value is the current load of the operating machinery determined by the model.
[0072] In this embodiment, by utilizing the current working condition information collected in real time and combining it with an efficient load weighing model, the load results can be output quickly and accurately to meet the real-time operation requirements of the operating machinery.
[0073] Figure 5 A schematic structural diagram of a training device for a load weighing model provided in another embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0074] Reference Figure 5 The training device 500 of the load weighing model may include: A first acquisition module 501 is used to acquire a training data set, the training data set includes a plurality of training samples, each training sample includes historical working condition information of the working machine and a load label of the working machine corresponding to the historical working condition information, wherein the historical working condition information includes at least one of angle information of a working device of the working machine, pressure information of a hydraulic system of the working machine, and speed-related information of the working device of the working machine; A construction module 502 is used to construct an initial model, wherein the basis functions of the initial model are rational functions; The training module 503 is used to train the initial model with historical operating condition information as input and load labels as output until the preset training stop condition is met, so as to obtain a load weighing model corresponding to the operating machinery. The load weighing model is used to load weigh the operating machinery.
[0075] Optionally, the construction module 502 is specifically configured to: Rational functions are used as basis functions in the KAN neural network, and absolute value operations and offset operations are performed on the denominators of the rational functions to obtain modified rational functions, so as to obtain an initial model based on the modified rational functions.
[0076] Optionally, the initial model is:
[0077] in, is the initial model, is the scaling factor, is the first coefficient, is the second coefficient, m and n are the degrees of the polynomial in the initial model, and x is the input quantity.
[0078] Optionally, the training module 503 includes: For each training sample, perform the following steps: An input submodule is used to input historical operating condition information into the initial model to obtain load prediction results corresponding to the historical operating condition information; A determination submodule is used to determine the loss function value of the initial model according to the load prediction result and the load label; The update submodule is used to update the initial model when the loss function value does not meet the preset training stop condition, obtain the updated initial model, and return the historical operating condition information to the initial model to obtain the load prediction result corresponding to the historical operating condition information until the preset training stop condition is met to obtain the load weighing model.
[0079] Optionally, update submodules, including: An adjustment unit, used to adjust the model parameters of the initial model to obtain an optimized initial model when the loss function value does not meet the training stop condition; A pruning unit, used for pruning neurons in the optimized initial model when the absolute value of the weight of the neurons corresponding to the hidden layer is less than a preset value, so as to obtain a pruned initial model; The training unit is used to accurately train the pruned initial model to obtain an updated initial model.
[0080] Optionally, the training unit comprises: A first adjustment subunit is used to perform function adjustment on the updated initial model using a preset objective function to obtain an adjusted initial model, wherein the objective function is used to standardize the expression form of the initial model; The training subunit is used to train the adjusted initial model using the training data set to obtain an updated initial model.
[0081] Optionally, the adjustment unit comprises: A determination subunit, used to determine the gradient of the model parameters using a pre-acquired gradient expression when the loss function value does not satisfy a preset training stop condition; The second adjustment subunit is used to adjust the model parameters of the initial model according to the gradient.
[0082] Optionally, determine a submodule, specifically for: Add a regularization term to the preset loss function to obtain the target loss function; The loss function value of the initial model is determined using the target loss function, load prediction results, and load labels.
[0083] Figure 6 A schematic structural diagram of a load weighing device provided in another embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0084] Reference Figure 6 , a load weighing device 600 for a working machine may include: The second acquisition module 601 is used to acquire the current working condition information of the operating machine; A determination module 602 is used to input the current working condition information into the load weighing model to determine the load of the working machine; The load weighing model is a load weighing model trained according to a training method of a load weighing model.
[0085] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0086] The device may include a processor 701 and a memory 702 storing program instructions.
[0087] When the processor 701 executes the program, the steps in any of the above method embodiments are implemented.
[0088] Exemplarily, the program may be divided into one or more modules / units, one or more modules / units are stored in the memory 702 and executed by the processor 701 to complete the present application. One or more modules / units may be a series of program instruction segments capable of completing a specific function, and the instruction segments are used to describe the execution process of the program in the device.
[0089] Specifically, the processor 701 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0090] The memory 702 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 702 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 702 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 702 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 702 is a non-volatile solid-state memory.
[0091] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, typically, the memory includes one or more tangible (non-transitory) readable storage media (e.g., memory devices) encoded with software including computer executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0092] The processor 701 implements any one of the methods in the above embodiments by reading and executing program instructions stored in the memory 702 .
[0093] In one example, the electronic device may further include a communication interface 703 and a bus 710. The processor 701, the memory 702, and the communication interface 703 are connected via the bus 710 and communicate with each other.
[0094] The communication interface 703 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0095] Bus 710 includes hardware, software or both, and the parts of online data flow billing equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front-end bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 710 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the present application considers any suitable bus or interconnection.
[0096] In addition, in combination with the method in the above embodiment, the embodiment of the present application can provide a storage medium for implementation. The storage medium stores program instructions; when the program instructions are executed by a processor, any one of the methods in the above embodiment is implemented.
[0097] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0098] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0099] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0100] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0101] The functional modules shown in the above block diagram can be implemented as hardware, software, firmware or their combination. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), suitable firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. Programs or code segments can be stored in machine-readable media, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable media" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. Code segments can be downloaded via computer grids such as the Internet, intranets, etc.
[0102] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0103] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0104] The above are only specific implementation methods of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A training method for a load weighing model, characterized in that: The method comprises: Acquire a training data set, the training data set comprising a plurality of training samples, each of the training samples comprising historical working condition information of a working machine and a load label of the working machine corresponding to the historical working condition information, wherein the historical working condition information comprises at least one of angle information of a working device of the working machine, pressure information of a hydraulic system of the working machine, and speed-related information of the working device of the working machine; Constructing an initial model, wherein the basis functions of the initial model are rational functions; The historical operating condition information is used as input and the load label is used as output to train the initial model until a preset training stop condition is met, thereby obtaining a load weighing model corresponding to the operating machinery. The load weighing model is used to perform load weighing on the operating machinery.
2. The method according to claim 1, characterized in that The constructing of the initial model comprises: The rational function is used as a basis function in a KAN neural network, and an absolute value operation and an offset operation are performed on the denominator of the rational function to obtain a modified rational function, so as to obtain the initial model based on the modified rational function.
3. The method according to claim 1 or 2, characterized in that: The initial model is: in, is the initial model, is the scaling factor, is the first coefficient, is the second coefficient, m and n are the degrees of the polynomial in the initial model, and x is the input quantity.
4. The method according to claim 1, characterized in that: The method uses the historical working condition information as input and the load label as output to train the initial model until a preset training stop condition is met to obtain a load weighing model corresponding to the working machine, including: For each training sample, perform the following steps: Inputting the historical operating condition information into the initial model to obtain a load prediction result corresponding to the historical operating condition information; Determining a loss function value of the initial model according to the load prediction result and the load label; When the loss function value does not meet the preset training stop condition, the initial model is updated to obtain an updated initial model, and the historical operating condition information is returned to the initial model to obtain the load prediction result corresponding to the historical operating condition information, until the preset training stop condition is met to obtain the load weighing model.
5. The method according to claim 4, characterized in that When the loss function value does not satisfy a preset training stop condition, updating the initial model to obtain an updated initial model includes: When the loss function value does not satisfy the training stop condition, adjusting the model parameters of the initial model to obtain an optimized initial model; In the optimized initial model, when the absolute value of the weight of the neuron corresponding to the hidden layer is less than a preset value, the neuron is pruned to obtain a pruned initial model; The pruned initial model is accurately trained to obtain an updated initial model.
6. The method according to claim 5, characterized in that The step of accurately training the pruned initial model to obtain an updated initial model includes: Performing function adjustment on the updated initial model using a preset objective function to obtain an adjusted initial model, wherein the objective function is used to standardize the expression form of the initial model; The adjusted initial model is trained using the training data set to obtain an updated initial model.
7. The method according to claim 4 or 5, characterized in that: When the loss function value does not satisfy a preset training stop condition, adjusting the model parameters of the initial model to obtain an updated initial model comprises: When the loss function value does not satisfy a preset training stop condition, determining the gradient of the model parameter using a pre-acquired gradient expression; According to the gradient, the model parameters of the initial model are adjusted.
8. The method according to claim 4, characterized in that Determining the loss function value of the initial model according to the load prediction result and the load label, further comprising: Add a regularization term to the preset loss function to obtain the target loss function; The loss function value of the initial model is determined using the target loss function, the load prediction result and the load label.
9. A load weighing method, characterized in that: The method comprises: Obtain the current working condition information of the operating machinery; Inputting the current working condition information into a load weighing model to determine the load of the working machine; The load weighing model is a load weighing model trained according to any one of claims 1-8.
10. A training device for a load weighing model, characterized in that: The device comprises: A first acquisition module is used to acquire a training data set, wherein the training data set includes a plurality of training samples, each of the training samples includes historical working condition information of a working machine and a load label of the working machine corresponding to the historical working condition information, wherein the historical working condition information includes at least one of angle information of a working device of the working machine, pressure information of a hydraulic system of the working machine, and speed-related information of the working device of the working machine; A construction module, used for constructing an initial model, wherein the basis functions of the initial model are rational functions; A training module is used to train the initial model using the historical operating condition information as input and the load label as output until a preset training stop condition is met, thereby obtaining a load weighing model corresponding to the operating machinery, and the load weighing model is used to perform load weighing on the operating machinery.
11. A load weighing device, characterized in that: The device comprises: The second acquisition module is used to acquire the current working condition information of the operating machine; A determination module, used for inputting the current working condition information into a load weighing model to determine the load of the working machine; The load weighing model is a load weighing model trained according to any one of claims 1-8.
12. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the load weighing model training method as described in any one of claims 1 to 8, or the load weighing method as described in claim 9.
13. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores instructions for causing a machine to execute the load weighing model training method according to any one of claims 1 to 8, or the load weighing method according to claim 9.
Citation Information
Patent Citations
Method and system for constructing rational function neural network, and readable storage medium
CN113537458A
Truck load estimation method and model training method and device thereof
CN116306254A
Method for determining the weight of a load of a mobile work machine, learning method for a data-based model, and mobile work machine
WO2022069467A1
Sensor physical quantity regression method based on ICS-BP neural network
WO2023231204A1