Agricultural Machinery Behavior Recognition Method Combining Interpolation Enhancement and Deep Spatiotemporal Feature Modeling

Through the combined combination of DBSCAN clustering and variational autoencoder VAE with improved BiLSTM network, the problems of data imbalance and spatial characteristics neglect in agricultural machinery behavior recognition are solved, and the classification accuracy is improved.

CN120180277BActive Publication Date: 2025-07-18QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510623075.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-18
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

There are problems in the existing agricultural machinery behavior recognition technology that data distribution is unbalanced and the spatial distribution characteristics of trajectory points are ignored, resulting in insufficient classification accuracy.

Method used

Data augmentation is performed using DBSCAN clustering, combining motion features and spatial distribution feature extraction, and spatiotemporal modeling is performed through variational autoencoder VAE and improved BiLSTM network to alleviate data imbalance and extract more discriminant features. Finally, Sigmoid function is used for classification.

Benefits of technology

It significantly improves the classification accuracy of agricultural machinery behavior recognition, reduces the misclassification rate of road points, and achieves more accurate field-road classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180277B_ABST
    Figure CN120180277B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of electrical digital data processing, and more specifically, relates to a method for identifying agricultural machinery behavior by integrating interpolation enhancement and depth spatio-temporal feature modeling. The method includes: obtaining GNSS trajectory data of agricultural machinery and preprocessing the trajectory data; using a local interpolation method based on DBSCAN clustering to perform data enhancement on road points in the dataset; designing a set of spatial distribution feature extraction methods; extracting features from the trajectory data after data enhancement by combining motion feature extraction and spatial distribution feature extraction; introducing a variational autoencoder VAE to perform latent feature extraction on the initial features; subsequently, using a ResBiLSTM network for spatio-temporal modeling; and finally, implementing end-to-end classification decision through a linear classifier, so as to realize intelligent identification of agricultural machinery operation behavior. The present invention solves two key problems existing in the existing trajectory recognition technology: one is the serious imbalance in data distribution, and the other is that the spatial distribution characteristics of trajectory points are often ignored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electrical digital data processing, and more specifically, relates to a method for identifying agricultural machinery behavior by integrating interpolation enhancement and deep spatio-temporal feature modeling. Background Art

[0002] With the acceleration of the agricultural modernization process and the popularization of large-scale planting, the integrated development of sensor technology, the Global Navigation Satellite System (GNSS) and Geographic Information System (GIS) provides key support for the intelligentization of agricultural machinery. Through high-precision and real-time trajectory tracking technology, the massive operation data generated by agricultural machinery has become the core resource of smart agriculture. Under this background, the identification of agricultural machinery behavior, that is, accurately determining whether the agricultural machinery is in the field operation or road driving state, has become an important research topic in intelligent agricultural management. As the core technical means to realize the identification of agricultural machinery behavior, the field-road classification technology essentially classifies GNSS trajectory points accurately through machine learning or deep learning methods to distinguish field points from road points, so as to realize the intelligent identification of agricultural machinery operation behavior. This technological breakthrough not only builds a technical bridge from trajectory points to behavior identification, but also significantly reduces the dependence on traditional manpower and hardware, and realizes the rapid measurement of cultivated area. Further analysis shows that the value of the field-road classification technology in the intelligent management of agricultural machinery is reflected in three levels: First, through the automatic classification of agricultural machinery trajectory points, the accurate identification of agricultural machinery operation behavior is realized, providing data support for agricultural production decision-making; Second, based on the statistical analysis of the classification results, the operation efficiency and use intensity of agricultural machinery can be quantitatively evaluated to optimize resource allocation; Third, combined with the precise positioning ability of GNSS, it provides a basis for path planning and operation scheduling, improving agricultural production efficiency. In addition, by developing an intelligent identification system for agricultural machinery behavior, real-time monitoring and analysis tools can be provided for farmers, farms and regulatory departments to promote the efficient allocation of agricultural resources. To sum up, the collaborative application of the field-road classification technology and multi-source perception technology not only reconstructs the data collection and management paradigm of modern agriculture, but also lays a technical foundation for improving agricultural production efficiency and achieving sustainable development goals through precise and intelligent decision-making support, showing broad application prospects.

[0003] The current field-road classification methods are mainly divided into two categories: traditional machine learning and deep learning. Traditional machine learning methods rely on manually designed motion features, such as speed, acceleration, and classical algorithms, such as decision trees and support vector machines, for classification. For example, Poteko et al. extracted 25 motion features and input them into a decision tree, but ignored the spatial distribution features, resulting in poor robustness. Deep learning models have shown better performance in this field due to their advantages in spatio-temporal relationship modeling and deep feature learning. For example, the ConvTEBiLSTM model proposed by Bian et al. realizes the fusion of local and global trajectory features through one-dimensional convolution and Transformer-Encoder networks; and the pixel-level feature extraction method developed by Chen et al. combines trajectory visual features with statistical features. Although deep learning performs excellently in the field-road classification task, its performance is still restricted by trajectory quality, data distribution imbalance, and insufficient feature utilization. First, the quality problem of GNSS trajectory data is prominent. Affected by climate, terrain, and device performance, problems such as noise, drift, and repetition are common, and the existing methods have limited effects on processing noise and offsets. Second, the problem of data distribution imbalance is serious. The number of field points is significantly more than that of road points, resulting in the model being biased towards the features of the majority class, increasing the risk of misclassifying road points, and the existing methods often ignore this problem. Finally, the spatial distribution features of trajectory points, such as parallel relationships, density, and distance, are often ignored. Compared with road points, field points show more changes in the moving direction, higher point density, and smaller spacing between adjacent points, but these key information has not been fully utilized. These factors jointly restrict the further improvement of classification accuracy.

[0004] Chinese patent document CN119761566A discloses a ship trajectory prediction method based on variational autoencoder and BiLSTM. First, by reading the MMSI number, longitude, latitude, speed over ground, and course over ground of the ship trajectory data collected by AIS as the feature sequence of the input data of the trajectory prediction model, and dividing the trajectory sequences of different lengths into sequence segments of fixed length based on the trajectory segmentation of the sliding window; then, performing variational inference on the original trajectory data through the variational autoencoder model, learning the latent space representation, establishing the correlation relationship between the trajectory information and the historical trajectory information, and calculating the reconstructed trajectory; then, using the bidirectional long short-term memory network model to capture the long-term dependence relationship of the trajectory data and learning the characteristics of the trajectory to predict the trajectory position at the next moment.

[0005] In view of this, the present invention designs an agricultural machinery behavior recognition method that combines interpolation enhancement and deep spatio-temporal feature modeling. Summary of the Invention

[0006] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a method for identifying agricultural machinery behavior by integrating interpolation enhancement and depth spatio-temporal feature modeling, so as to solve the problems of unbalanced data distribution and often ignored spatial distribution characteristics of trajectory points in the existing trajectory recognition technology.

[0007] The detailed technical solution of the present invention is as follows:

[0008] A method for identifying agricultural machinery behavior by integrating interpolation enhancement and depth spatio-temporal feature modeling, the method comprising:

[0009] S1. Obtain the GNSS trajectory data of the agricultural machinery and preprocess the trajectory data;

[0010] S2. Data enhancement: Adopt a local interpolation method based on DBSCAN clustering to perform data enhancement on the road points in the preprocessed trajectory dataset, effectively solving the problem of class imbalance commonly existing in agricultural machinery trajectory data;

[0011] S3. Feature extraction: Design a set of spatial distribution feature extraction methods to quantitatively analyze the spatial attributes such as the density feature, parallel relationship and distance relationship of trajectory points;

[0012] Extract features from the trajectory data after data enhancement by combining motion feature extraction and spatial distribution feature extraction, extract the initial features of each trajectory point, each trajectory point consists of 14 initial features, and construct a multi-dimensional feature input space;

[0013] S4. Feature optimization: Introduce a variational autoencoder VAE to perform latent feature extraction on the initial features, and obtain a more discriminative embedded feature representation through non-linear transformation; Subsequently, adopt an improved BiLSTM network - ResBiLSTM network for spatio-temporal modeling: on the one hand, use its bidirectional structure to fully capture the temporal dependence relationship, and on the other hand, optimize the information flow transmission of the network through the residual connection mechanism, effectively alleviating the gradient disappearance problem of deep networks;

[0014] S5. Finally, the deep features extracted by the ResBiLSTM network are input into the fully connected layer for class mapping. Subsequently, use the Sigmoid function to calculate the probability distributions of the field and road classes, and select the class with the highest probability as the classification result, so as to realize the intelligent identification of agricultural machinery operation behavior.

[0015] Preferably according to the present invention, the obtaining of the GNSS data of the agricultural machinery in step S1 refers to:

[0016] Obtain the GNSS trajectory data of the agricultural machinery, wherein each GNSS trajectory point is manually labeled as "1" for the field and "0" for the road class, and a trajectory can be expressed as:

[0017] (1)

[0018] Among them, represents a trajectory, represents the trajectory points that make up the trajectory, represents the trajectory the number of trajectory points on;

[0019] The preprocessing of the trajectory data refers to:

[0020] 1.1 Removing duplicate values: Removing consecutive duplicate trajectory points with the same longitude and latitude coordinates;

[0021] 1.2 Removing drift points: Detecting deviation points by calculating the speed and direction differences between adjacent points, smoothing the calculated difference data, using the exponential weighted average method, which assigns higher weights to the latest data points and lower weights to the older data points to reduce noise and unnecessary fluctuations, thereby eliminating the influence of drift points; At the same time, set a threshold, when the difference value exceeds the threshold, mark these points as drift points and remove them from the trajectory data;

[0022] 1.3 Smoothing noise points: Using the Savitzky-Golay smoothing filter to optimize the trajectory data after removing duplicate values and drift points. This filter is based on the principle of local polynomial least squares fitting and realizes data optimization through the following processing flow: First, perform low-order polynomial fitting on the longitude and latitude coordinates of the trajectory points within a preset sliding window respectively; Then, replace the original observed value of the window center point with the fitting calculated value; Finally, complete the iterative processing of the entire trajectory through window sliding.

[0023] Preferably according to the present invention, in step S2, the specific steps of the data enhancement are as follows:

[0024] According to the spatial distribution characteristics and domain knowledge of the agricultural machinery trajectory data, initially determine the value of the neighborhood radius : Adopt the K-Distance Graph method, calculate the distance from each road point to the Kth nearest neighbor point, draw a distance sorting graph, and find the "inflection point" with a sudden increase. The distance corresponding to this inflection point is the candidate value of the neighborhood radius ;

[0025] Cluster according to the initially determined value of the neighborhood radius and ensure that the points on the same road segment are clustered into one cluster and the points on different road segments belong to different clusters through the visualization evaluation results of the field points and road points;

[0026] Perform linear interpolation on each cluster separately to ensure that the interpolation points are only within a range where the geographical distance is relatively close and the structure is similar to avoid incorrect connection and unreasonable interpolation.

[0027] Preferably according to the present invention, the linear interpolation for each cluster is specifically as follows:

[0028] First, the road points in each cluster are sorted by timestamp;

[0029] Then, check the road point and its adjacent road point for the distance therebetween, and skip it if the distance is greater than a threshold to ensure the spatial rationality of the interpolation points:

[0030] (2)

[0031] Further check whether the directions of these two road points and are consistent, that is, whether the direction angle difference satisfies the following formula to ensure that the interpolation points do not deviate from the driving direction of the road points:

[0032] (3)

[0033] or (4)

[0034] Wherein, represents the direction angle of the road point , represents the direction angle of the road point , is the maximum allowable direction angle difference;

[0035] Finally, perform linear interpolation on the road points that simultaneously meet the distance and direction conditions, and the interpolation formula is defined as follows:

[0036] (5)

[0037] (6)

[0038] Wherein, is the interpolation point, is the interpolation interval, is the number of interpolation points;

[0039] To ensure the rationality of , calculate the distances between all road points in the preprocessed dataset and their nearest neighbor points, and obtain the average value of these distances, which is used as the value of .

[0040] Preferably according to the present invention, the motion feature extraction in step S3 includes nine motion features, which are defined by formulas (7)-(12), and each trajectory point , and are the longitude and latitude, is the timestamp, is the speed, is the direction, , where the speed and direction are directly derived from the point record.

[0041] (7)

[0042] (8)

[0043] (9)

[0044] (10)

[0045] (11)

[0046] , (12)

[0047] wherein, , , , , , , , and respectively represent the speed change, acceleration, direction change, angular velocity, angular acceleration, latitude change, longitude change, distance and curvature of the -th GNSS trajectory point; , , , and respectively represent the speed, timestamp, direction angle, latitude and longitude of the -th GNSS trajectory point; , , , , , and respectively represent the speed, timestamp, direction angle, angular velocity, latitude, longitude and distance of the -th GNSS trajectory point; represents the average radius of the earth and = 6371000; represents the modulo operation, 。

[0048] Preferably according to the present invention, the extraction of the spatial distribution features in step S3 is specifically as follows:

[0049] Define a rectangular area around the current trajectory point, with the width direction consistent with the moving direction and the length direction perpendicular to the moving direction; within the rectangular area, calculate the number of trajectory points around the current trajectory point as its density feature. The density of field points is usually higher than that of road points, and the definition of the density feature is as follows:

[0050] (13)

[0051] Wherein, is the number of trajectory points within the rectangular area centered at point , represents the rectangular area of point , is the number of trajectory points within the trajectory, is an indicator function that returns 1 when point is within the rectangle and returns 0 otherwise;

[0052] Within the rectangular area, calculate the mean and standard deviation of the direction angle differences as the parallel relationship feature. The definition of the parallel relationship is as follows:

[0053] (14)

[0054] (15)

[0055] Wherein, and respectively represent the mean and standard deviation of the direction angle differences of the current point , represents other trajectory points within the rectangular frame, represents the trajectory point and the angular difference between them, is the number of trajectory points within the rectangular area centered at point , represents the rectangular area of point ;

[0056] Extract the distance feature based on the spatial distribution, using a method that combines short-term and long-term sliding windows: the short-term window is used to capture motion fluctuations, and the long-term window is used to reflect the overall trend; specifically, for each trajectory point, calculate its distance within the corresponding window length with two sliding window lengths. In the present invention, the short-term window is set to 5 and the long-term window is set to 10 to effectively highlight the distance difference between road points and field points.

[0057] Preferably according to the present invention, in step S4, the variational autoencoder VAE is introduced to extract latent features from the initial features, and a more discriminative embedded feature representation is obtained through non-linear transformation as follows:

[0058] The 14 initial features extracted in the feature extraction stage are used as the input of the VAE encoder. After passing through the VAE encoder, the input features output two independent vectors: the mean of the latent space distribution and the logarithm of the variance , which are defined as follows:

[0059] (16)

[0060] where I is the input of the VAE encoder, representing 14 initial features; represents the variance;

[0061] Subsequently, the VAE encoder uses the reparameterization technique to sample the latent space. This mechanism enables the gradient to be backpropagated through random nodes. By reconstructing the parameters of the latent vector distribution, the network can better understand the global structure of the data. The reparameterization process is defined as follows:

[0062] (17)

[0063] where, represents the random noise sampled from the standard normal distribution, represents the encoded representation of the initial data in the latent space, represents the variance.

[0064] The reparameterization technique may amplify the gradient, especially when the standard deviation is small. To alleviate the problem of gradient explosion that may occur in VAE training, a gradient clipping strategy is adopted to set the upper limit of the gradient to prevent it from exceeding the predefined threshold; based on gradient distribution analysis and multiple training tests, a process of clipping to the range of [-2, 2] is constructed, which is specifically defined as follows:

[0065] (18)

[0066] where, The function limits

[0067] Preferably according to the present invention, in step S4, the improved BiLSTM network - ResBiLSTM network is used for spatio-temporal modeling as follows:

[0068] After extracting the embedded features through VAE, these features are input into the BiLSTM to capture the deep temporal relationships. Different from the standard LSTM that can only learn forward information, the BiLSTM can capture the information between trajectory points bidirectionally by dividing the hidden neurons into two parts: forward and backward propagation, making full use of the forward and backward information. The BiLSTM network contains two LSTM modules, which are described as follows:

[0069] (19)

[0070] Among them, is the input feature vector of the -th point in the BiLSTM, is the hidden state, is the hidden state of the (t - 1)-th point;

[0071] (20)

[0072] (21)

[0073] (22)

[0074] Among them, captures the temporal relationship from past points and is the forward hidden state, then learns the temporal dependence from future points and is the backward hidden state. The forward and backward hidden states are concatenated to obtain ;

[0075] The stacking of multiple layers of networks in LSTM may lead to the loss of feature information and gradient problems, affecting network training. Adding residual connections can enable the deep network to retain low-level features, capture high-level temporal features at the same time, and provide a shorter transmission path for gradients, alleviating the gradient problem and accelerating network convergence. Therefore, a residual structure is added to the BiLSTM, and the output of the current point is added to the original input, which is defined as follows:

[0076] (23)

[0077] Among them, is the input of the -th point in the BiLSTM,

[0078] Compared with the prior art, the beneficial effects of the present invention are:

[0079] (1) By introducing the DBSCAN-guided data augmentation technique, the present invention effectively alleviates the problem of imbalanced class distribution commonly existing in trajectory datasets, significantly reduces the misclassification rate of road points, i.e., the minority class, thereby improving the overall classification accuracy.

[0080] (2) Based on the essential feature of the spatio-temporal continuity of trajectory data, the present invention integrates the spatial distribution feature and the motion feature to construct a more discriminative multi-dimensional feature representation system.

[0081] (3) In terms of the model architecture design, the present invention organically combines the residual connection mechanism, the bidirectional long short-term memory network BiLSTM, and the variational autoencoder VAE. This multi-module collaborative deep spatio-temporal modeling method significantly improves the feature expression ability, effectively prevents information loss, and finally achieves a more accurate field-road classification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 is a flowchart of the agricultural machinery behavior recognition method that integrates interpolation enhancement and deep spatio-temporal feature modeling according to the present invention.

[0083] Figure 2 is the data imbalance distribution diagram in Embodiment 1 of the present invention.

[0084] Figure 3 is the K-Distance Graph of the wheat dataset in Embodiment 1 of the present invention.

[0085] Figure 4 is the visualization distribution diagram of field points and road points in Embodiment 1 of the present invention.

[0086] Figure 5 is the clustering result diagram of road points in Embodiment 1 of the present invention.

[0087] Figure 6 is the original trajectory diagram before data augmentation in Embodiment 1 of the present invention.

[0088] Figure 7 is the trajectory diagram after data augmentation in Embodiment 1 of the present invention.

[0089] Figure 8 is the feature diagram of density and parallel relationship based on spatial distribution in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0090] The following further describes the present disclosure in conjunction with the drawings and embodiments.

[0091] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.

[0092] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0093] In the absence of conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.

[0094] In order to solve the shortcomings of the prior art, the present invention provides an agricultural machinery behavior recognition method that integrates interpolation enhancement and deep spatiotemporal feature modeling. The method designs an innovative deep neural network framework DBSCAN-VE-ResBiLSTM model, which integrates data enhancement, spatial distribution feature extraction and deep spatiotemporal representation learning for field-road classification. First, the trajectory data is preprocessed using a data smoothing function to improve data quality. Secondly, a DBSCAN-based clustering method is used to perform local interpolation on road points to alleviate the problem of unbalanced data distribution. In addition, the motion features and spatial distribution features are combined to improve the data representation capability. Then, in order to fully explore the potential features in the trajectory data, the BiLSTM integrated into the residual structure is combined with the variational autoencoder VAE, and a VE-ResBiLSTM network is proposed to learn deeper feature representations from the latent space. Finally, a linear classifier is used to complete the field-road classification task.

[0095] The agricultural machinery behavior recognition method integrating interpolation enhancement and deep spatiotemporal feature modeling of the present invention is further described below in conjunction with specific embodiments.

[0096] Embodiment 1,

[0097] Ginseng Figure 1 This embodiment provides an agricultural machinery behavior recognition method integrating interpolation enhancement and deep spatiotemporal feature modeling, the method comprising:

[0098] S1, obtaining GNSS trajectory data of agricultural machinery and preprocessing the trajectory data;

[0099] To prove the effectiveness of the present invention, 150 sets of existing GNSS trajectory data of rice harvesters were selected. These data were collected by the Beidou team of China Agricultural University, and the labels of these trajectory data were manually annotated by the team. Among them, each GNSS trajectory point was labeled as "1" for the field and "0" for the road category. A trajectory can be expressed as:

[0100] (1)

[0101] Wherein, represents a trajectory, represents the trajectory points that make up the trajectory, represents the trajectory the number of points on.

[0102] The GNSS receiver is usually installed at a fixed position of the agricultural machinery, such as the top of the cab or the center position of the vehicle, to receive satellite signals in real time and calculate the longitude and latitude coordinates of the agricultural machinery. However, GNSS signals are vulnerable to climate conditions, terrain, and equipment performance limitations in an open environment. At the same time, frequent braking operations of agricultural machinery, such as stopping and turning, will also affect the performance of the GNSS receiver. These factors together lead to problems such as noise, drift, and repetition in the collected GNSS agricultural machinery trajectory data. These data quality problems will not only reduce the stability of the trajectory point classification results but also weaken the robustness and classification accuracy of the model. The following steps are specific data preprocessing operations:

[0103] 1.1 Removing duplicate values: When the agricultural machinery pauses operation in the field but remains powered on, the GNSS receiver will continuously record coordinates. However, since the agricultural machinery does not actually move, a large number of duplicate points will be generated. Therefore, consecutive duplicate trajectory points with the same longitude and latitude coordinates are removed.

[0104] 1.2 Removing drift points: Drift points refer to abnormal points that deviate significantly from the rest of the data points in the data, usually caused by instantaneous errors or other abnormal situations. Compared with the points around them, the distance of these points is significantly larger, which does not conform to the trend of the overall data. These deviated points are detected by calculating the speed and direction differences between adjacent points, and the calculated difference data is smoothed using the exponential weighted average method. This method assigns a higher weight to the latest data points and a lower weight to the older data points to reduce noise and unnecessary fluctuations, thereby eliminating the influence of drift points. A threshold of 2 is set. When the difference value exceeds this threshold, these points are marked as drift points and removed from the trajectory data.

[0105] 1.3 Smooth noise points: During the GNSS trajectory acquisition process, signal noise interference is inevitable, and this noise will significantly affect the accuracy of trajectory data and the reliability of subsequent analysis. To address this issue, in this embodiment, a Savitzky-Golay smoothing filter is used to optimize the trajectory data after removing duplicate values and drift points. This filter is based on the principle of local polynomial least squares fitting and realizes data optimization through the following processing flow: First, low-order polynomial fitting is performed on the longitude and latitude coordinates of the trajectory points within a preset sliding window; then, the original observed value at the center point of the window is replaced with the fitting calculated value; finally, the entire trajectory is iteratively processed by sliding the window. This processing method has two advantages: on the one hand, it can effectively eliminate the interference of random noise and outliers on the data, making the trajectory smoother and more continuous; on the other hand, it can better retain the main movement trend and spatial pattern characteristics of the original data, ensuring that the processed trajectory data not only improves the quality but also maintains a high degree of consistency with the actual movement trajectory.

[0106] S2. Data augmentation: A local interpolation method based on DBSCAN clustering is used to augment the road points in the preprocessed trajectory dataset, effectively solving the problem of class imbalance commonly existing in agricultural machinery trajectory data.

[0107] Specifically, the DBSCAN-guided method is used for data augmentation:

[0108] The imbalance problem of GNSS points significantly affects the classification accuracy, and the data is unevenly distributed as Figure 2 shown. To solve this problem, the present invention proposes a local interpolation method based on DBSCAN clustering for data augmentation, which improves the traditional oversampling technique of synthesizing minority class samples. By supplementing samples in the sparse areas of road points in each cluster, it not only alleviates the imbalance of data distribution but also retains the original distribution characteristics, effectively avoiding the problem that the traditional SMOTE method is prone to generating noise samples when the data is imbalanced. DBSCAN is a density-based clustering algorithm, and its core idea is that clusters are composed of core points with higher density. The determination of core points depends on two parameters: the neighborhood radius and the minimum number of points within the neighborhood .

[0109] The so-called core point means that if the number of points in the neighborhood of point is greater than , then point is a core point.

[0110] The neighborhood of the point refers to: the neighborhood of point , denoted as , is defined as: , where is a set of points in the trajectory, is the point and the point is the Euclidean distance between them.

[0111] DBSCAN completes the clustering process by successively extracting clusters. Starting from an arbitrary point , its neighborhood is checked. If is a core point, a new cluster is created with the points in . Then, for each point in whose neighborhood has not been checked yet, if is a core point, the points in that have not been included in are added to the cluster. The expansion of the cluster continues until no new points can be added to the cluster. The clustering process terminates when no new clusters are created. Points that are neither core points nor within the neighborhood of any core point are labeled as outliers.

[0112] The agricultural machinery trajectory contains multiple road segments or turning areas, and clustering needs to effectively distinguish different road segments. Due to the spatial separation between segments, trajectory points should not be wrongly clustered, so as not to affect the data accuracy when connecting errors occur during interpolation. Therefore, the goal of clustering is to cluster the points on the same road segment into the same cluster and avoid cross-segment interference. When performing DBSCAN clustering, the selection of the value is crucial.

[0113] If is too small, it may cause the trajectory points on the same segment to be wrongly segmented or labeled as outliers. If it is too large, it may wrongly cluster different road segments into one cluster, especially in areas with close distances or intersections, affecting the interpolation and classification performance. value, this embodiment adopts the K-Distance Graph method to select appropriate parameters by showing the change of data density. Calculate the distance from each road point to its Kth nearest neighbor, draw a distance sorting graph, and find the sudden increase "inflection point". The distance corresponding to this inflection point is the candidate value of . Because the points before the inflection point belong to the high-density area and the distances are small, while the points after the inflection point belong to the low-density area or outliers and the distances are large.Figure 3 shows the K-Distance Graph of the wheat dataset, and the points marked with red circles are candidate values.

[0114] In this embodiment, clustering is first performed according to the preliminarily determined value, and the results are evaluated visually to ensure that the points on the same road segment are clustered into one cluster, and the points on different road segments belong to different clusters. Since there is a spatial distance separation between different road segments, which is usually connected by field areas. Therefore, in this embodiment, each road segment, that is, each cluster, is interpolated separately to maintain the original spatial distance and avoid incorrect connections. If the clustering effect is not ideal, then is fine-tuned and optimized by combining domain knowledge and data density , and finally the optimal parameter combination is determined through multiple experiments. Figure 4 shows the visual distribution of field points and road points, Figure 5 shows the clustering results of road points. Different road segments are correctly divided into different clusters, and isolated points are marked in black.

[0115] In this embodiment, linear interpolation is performed separately for each cluster to ensure that the interpolation points are only within a range where the geographical distance is relatively close and the structure is similar, so as to avoid incorrect connections and unreasonable interpolation. The local interpolation calculation cost is lower than that of global interpolation, especially more efficient when the data volume is large.

[0116] First, the road points in each cluster are sorted by timestamp. Then, check the road points and its adjacent road points to see if the distance between them is less than the threshold . If it is greater than the threshold, skip it to ensure the spatial rationality of the interpolation points.

[0117] (2)

[0118] Further check whether the directions of these two road points and are the same, that is, whether the direction angle difference satisfies the following formula to ensure that the interpolation points do not deviate from the driving direction of the road points, where is the maximum allowable direction angle difference.

[0119] (3)

[0120] or (4)

[0121] where, represents the direction angle of the road point , represents the road point The direction angle is the maximum allowable direction angle difference;

[0122] Finally, linear interpolation is performed on the road points that simultaneously meet the distance and direction conditions. The interpolation formula is defined as follows:

[0123] (5)

[0124] (6)

[0125] Where is the interpolation point, is the interpolation interval, is the number of interpolation points. To ensure is reasonable, calculate the distances between all road points in the preprocessed dataset and their nearest neighbors, and obtain the average of these distances, which is used as the value of .

[0126] In this embodiment, is set to be the same as the value of DBSCAN clustering to ensure that the interpolation follows the spatial constraints of the clustering and avoid cross-cluster interpolation. In trajectory data, the direction angles of adjacent points are usually consistent. If the direction difference is too large, it may indicate that they belong to different trajectory segments. Therefore, in this embodiment, is set to to cover common agricultural machinery steering situations and avoid excessive direction changes, ensuring that the interpolation points accurately reflect the trajectory trend.

[0127] Figure 6 and Figure 7 show the change in the distribution density of trajectory points before and after data augmentation, where Figure 6 is the original trajectory map, Figure 7 is the trajectory map after data augmentation; The DBSCAN-guided data augmentation method effectively generates road points, helping the model learn more road point features, thereby more accurately distinguishing field points from road points.

[0128] S3. Feature extraction: Design a set of spatial distribution feature extraction methods to quantitatively analyze the spatial attributes such as the density feature, parallel relationship, and distance relationship of trajectory points;

[0129] By combining motion feature extraction and spatial distribution feature extraction, feature extraction is performed on the trajectory data after data augmentation, and the initial features of each trajectory point are extracted. Each trajectory point consists of 14 initial features, constructing a multi-dimensional feature input space;

[0130] In the feature extraction module of the present invention, a statistical method is designed to extract the motion and spatial distribution features of trajectory points from trajectory data. For a set of trajectory points , each GNSS point records , and as the latitude and longitude, as the timestamp, as the speed, as the direction, is the number of points of the trajectory.

[0131] Motion feature extraction: Agricultural machines usually drive straight at a relatively fast speed on roads, while they turn frequently and adjust directions during field operations, and the speed is low and tends to be uniform. Therefore, motion features can effectively distinguish the behaviors of agricultural machines in the field and on the road. According to the trajectory point records, nine motion features are constructed in this embodiment and defined by formulas (7)-(12), where the speed and direction directly come from the point records.

[0132] (7)

[0133] (8)

[0134] (9)

[0135] (10)

[0136] (11)

[0137] , (12)

[0138] Among them, , , , , , , , and respectively represent the speed change, acceleration, direction change, angular velocity, angular acceleration, latitude change, longitude change, distance, and curvature of the th GNSS trajectory point; , , , and respectively represent the speed, timestamp, direction angle, latitude, and longitude of the th GNSS trajectory point; , , , , , and respectively represent the velocity, timestamp, direction angle, angular velocity, latitude, longitude, and distance of the th GNSS trajectory point; represents the average radius of the Earth and = 6371000; represents the modulo operation, .

[0139] Spatial distribution characteristics: As shown on the Figure 8 left side, in this embodiment, a rectangular area is defined around the current GNSS point, with the width direction consistent with the moving direction and the length direction perpendicular to the moving direction.

[0140] As shown on the Figure 8 left side, field points are usually more densely distributed, while road points are relatively sparse. Within the rectangular area, the density feature is calculated by counting the number of GNSS points around the current point. The density of field points is usually higher than that of road points. The definition of the density feature is as follows:

[0141] (13)

[0142] where, is the number of GNSS points within the rectangular area centered at point , represents the rectangular area of point , is the number of GNSS points within the trajectory, is the indicator function that returns 1 when point is inside the rectangle and 0 otherwise.

[0143] Figure 8 The right side shows the difference in direction angles between the current point and adjacent points within the rectangular area. Compared with road points, the difference in direction angles between field points is larger, and the fluctuation of the parallel relationship is more obvious. Within the rectangular area, the mean and standard deviation of the direction angle difference are calculated as the parallel relationship feature. The definition of the parallel relationship is as follows:

[0144] (14)

[0145] (15)

[0146] where, and respectively represent the mean and standard deviation of the direction angle difference of the current point , represents other GNSS points within the rectangular frame, represents the GNSS point and The angular difference between and . Since the distance between adjacent GNSS points is about 5 meters, the width of the rectangle is set to 20 meters and the length is set to 40 meters.

[0147] Continue to extract distance features based on spatial distribution because the spacing of road points is usually greater than that of field points. For this purpose, this embodiment adopts a method combining short-term and long-term sliding windows: the short-term window is used to capture motion fluctuations such as acceleration or deceleration, and the long-term window is used to reflect the overall trend. Specifically, for each trajectory point, its corresponding distance within the window length is calculated with two sliding window lengths. In the present invention, the short-term window is set to 5 and the long-term window is set to 10 to effectively highlight the distance difference between road points and field points.

[0148] S4. Feature optimization: Introduce a variational autoencoder VAE to extract latent features from the initial features, and obtain a more discriminative embedded feature representation through non-linear transformation; subsequently, adopt an improved BiLSTM network - ResBiLSTM network for spatio-temporal modeling: on the one hand, use its bidirectional structure to fully capture temporal dependencies, and on the other hand, optimize the information flow transmission of the network through the residual connection mechanism, effectively alleviating the problem of gradient disappearance in deep networks.

[0149] Previous studies mainly used BiLSTM to directly process the initial features for field-road classification, ignoring the importance of latent spatial features, resulting in low classification accuracy. Compared with the initial features, latent spatial features can provide richer semantic information and more continuous data representation. For this purpose, this embodiment uses a variational autoencoder VAE to enhance the representation ability of the initial features. Specifically, the proposed VE-ResBiLSTM network consists of three modules: a VAE encoder, a BiLSTM with a fusion residual structure, and a linear classifier.

[0150] VAE encoder module: The variational autoencoder VAE consists of an encoder and a decoder, and can learn the latent representation of data and generate new samples. In the present invention, only the encoder is used, focusing on feature extraction rather than data reconstruction. The latent space of VAE is a low-dimensional representation of data, capturing core features and latent structures, providing a compact and continuous embedding method, making similar data points close to each other in the latent space. As Figure 1 shown, the input feature of VAE has a dimension of 14, corresponding to the number of initial features of the feature extraction module. After the input feature passes through the encoder, two independent vectors are output: the mean of the latent space distribution and the logarithm of the variance . Defined as follows:

[0151] (16)

[0152] Among them, I is the input of the encoder VAE, representing 14 initial features; represents variance;

[0153] Subsequently, the VAE encoder uses the reparameterization technique to sample the latent space. This mechanism enables the gradient to be backpropagated through random nodes. By reconstructing the parameters of the latent vector distribution, the network can better understand the global structure of the data. The reparameterization process is defined as follows:

[0154] (17)

[0155] Among them, represents the random noise sampled from the standard normal distribution, represents the encoded representation of the initial data in the latent space, represents variance.

[0156] The reparameterization technique may amplify the gradient, especially when the standard deviation is small. To alleviate the problem of gradient explosion that may occur in VAE training, the present invention adopts a gradient clipping strategy, sets an upper limit for the gradient, and prevents it from exceeding a predefined threshold. Based on gradient distribution analysis and multiple training tests, this paper constructs a process of clipping to the range of [-2, 2], which is specifically defined as follows:

[0157] (18)

[0158] Among them, The function limits

[0159] to the range of [-2, 2]. This function helps to maintain the stability of the gradient and reduce the impact of gradient explosion on model training.

[0160] (19)

[0161] Among them, is the input feature vector of the th point in the BiLSTM, is the hidden state, is the hidden state of the (t - 1)th point;

[0162] (20)

[0163] (21)

[0164] (22)

[0165] Among them, Capture the time relationship from past points, which is the forward hidden state. Then learn the time dependence from future points, which is the backward hidden state. Connect the forward and backward hidden states to obtain .

[0166] The stacking of multiple layers of networks in LSTM may lead to feature information loss and gradient problems, affecting network training. Adding residual connections can enable deep networks to retain low-level features, capture high-level temporal features at the same time, and provide a shorter transmission path for gradients, alleviating gradient problems and accelerating network convergence. Therefore, the present invention adds a residual structure to BiLSTM, adding the output of the current point to the original input, which is defined as follows:

[0167] (23)

[0168] Among them, is the input at the t-th point in BiLSTM. is the state after adding the residual. The residual connection is located after the output of each layer of LSTM. That is, after each layer of LSTM processes the sequence, its output is added to the residual feature to obtain a new feature and input it into the next layer of LSTM.

[0169] S5. Finally, the deep features extracted by the ResBiLSTM network are input into the fully connected layer for class mapping. Subsequently, the Sigmoid function is used to calculate the probability distributions of the field and road classes. The model selects the class with the highest probability as the classification result and outputs the classification results of each trajectory point, thereby realizing the intelligent recognition of agricultural machinery operation behaviors.

[0170] Experimental example

[0171] The method of the present invention proposes an innovative deep neural network framework, the DBSCAN-VE-ResBiLSTM model, which integrates data augmentation, spatial distribution feature extraction, and deep spatio-temporal representation learning for field-road classification. Experimental results show that the model has significant advantages in classification performance:

[0172] The model training is based on the PyTorch framework and is carried out in the NVIDIA A100 GPU environment. The wheat dataset is randomly divided into three subsets: 80% for training, 10% for validation, and 10% for testing. Referring to previous studies, the model evaluation adopts 10-fold cross-validation. In the model training stage, = 0.027, = 3, the maximum number of training epochs is set to 400, the optimizer is AdamW, the learning rate is 0.0005, and the weight decay is 0.00001. In addition, the focal loss FL is used to train the model, and the sample weight is 0.25, is an adjustable parameter, set to 2, to make the model pay more attention to difficult-to-classify samples and improve the accuracy of difficult-to-predict points. Subsequently, L2 regularization RT is introduced into the focal loss to prevent the model from overfitting on the training data. The total loss TL is defined as follows:

[0173] (24)

[0174] (25)

[0175] (26)

[0176] Among them, represents the predicted probability of the true label of the model, represents the regularization of the model parameters regularization, as the regularization coefficient, by default is set to 0.00001.

[0177] After every 10 epochs of training, the best hyperparameters of the model are evaluated and saved using the validation data. An early stopping mechanism is introduced. When the performance metric on the validation set has not improved for 10 consecutive epochs, the training automatically terminates to prevent overfitting.

[0178] In the model validation stage, the optimal model obtained from training is used to predict the test dataset, and the classification result of each trajectory point is output, that is, the field or road category. To comprehensively evaluate the model performance, four standard evaluation metrics are selected in the present invention: precision Pre, recall Rec, F1-score F1, and accuracy Acc. The calculation of these metrics is based on the four basic elements of the confusion matrix: true positive TP, true negative TN, false positive FP, and false negative FN. Specifically, precision reflects the accuracy of the model in predicting positive class samples, and recall measures the completeness of the model in identifying positive class samples. The F1-score, as the harmonic mean of precision and recall, can comprehensively evaluate the classification performance of the model; accuracy directly reflects the correct rate of the overall prediction of the model. To ensure the statistical significance of the experimental results, a ten-fold cross-validation method is adopted in this study, and the average performance metrics of ten independent experiments are finally reported. The calculation formulas of each metric are defined as follows:

[0179] (27)

[0180] (28)

[0181] (29)

[0182] (30)

[0183] Here, the "field" points are regarded as "P (Positive)", and the "road" points are regarded as "N (Negative)". Therefore, TP represents the field points correctly classified as "field"; TN represents the road points correctly classified as "road"; FP represents the road points misclassified as "field"; FN represents the field points misclassified as "road".

[0184] To prove the effectiveness of the invention, the proposed method DBSCAN-VE-ResBiLSTM was evaluated against four existing field-road pattern classification methods, including DBSCAN + Rules, Decision Tree DT, Graph Convolutional Network GCN, and STF+VFAU+BiLSTM. The experimental results of the five methods are shown in Table 1. The field-road classification method DBSCAN-VE-ResBiLSTM proposed in the present invention achieved the best results in all four evaluation metrics. Specifically, on the wheat dataset, the method DBSCAN-VE-ResBiLSTM of the present invention obtained an accuracy of 99.09% and an F1 score of 98.61%, which were 8.16% and 15.03% higher than the best baseline STF+VFAU+BiLSTM, respectively. These results fully confirm the excellent effectiveness of the invention in field-road classification.

[0185] Table 1 Overall performance of various methods on the wheat dataset

[0186]

[0187] The present invention proposes a solution to the problems existing in the field-road classification task. First, by introducing the DBSCAN-guided data augmentation technique, the problem of class distribution imbalance commonly existing in the trajectory dataset is effectively alleviated, and the misclassification rate of road points is significantly reduced, thereby improving the overall classification accuracy. Second, based on the essential characteristics of the spatio-temporal continuity of trajectory data, the spatial distribution characteristics and motion characteristics are fused to construct a more discriminative multi-dimensional feature representation system. In terms of model architecture design, the present invention organically combines the residual connection mechanism, the bidirectional long short-term memory network BiLSTM, and the variational autoencoder VAE. This multi-module collaborative deep spatio-temporal modeling method significantly improves the feature expression ability, effectively prevents information loss, and finally achieves a more accurate field-road classification effect.

[0188] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An agricultural machinery behavior recognition method integrating interpolation enhancement and depth spatio-temporal feature modeling, characterized in that, The method includes: S1. Obtain the GNSS trajectory data of the agricultural machinery and preprocess the trajectory data; S2. Data augmentation: Adopt a local interpolation method based on DBSCAN clustering to perform data augmentation on the road points in the preprocessed trajectory dataset; S3. Feature extraction: Design a set of spatial distribution feature extraction methods to quantitatively analyze the spatial attributes such as the density feature, parallel relationship, and distance relationship of the trajectory points; Extract features from the data-augmented trajectory data by combining motion feature extraction and spatial distribution feature extraction, extract the initial features of each trajectory point, each trajectory point consists of 14 initial features, and construct a multi-dimensional feature input space; S4. Feature optimization: Introduce a variational autoencoder (VAE) to perform latent feature extraction on the initial features, and obtain a more discriminative embedded feature representation through non-linear transformation; Subsequently, adopt an improved BiLSTM network - ResBiLSTM network for spatio-temporal modeling: on the one hand, use its bidirectional structure to fully capture the temporal dependence relationship, and on the other hand, optimize the information flow transmission of the network through the residual connection mechanism; S5. Finally, the deep features extracted by the ResBiLSTM network are input into the fully connected layer for class mapping. Subsequently, use the Sigmoid function to calculate the probability distributions of the field and road categories, and select the category with the highest probability as the classification result, so as to realize the intelligent identification of the agricultural machinery operation behavior.

2. The agricultural machinery behavior recognition method that combines interpolation enhancement and depth spatio-temporal feature modeling according to claim 1, characterized in that, The obtaining of the GNSS data of the agricultural machinery in step S1 refers to: Collect the GNSS trajectory data of the agricultural machinery, where each GNSS trajectory point is labeled as the "1" field and "0" road categories, and a trajectory can be expressed as: (1) Among them, represents a trajectory, represents the trajectory points that make up the trajectory, represents the trajectory the number of trajectory points on the trajectory; The preprocessing of the trajectory data refers to: 1.1 Remove duplicate values: Remove consecutive duplicate trajectory points with the same longitude and latitude coordinates; 1.2 Remove drift points: Detect deviation points by calculating the speed and direction differences between adjacent points, smooth the calculated difference data, use the exponential weighted average method to eliminate the influence of drift points; At the same time, set a threshold, when the difference value exceeds the threshold, mark these points as drift points and remove them from the trajectory data; 1.3 Smooth noise points: Adopt the Savitzky-Golay smoothing filter to optimize the trajectory data after removing duplicate values and drift points. This filter is based on the principle of local polynomial least squares fitting and realizes data optimization through the following processing flow: First, perform low-order polynomial fitting on the longitude and latitude coordinates of the trajectory points within a preset sliding window respectively; Then, replace the original observed value of the window center point with the fitting calculated value; Finally, complete the iterative processing of the entire trajectory through window sliding.

3. The agricultural machinery behavior recognition method integrating interpolation enhancement and depth spatio-temporal feature modeling according to claim 1, characterized in that, In step S2, the specific steps of the data augmentation are as follows: According to the spatial distribution characteristics of agricultural machinery trajectory data and domain knowledge, preliminarily determine the value of the neighborhood radius: Adopt the K-Distance Graph method to calculate the distance from each road point to the Kth nearest neighbor point, draw a distance sorting graph, and find the sudden "inflection point". The distance corresponding to this inflection point is the candidate value of the neighborhood radius; ​ Cluster according to the initially determined neighborhood radius and ensure that points on the same road segment are clustered into one cluster and points on different road segments belong to different clusters through the visualization evaluation results of field points and road points; Perform linear interpolation on each cluster separately, ensuring that the interpolation points are only within the range of relatively close geographical distance and similar structure.

4. The agricultural machinery behavior recognition method integrating interpolation enhancement and depth spatio-temporal feature modeling according to claim 3, wherein The specific linear interpolation on each cluster is as follows: First, sort the road points in each cluster by timestamp; Then, check the road points and the adjacent road points to see if the distance between them is less than the threshold . If it is greater than the threshold, skip it to ensure the spatial rationality of the interpolation points: (2) Further check the directions of these two road points and to see if they are the same, i.e., whether the difference in direction angles satisfies the following formula to ensure that the interpolation point does not deviate from the driving direction of the road point: (3) or (4) Among them, represents the direction angle of the road point , represents the direction angle of the road point , is the maximum allowable direction angle difference; Finally, perform linear interpolation on the road points that meet both the distance and direction conditions, and the interpolation formula is defined as follows: (5) (6) Among them, is the interpolation point, is the interpolation interval, is the number of interpolation points; calculate the distances between all road points in the pre-processed dataset and their nearest neighbor points, and obtain the average value of these distances, and use it as the value.

5. The agricultural machinery behavior recognition method integrating interpolation enhancement and depth spatio-temporal feature modeling according to claim 1, characterized in that The motion feature extraction described in step S3 includes nine motion features, which are defined by formulas (7)-(12), and each trajectory point , and are the longitude and latitude, is the timestamp, is the speed, is the direction, , where the speed and direction are directly derived from the point records: (7) (8) (9) (10) (11) , (12) Among them, , , , , , , , and respectively represent the velocity change, acceleration, direction change, angular velocity, angular acceleration, latitude change, longitude change, distance, and curvature of the th GNSS trajectory point; , , , and respectively represent the velocity, timestamp, direction angle, latitude, and longitude of the th GNSS trajectory point; , , , , , and respectively represent the velocity, timestamp, direction angle, angular velocity, latitude, longitude, and distance of the th GNSS trajectory point; represents the average radius of the Earth and = 6371000; represents the modulo operation, .

6. The agricultural machinery behavior recognition method integrating interpolation enhancement and depth spatio-temporal feature modeling according to claim 1, characterized in that The spatial distribution feature extraction in step S3 is specifically as follows: Define a rectangular area around the current trajectory point, with the width direction consistent with the moving direction and the length direction perpendicular to the moving direction; within the rectangular area, calculate the number of trajectory points around the current trajectory point as its density feature. The density of field points is usually higher than that of road points, and the definition of the density feature is as follows: (13) Among them, is the number of trajectory points within the rectangular area centered at point . represents the rectangular area of point . is the number of trajectory points within the trajectory. is the indicator function, which returns 1 when point is within the rectangle and returns 0 otherwise. Within the rectangular area, calculate the mean and standard deviation of the direction angle differences as the parallel relationship feature, and the definition of the parallel relationship is as follows: (14) (15) Among them, and respectively represent the mean and standard deviation of the direction angle difference of the current point ; represents other trajectory points within the rectangular box, represents the trajectory point and the angular difference between them, is the number of trajectory points within the rectangular area centered at the point ; represents the rectangular area of the point . Extract the distance feature based on spatial distribution, using a method that combines short-term and long-term sliding windows: the short-term window is used to capture motion fluctuations, and the long-term window is used to reflect the overall trend; specifically, for each trajectory point, calculate the distance within its corresponding window length with two sliding window lengths.

7. The agricultural machinery behavior recognition method integrating interpolation enhancement and depth spatio-temporal feature modeling according to claim 1, wherein In step S4, the variational autoencoder VAE is introduced to extract latent features from the initial features, and a more discriminative embedded feature representation is obtained through non-linear transformation as follows: The 14 initial features extracted in the feature extraction stage are used as the input of the VAE encoder. After passing through the VAE encoder, the input features output two independent vectors: the mean of the latent space distribution and the logarithm of the variance , which are defined as follows: (16) Among them, I is the input of the encoder VAE, representing 14 initial features; represents the variance; Subsequently, the VAE encoder uses the reparameterization technique to sample the latent space, and the reparameterization process is defined as follows: (17) wherein, represents random noise sampled from a standard normal distribution, represents the encoded representation of the initial data in the latent space, represents the variance; Adopt a gradient clipping strategy, set an upper limit for the gradient to prevent it from exceeding a predefined threshold. Specifically, construct a process that clips to the range [-2, 2], which is defined as follows: (18) Among them, The function will be restricted to the range of [-2, 2].

8. The agricultural machinery behavior recognition method integrating interpolation enhancement and depth spatio-temporal feature modeling according to claim 1, characterized in that, In step S4, the improved BiLSTM network - ResBiLSTM network is used for spatio-temporal modeling, as follows: After extracting the embedded features through VAE, these features are input into BiLSTM to capture deep time relationships; the BiLSTM network contains two LSTM modules, which are described as follows: (19) Among them, is the input feature vector of the th point in the BiLSTM, is the hidden state, is the hidden state of the (t - 1)th point; (20) (21) (22) Among them, capture the temporal relationship from past points as the forward hidden state, then learn the temporal dependence from future points as the backward hidden state, and connect the two forward and backward hidden states to obtain ; Add a residual structure to BiLSTM, and add the output of the current point to the original input, as defined below: (23) Among them, is the input at the t-th point in the BiLSTM, is the state after adding the residual. The residual connection is located after the output of each layer of LSTM. That is, after each layer of LSTM processes the sequence, its output is added to the residual feature to obtain a new feature and input it into the next layer of LSTM.

Citation Information

Patent Citations

  • Agricultural machinery motion behavior recognition method, device and equipment based on multi-feature enhancement and storage medium

    CN118520371A

  • Ship trajectory prediction method based on variational autoencoder and BiLSTM

    CN119761566A