Regional scale soil classification method and system based on multi-modal data
By combining the sliding window mechanism and the Transformer architecture, the characteristics of soil images and structured data are extracted, and data fusion and reconstruction are carried out through the cross-co-attention mechanism and multimodal autoencoder, the efficiency and accuracy problems of traditional soil classification methods are solved when processing large-scale complex data, and efficient and accurate soil classification is achieved.
Patent Information
- Application Number
- CN202510431980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional soil classification methods rely on manual experience and rule-based feature extraction, and cannot effectively process large-scale complex soil data, and have high computational overhead and dependence on large amounts of labeled data.
The regional scale soil classification method based on multimodal data is adopted, and the characteristics of images and structured data are extracted through innovative technologies combining sliding window mechanism and Transformer architecture, and data fusion and reconstruction are carried out through cross-co-attention mechanism and multimodal autoencoder, and shared representations are optimized to build a soil classification model.
It improves the accuracy and efficiency of soil classification, reduces calculation overhead, and enhances the ability to identify soil species, providing more accurate and reliable decision-making support for land resource management, agricultural production and environmental protection.
Smart Images

Figure CN119939364A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multimodal deep learning, and in particular relates to a regional scale soil classification method and system based on multimodal data. Background Art
[0002] Accurate identification of soil types has important application value in the fields of land resource management, agricultural production, environmental protection, etc. By accurately identifying soil types, reliable data support can be provided for land use planning, agricultural production management, and soil pollution control, thereby effectively promoting sustainable agricultural production and environmental protection.
[0003] In the agricultural field, identifying soil types helps optimize crop planting decisions and improve crop yields and quality; in the field of environmental protection, soil type identification can help monitor soil quality, assess the degree of pollution, and provide a scientific basis for soil remediation.
[0004] However, traditional soil classification methods face many challenges:
[0005] Traditional methods rely on manual experience and rule-based feature extraction, which are often unable to handle large-scale, complex soil data, and have low accuracy and efficiency. With the rapid development of data science and deep learning technology, soil type recognition methods based on images and other data modalities have gradually become a research hotspot. However, most existing deep learning methods rely on architectures such as convolutional neural networks (CNNs), which face high computational overhead and dependence on a large amount of labeled data. Summary of the invention
[0006] In view of the shortcomings of the existing technology, the present invention proposes a regional-scale soil classification method and system based on multimodal data, aiming to improve the accuracy and efficiency of soil classification through multimodal learning technology.
[0007] A first aspect of the present invention provides a regional scale soil classification method based on multimodal data, comprising the following steps:
[0008] Dataset construction: Collect soil image data and survey data within the region to build the initial dataset;
[0009] Data preprocessing: Clean and structure the collected image data and survey data to ensure data quality and consistency;
[0010] Feature extraction: Extract features from image data and structured survey data to obtain image features and structured features;
[0011] Multimodal data fusion: Fusion of image features and structural features to generate shared representations;
[0012] Model reconstruction: reconstruct the fused data through the autoencoder to optimize the shared representation;
[0013] Soil classification model construction: Based on the optimized shared representation, a soil classification model is constructed;
[0014] Model application: The constructed soil classification model is applied to actual scenarios to classify regional-scale soils in different plots.
[0015] A second aspect of the present invention provides a regional scale soil classification system based on multimodal data, comprising:
[0016] Data acquisition module, used to collect soil image data and survey data within the area;
[0017] Data preprocessing module, used to clean and structure the collected image data and survey data;
[0018] A feature extraction module is used to extract features from image data and structured survey data to obtain image features and structured features;
[0019] Data fusion module, used to fuse image features and structural features to generate shared representation;
[0020] The model reconstruction module is used to reconstruct the fused data through the autoencoder and optimize the shared representation;
[0021] A soil classification model building module, used to build a soil classification model based on the optimized shared representation;
[0022] The model application module is used to apply the constructed soil classification model to actual scenarios and classify regional-scale soils in different plots.
[0023] A third aspect of the present invention provides a regional scale soil classification device based on multimodal data, comprising:
[0024] Memory, used to store computer programs and data;
[0025] A processor is used to execute the computer program to implement the regional scale soil classification method based on multimodal data.
[0026] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program for executing the regional-scale soil classification method based on multimodal data.
[0027] Beneficial effects of the present invention:
[0028] Different from the traditional method, the present invention does not rely on traditional neural networks such as CNN, but adopts an innovative technology combining the sliding window mechanism and the Transformer architecture. Specifically:
[0029] Image data and structured survey data are respectively subjected to feature extraction by a sliding window transformer (SWIN Transformer) and a multi-layer perceptron (MLP). SWIN Transformer uses a sliding window mechanism and a window size that increases layer by layer to establish an effective connection between the local and the global, and is particularly suitable for processing high-resolution soil images. At the same time, the survey data is processed by MLP to extract meaningful features from structured data. In order to further improve the accuracy and robustness of soil classification, the present invention introduces a cross-co-attention mechanism to capture and strengthen the relationship between different modalities. The cross-co-attention mechanism can help the model more effectively align information from different modalities by weighted fusion of image features and survey data features, thereby improving classification performance.
[0030] In addition, the present invention also combines multimodal model reconstruction technology to reconstruct images and structured data through autoencoders. The reconstruction process can not only maintain the individual characteristics of each data modality, but also ensure the distinguishability of shared representation in the latent space. By minimizing the reconstruction loss, the shared representation of the model is optimized, thereby enhancing the recognition ability of soil types.
[0031] In summary, the present invention successfully improved the accuracy and efficiency of the soil classification model through an innovative multimodal learning framework, combined with SWIN Transformer, MLP, cross co-attention mechanism and autoencoder reconstruction, and provided more accurate and reliable decision support for land resource management, agricultural production and environmental protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Flowchart of regional-scale soil classification method based on multimodal data;
[0033] Figure 2 Schematic diagram of survey data structuring and feature extraction;
[0034] Figure 3 It is a schematic diagram of image data preprocessing and feature extraction;
[0035] Figure 4 A schematic diagram for representing multimodal data sharing;
[0036] Figure 5 Optimize schematics for model reconstruction and shared representation. DETAILED DESCRIPTION
[0037] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0038] like Figure 1 As shown, the embodiment of the present application includes the following steps:
[0039] S1: Construction of initial data set: By inspecting soil survey points within the region, high-resolution field sampling point image data and land quality survey record data are extracted to construct the initial data set. In the data preprocessing stage, the image data and survey data are cleaned and structured to ensure data quality and consistency. The image data uses a sliding window mechanism to extract local features, and the survey data is processed through standardization and other technologies.
[0040] S2: Feature extraction module construction: The feature extraction module consists of a structured survey data module and an image encoding module; construct a feature extraction module, extract structured features from survey data through a multi-layer perceptron (MLP), and extract local and global features from image data through a sliding window transformer (SWIN Transformer), ensuring that the feature representation of image and survey data is discriminative and has rich semantic information. Significantly improve the processing efficiency of regional soil multimodal data and ensure that the model can effectively process complex sampling point information.
[0041] S3: Shared representation of multimodal data: Based on the extracted image and survey data features, a soil classification prediction model is constructed. The image and survey data features are weighted and fused through the cross-co-attention mechanism to capture the correlation between the two. Then, the multimodal deep autoencoder is used to denoise and reconstruct the image and structured data. Combined with the multimodal deep generative model, high-quality feature representation is generated to ensure the consistency and integrity of different modal data in the shared space.
[0042] S4: Multimodal model reconstruction: Through multimodal model reconstruction, the image and structured data are reconstructed in combination with the autoencoder to narrow the distance between the matching image and text data pairs, and the decoding structure is used to enable the representation in the corresponding representation space to reconstruct the image and structured data, maintain the individuality of the data and ensure that its representation has strong distinguishability. The model is shared based on the reconstruction loss. This process enhances the robustness of the features by optimizing the shared representation and improves the classification accuracy of the model for soil types.
[0043] S5: Model application and verification: Apply the trained soil classification model to actual scenarios, classify regional-scale soils in different plots through soil sample collection and classification prediction, verify the adaptability and accuracy of the model in different soil environments, and provide effective support for soil management, agricultural production, and environmental monitoring.
[0044] In one embodiment, the initial data set construction comprises the following steps:
[0045] S11 Soil data collection: Conduct field surveys of soil quality points, refer to national sampling record standards, and use drones, industrial cameras and other photographic equipment to obtain high-resolution image data of sampling points. The field image data of each sampling point in the shooting area should cover the surface characteristics of the soil, vegetation coverage, topography and other visual information; at the same time, fill in the soil geochemical survey sampling record card, which should include: color, origin, landform, slope, texture, land use, crop type and other information, record the photo number and longitude and latitude of the survey point for subsequent coding of survey data, and form a preliminary data set.
[0046] S12 Image data preprocessing: Crop the image data to ensure that the image size is consistent and normalize it to facilitate input into the subsequent model. Apply denoising algorithms (median filtering) and enhancement techniques (rotation, flipping, etc.) to the image data to improve image quality and increase data diversity, helping the model to better process image information in different environments.
[0047] S13 Survey data structuring: Clean the survey data, including removing duplicate data, correcting erroneous values, filling missing values, etc. For all discrete survey data, use One-Hot Encoding technology to convert them into numerical data. One-Hot Encoding Structured Survey Data: In it, each category is converted into a binary vector, each category occupies a position in the vector, the value is 1 to represent the category, and the other positions are 0. By converting each category variable into a binary vector, it is convenient for subsequent model processing. It can be spliced with image features to form a complete data set containing all features for subsequent MLP models or Transformer architectures. It not only ensures the structuring of the data, but also ensures alignment and fusion with image data.
[0048] S14 Data division: When constructing the initial data set, pair the collected image data with the corresponding soil survey data. Ensure that each image data can correspond one-to-one with the corresponding soil survey data, and provide high-quality labeled data for subsequent model training. Divide the constructed data set into training set, validation set, and test set. The training set is used to train the model, the validation set is used for model tuning and selecting the optimal model, and the test set is used to finally evaluate the performance of the model.
[0049] In one embodiment, the feature extraction module construction process includes:
[0050] The constructed feature extraction modules include survey data feature extraction and image data feature extraction modules. In the process of structured survey data feature extraction, MLP (multi-layer perceptron) is usually good at processing structured numerical data, and the dense matrix or vector generated by one-hot encoding just meets the requirements of MLP for input data. In the process of image data feature extraction, SWINTransformer uses a sliding window mechanism and a layer-by-layer increasing window size to establish a connection between the local and the global and extract multi-scale features. It performs well in high-resolution image processing, can efficiently extract local and global features of images, has low computational overhead, and is particularly suitable for data processing of high-resolution images; it includes the following steps:
[0051] S21 Survey data feature extraction (see Figure 2 ): The constructed survey data feature extraction uses a multi-layer perceptron as an encoder to effectively encode the input categorical variables. After the S13 survey data structuring, the soil survey data is converted into a dense matrix suitable for further processing through MLP.
[0052] Assume that the input survey data set is D={d1,d2,…,d n}, where n is the number of data samples, d i is the feature vector of each sample. Before entering the encoding module, each survey data sample is first standardized to ensure the uniformity and stability of the input data. The standardization operation can be expressed by the following formula (1):
[0053] (1)
[0054] Where: D input is the standardized data set, μ(D) and σ(D) represent the mean and standard deviation of the data set D, respectively. The purpose of standardization is to reduce the deviation of the data and make the data distribution more uniform, thereby improving the convergence speed and stability of the subsequent model.
[0055] Next, the data is input into the multi-layer perceptron (MLP) encoding module for processing. MLP consists of multiple fully connected layers, each layer contains several neurons and has a nonlinear activation function. Assume that the input data is D input , the multi-layer perceptron encoding process can be achieved through the following steps:
[0056] Each layer uses a fully connected (dense) operation to map the input features to a higher dimensional space, and the output is H1. Use SELU (activation function) for nonlinear transformation after each fully connected layer. SELU activation function can automatically adjust the mean and variance of the output so that the output of each layer remains within a stable range, which helps to speed up the training process. In order to prevent overfitting, a regularization layer is added. After multiple layers of nonlinear transformation, the final output D feature This is the feature representation of structured data. Through this encoding process, the structured survey data is converted into a feature representation with high-dimensional information, which can effectively provide rich input features for subsequent soil classification prediction. The formula is as follows:
[0057] (2)
[0058] (3)
[0059] (4)
[0060] (5)
[0061] Among them, W1 and b1 are the weight matrix and bias of the first fully connected layer respectively; λ and α are the hyperparameters of the SELU function; p is the dropout rate, and the regularization layer can effectively prevent overfitting. output It is the output after the Dropout layer is processed. Dropout is a regularization technique used to prevent overfitting of neural networks.
[0062] S22 Image data feature extraction (see Figure 3 ): The goal of image data feature extraction is to efficiently extract features from the collected soil images to ensure that local and global information in the image is captured. The image data is processed by SWIN Transformer, which combines a sliding window mechanism with a layer-by-layer increase in window size to establish a connection between the local and global and extract multi-scale features.
[0063] Assume the input image is I , first split it into multiple small blocks and input it into SWIN Transformer for encoding; each block will be sent to SWIN Transformer for processing. Through the sliding window mechanism, SWIN Transformer can effectively process the local area features of the image, and at the same time gradually capture the global features by increasing the window size layer by layer. In order to map the high-dimensional features of the image to the low-dimensional space, the linear projection method is used to process each image block. Assume that the projection result of the jth block is I' j , which can be expressed as the following formula (6):
[0064] (6)
[0065] Among them, W j and b j are weight matrices and bias terms respectively, and the projection result I' j It will be used as input for subsequent self-attention calculations.
[0066] SWIN Transformer uses the self-attention mechanism to weight the features of image blocks. By calculating the similarity relationship between blocks, the model can focus on more important areas in the image. The calculation formula (7) of self-attention is:
[0067] (7)
[0068] Among them, Q, K, V represent query, key and value respectively, dk is the dimension of the key, and the self-attention mechanism further enhances the ability to capture important features in the image by calculating the relationship between the query and the key.
[0069] After self-attention, the features are further processed by a feedforward neural network to enhance the expressiveness of the features. The feedforward layer transforms the input features nonlinearly to produce a richer feature representation:
[0070] (8)
[0071] Through this process, SWIN Transformer can extract multi-scale features of the image, covering image information from local to global; the final image features I feature It is a collection of multiple feature vectors that contain rich information about key areas in the image and provide a powerful input for the subsequent soil type classification task.
[0072] In one embodiment, the process of constructing a multimodal data sharing representation (see Figure 4 )include:
[0073] This step constructs a shared low-dimensional latent space by introducing a cross-attention mechanism, a multimodal deep autoencoder, and a multimodal deep generative model, aiming to effectively fuse image data and structured survey data and improve the ability to identify soil types. First, the cross-attention mechanism is used to weightedly fuse the features of image and survey data to capture the relationship between the two modalities. Then, the image and structured data are mapped to the shared latent space through a multimodal deep autoencoder, and discriminative feature representations are further learned. Finally, combined with a multimodal deep generative model, high-quality image and structured data features are generated to ensure the consistency and information completeness of different modal data in the shared space. This process improves the accuracy and robustness of the model in the soil classification task by optimizing the shared representation.
[0074] The specific steps include:
[0075] S31 Cross-co-attention fusion module: The features of image data and structured survey data are fused through the cross-co-attention mechanism. The cross-co-attention mechanism captures and strengthens the interdependence between different modalities by calculating the relationship between image features and structured data features, thereby improving the alignment and fusion effect of information, and generating queries, keys, and values for image features I and survey data features D respectively. Image data I generates query Q image , survey data D to generate key K survey Sum value V survey .
[0076] S32 Multimodal Deep Autoencoder: Use a multimodal deep autoencoder to further process the fused features. feature and survey data characteristics D feature Encode and map to a shared latent space. This step helps to extract latent features in image data and structured data and generate shared representations. The shared representation is mapped back to the original data space through the decoder for reconstruction optimization to ensure the validity of the shared representation. This process can be optimized by minimizing the reconstruction error:
[0077] (9)
[0078] Among them, f enc and f dec are the encoder and decoder functions respectively, x is the input data, and minimizing the reconstruction error helps to generate a shared representation with high-dimensional information.
[0079] S33 Multimodal Deep Generative Model: Finally, through the weighted fusion of formulas (7) and (8), the features of image data and survey data are effectively combined in a shared space, enhancing the relationship between the two modalities.fused Through further transformation by multiple deep Boltzmann machines, high-quality images and structured data are generated from the low-dimensional latent space and a new shared representation is formed. shared .
[0080] In one embodiment, the multimodal model reconstruction process (see Figure 5 )include:
[0081] Through multimodal model reconstruction, the image and structured data are reconstructed in combination with the constructed autoencoder structure, and the shared representation is optimized by minimizing the reconstruction loss. This process ensures their consistency in the shared representation space by shortening the distance between the matching pairs of image and structured data. Using the decoder structure, the model can effectively reconstruct the original features of the image and structured data while maintaining the personalized information of the data and ensuring that the reconstructed representation has strong distinguishability. By optimizing the reconstruction loss function and calculating the reconstruction loss, the model can enhance the robustness of the features and optimize the shared representation in the process of learning the shared representation, and further improve the accuracy and reliability of soil type classification. Specifically, it includes the following steps:
[0082] S41 Investigation and Image Data Reconstruction: The features of image data and structured data are input into the encoder part of the multimodal autoencoder respectively and mapped to a shared low-dimensional latent space H shared , this representation contains the common information of images and structured data.
[0083] (10)
[0084] where f enc is the encoder function.
[0085] Through the decoder part, the shared representation H shared Converted into reconstructed image data I reconstructed ; Shared representation H shared Converted into reconstructed image data D reconstructed .
[0086] (11)
[0087] (12)
[0088] g dec and h dec These are the decoder functions for images and structured data, respectively.
[0089] S42 calculates the reconstruction loss: The reconstruction loss is the error between the original input data and the data reconstructed by the model. For multimodal models, the reconstruction loss involves two modal data: image data and structured data. The shared representation generated by the model can reconstruct image data and structured data at the same time and make them as close to the original data as possible. The difference between the input data and the reconstructed data is calculated to obtain the reconstruction loss of the image and structured data. Use the mean square error (MSE) as the reconstruction loss function:
[0090] (13)
[0091] Among them, I input and D input are the original images and structured data, I reconstructed and D reconstructed It is the reconstructed image and structured data.
[0092] S43 Optimizing shared representations: by minimizing reconstruction loss reconstruction , the model can optimize the shared representation H shared , making the representation of images and structured data in the shared space more compact and consistent. This process uses the back-propagation algorithm and gradient descent optimization for training. In the process of minimizing the reconstruction loss, the model continuously adjusts the parameters of the encoder and decoder to enhance the similarity of image features and structured data features in the shared representation space, ensuring that the distribution of the two in the latent space is more consistent.
[0093] The optimized shared representation can better represent the joint features of image and structured data, thereby improving the classification model's recognition accuracy for soil types. The optimized shared representation can also enhance the model's generalization ability under different data distributions and improve the accuracy and robustness in soil classification tasks.
[0094] Embodiment: This embodiment includes an initial data set based on previous soil survey data and image data, feature extraction and shared representation of structured survey data and image data, and then model reconstruction, and finally obtaining a soil classification prediction model with optimized shared representation, and applying this optimized model to soil multimodal data analysis to achieve accurate distinction of regional soil categories;
[0095] The specific steps include:
[0096] S51 Soil data collection: Through land quality surveys within the region, obtain rich image information and survey data, fill in soil geochemical survey sampling record cards, and provide original data support for subsequent survey data structuring and feature extraction;
[0097] S52 Survey data structuring: Soil geochemistry survey sampling record cards were sorted, and a python crawler program was written to extract discrete variable data from each record card to meet the requirements of the one-hot encoding input paradigm, and the survey data was sorted and structured with one-hot encoding to facilitate subsequent feature extraction;
[0098] S53 Multimodal data shared representation: Based on the extracted image and survey data features, a soil classification prediction model is constructed. The image and survey data features are weighted and fused through the cross-co-attention mechanism to capture the correlation between the two. Multimodal deep autoencoders are used to denoise and reconstruct image and structured data. Combined with multimodal deep generative models, high-quality feature representations are generated to ensure the consistency and integrity of different modal data in the shared space, maintain the individuality of the data and have strong distinguishability;
[0099] S54 Regional soil classification prediction results: The model that has been reconstructed and shared and optimized is applied to the existing regional soil data to verify the adaptability and accuracy of the model in different regional soil environments, providing effective support for soil management, agricultural production, and environmental monitoring.
[0100] Based on the same concept as the above method, the embodiment of the present application also provides a regional scale soil classification system based on multimodal data, including:
[0101] Data acquisition module, used to collect soil image data and survey data within the area;
[0102] Data preprocessing module, used to clean and structure the collected image data and survey data;
[0103] A feature extraction module is used to extract features from image data and structured survey data to obtain image features and structured features;
[0104] Data fusion module, used to fuse image features and structural features to generate shared representation;
[0105] The model reconstruction module is used to reconstruct the fused data through the autoencoder and optimize the shared representation;
[0106] A soil classification model building module, used to build a soil classification model based on the optimized shared representation;
[0107] The model application module is used to apply the constructed soil classification model to actual scenarios and classify regional-scale soils in different plots.
[0108] Based on the same concept as the above method, the embodiment of the present application also provides a regional scale soil classification device based on multimodal data, including:
[0109] Memory, used to store computer programs and data;
[0110] A processor is used to execute the computer program to implement the regional scale soil classification method based on multimodal data.
[0111] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0112] The memory may be used to store the computer programs and / or modules, and the processor implements various functions of the device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.
[0113] Based on the same concept as the above method, an embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the regional-scale soil classification method based on multimodal data.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.
Claims
1. A regional scale soil classification method based on multimodal data, characterized in that: The following steps are involved: Dataset construction: Collect soil image data and survey data within the region to build the initial dataset; Data preprocessing: Clean and structure the collected image data and survey data to ensure data quality and consistency; Feature extraction: Extract features from image data and structured survey data to obtain image features and structured features; Multimodal data fusion: Fusion of image features and structural features to generate shared representations; Model reconstruction: reconstruct the fused data through the autoencoder to optimize the shared representation; Soil classification model construction: Based on the optimized shared representation, a soil classification model is constructed; Model application: The constructed soil classification model is applied to actual scenarios to classify regional-scale soils in different plots.
2. The regional scale soil classification method based on multimodal data according to claim 1, characterized in that: The data preprocessing step includes: Crop and normalize image data, and apply denoising algorithms and enhancement techniques to improve image quality and increase data diversity; The survey data was cleaned, including removing duplicate data, correcting erroneous values, and filling missing values; all discrete survey data were converted into numerical data using one-hot encoding technology.
3. The regional scale soil classification method based on multimodal data according to claim 1 or 2, characterized in that: In the feature extraction step, local and global features are extracted from image data through a sliding window transformer, and structured features are extracted from structured survey data through a multi-layer perceptron.
4. The regional scale soil classification method based on multimodal data according to claim 3 is characterized in that: In the multimodal data fusion step, the image features and structural features are weighted and fused through a cross-co-attention mechanism to capture the correlation between the two.
5. The regional scale soil classification method based on multimodal data according to claim 1, characterized in that: In the model reconstruction step, the image and structured data are denoised and reconstructed through a multimodal deep autoencoder to optimize the shared representation.
6. The regional scale soil classification method based on multimodal data according to claim 1 or 5, characterized in that: In the soil classification model construction step, a deep Boltzmann machine is combined to extract high-order features from each modality, and the generated features are used to perform multimodal feature fusion to ensure the consistency and integrity of different modal data in a shared space.
7. A regional scale soil classification system based on multimodal data, characterized in that: include: Data acquisition module, used to collect soil image data and survey data within the area; Data preprocessing module, used to clean and structure the collected image data and survey data; A feature extraction module is used to extract features from image data and structured survey data to obtain image features and structured features; Data fusion module, used to fuse image features and structural features to generate shared representation; The model reconstruction module is used to reconstruct the fused data through the autoencoder and optimize the shared representation; A soil classification model building module, used to build a soil classification model based on the optimized shared representation; The model application module is used to apply the constructed soil classification model to actual scenarios and classify regional-scale soils in different plots.
8. A regional scale soil classification device based on multimodal data, characterized in that: include: Memory, used to store computer programs and data; A processor, configured to execute the computer program to implement the regional scale soil classification method based on multimodal data as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the regional scale soil classification method based on multimodal data as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-source data fusion ground feature classification method based on multi-scale convolution auto-encoder
CN117576483A
Soil microorganism classification method based on deep learning
CN119251836A