Target function type high-precision remote sensing recognition method fusing multi-modal features

By integrating features from multimodal remote sensing data sources, the area, radiation, and structural features of targets are extracted, solving the problem of insufficient accuracy and robustness of single-source remote sensing technology in identifying geographic target functional types, and achieving high-precision target functional type identification.

CN121811274APending Publication Date: 2026-04-07SHANDONG JIANZHU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

When identifying the functional types of geographic targets, existing remote sensing technologies suffer from insufficient classification accuracy, robustness, and reliability due to the feature mining mode of a single data source. This is especially true when target features are weak, targets have similar features, or the environment is complex, leading to frequent misclassification and missed classification.

Method used

A multimodal feature fusion method is adopted, which combines high-resolution optical imagery, nighttime light imagery, and synthetic aperture radar imagery. By extracting the area size, light radiation intensity, and structural reflection intensity features of the target, a multimodal feature vector is constructed, and a random forest model is used for classification.

Benefits of technology

It improves the accuracy and robustness of target function type identification, enabling accurate identification of target function types in complex environments and significantly reducing the false positive rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811274A_ABST
    Figure CN121811274A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal feature fused target function type high-precision remote sensing recognition method, and belongs to the technical field of remote sensing image processing and target recognition, and the method comprises the following steps: S1, obtaining a multi-modal remote sensing data set covering an area to be recognized, and carrying out the preprocessing; s2, performing fine identification and area size feature extraction on a target in the to-be-identified region; s3, calculating light radiation intensity characteristics of each target; s4, extracting structural reflection intensity characteristics of each target; s5, fusing the three extracted features of each target to obtain a multi-modal feature vector of the target, and obtaining multi-modal feature vectors of all targets; and S6, training a machine learning classification model, inputting the multi-modal feature vectors of all targets into the machine learning classification model, and carrying out function type classification and precision evaluation.By adopting the method, the advantages of different data sources are fully utilized, the limitation of a single data source is effectively overcome, and the precision and robustness are higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing and target recognition technology, and in particular to a high-precision remote sensing recognition method for target functional types that integrates multimodal features. Background Technology

[0002] Geographic space contains numerous man-made targets serving various socio-economic functions, such as industrial production facilities, energy infrastructure (oil and gas platforms, power plants), transportation hubs, and resource extraction sites. Accurately and efficiently identifying and clarifying the functional types of these geographic targets is of irreplaceable importance for assessing regional development levels and monitoring energy consumption and environmental pollution.

[0003] Remote sensing technology, as a core means of Earth observation, has become a major data source for large-scale, dynamic target identification and monitoring research due to its unique advantages such as wide coverage, periodic repeatability, and minimal geographical limitations in data acquisition. With the rapid development of remote sensing technology, the available data sources have become increasingly abundant, forming a multi-modal observation system including high-resolution optical remote sensing, nighttime light remote sensing, and synthetic aperture radar (SAR) remote sensing, providing multi-faceted information support for precise target identification.

[0004] However, existing technologies mostly rely on feature mining from a single data source. While each has its own emphasis, they all suffer from inherent limitations that are difficult to overcome: optical images are "affected by cloudy skies," nighttime lights are "too blurry," and radar images are "difficult to interpret." This single-data-source feature mining model results in the classifier relying on a feature set with limited dimensions, leading to an insufficiently comprehensive and three-dimensional description of the target. When faced with challenging scenarios such as targets with weak features (e.g., small facilities), different types of targets with similar features (e.g., factories with different functions), or complex environmental backgrounds (e.g., interference from ships in near-shore areas), classification based on a single feature often fails to guarantee the accuracy, robustness, and reliability of the results, leading to frequent misclassifications and missed classifications. Summary of the Invention

[0005] The purpose of this invention is to provide a high-precision remote sensing identification method for target functional types that integrates multimodal features. It uses a multi-source information complementarity approach, which fully utilizes the advantages of different data sources by integrating features from optical, night light, and radar modes. Morphological, radiation, and structural information corroborate each other, effectively overcoming the limitations of a single data source, and achieving higher accuracy and stronger robustness.

[0006] To achieve the above objectives, this invention provides a high-precision remote sensing identification method for target functional types that integrates multimodal features, comprising the following steps: S1. Obtain the multimodal remote sensing dataset of the area to be identified, and preprocess the data in the dataset; S2. Based on relevant data in the multimodal remote sensing dataset, perform fine identification and size feature extraction of targets within the area to be identified; S3. Calculate the light radiation intensity characteristics of each target based on the relevant data in the multimodal remote sensing dataset and the target size characteristics obtained in S2; S4. Extract the structural reflection intensity features of each target based on the relevant data in the multimodal remote sensing dataset and the refined identification results obtained in S2; S5. The extracted size features, light radiation intensity features, and structural reflection intensity features of each target are fused to obtain the multimodal feature vector of the target, and the multimodal feature vectors of all targets are obtained. S6. Train the machine learning classification model by inputting the multimodal feature vectors of all targets into the machine learning classification model to perform functional type classification and accuracy evaluation.

[0007] Preferably, the process of S1 is as follows: S11. Determine the area to be identified; S12. Collect multimodal remote sensing data within the area, including high-resolution optical images, nighttime light images, and synthetic aperture radar images; S13. Preprocess the multimodal remote sensing data, and the preprocessed multimodal remote sensing data form a multimodal remote sensing dataset.

[0008] Preferably, the process of S2 is as follows: S21. Extract high-resolution optical images from the multimodal remote sensing dataset; S22. Use an object-oriented segmentation method to perform multi-scale segmentation on high-resolution optical images; S23. Perform detailed identification of targets within the image and obtain the precise contour vector boundary of each target; S24. Calculate the area of ​​each target based on the vector boundary, that is, the area size characteristics of each target.

[0009] Preferably, the process of S3 is as follows: S31. Extract nighttime light images from the multimodal remote sensing dataset. Based on the multi-scale segmentation results in S2, obtain the total radiance value of the target nighttime light area from the nighttime light images, denoted as... ; S32. Based on the area size characteristics of each target in S2, traverse all target vector polygons within the night-light target area and calculate the vector polygon of each target. The actual area is denoted as . ; S33. Add up the actual areas of all target vector polygons within the luminous target area to obtain the total area of ​​the target within the target area. ; S34. Calculate the proportion of the actual area of ​​each target in the total area of ​​the target area. That is, the area weight of each target, the process is as follows: ; S35. The total radiance value of the target's night-light target area is allocated according to the area weight of each target based on the night-light hybrid pixel decomposition model, to obtain the radiance intensity feature value allocated to each target. The process is as follows: ; in This represents the radiation intensity characteristic value assigned to each target, i.e., the light radiation intensity characteristic.

[0010] Preferably, the process of S4 is as follows: S41. Set the expansion distance, and build a buffer by expanding outward based on the precise contour vector boundary of each target obtained in S2; S42. Extract the backscattering coefficients of all pixels in the synthetic aperture radar image from the buffer. S43. Take the maximum value of the backscattering coefficients of all pixels of each target as the structural reflection intensity characteristic of the target.

[0011] Preferably, the process of obtaining the multimodal feature vectors of all targets in S5 is as follows: the area size features, light radiation intensity features and structural reflection intensity features of each target extracted in S2, S3 and S4 are combined to form the three-dimensional feature vector of each target.

[0012] Preferably, the machine learning classification model used in S6 is a random forest model. The pre-training process of the random forest model is to train the random forest model using target samples with known functional type labels and their corresponding multimodal feature vectors to obtain a trained random forest model.

[0013] Preferably, the process of S6 is as follows: S61. Input the multimodal feature vectors of all targets obtained in S5 into the trained random forest model to obtain the functional type of each target; S62. Calculate the precision, recall, and F1 score for all target classification results to evaluate the performance of the random forest model and complete the process of identifying target function types.

[0014] Therefore, the present invention adopts a high-precision remote sensing identification method for target functional types that integrates multimodal features using the above-mentioned structure. It uses a multi-source information complementarity approach, and fully utilizes the advantages of different data sources by integrating the features of optical, night light, and radar modes. The morphological, radiation, and structural information corroborate each other, effectively overcoming the limitations of a single data source, and has higher accuracy and stronger robustness.

[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0016] Figure 1 This is an overall flowchart of a high-precision remote sensing identification method for target function types that integrates multimodal features according to the present invention. Figure 2 In the example, (a) represents the luminescent target region extracted from the target area size features; Figure 2 In the example, (b) represents the actual boundary for the extraction of the target area size features; Figure 2 In the example, (c) represents the target area size feature; Figure 3 In the example, (a) represents the night-light target area for the distribution of nighttime light radiation intensity; Figure 3 In the example, (b) represents the light radiation intensity characteristics in the embodiment; Figure 4 In the example, (a) represents the luminescent target region with structural reflection intensity characteristics in the embodiment; Figure 4 (b) in the example represents the construction graph of the buffer; Figure 4 In the example, (c) represents the reflection intensity characteristics of the target structure; Figure 5 In the example, (a) represents the area feature; Figure 5 In the example, (b) represents the light radiation intensity characteristics in the embodiment; Figure 5 In the example, (c) represents the structural reflection intensity characteristics; Figure 5 In the example, (d) represents the feature synthesis result; Figure 6 (a) in the figure represents the target classification result of a certain place in a certain year using this method; Figure 6 (b) in the figure represents the verification result of the classification accuracy of a certain place in a certain year using this method. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0018] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0019] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0020] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed when in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0021] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," and "connect" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0022] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0023] Example 1 like Figure 1 and Figure 6 As shown, the present invention provides a high-precision remote sensing identification method for target functional types that integrates multimodal features, comprising the following steps: S1. Obtain the multimodal remote sensing dataset of the area to be identified, and preprocess the data in the dataset; S11. Determine the area to be identified; S12. The acquisition device acquires multimodal remote sensing data within the area, including high-resolution optical images, nighttime light images, and synthetic aperture radar images. S13. Preprocess the multimodal remote sensing data, and the preprocessed multimodal remote sensing data form a multimodal remote sensing dataset.

[0024] S2. Based on relevant data in the multimodal remote sensing dataset, perform fine identification and area size feature extraction of targets within the area to be identified; S21. Extract high-resolution optical images from the multimodal remote sensing dataset; S22. Use an object-oriented segmentation method to perform multi-scale segmentation on high-resolution optical images; S23. Perform detailed identification of targets within the image and obtain the precise contour vector boundary of each target; S24. Calculate the area of ​​each target based on the vector boundary, that is, the area size characteristics of each target.

[0025] S3. Calculate the light radiation intensity characteristics of each target based on the relevant data in the multimodal remote sensing dataset and the target size characteristics obtained in S2; S31. Extract nighttime light images from the multimodal remote sensing dataset. Based on the multi-scale segmentation results in S2, obtain the total radiance value of the target nighttime light area from the nighttime light images, denoted as... The target luminescent area is a single pixel or a region composed of multiple consecutive pixels, which appears as a connected bright spot in the image. Based on high spatial resolution optical images (such as Sentinel-2), the vector boundaries of all independent target contours within the luminescent target area are accurately identified and drawn using object-oriented image segmentation methods. S32. Based on the area size characteristics of each target in S2, traverse all target vector polygons within the night-light target area and calculate the vector polygon of each target. The actual area is denoted as . ; S33. Add up the actual areas of all target vector polygons within the luminous target area to obtain the total area of ​​the target within the target area. ; S34. Calculate the proportion of the actual area of ​​each target in the total area of ​​the target area. That is, the area weight of each target, the process is as follows: ; S35. The total radiance value of the target's night-light target area is allocated according to the area weight of each target based on the night-light hybrid pixel decomposition model, to obtain the radiance intensity feature value allocated to each target. The process is as follows: ; in This represents the radiation intensity characteristic value assigned to each target, i.e., the light radiation intensity characteristic.

[0026] The luminescent mixed pixel decomposition model reasonably assumes that the radiation contribution of a target is roughly proportional to its physical size.

[0027] S4. Extract the structural reflection intensity features of each target based on the relevant data in the multimodal remote sensing dataset and the refined identification results obtained in S2; S41. Set the expansion distance, and build a buffer by expanding outward based on the precise contour vector boundary of each target obtained in S2; S42. Extract the backscattering coefficients of all pixels in the synthetic aperture radar image from the buffer. S43. Take the maximum value of the backscattering coefficients of all pixels of each target as the structural reflection intensity characteristic of the target.

[0028] S5. The area size features, light radiation intensity features, and structural reflection intensity features of each extracted target are fused to obtain the multimodal feature vector of the target, and the multimodal feature vectors of all targets are obtained. The area size features, light radiation intensity features, and structural reflection intensity features of each target extracted from S2, S3, and S4 are combined to form a three-dimensional feature vector for each target.

[0029] Thus, the fusion process of mapping multimodal features with different spatial resolutions and physical meanings onto the same target has been completed, constructing a feature set that can comprehensively describe the target's attributes from three dimensions: morphology, nighttime activity intensity, and macroscopic structure. This provides an information foundation for subsequent machine learning models to perform high-precision functional classification.

[0030] S6. Train the machine learning classification model by inputting the multimodal feature vectors of all targets into the machine learning classification model to perform functional type classification and accuracy evaluation.

[0031] The machine learning classification model used is the random forest model. The pre-training process of the random forest model is to train the random forest model using target samples with known functional type labels and their corresponding multimodal feature vectors to obtain a trained random forest model.

[0032] S61. Input the multimodal feature vectors of all targets obtained in S5 into the trained random forest model to obtain the functional type of each target; S62. Calculate the precision, recall, and F1 score for all target classification results to evaluate the performance of the random forest model and complete the process of identifying target function types.

[0033] Example 2 like Figures 2-5 The method was verified as shown below. Step 1: Data Acquisition and Preprocessing.

[0034] Three types of remote sensing data covering the same area and at similar times were acquired: Sentinel-2 high-resolution multispectral optical image (spatial resolution 10 meters), NPP / VIIRS nighttime light monthly composite image (spatial resolution approximately 500 meters), and Sentinel-1 C-band SAR radar image (interferometric wide-swath mode, VV polarization, GRD product).

[0035] All data was acquired through the Google Earth Engine platform and underwent necessary preprocessing, including radiometric calibration, geometric correction, image registration, image cropping, and projection transformation, to ensure spatial reference consistency. Optical imagery also underwent cloud masking.

[0036] Step 2: Target fine recognition and area size feature extraction.

[0037] like Figure 2 In Figure (a), the target area size feature is extracted for nighttime illumination, where the red border represents the target area boundary and the blue area represents the target object. Leveraging the high spatial resolution of Sentinel-2 optical imagery, an object-oriented segmentation method is employed to segment the study area at multiple scales. Through merging and filtering operations, the contour vector polygon of each target is accurately extracted. Then, the area (unit: square meters) of each vector polygon is calculated as the target's area size feature (denoted as Feature_Area).

[0038] Step 3, Radiation intensity feature extraction.

[0039] like Figure 3 In the diagram, (a) represents the nighttime light radiation intensity allocation target area, where the red border indicates the target area boundary and the blue circle indicates the target object. Due to the low spatial resolution of VIIRS nighttime light images, a single pixel (target area) often contains multiple targets. First, the total radiance value (DN_total) of each target area is determined. Then, using the target area obtained in step 2, an area-based allocation model is used to calculate the radiance characteristics of each target. For example, if a target area contains 3 targets with areas of... , , Total area The total radiation value of the target area is Then the characteristic value of the radiation intensity of target 1. Similarly, calculate... , This feature is denoted as .

[0040] Step 4: Extraction of structural reflection intensity features.

[0041] like Figure 4 In the example, (a) represents the luminescent target region of the structure's reflection intensity characteristics, where the area within the red boundary is the target luminescent target region, and the area within the blue dashed line is... Figure 4 The magnified region of (c) in the image; Figure 4 (b) in the diagram represents the construction of the buffer zone, where the yellow ellipse represents the target buffer zone range. Based on the Sentinel-1 SAR image, and using the target vector boundary obtained in step 2 as a basis, a buffer zone with a range of 2 kilometers is constructed outward in the GEE. The backscattering coefficients of all pixels within this buffer zone on the SAR image are extracted. The maximum value of these values ​​is taken as the structural reflection intensity feature of the target (denoted as Feature_SAR). This is intended to capture the strongest structural reflection signal of the target, reflecting the complexity of its facilities.

[0042] Step 5: Multimodal feature fusion.

[0043] For each target, its three feature values ​​(Feature_Area, Feature_NTL, Feature_SAR) are combined into a three-dimensional feature vector, as shown below: The feature vectors of all targets constitute the training or classification dataset.

[0044] Step 6: Functional type classification and accuracy evaluation.

[0045] Random forest is used as the classifier. First, the random forest model is trained using target samples with known functional type labels (e.g., obtained through manual interpretation of high-resolution images or field surveys) and their corresponding multimodal feature vectors. The trained model can then be used to classify and predict targets of unknown types. Finally, using reserved test samples, the precision, recall, and F1 score of the classification results are calculated to evaluate the model's performance.

[0046] Therefore, this method achieves high-precision automatic identification of target functional types. The application of multimodal features enables the identification results to take into account the target's size, optical radiation, and structural features, significantly improving the ability to distinguish targets of different functional types. When the data of one modality is damaged due to noise or occlusion, other modal information can still provide effective feature supplementation, making the identification more robust and reliable.

[0047] Using multimodal features provides a richer and more comprehensive description of the target, enabling the classification model to learn more fundamental class distinction rules. Even when some features are disturbed (such as cloud cover causing poor optical images), the features of other modalities can still support making correct judgments, significantly improving classification accuracy and stability, and has the advantages of high accuracy and strong robustness. The core idea of ​​this method is not limited to specific types of targets. By adjusting the specific parameters of feature extraction and the training samples of the classification model, it can be widely applied to various geospatial target recognition scenarios that require functional type differentiation based on remote sensing images. At the same time, the method has a clear process, is easy to automate, and is suitable for batch and efficient functional type recognition of targets in a large area.

[0048] Therefore, the present invention adopts a high-precision remote sensing identification method for target functional types that integrates multimodal features using the above-mentioned structure. It uses a multi-source information complementarity approach, and fully utilizes the advantages of different data sources by integrating the features of optical, night light, and radar modes. The morphological, radiation, and structural information corroborate each other, effectively overcoming the limitations of a single data source, and has higher accuracy and stronger robustness.

[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A high-precision remote sensing identification method for target functional types that integrates multimodal features, characterized in that, Includes the following steps: S1. Obtain the multimodal remote sensing dataset of the area to be identified, and preprocess the data in the dataset; S2. Based on relevant data in the multimodal remote sensing dataset, perform fine identification and area size feature extraction of targets within the area to be identified; S3. Calculate the light radiation intensity characteristics of each target based on the relevant data in the multimodal remote sensing dataset and the target size characteristics obtained in S2; S4. Extract the structural reflection intensity features of each target based on the relevant data in the multimodal remote sensing dataset and the refined identification results obtained in S2; S5. The extracted size features, light radiation intensity features, and structural reflection intensity features of each target are fused to obtain the multimodal feature vector of the target, and the multimodal feature vectors of all targets are obtained. S6. Train the machine learning classification model by inputting the multimodal feature vectors of all targets into the machine learning classification model to perform functional type classification and accuracy evaluation.

2. The high-precision remote sensing identification method for target function type fusion based on multimodal features according to claim 1, characterized in that, The process of S1 is as follows: S11. Determine the area to be identified; S12. Collect multimodal remote sensing data within the area, including high-resolution optical images, nighttime light images, and synthetic aperture radar images; S13. Preprocess the multimodal remote sensing data, and the preprocessed multimodal remote sensing data form a multimodal remote sensing dataset.

3. The high-precision remote sensing identification method for target function type fusion based on multimodal features according to claim 2, characterized in that, The process of S2 is as follows: S21. Extract high-resolution optical images from the multimodal remote sensing dataset; S22. Use an object-oriented segmentation method to perform multi-scale segmentation on high-resolution optical images; S23. Perform detailed identification of targets within the image and obtain the precise contour vector boundary of each target; S24. Calculate the area of ​​each target based on the vector boundary, that is, the area size characteristics of each target.

4. The high-precision remote sensing identification method for target function type fusion based on multimodal features according to claim 3, characterized in that, The process of S3 is as follows: S31. Extract nighttime light images from the multimodal remote sensing dataset. Based on the multi-scale segmentation results in S2, obtain the total radiance value of the target nighttime light area from the nighttime light images, denoted as... ; S32. Based on the area size characteristics of each target in S2, traverse all target vector polygons within the night-light target area and calculate the vector polygon of each target. The actual area is denoted as . ; S33. Add up the actual areas of all target vector polygons within the luminous target area to obtain the total area of ​​the target within the target area. ; S34. Calculate the proportion of the actual area of ​​each target in the total area of ​​the target area. That is, the area weight of each target, the process is as follows: ; S35. The total radiance value of the target's night-light target area is allocated according to the area weight of each target based on the night-light hybrid pixel decomposition model, to obtain the radiance intensity feature value allocated to each target. The process is as follows: ; in This represents the radiation intensity characteristic value assigned to each target, i.e., the light radiation intensity characteristic.

5. The high-precision remote sensing identification method for target function type fusion based on multimodal features according to claim 4, characterized in that, The process of S4 is as follows: S41. Set the expansion distance, and build a buffer by expanding outward based on the precise contour vector boundary of each target obtained in S2; S42. Extract the backscattering coefficients of all pixels in the synthetic aperture radar image from the buffer. S43. Take the maximum value of the backscattering coefficients of all pixels of each target as the structural reflection intensity characteristic of the target.

6. The high-precision remote sensing identification method for target function type fusion based on multimodal features according to claim 5, characterized in that, The process of obtaining the multimodal feature vectors of all targets in S5 is as follows: the area size features, light radiation intensity features, and structural reflection intensity features of each target extracted from S2, S3, and S4 are combined to form the three-dimensional feature vector of each target.

7. A high-precision remote sensing identification method for target function type fusion based on multimodal features according to claim 6, characterized in that, The machine learning classification model used in S6 is the random forest model. The pre-training process of the random forest model is to train the random forest model using target samples with known functional type labels and their corresponding multimodal feature vectors to obtain a trained random forest model.

8. A high-precision remote sensing identification method for target function type fusion based on multimodal features according to claim 7, characterized in that, The process of S6 is as follows: S61. Input the multimodal feature vectors of all targets obtained in S5 into the trained random forest model to obtain the functional type of each target; S62. Calculate the precision, recall, and F1 score for all target classification results to evaluate the performance of the random forest model and complete the process of identifying target function types.