Farmland boundary extraction method based on multimodal reasoning model
Through the multimodal inference model, the reconstructed network and gated network module are used to adaptively suppress interference, which solves the problem of insufficient precision of cultivating land boundary extraction in the traditional method, and achieves higher accuracy and robust farmland boundary recognition.
Patent Information
- Application Number
- CN202510660232.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-22
AI Technical Summary
Traditional arable land boundary extraction methods rely on a single data source to distinguish arable land from similar land objects, and have weak anti-interference ability. Multimodal data has inconsistent feature distribution due to differences in resolution and imaging principles. Direct fusion is prone to conflicts, and some modal data are severely affected by environmental or equipment factors.
The multimodal inference model is adopted, and the reconstruction network, gated network module and contrast learning network module are constructed, combined with the autoencoder network and multi-scale feature extraction, and the interference information is adaptively suppressed, and the accuracy and robustness of the farmland boundary extraction are improved.
It improves the accuracy and robustness of farmland boundary extraction, effectively integrates multimodal information, and reduces the interference influence of environmental and equipment factors.
Smart Images

Figure CN120182629B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for extracting farmland boundaries based on a multimodal reasoning model. Background Art
[0002] Traditional methods for extracting cultivated land boundaries primarily rely on visible light remote sensing imagery or single-spectral band data. This single data source makes it difficult to distinguish cultivated land from similar features such as bare land and vegetation-covered areas, leading to boundary misjudgments. The highly similar spectral reflectance characteristics of cultivated land, bare land, and vegetation-covered areas in visible light or a single near-infrared band make it difficult for traditional algorithms to distinguish between them. Furthermore, these methods have weak anti-interference capabilities and are significantly affected by cloud cover, varying lighting conditions, and sensor noise.
[0003] In recent years, although multimodal fusion technology has improved feature richness, for example, hyperspectral data can provide dozens to hundreds of narrow band information, different modal data have inconsistent feature distribution due to differences in resolution and imaging principles. Direct fusion is prone to conflicts. If the feature space is not aligned, the classification accuracy will decrease after fusion, or some modal data may have poor quality due to environmental or equipment factors, such as hyperspectral bands interfered by clouds. Traditional average weighted or fixed weight fusion strategies are difficult to dynamically suppress noise. Summary of the Invention
[0004] The present invention provides a farmland boundary extraction method based on a multimodal reasoning model to solve existing problems: since different modal data have inconsistent feature distributions due to differences in resolution and imaging principles, direct fusion is prone to cause conflicts, and is also affected by the poor quality of some modal data due to environmental or equipment factors, resulting in farmland boundary extraction being seriously affected by noise.
[0005] The method for extracting cultivated land boundaries based on a multimodal reasoning model of the present invention adopts the following technical solutions:
[0006] The present invention proposes a method for extracting cultivated land boundaries based on a multimodal reasoning model, which includes the following steps:
[0007] Acquire multimodal data of cultivated land, construct a multimodal dataset for the multimodal data, construct a reconstruction network, use the multimodal dataset to train to obtain a pre-trained reconstruction network, and use the pre-trained reconstruction network to extract multi-scale features of the multimodal dataset; splice the multi-scale features of the multimodal dataset to obtain splicing features of the multimodal dataset, and group the multimodal dataset; construct a cultivated land boundary extraction network, which includes an encoder, a gating network module, a contrastive learning network module, and a decoder module of the pre-trained reconstruction network; construct a cultivated land boundary extraction network loss function, use the cultivated land boundary extraction network loss function and the grouped multimodal dataset to train the cultivated land boundary extraction network, obtain a trained cultivated land boundary extraction network, input the collected multimodal data into the cultivated land boundary extraction network, and obtain a cultivated land boundary extraction result.
[0008] Preferably, the multimodal data set includes: collecting hyperspectral data using a drone, where the drone carries a hyperspectral device to obtain hyperspectral data of cultivated land.
[0009] Preferably, the pre-trained reconstruction network includes: selecting an autoencoder network as a reconstruction network for multimodal data, freezing network learning parameters starting from the fourth feature layer of the encoder, and freezing each layer; the decoder in the autoencoder network structure is a normal convolution layer, no parameter freezing is performed, the loss function is the loss function of the autoencoder network, each multimodal data in the multimodal data set is used as the input of the autoencoder network, and the multimodal data set is used to implement autoencoder network training to obtain a trained autoencoder network as a pre-trained reconstruction network.
[0010] Preferably, the gated network module includes:
[0011] The gated network module is a fully connected network. A softmax layer is added after the fully connected layer, and the output of the softmax layer is used as the weight coefficient of the single modality data in the same group of multimodal data. In order to prevent the weight coefficient of the single modality data in the same group of multimodal data from being uniform and unchanged, the information entropy calculation formula is used to calculate the information entropy value of the output result of the softmax layer as the loss function.
[0012] Preferably, the contrastive learning network module includes: when the contrastive learning network module extracts the farmland boundary, the farmland boundary pixel points are used as positive samples, and the non-farmland boundary pixel points are used as negative samples to perform pixel-level contrastive learning, and the contrastive loss function of the contrastive learning network module is combined with the weight coefficient of the gated network module to obtain the weighted contrastive loss function of the contrastive learning network module. .
[0013] Preferably, the decoder module includes: the decoder is a multi-layer convolution and pooling structure, the size of the last feature layer of the decoder is consistent with the same size of each multimodal data, and the output result of the last feature layer of the decoder is mapped to 0 or 1 using a 0-1 activation function.
[0014] Preferably, the farmland boundary extraction network includes: inputting multimodal data collected by the drone in real time into the trained farmland boundary extraction network, extracting multi-scale features through the pre-trained reconstruction network encoder in the farmland boundary extraction network, splicing the multi-scale features to obtain splicing features, inputting the splicing features into the gated network module, obtaining the weight coefficient output by the gated network module, inputting the splicing features into the decoder, calculating the total loss function loss according to the decoder output result combined with the weight coefficient output by the gated network module, and using the total loss function loss to train the farmland boundary extraction network to obtain a trained farmland boundary extraction network.
[0015] The beneficial effects of the technical solution of the present invention are as follows: the present invention utilizes the effectiveness of effective multimodal information in farmland boundary extraction through multimodal data fusion and multi-scale feature extraction of the autoencoder network, combined with the weight coefficient and contrast learning mechanism of the gated network module, and adaptively suppresses interference model information, thereby improving the accuracy and robustness of farmland boundary extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is a flowchart of the steps of the farmland boundary extraction method based on the multimodal reasoning model of the present invention. DETAILED DESCRIPTION
[0018] To further illustrate the technical means and effectiveness of the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of the farmland boundary extraction method based on a multimodal reasoning model proposed by the present invention. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0019] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0020] The specific scheme of the cultivated land boundary extraction method based on the multimodal reasoning model provided by the present invention is described in detail below with reference to the accompanying drawings.
[0021] See also Figure 1 , which shows a flowchart of a method for extracting cultivated land boundaries based on a multimodal reasoning model according to an embodiment of the present invention. The method includes the following steps:
[0022] Step S1: Acquire multimodal data of cultivated land, construct a multimodal dataset for the multimodal data, construct a reconstruction network, train the multimodal dataset to obtain a pre-trained reconstruction network, and extract multi-scale features of the multimodal dataset using the pre-trained reconstruction network.
[0023] Step S101: Acquire multimodal data of cultivated land, and construct a multimodal dataset based on the multimodal data.
[0024] Hyperspectral data is collected using a drone with a planned path. The drone carries a hyperspectral device and collects hyperspectral data of cultivated land while flying along the planned path, thereby obtaining hyperspectral data of the cultivated land. Since hyperspectral data is a combination of multi-wavelength data, the collected hyperspectral data of cultivated land is multimodal data.
[0025] Multimodal data contains more cultivated land boundary information, but not all multimodal data are conducive to cultivated land boundary extraction. When the UAV is flying, the field of view is different due to different flight altitudes or the replacement of hyperspectral equipment, which makes the collected cultivated land multimodal data have field of view differences. Therefore, the multimodal data needs to be preprocessed to obtain multimodal data of uniform size. Moreover, when collecting multimodal data, different hyperspectral devices are affected differently by factors such as clouds and vegetation, resulting in some multimodal data interfering with the extraction of cultivated land boundaries.
[0026] To obtain uniform multimodal data, we used image scaling to resize each hyperspectral data element to a uniform size, achieving uniform multispectral data size and completing multispectral data preprocessing. This uniform size was achieved by calculating the mode value of the overall size of all hyperspectral data elements in the multispectral data set and using that as the uniform size.
[0027] Step S102: constructing a reconstruction network, using the multimodal dataset to train a pre-trained reconstruction network, and using the pre-trained reconstruction network to extract multi-scale features of the multimodal dataset.
[0028] When using multimodal data to extract cultivated land boundaries, in order to obtain better cultivated land boundary extraction results, it is necessary to fuse multimodal data. When fusing multimodal data, since the collected multimodal data have different fields of view, the spatial distribution is not uniform. The reconstruction network is used to extract features from the multimodal data, and the feature extraction results are anchored using contrastive learning to solve the problem of inconsistent spatial distribution and realize cultivated land boundary extraction.
[0029] To realize feature extraction of multimodal data, the present invention chooses to use a reconstruction network to extract features of multimodal data, and uses the autoencoder network as the reconstruction network of multimodal data. Implementers can choose other reconstruction networks for multimodal feature extraction according to specific implementation scenarios, such as U-net network.
[0030] The autoencoder network has an encoder and a decoder part. If the autoencoder network feature extraction is performed separately for each multimodal data, the feature extraction results will be scattered in different feature spaces, which is not conducive to the fusion of effective features when extracting cultivated land boundaries.
[0031] For the encoder part in the autoencoder network structure, the present invention chooses to freeze the network learning parameters starting from the fourth feature layer of the encoder, and freezes each layer, that is: the network learning parameters of the first, second, and third feature layers are not frozen, the network learning parameters of the fourth feature layer are frozen, the network learning parameters of the fifth feature layer are not frozen, and the network learning parameters of the third feature layer are frozen, and so on to implement the encoder of the autoencoder network. In this embodiment, the number of feature layers of the encoder is set to 7 layers, and the implementer can adjust the number of feature layers of the encoder according to the specific implementation scenario.
[0032] The decoder in the autoencoder network structure is a normal convolutional layer, and parameter freezing is not performed. The loss function is the loss function of the autoencoder network. The loss function of the autoencoder network is well known and will not be described in detail in the present invention.
[0033] Each multimodal data in the multimodal data set is used as the input of the autoencoder network, and the multimodal data set is used to implement autoencoder network training to obtain a trained autoencoder network. The training parameters and training process of the autoencoder network are well-known means and will not be described in detail in the present invention. Among them, each multimodal data in the multimodal data set is single-band spectral acquisition data in the hyperspectral image data, and the hyperspectral data is composed of acquired multi-band spectral data.
[0034] Each multimodal data in the multimodal dataset is input into the trained autoencoder network to obtain the feature layer data encoded by the encoder of the autoencoder network. The multi-scale feature layer is extracted from the feature layer data as the feature extraction result of the autoencoder network input data. Through multi-scale feature layer extraction, the feature extraction results under multiple receptive fields can be retained.
[0035] When extracting multi-scale feature layers from feature layer data, the present invention chooses to extract three feature layers, which are the first feature layer, the middle feature layer, and the last feature layer of the encoder, as the feature extraction results of the autoencoder network input data. Among them, the selection of three feature layers is an empirical value selection, and the implementer can adjust the number of layers and the position of the selected feature layer according to the specific implementation method.
[0036] The multi-scale feature layer extracted from the feature layer data is the multi-scale feature of each multimodal data. Similarly, the multi-scale feature belonging to each multimodal data is obtained.
[0037] Step S2: Splice the multi-scale features of the multimodal dataset to obtain the spliced features of the multimodal dataset and group the multimodal dataset.
[0038] By step S1, the multi-scale features of each multimodal data can be obtained. Since the multi-scale features are matrices of different scales, they cannot be directly used as input data of the cultivated land boundary extraction network. Multi-scale feature splicing is required, wherein the multi-scale feature splicing is to splice the multi-scale features into the same matrix by matrix splicing. The specific matrix splicing process is as follows: each feature matrix in the multi-scale feature is expanded into a one-dimensional vector by column, and the remaining feature matrices are expanded into one-dimensional vectors by column. The multiple one-dimensional vectors corresponding to the multi-scale features are connected end to end according to the scale size data, and the multiple one-dimensional vectors are spliced into a single one-dimensional vector. The single one-dimensional vector is reorganized into a new matrix by sorting the columns. The length and width of the new matrix are: the length of the new matrix is the square root of the number of single one-dimensional vector data, rounded up. The new matrix is a square matrix. The null value data in the new matrix is padded with zeros. The new matrix is the splicing matrix of the multi-scale features, and the splicing matrix of each multimodal data is obtained.
[0039] The multimodal data collected in the same shooting area at the same time is hyperspectral data, which is composed of data from multiple spectral bands. Therefore, the multimodal data collected in the same shooting area at the same time is a group of multimodal data, and a single group of multimodal data can be used to identify cultivated land boundaries.
[0040] Step S3: constructing a farmland boundary extraction network, which includes an encoder of a pre-trained reconstruction network, a gating network module, a contrastive learning network module, and a decoder module; constructing a farmland boundary extraction network loss function, using the farmland boundary extraction network loss function and the grouped multimodal data set to train the farmland boundary extraction network, obtaining a trained farmland boundary extraction network, and inputting the collected multimodal data into the farmland boundary extraction network to obtain a farmland boundary extraction result.
[0041] Since hyperspectral data in multimodal data contain multiple spectral bands, not all spectral bands are conducive to the extraction of cultivated land boundaries. In order to make the hyperspectral data in the same set of multimodal data be able to use effective modal data for cultivated land boundary extraction, an efficient method is proposed to extract cultivated land boundaries.
[0042] The present invention proposes a farmland boundary extraction network, which includes an encoder of a pre-trained reconstruction network, a gating network module, a contrastive learning network module and a decoder module.
[0043] Input a single group of multimodal data into the encoder of the reconstruction network trained in step S1 to obtain the multi-scale features of each modal data in the single group of multimodal data, splice the multi-scale features of each modal data to obtain the splicing features of each modal data in the single group of multimodal data, and use each modal data in the single group of multimodal data as the input of the gating network module and the decoder module respectively.
[0044] The gated network module is a fully connected network. A softmax layer is added after the fully connected layer, and the output of the softmax layer is used as the weight coefficient of the single modal data in the same group of multimodal data. In order to prevent the weight coefficient of the single modal data in the same group of multimodal data from being uniform and unchanged, the information entropy calculation formula is used to calculate the information entropy value of the output result of the softmax layer as the loss function. The smaller the entropy value of the output result of the softmax layer, the more dispersed the values between the weight coefficients of the single modal data are. Since the output result of the softmax layer is used as the weight coefficient of the single modal data in the same group of multimodal data, the sum of the weight coefficients of the single modal data in the same group of multimodal data is 1. The information entropy calculation formula is well known and will not be described in detail in this invention.
[0045] The gating network module can be used to obtain the weight coefficient of a single modal data in the same group of multimodal data. Contrastive learning can be used to anchor the feature space in the same group of multimodal data while also realizing the extraction of cultivated land boundaries. When contrastive learning is used to extract cultivated land boundaries from the same group of multimodal data, the corresponding weight coefficient is combined to improve the role of effective modal information in extracting cultivated land boundaries through contrastive learning, while reducing the interference effect of interfering modal information on the extraction of cultivated land boundaries.
[0046] Since the cultivated land boundary is extracted at the pixel level, when using the contrastive learning network module to extract the cultivated land boundary, it is necessary to use the cultivated land boundary pixels as positive samples and the non-cultivated land boundary pixels as negative samples for contrastive learning. Since the cultivated land boundary is extracted at the pixel level, it is necessary to annotate the multimodal data set. By manually annotating, the pixels belonging to the cultivated land boundary in each multimodal data set are marked as 1, and the remaining non-cultivated land boundary pixels are marked as 0. The standardized multimodal data set is obtained for contrastive learning to extract the cultivated land boundary.
[0047] Since the multimodal data in the same group are collected from the same shooting area at the same time, the cultivated land boundary data in the same group of multimodal data are the same cultivated land boundary, and thus the positive samples in the same group of multimodal data are the same positive samples, and the negative samples are the same negative samples. However, since the contrastive learning input is the splicing feature of each multimodal data in the same group of multimodal data, pixel-level recognition cannot be achieved, and a decoder needs to be added. The decoder is used to generate a cultivated land boundary mask image for the splicing feature of each multimodal data to complete the cultivated land boundary extraction. The decoder is a multi-layer convolution and pooling structure. The size of the last feature layer of the decoder is consistent with the same size of each multimodal data, and the 0-1 activation function is used to map the output result of the last feature layer of the decoder to 0 or 1. In this implementation, the Heaviside function is selected as the 0-1 activation function, and the implementer can adjust it according to the specific implementation method.
[0048] Then the total loss function of the farmland boundary extraction network is:
[0049]
[0050] 、 、 is a hyperparameter, adjust 、 、 The weights between and can be obtained by negatively adapting the network to obtain the parameter values.
[0051] The information entropy value of the weight coefficient corresponding to each multimodal data in the same set of multimodal data. The smaller the entropy value, the greater the difference between the values of the weight coefficients, thereby suppressing the influence of the interfering mode.
[0052] This is a weighted contrast loss function for the same set of multimodal data combined with the corresponding weight coefficient. To ensure that the larger the weight coefficient, the more important the corresponding multimodal data, the contrast loss of each modal data in the same set of multimodal data is multiplied by the corresponding weight coefficient to select valid multimodal data for farmland boundary extraction.
[0053] in, The specific calculation method is:
[0054]
[0055] is the number of samples in a multimodal data set, Express The traversal value of , is the temperature coefficient, which is used to adjust the smoothness of the feature similarity distribution. The empirical value of the temperature coefficient is 1.2. When the present invention is implemented, the input of each round of cultivated land recognition network training is a group of multimodal data.
[0056] A set of multimodal data The weight coefficient corresponding to the modal data represents the The effectiveness of the modal data in farmland boundary extraction is learned through parameter update in farmland boundary extraction network model training.
[0057] is a set of multimodal data The specific calculation process is as follows: in the output result of the decoder, extract the farmland boundary mask in the output result, and calculate the positive sample intersection and union ratio of the modal data. The intersection-over-union ratio between the masks composed of farmland boundary labels of the two modal data , when the intersection-union ratio is closer to 1, the better the farmland boundary extraction effect is, so Using the exponential function to map, we get , as the first Similarity of positive samples of modal data.
[0058] is a set of multimodal data The specific calculation process is as follows: in the output result of the decoder, extract the non-cultivated land boundary mask in the output result, and calculate the negative sample intersection and union ratio of the modal data. The intersection-over-union ratio between the non-cultivated land boundary label components of the two modal data When the intersection-union ratio is closer to 1, the non-cultivated land boundary extraction effect is better. Using the exponential function to map, we get , as the first Similarity of negative samples of data from different modalities.
[0059] is the sum of the similarities of the current samples, which is used to normalize the probability distribution.
[0060] As the farmland boundary extraction network is trained, The value will decrease, thereby maximizing the similarity between the positive sample pairs and minimizing the similarity between the positive sample pairs and the negative sample pairs, wherein the intersection-over-union ratio between the image masks is calculated as a well-known technique and will not be described in detail in the present invention.
[0061] For the same set of multimodal data The farmland boundary segmentation loss of modal data is calculated as follows: Indicates that in the decoder output, the farmland boundary mask in the output is extracted and the The consistency between the masks composed of farmland boundary labels of the modal data is then obtained by exponential function mapping. , represents the loss of farmland boundary segmentation, when The larger the value of The results of farmland boundary extraction for multimodal data are very poor, and the mean of all Z values in the same group of multimodal data is obtained as , as the farmland boundary segmentation loss for the same set of multimodal data as a whole, where The empirical value is 0.5, which can be adjusted by the implementer according to the specific implementation scenario.
[0062] The above-mentioned loss function and network search method are used to perform hyperparameter adaptation to realize the training of the cultivated land boundary extraction network, and the trained cultivated land boundary extraction network is obtained. The multimodal data collected by the UAV in real time is input into the trained cultivated land boundary extraction network to obtain the cultivated land boundary extraction result.
[0063] Through the above steps, the cultivated land boundary extraction method based on the multimodal reasoning model is completed.
[0064] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. The cultivated land boundary extraction method based on the multimodal reasoning model is characterized by: The method comprises the following steps: acquiring multimodal data of cultivated land, constructing a multimodal dataset for the multimodal data, constructing a reconstruction network, training the multimodal dataset to obtain a pre-trained reconstruction network, and extracting multi-scale features of the multimodal dataset using the pre-trained reconstruction network; Splicing multi-scale features of the multimodal dataset to obtain the splicing features of the multimodal dataset and grouping the multimodal dataset; A farmland boundary extraction network is constructed, which includes an encoder of a pre-trained reconstruction network, a gating network module, a contrastive learning network module, and a decoder module. A farmland boundary extraction network loss function is constructed, and the farmland boundary extraction network is trained using the farmland boundary extraction network loss function and the grouped multimodal dataset to obtain a trained farmland boundary extraction network. The collected multimodal data is input into the farmland boundary extraction network to obtain a farmland boundary extraction result. The contrastive learning network module includes: when the contrastive learning network module extracts the farmland boundary, the farmland boundary pixel points are used as positive samples, and the non-farmland boundary pixel points are used as negative samples to perform pixel-level contrastive learning, and the contrastive loss function of the contrastive learning network module is combined with the weight coefficient of the gated network module to obtain the weighted contrastive loss function of the contrastive learning network module. ; The specific calculation method is: ; Where, is the number of samples in a multimodal data set; is the temperature coefficient; A set of multimodal data The weight coefficient corresponding to each modal data; A set of multimodal data The negative sample orthogonal ratio of the modal data; is the total number of data in a set of sample multimodal data sets, is the positive sample intersection-union ratio of the modal data in a set of multimodal data; The collected hyperspectral data of cultivated land are multimodal data; The gated network module includes: the gated network module is a fully connected network, by adding a softmax layer after the fully connected layer, the output result of the softmax layer is used as the weight coefficient of the single modal data in the same group of multimodal data, in order to prevent the weight coefficient of the single modal data in the same group of multimodal data from being uniform and unchanged, the information entropy calculation formula is used to calculate the information entropy value of the output result of the softmax layer as the loss function .
2. The method for extracting cultivated land boundaries based on a multimodal reasoning model according to claim 1, characterized in that: The multimodal data set includes: using a drone to collect hyperspectral data, the drone carries a hyperspectral device to obtain hyperspectral data of cultivated land.
3. The method for extracting cultivated land boundaries based on a multimodal reasoning model according to claim 1, characterized in that: The pre-trained reconstruction network includes: selecting an autoencoder network as a reconstruction network for multimodal data, freezing network learning parameters starting from the fourth feature layer of the encoder, freezing each layer, the decoder in the autoencoder network structure is a normal convolution layer, no parameter freezing is performed, the loss function is the loss function of the autoencoder network, each multimodal data in the multimodal data set is used as the input of the autoencoder network, the multimodal data set is used to implement autoencoder network training, and a trained autoencoder network is obtained as the pre-trained reconstruction network.
4. The method for extracting cultivated land boundaries based on a multimodal reasoning model according to claim 1, characterized in that: The decoder module includes: a decoder with a multi-layer convolution and pooling structure, the size of the last feature layer of the decoder is consistent with the same size of each multimodal data, and a 0-1 activation function is used to map the output result of the last feature layer of the decoder to 0 or 1.
5. The method for extracting cultivated land boundaries based on a multimodal reasoning model according to claim 1, characterized in that: The farmland boundary extraction network includes: inputting multimodal data collected by a drone in real time into a trained farmland boundary extraction network, extracting multi-scale features through a pre-trained reconstruction network encoder in the farmland boundary extraction network, splicing the multi-scale features to obtain splicing features, inputting the splicing features into a gated network module to obtain a weight coefficient output by the gated network module, inputting the splicing features into a decoder, calculating a total loss function loss based on the decoder output result combined with the weight coefficient output by the gated network module, and using the total loss function loss to train the farmland boundary extraction network to obtain a trained farmland boundary extraction network.
Citation Information
Patent Citations
End-to-end speech recognition method and system and storage medium
CN114596839A
Method, apparatus and program for generating composition map and for detecting source gas using artificial intelligence
KR102407251B1