Information identification method and device and storage medium

By keyframe sorting and time stamping determination of remote sensing images within the preset time period, combined with the trained abandoned identification network model, the problem of low accuracy of abandoned land detection in remote sensing image change detection is solved, and more accurate abandoned land detection is achieved.

CN120107773APending Publication Date: 2025-06-06CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311659948.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing remote sensing image change detection methods have low accuracy in the detection of abandoned land, and the long-term change trend of inland objects during time intervals is not fully considered.

Method used

By obtaining the remote sensing image set within the preset time period, selecting at least two key images, sorting them in chronological order, determining the acquisition time of each frame of images, and finally using the trained abandoned recognition network model for abandoned recognition.

Benefits of technology

The accuracy of the detection results of abandoned land is improved, and the reliability of the detection is enhanced by considering the long-term trend of land objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107773A_ABST
    Figure CN120107773A_ABST
Patent Text Reader

Abstract

The invention discloses an information identification method. The method comprises the steps of obtaining a to-be-analyzed remote sensing image set corresponding to a to-be-identified region in a preset time period; sorting at least two frames of key images selected from the to-be-analyzed remote sensing image set according to a time sequence to obtain a continuous to-be-processed image set; determining the image acquisition time of each frame of the key image in the to-be-processed image set to obtain at least two target acquisition times; and through a trained abandoned land identification network model, carrying out abandoned land identification on the image set to be processed and the at least two target acquisition times to obtain a target identification result of the region to be identified, so that the method for realizing abandoned land detection by using the remote sensing images of the continuous time sequence fully considers the long-term change trend of the ground features, and improves the detection accuracy of the abandoned land. And the accuracy of the abandoned land detection result is improved. The invention further discloses information identification equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer application technology, and in particular to an information identification method, device and storage medium. Background Art

[0002] With the rapid development of remote sensing technology, remote sensing image detection technology has been further improved and has been widely used in many fields such as agriculture, military, and national defense security. In the agricultural field, the retention of cultivated land is closely linked to food security, and the abandonment of cultivated land will seriously affect food production. Cultivated land resources have been in a serious vicious cycle for a long time, that is, on the one hand, it is imperative to expand the area of ​​cultivated land and reclaim wasteland, and on the other hand, due to various reasons, it is abandoned. Due to the great impact of abandoned cultivated land, people are paying more and more attention to it. Therefore, conducting abandoned land surveys and monitoring is of great significance to help identify non-agriculturalization and ensure the red line of food security. At present, the method of using remote sensing image change detection to realize abandoned land inspection can reduce manpower and material resources.

[0003] However, the current remote sensing image change detection process usually analyzes a single remote sensing image to implement abandoned land inspection, without considering the correlation between time and ignoring the changes of some vegetation during the growth cycle, resulting in low accuracy of abandoned land detection results.

[0004] Application Contents

[0005] In order to solve the above-mentioned technical problems, the present application hopes to provide an information identification method, device and storage medium, which solves the problem of low accuracy of abandoned land detection results in remote sensing image change detection, and proposes a method for abandoned land detection using continuous time series remote sensing images, which fully considers the long-term change trend of the ground objects and improves the accuracy of abandoned land detection results.

[0006] The technical solution of this application is implemented as follows:

[0007] In one aspect, an information identification method comprises:

[0008] Acquire a set of remote sensing images to be analyzed corresponding to the area to be identified within a preset time period;

[0009] Sorting at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous image set to be processed;

[0010] Determine the image acquisition time of each key image in the image set to be processed to obtain at least two target acquisition times;

[0011] Through the trained abandonment identification network model, abandonment identification is performed on the image set to be processed and at least two target acquisition times to obtain the target identification result of the area to be identified.

[0012] In one aspect, an information identification device comprises at least: a memory, a processor and a communication bus; wherein:

[0013] The memory is used to store executable instructions;

[0014] The communication bus is used to realize the communication connection between the processor and the memory;

[0015] The processor is used to execute the information identification program stored in the memory to implement the steps of the information identification method as described in any one of the above items.

[0016] On the one hand, a storage medium stores an information identification program, and when the information identification program is executed by a processor, the steps of any of the above-mentioned information identification methods are implemented.

[0017] The embodiment of the present application provides an information identification method, device and storage medium, wherein the information identification device obtains a set of ground analysis remote sensing images corresponding to the area to be identified within a preset time period, sorts at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous set of images to be processed, and determines the image acquisition time of each key image in the image set to be processed to obtain at least two target acquisition times, and finally uses a trained abandonment identification network model to perform abandonment identification on the image set to be processed and at least two target acquisition times to obtain a target identification result of the area to be identified. In this way, key frames are extracted from remote sensing images collected within a period of time, sorted in chronological order to obtain a continuous set of images to be processed, and abandonment identification is performed on the image set to be processed and the corresponding target acquisition time through the trained abandonment identification network model to obtain a target identification result of the area to be identified, which solves the problem of low accuracy of abandoned land detection results of remote sensing image change detection, and proposes a method for realizing abandoned land detection using remote sensing images with continuous time series, which fully considers the long-term change trend of ground objects and improves the accuracy of abandoned land detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of the process of the information identification method provided in the embodiment of the present application Figure 1 ;

[0019] Figure 2 Schematic diagram of the process of the information identification method provided in the embodiment of the present application Figure 2 ;

[0020] Figure 3A schematic diagram of a system architecture structure for implementing an information identification method provided in an embodiment of the present application;

[0021] Figure 4 A schematic diagram of the implementation process of image data set selection and labeling provided in an embodiment of the present application;

[0022] Figure 5 A schematic diagram of a system architecture of a data encoding module provided in an embodiment of the present application;

[0023] Figure 6 A schematic diagram of the structure of an improved patch embedding system provided in an embodiment of the present application;

[0024] Figure 7 A schematic diagram of an implementation of an improved window calculation range provided in an embodiment of the present application;

[0025] Figure 8 A schematic diagram of an improved attention calculation method provided in an embodiment of the present application;

[0026] Fig. 9 A schematic diagram of a system architecture for implementing coding feature learning based on an improved W-MSA and an improved VSA module provided in an embodiment of the present application;

[0027] Fig.10 A schematic diagram of the system architecture of a data decoding module provided in an embodiment of the present application;

[0028] Fig.11 A flowchart of an application implementation of an information identification method provided in an embodiment of the present application;

[0029] Fig.12 A schematic diagram of the structure of an information identification device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0031] The embodiment of the present application provides an information identification method, referring to Figure 1 As shown, the method is applied to an information recognition device, and the method comprises the following steps:

[0032] Step 101: Obtain a set of remote sensing images to be analyzed corresponding to a region to be identified within a preset time period.

[0033] In the embodiment of the present application, the preset time period may be a preset time period, which may be an empirical time period value obtained from a large number of experiments, or an empirical time period value set by the user based on his or her own experience. The area to be identified is an area where wasteland identification is required, such as a cultivated land area that needs to be monitored and identified. The information identification device downloads and obtains a set of remote sensing images to be analyzed within a preset time period corresponding to the area to be identified from a data source that can provide remote sensing images taken of the earth's surface.

[0034] Step 102: sort at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous image set to be processed.

[0035] In an embodiment of the present application, the information recognition device performs image analysis on the images in the acquired remote sensing image set to be analyzed, selects at least two key image frames therefrom, and then sorts and processes the selected at least two key image frames in chronological order to obtain a set of images to be processed with temporal continuity.

[0036] Step 103: Determine the image acquisition time of each key image in the image set to be processed, and obtain at least two target acquisition times.

[0037] In an embodiment of the present application, the image acquisition time of each key image in the image set to be processed is obtained to obtain at least two target acquisition times.

[0038] Step 104: Using the trained abandonment identification network model, perform abandonment identification on the image set to be processed and at least two target acquisition times to obtain target identification results for the area to be identified.

[0039] In the embodiment of the present application, the target recognition result may be the map image information corresponding to the region to be recognized, wherein identification information of whether it is an abandoned area may be provided for different blocks in the region to be recognized. A trained abandonment recognition network model is obtained, and the image set to be processed and the corresponding at least two target acquisition times are input into the trained abandonment recognition network model. The trained abandonment recognition network model is used to perform abandonment recognition on the image set to be processed and the at least two target acquisition times, thereby obtaining a target recognition result on whether there is abandonment in the region to be recognized.

[0040] In some application scenarios, after the information identification device obtains the target identification result by performing abandonment identification on the area to be identified through the abandonment identification network model, the target identification result can be output and displayed. For example, it can be displayed directly on the display screen corresponding to the information identification device, or the target identification result can be sent to the third-party display device corresponding to the monitoring personnel, such as a smart mobile terminal or other display terminal for display. In this way, the monitoring personnel can manage the cultivated land in the area to be identified according to the target identification result, so as to improve the land resource utilization rate of the area to be identified.

[0041] The embodiment of the present application provides an information identification method, which obtains a set of ground analysis remote sensing images corresponding to a region to be identified within a preset time period through an information identification device, sorts at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous set of images to be processed, and determines the image acquisition time of each key image in the image set to be processed to obtain at least two target acquisition times, and finally uses a trained abandonment identification network model to perform abandonment identification on the image set to be processed and at least two target acquisition times to obtain a target identification result of the region to be identified. In this way, key frames are collected from remote sensing images collected within a period of time, sorted in chronological order to obtain a continuous set of images to be processed, and abandonment identification is performed on the image set to be processed and the corresponding target acquisition time through a trained abandonment identification network model to obtain a target identification result of the region to be identified, which solves the problem of low accuracy of abandoned land detection results of remote sensing image change detection, and proposes a method for realizing abandoned land detection using remote sensing images with continuous time series, which fully considers the long-term change trend of ground objects and improves the accuracy of abandoned land detection results.

[0042] Based on the foregoing embodiments, an embodiment of the present application provides an information identification method, which is applied to an information identification device, and the method includes the following steps:

[0043] Step 201: Obtain a set of remote sensing images to be analyzed corresponding to a region to be identified within a preset time period.

[0044] In the embodiment of the present application, the elements in the remote sensing image set to be analyzed are images obtained by remote sensing the ground content of the area to be identified using a satellite as an example for explanation, and the satellite remote sensing images of the area including the area to be identified within the preset time period are downloaded from the data source corresponding to the satellite to obtain the remote sensing image set to be analyzed corresponding to the area to be identified within the preset time period. In order to ensure accurate identification of whether there is abandonment in the area to be identified, the preset time period is usually as long as possible, for example, it can be in years, so that it can be determined whether there is abandonment based on the growth changes of crops in the area to be identified.

[0045] Step 202: sort at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous image set to be processed.

[0046] In the embodiment of the present application, since the number of images included in the remote sensing image set to be analyzed is too large, in order to reduce the processing and calculation pressure of the information recognition device, and within the longer time range corresponding to the same season, the changes in the vegetation and other crops in the area to be identified are small, therefore, at least one image can be selected as a key image within the same season time range. In this way, the information recognition device can select at least two frames of key images from the remote sensing image set to be analyzed. The information recognition device can sort the selected at least two frames of key images in the time sequence of the acquisition time to obtain the image set to be processed, and the image set to be processed can be in the form of a video.

[0047] The Digital Surface Model (DSM) represents the surface of the earth and includes the elevation information of all objects on it, including slope, height, coordinates, etc. The Digital Terrain Model (DTM) represents the bare ground without any objects such as plants and buildings, and only contains the elevation data of natural terrain. These two types of data are indispensable in many projects involving terrain, and can be used for three-dimensional establishment, slope and aspect analysis, engineering construction applications, geological applications, land use detection and many other fields. In the implementation process, in order to reduce the amount of data calculation, in the implementation process, only the elevation information corresponding to the first frame and the last frame included in the image set to be processed can be obtained, and the corresponding elevation change label can be determined.

[0048] Normalized Difference Vegetation Index (NDVI) is one of the important parameters that reflects the growth and nutritional information of crops. It is used in remote sensing images to detect vegetation growth status, vegetation coverage, and eliminate some radiation errors. NDVI can reflect the background influence of plant canopies, such as soil, wet ground, snow, dead leaves, roughness, etc., and is related to vegetation coverage. The NDVI calculation formula can be recorded as NDVI = (NIR-Red) / (NIR + Red), where NIR is the near-infrared band of the spectrum and Red is the red band of the spectrum.

[0049] Step 203: Determine the image acquisition time of each key image in the image set to be processed, and obtain at least two target acquisition times.

[0050] In the embodiment of the present application, each key image frame in the image set to be processed has a shooting time. Therefore, the image acquisition time of each key image frame can be obtained from each key image frame to obtain at least two target acquisition times.

[0051] Step 204: Using the trained abandonment identification network model, perform abandonment identification on the image set to be processed and at least two target acquisition times to obtain target identification results for the area to be identified.

[0052] In an embodiment of the present application, the abandonment identification network model is a trained abandonment identification network model obtained by improving the transformer model and using model training.

[0053] Based on the above embodiments, in other embodiments of the present application, refer to Figure 2 As shown, before the information recognition device executes step 204, it is also used to execute the following steps:

[0054] Step 205: Determine a first type of sample data set.

[0055] In the embodiment of the present application, a first type of sample data set is obtained from a first type of sample data source.

[0056] Step 206: Determine a second type of sample data set.

[0057] The data in the first type sample data set and the data in the second type sample data set are different.

[0058] In an embodiment of the present application, a second type of sample data set is obtained from a second type of sample data source. The sample data in the first type of sample data set and the second type of sample data set are both remote sensing image data, which are collected by remote sensing of the earth's surface through remote sensing methods with different remote sensing accuracies. Generally, the data accuracy of the sample data in the first type of sample data set is better than the data accuracy of the sample data in the second type of sample data set.

[0059] Step 207: Determine the network model to be trained.

[0060] In the embodiment of the present application, the network model to be trained is a pre-set initial network model that requires model training.

[0061] Step 208: Use the first type sample data set and the second type sample data set to perform model training on the network model to be trained, so as to obtain a trained abandonment recognition network model.

[0062] In an embodiment of the present application, the first type of sample data set and the second type of sample data set are used as training samples to perform model training on the network model to be trained until a wasteland recognition network model with a model loss value that meets the requirements is obtained through training.

[0063] It should be noted that steps 205 to 208 can also be implemented as a separate embodiment, that is, steps 205 to 208 can also be that the information recognition device pre-trains the sample data to obtain the abandonment identification network model. In this way, when the abandonment identification network model is needed in the future, the trained abandonment identification network model can be directly called.

[0064] Based on the above embodiments, in other embodiments of the present application, step 205 can be implemented by steps 205a to 205c:

[0065] Step 205a: Obtain, from a first data source, a first remote sensing image set collected for the photographed object within a first time period.

[0066] Among them, each element in the first remote sensing image set includes at least: image elements of at least two frames of images for the same photographed object, regional feature labels about regional features corresponding to each frame of image in each image element, image acquisition time of each frame of image in each image element, and elevation data information in each frame of image.

[0067] In the embodiment of the present application, the first data source may be a data source for providing remote sensing images specified by a user, and the first time period may be a time period specified by a user, usually a historical time period. In order to ensure that the trained model is optimal, the longer the first time period is, the better. In this way, the first remote sensing image set obtained includes more sample data. The photographed object is usually the surface of the earth.

[0068] Step 205b: Determine a first sub-image set and a second sub-image set based on the first remote sensing image set.

[0069] Each image element in the first sub-image set includes: at least two key image frames corresponding to the same photographed object within a preset time period, and the second sub-image set includes the collective elements in the first remote sensing image set except the first sub-image set.

[0070] In an embodiment of the present application, images in the first remote sensing image set are selected and processed, and image selection is performed on images corresponding to each photographed object at a time interval of a preset length to obtain at least two frames of key images corresponding to the same photographed object, to obtain a first sub-image set, and the image set in the first remote sensing image set excluding the elements in the first sub-image set is determined as a second sub-image set.

[0071] For example, it is assumed that in a first time period, remote sensing images are collected for four areas corresponding to a, b, c and d on the surface of the earth, and the first remote sensing image set is recorded as ((a0, a1, a2, ..., a190), (b0, b1, b2, ..., b190), (c0, c1, c2, ..., c190), (d0, d1, d2, ..., d190)), and the preset time length is the time length for taking 100 remote sensing images. Correspondingly, when selecting key frames, every 10 remote sensing images are selected. When an image is used as a key image, the corresponding first sub-image set can be recorded as ((a0, a10, a20, a30, a40, a50, a60, a70, a80, a90), (a100, a110, a120, a130, a140, a150, a160, a170, a180, a190), (b0, b10, b20, b30, b40, b50, b60, b70, b80, b90), (b100, b110, b120 , b130, b140, b150, b160, b170, b180, b190), (c0, c10, c20, c30, c40, c50, c60, c70, c80, c90), (c100, c110, c120, c130, c140, c150, c160, c170, c180, c190), (d0, d10, d20, d30, d40, d50, d60, d70, d80, d90), (d100, d 110, d120, d130, d140, d150, d160, d170, d180, d190)), and the corresponding second sub-image set is ((a1, a2, … , a9, a11, a12, … , a189), (b1, b2, … , b9, b11, b12, … , b189), (c1, c2, … , c9, c11, c12, … , c189), (d1, d2, … , d9, d11, d12, … , d189)).

[0072] Step 205c: Determine a first type of sample data set based on the first sub-image set.

[0073] In the embodiment of the present application, the image elements in the first sub-image set are identified and the like to obtain a first type of sample data set including abandoned results.

[0074] Based on the above embodiment, in other embodiments of the present application, step 205c can be implemented by steps a11 to a16:

[0075] Step a11: Use preset identification information to identify the non-cultivated land area in each image element in the first sub-image set to obtain a third sub-image set.

[0076] In an embodiment of the present application, the preset identification information may be information pre-adopted for image identification processing. For example, the preset identification information 0 may be used as mask information corresponding to the non-arable land area. In this way, only the cultivated land information in each image element in the first sub-image set is retained to obtain the third sub-image set, thereby reducing the possibility of subsequent analysis of the non-arable land area, reducing two calculation analyses, and improving the efficiency of the calculation analysis.

[0077] Step a12: determine the cultivated land change in each image element in the third sub-image set, and obtain the cultivated land change label of each image element.

[0078] In the embodiment of the present application, the cultivated land change in each image element is determined based on the change of the cultivated land area in the two previous and next frames of image elements in the third sub-image set, and the cultivated land change label of each image element is obtained.

[0079] Step a13: Based on the elevation data information of each frame of image, determine the elevation change label of each image element in the third sub-image set.

[0080] In an embodiment of the present application, the elevation data information of each image element in the third sub-image set is calculated and analyzed. For example, for two adjacent frames of image elements, the elevation data information of the latter frame of image elements and the elevation data information of the previous frame of image elements are used to perform calculations to obtain an elevation change label for each frame of image.

[0081] Step a14: determine the normalized vegetation index of each image element in the third sub-image set.

[0082] In the embodiment of the present application, the normalized vegetation index calculation formula is used to calculate the normalized vegetation index of each image element in the third sub-image set.

[0083] Step a15, take the third sub-image set, the normalized vegetation index corresponding to each image element in the third sub-image set, the elevation change label corresponding to each image element in the third sub-image set, the image acquisition time of each image element in the third sub-image set, and the cultivated land change label corresponding to each image element in the third sub-image set as the fourth sub-image set.

[0084] In an embodiment of the present application, it is determined that the fourth sub-image set includes the third sub-image set, the normalized vegetation indication corresponding to each image element in the third sub-image set, the elevation change label corresponding to each image element in the third sub-image set, the image acquisition time of each image element in the third sub-image set, and the cultivated land change label corresponding to each image element in the third sub-image set, wherein the above elements in the fourth sub-image set can be recorded and stored in the form of channels.

[0085] Step a16: perform image segmentation processing on each frame of image corresponding to each image element in the fourth sub-image set according to a preset size to obtain a first type of sample data set.

[0086] In an embodiment of the present application, the preset size may be an empirical value obtained from a large number of experiments, and is used as the cutting size for each group of image elements when performing image cutting. For example, it may be recorded as M*N, and the values ​​of M and N may be the same or different, and may be determined specifically according to actual conditions. Exemplarily, for the image elements (a0, a10, a20, a30, a40, a50, a60, a70, a80, a90) in the fourth sub-image set, a preset size of 896*896 is used to perform image segmentation processing on (a0, a10, a20, a30, a40, a50, a60, a70, a80, a90), and the image elements in the first type of sample data set corresponding to (a0, a10, a20, a30, a40, a50, a60, a70, a80, a90) are obtained.

[0087] Based on the above embodiment, in other embodiments of the present application, step 206 can be implemented by steps 206a to 206d:

[0088] Step 206a: Obtain, from a second data source, a second remote sensing image set collected for the photographed object within a second time period.

[0089] Each element in the second remote sensing image set includes at least: an image, an image acquisition time of each image, and a ground object segmentation label.

[0090] In an embodiment of the present application, the information identification device obtains from the second data source a second remote sensing image set obtained by remote sensing photography of the earth's surface by a remote sensing device corresponding to the second data source during a second time period.

[0091] Step 206b: Determine the ground object segmentation label corresponding to each image element in the second sub-image set.

[0092] In the embodiment of the present application, each image element in the second sub-image set corresponding to the first data source is analyzed to determine the ground object segmentation label corresponding to each image element in the second sub-image set.

[0093] Step 206c: perform fusion processing on the second remote sensing image set, the second sub-image set, and the elements in the ground object segmentation label corresponding to the second sub-image set to obtain a fifth sub-image set.

[0094] In an embodiment of the present application, the elements in the second sub-image set and the ground object segmentation labels corresponding to the second sub-image set are fused with the images in the second remote sensing image set, that is, the second sub-image set and the ground object segmentation labels corresponding to the second sub-image set are merged with the second remote sensing image set to obtain the fifth sub-image set.

[0095] Step 206d: perform image segmentation processing on the image elements in the fifth sub-image set according to a preset size to obtain a second type of sample data set.

[0096] In the embodiment of the present application, a preset size is used to perform image segmentation processing on each image element in the fifth sub-image set, so that a second type of sample data set can be obtained.

[0097] Based on the above embodiment, in other embodiments of the present application, step 208 can be implemented by steps 208a to 208g:

[0098] Step 208a: Acquire first training sample data from the first type sample data set.

[0099] Step 208b: Acquire second training sample data from the second type sample data set.

[0100] Step 208c: input the first training sample data and the second training sample data into the network model to be trained to obtain a first abandoned land recognition result, a first elevation recognition result and a first land feature segmentation recognition result.

[0101] In an embodiment of the present application, first training sample data selected from a first type of sample data set and second training sample data selected from a second type of sample data set are input into a network model to be trained, and the network model to be trained outputs a first abandonment recognition result, a first elevation recognition result, and a first land feature segmentation recognition result.

[0102] Step 208d: Determine a target loss value based on the first abandonment identification result, the first elevation identification result, and the first land feature segmentation identification result.

[0103] In the embodiment of the present application, the first abandonment identification result, the first elevation identification result and the first land feature segmentation identification result are calculated using corresponding loss functions to obtain a target loss value.

[0104] Among them, after the information identification device executes step 208d, it can choose to execute step 208e or steps 208f~208g. If the target loss value is less than or equal to the preset threshold, choose to execute step 208e; if the target loss value is greater than the preset threshold, choose to execute steps 208f~208g.

[0105] Step 208e: If the target loss value is less than or equal to the preset threshold, the abandonment identification network model is determined to be the network model to be trained.

[0106] In the embodiment of the present application, the preset threshold is an empirical value obtained from a large number of experiments, or an empirical value set according to the user's own actual use experience. When the target loss value is less than or equal to the preset threshold, it indicates that the first abandonment identification result, the first elevation identification result, and the first land object segmentation identification result are slightly different from the actual corresponding results, indicating that the current network model to be trained has met the requirements. Therefore, it can be determined that the network model to be trained is the abandonment identification network model finally used for prediction.

[0107] Step 208f: If the target loss value is greater than the preset threshold, update the parameters in the network model to be trained based on the target loss value to obtain a reference network model.

[0108] In an embodiment of the present application, when the target loss value is greater than a preset threshold, it indicates that the predicted result of the network model to be trained is significantly different from the actual result, and the network model to be trained needs to be trained. Therefore, the target loss value is transmitted in a direction and input into the network model to be trained, so that the adjustable weight parameters in the network model to be trained are adjusted according to the target loss value. In this way, after the adjustable weight parameters in the network model to be trained are adjusted according to the target loss value, a reference network model is obtained.

[0109] Step 208g: After updating the model to be trained as the reference network model, repeat the step of “obtaining the first training sample data from the first type of sample data set” until the abandonment identification network model is obtained.

[0110] In an embodiment of the present application, the model to be trained is updated to a reference network model, and then steps 208a to 208g are re-executed until the target loss value of the model to be trained is less than or equal to a preset threshold value, and the model to be trained is determined to be an abandonment identification network model.

[0111] Based on the above embodiment, in other embodiments of the present application, step 208c can be implemented by steps c11 to c14:

[0112] Step c11: perform temporal information integration processing on the first training sample data and the second training sample data respectively through the patch embedding layer in the encoding module of the network model to be trained to obtain corresponding first fused sample data and second fused sample data.

[0113] In an embodiment of the present application, the network model to be trained includes an encoding module, in which a patch embedding layer is provided, and the patch embedding layer is used to fuse the target acquisition time corresponding to each image element in the first training sample data and the second training sample data into the corresponding image element to reduce the number of parameters of image calculation. The first fused sample data corresponds to the first training sample data, and the second fused sample data corresponds to the second training sample data.

[0114] Step c12: extract the abandonment feature information and elevation feature information in the first fused sample data through the first network including a multi-head self-attention mechanism and a rotational self-attention mechanism in the encoding module to perform attention calculation.

[0115] Among them, the multi-head self-attention mechanism includes at least a spatial offset relative to a reference window, a spatial scale scaling factor, and a temporal offset for learning three-dimensional window information of variable size and position. The rotational self-attention mechanism is used to implement a temporal feature learning mechanism that learns temporal information and spatial information separately.

[0116] In an embodiment of the present application, a first network including a multi-head self-attention mechanism and a rotational self-attention mechanism for attention calculation performs feature extraction on the abandonment feature information and elevation feature information in the first fused sample data to obtain the abandonment feature information and elevation feature information in the first fused sample data. The first network may be a Swin transformer network module including a multi-head self-attention mechanism (W-MSA) and a rotational self-attention mechanism (VSA).

[0117] Step c13: extracting the ground object segmentation feature information in the second fused sample data through the second network in the encoding module.

[0118] In the embodiment of the present application, the second network in the encoding module is used to perform learning and analysis on the second fused sample data, and extract corresponding ground feature segmentation feature information from the second fused sample data. The second network may be a Segformer encoder network module.

[0119] Step c14: Decode the abandonment feature information, the elevation feature information and the land feature segmentation feature information through a decoding module to obtain a first abandonment identification result, a first elevation identification result and a first land feature segmentation identification result.

[0120] In an embodiment of the present application, the abandonment feature information, elevation feature information and land feature segmentation feature information are decoded and analyzed respectively through corresponding decoding layers to obtain the first abandonment identification result, the first elevation identification result and the first service segmentation identification result corresponding to the first sample data.

[0121] Based on the above embodiment, in other embodiments of the present application, step 208d can be implemented by steps d11 to d14:

[0122] Step d11: Determine an abandonment loss value based on the first abandonment identification result and the abandonment identification label in the first training sample data.

[0123] In an embodiment of the present application, a pre-set abandonment loss function is used to calculate the loss value of the first abandonment identification result and the actual abandonment identification label set in the first training sample data to obtain the abandonment loss value.

[0124] Step d12: determining an elevation loss value based on the first elevation recognition result and the elevation change label in the first training sample data.

[0125] In an embodiment of the present application, a preset elevation loss function is used to calculate the loss value of the first elevation recognition record and the actual elevation change label set in the first training sample data to obtain the elevation loss value.

[0126] Step d13: Determine the ground object loss value based on the first ground object segmentation and recognition result and the ground object segmentation label in the second training sample data.

[0127] In the embodiment of the present application, a preset fifth segmentation loss function is used to calculate the loss value of the first ground object segmentation and recognition result and the actual ground object segmentation label set in the second training sample data to obtain the ground object loss value.

[0128] Step d14: perform weighted sum calculation on the abandonment loss value, elevation loss value and ground feature loss value to obtain a target loss value.

[0129] In an embodiment of the present application, the weighted weight coefficients corresponding to the abandonment loss value, elevation loss value and terrain loss value are determined. For example, the weighted weight coefficient of the abandonment loss value is recorded as x1, the weighted weight coefficient of the elevation loss value is recorded as x2, and the weighted weight coefficient of the terrain loss value is recorded as x3. The calculation formula of the corresponding target loss value can be recorded as "target loss value = x1*abandonment loss value + x2*elevation loss value + x3*terrain loss value".

[0130] Based on the above embodiments, in other embodiments of the present application, step 204 can be implemented by steps 204a to 204c:

[0131] Step 204a: perform image segmentation processing on the images in the to-be-processed image set according to a preset size to obtain an image set to be input.

[0132] The image set to be input includes a preset number of sub-image groups.

[0133] In the embodiment of the present application, the preset number in the image set to be input is an actual number value determined after the images in the image set to be processed are segmented according to the image segmentation method according to the preset size. The image set to be processed is a group of images, so by cutting this group of images according to the preset size, a preset number of sub-image groups can be obtained.

[0134] Step 204b: using the abandonment recognition network model to perform recognition processing on the input image set and at least two target acquisition times, and obtaining the ground object prediction results, elevation prediction results and abandonment prediction results corresponding to a preset number of sub-images.

[0135] In the embodiment of the present application, the wasteland identification network model performs identification and prediction processing on each sub-image group in the input image set, and obtains the ground object prediction result, elevation prediction result and wasteland prediction result corresponding to each sub-image group. The preset number of sub-images can be the images obtained by segmenting the images in the image set to be processed that were collected most recently from the current time.

[0136] Step 204c: splice and fuse the ground feature prediction results, elevation prediction results and land abandonment prediction results corresponding to a preset number of sub-images to obtain a target recognition result.

[0137] In an embodiment of the present application, data splicing and fusion processing is performed on the ground feature prediction results, elevation prediction results and abandonment prediction results corresponding to a preset number of sub-images to obtain the final target recognition result.

[0138] Based on the above embodiment, in other embodiments of the present application, step 204c can be implemented by steps e11 to e12:

[0139] Step e11: Use weighted convolution to concatenate the land object prediction results and land abandonment prediction results corresponding to a preset number of sub-images to obtain land object detection results and land abandonment detection results of the area to be identified.

[0140] Among them, the target recognition results include the ground object detection results and abandonment detection results of the area to be identified.

[0141] Step e12: Use normalized weighted convolution to concatenate the elevation prediction results corresponding to a preset number of sub-images to obtain an elevation detection result.

[0142] Among them, the target recognition result includes the elevation detection result.

[0143] Based on the above embodiments, the present application embodiment provides a schematic diagram of a system architecture structure for implementing the above information identification method. Figure 3As shown, it includes: a data input module, which is used to obtain a pseudo video frame sequence, a data preprocessing module, which is used to preprocess the sample data, and then input the preprocessed pseudo video frame sequence to the data encoding module, the result of the encoding processing by the data encoding module is input to the decoding module, and finally the decoding module performs multi-task coupled decoding. It also has a splicing module, which is at least used to splice images, and can also splice the prediction results output by the decoding module based on the splicing of the images. The decoding module and the splicing module output the results to the output module, and the final result is output through the output module for subsequent display light processing.

[0144] Based on the foregoing embodiments, the present application provides an implementation process for implementing model training in an information recognition method, including the following implementation steps:

[0145] Step 1: Remote sensing image dataset selection and labeling.

[0146] Among them, the general implementation process of step one can refer to Figure 4 As shown, it includes selecting a remote sensing image data set, and performing a training data processing process and a pre-training data processing process on the data in the remote sensing image data set, wherein the training data processing process includes: obtaining a selected pseudo video sequence and unselected data from a first data set; obtaining the time sequence information of the pseudo video sequence; obtaining the elevation information of the pseudo video sequence, and then determining the elevation change label of the pseudo video sequence; obtaining the abandoned change label of the pseudo video sequence; and obtaining the image semantic label of the pseudo video sequence. The pre-training data processing process includes: obtaining the image semantic label of the unselected data; obtaining a second data set; obtaining the image semantic label of the second data set; fusing the image semantic label of the unselected data with the image semantic label of the second data set to obtain the fused image semantic label; fusing the image data in the second data set with the image data in the unselected data to obtain the fused image data. The training data processing process corresponds to the aforementioned process of determining the first type of sample data set, and the pre-training data processing process corresponds to the aforementioned process of determining the second type of sample data set. The data obtained in the training data processing process is used for network training, verification and testing, and the data obtained in the pre-training data processing process is used for network pre-training, so that the model training process can achieve the best accuracy at a faster speed.

[0147] The specific implementation of step one may include the following steps:

[0148] Step 1.1: Download the first data set and corresponding semantic labels from the official website of the first data source.

[0149] For example, the first data set may be a first data source DynamicEarthNet data set X IMG and the corresponding semantic label Ytype-dy ; Semantic labels can refer to the types of remote sensing ground, such as cultivated land, forest land and grassland, wetland, bare land, water area and ice and snow area.

[0150] Step 1.2: Select T-frame temporal remote sensing images of multi-view data according to actual needs to form multi-view pseudo video sequence frames Unselected images Used for subsequent pre-training,

[0151] There are two ways to select T-frame time remote sensing images of multi-view data:

[0152] One is to select key frames for a scene of time series data based on artificial experience, statistical analysis, video key frame selection and other ideas, based on factors such as crop phenological changes, seasonal cycle changes, and cloud cover, and select data based on the principle of "cloud-free data first". For example, assuming that the main crops in a certain area are corn, potatoes, soybeans, millet and other crops, then the T-period images can be selected based on NDVI, the number of crops per year, and manual judgment. For example, for areas with the above crops, remote sensing images in spring, summer and autumn can be selected every year to effectively determine abandonment.

[0153] The other is to use software to automatically extract key frames. When selecting "key frames" of actual data, in order to reduce manual dependence, key frames can be extracted based on the key frame extraction module set in the software, such as selecting a key frame extraction method based on spectral information difference method or a key frame extraction method based on principal component analysis method to extract key frames.

[0154] Note that when selecting remote sensing frame sequences for data from different regions, the step sizes between frames may be inconsistent.

[0155] Step 1.3: Get The elevation data corresponding to the data X DSM and X DTM , so that the elevation change label Y can be calculated DSM and Y DTM ; Get Data capture time encoding data X T-fea ; Get Semantic segmentation labels of the first and last frames corresponding to different scenes of data Get Semantic segmentation labels corresponding to the data

[0156] in, Where H, W, and B are image height, width, and number of spectral channels respectively; the corresponding elevation information X can be obtained by processing the image data based on specific application software DSM, X DTM It should be noted that usually, only the elevation information of the first frame and the last frame of the same scene is needed to obtain the subsequent elevation change label Y DSM and Y DTM The time coding data can be based on the name of the image after downloading or the tag image file format (Tag Image File Format, TIFF) file annotation information to obtain the corresponding shooting time X T-fea Based on data capture time X T Perform one-hot encoding on each image. Specifically, the 12 months of a year can be evenly divided into 24 time periods. T If the data is in the current period, the 24-dimensional one-hot encoding value of the current period is set to 1, and the rest are 0, thus obtaining the time series encoding representation X T-fea ; Based on single-view data, the public dataset has the semantic segmentation label Y of each remote sensing image type-dy ; Based on a scene of time series data, the selected abandoned area has a change detection label from to maps, that is, the change result mask map Y from one category to which category change , since we only focus on the change of cultivated land, Y change In the segmentation mask of non-cultivated land change type, it is set to 0, and different values ​​are assigned as labels for different types of abandoned cultivated land change. If customized development based on other types of changes is required, the label can be re-labeled. DSM , X DTM The difference method is used to obtain the elevation change label Y of the first frame and the last frame of a scene data. DSM , Y DTM .

[0157] Step 1.4: Determine the pseudo video frame based on the definition of abandoned farmland and the comparison on the map Include the changes in abandoned farmland and generate farmland change label Y change .

[0158] Among them, since we only focus on the change of cultivated land, we can use Y change The segmentation mask for non-cultivated land change type is set to 0. The implementation process of selecting abandoned areas can be: multi-view time series data After selection, the cultivated land information is extracted according to the pixel classification results of the existing data. Through overlay analysis, it is found that the cultivated land has decreased, and it can be determined that the area may be an abandoned area. The corresponding area is found on the map of the earth's surface, and the historical image of the map is compared to confirm the actual area, and then the change of abandoned cultivated land in the area is determined. At the same time, the historical image of the current area on the map for a period of time, such as 2 years, is used to determine whether the current area is cultivated land in a changing period, and the area is selected according to the definition of abandonment.

[0159] Step 1.5: Get pre-training data and perform fusion. Download the free Sentinel 2 data from the second data source official website. esa And the corresponding segmentation mask label Y type-esa , Fusion X esa and Get X seg_pretrain , fusion Y typr-esa and Corresponding semantic segmentation labels Data obtained Y seg_ptryrain , so as to pre-train the segmentation network later.

[0160] In order to reduce the amount of model calculation, only the segmentation labels of the first and last frames of the current foreground data are used during model training. Train the segmentation network and the remaining segmentation labels The second data source data is integrated and used for pre-training of the model segmentation sub-network. The principle of image data fusion is the same as that of segmentation label fusion. By pre-training the model segmentation sub-network with preset training data, the algorithm is prompted to initially learn the texture features of the image, helping to improve the model accuracy.

[0161] Step 2: Time series data segmentation and change detection network training module.

[0162] Among them, step 2 is specifically implemented by the following steps:

[0163] Step 2.1: Data preprocessing.

[0164] The data preprocessing operations include at least adding NDVI channels, image cropping and patch segmentation, and data standardization.

[0165] The process of adding NDVI channel is: adding NDVI index display to training sample data, which can be realized by directly adding channels to the original data. In this way, sample data can be changed from (time T, width W, image height H, number of spectral channels B) to (T, W, H, C), where C = B + 1.

[0166] Based on the large-scale, high-resolution, intensive remote sensing prediction task with time series information, if the data is directly used as the model input in the original four-dimensional form of (T, W, H, C) to realize change detection, the computational cost of the model is too high and difficult to bear. Therefore, the image is cropped and patch segmented. The specific image cropping and patch segmentation process can be: crop the remote sensing image to M*M size, so that the (T, W, H, B) dimension data will be divided into W / M]*[H / M] (T, M, M, C) dimension data, where [] represents rounding down. If it cannot be divided evenly, the last cropped data can be slid to a sufficient window based on the original image and cropped. For each cropped M*M image, it is divided into overlapping S*S size patches, so that the (T, M, M, C) dimensional image will be divided into multiple (T, S, S, C) dimensional patch data. Exemplarily, the value of S can be 4. It should be noted that the value of M should not be too large or too small, because if the scale is too large, it is easy to fail to learn the detailed information of the image, and if it is too small, it is easy to produce a salt and pepper effect during the prediction splicing process. For example, the value of M can be 896.

[0167] Data standardization: There are large differences in data between channels of remote sensing images, so the data of each channel needs to be standardized one by one.

[0168] Step 2.2: Data encoding processing.

[0169] Among them, a system architecture diagram of the data encoding module can be referred to Figure 5 As shown, Figure 5 A schematic diagram of the improved patch embedding system architecture can be found in Figure 6 As shown. Among them, the data encoding module is divided into two sub-networks for learning according to different learning tasks, namely the improved Swin transformer network and the Segformer encoder network. The improved Swin transformer network is mainly used to learn abandonment recognition and elevation information features, and the Segformer encoder network is mainly used to learn ground object segmentation features. The improved part of the improved Swin transformer network is mainly to improve the patch embedding part shared by the network by integrating time series information, and the improved Segformer encoder network is mainly based on the abandonment recognition and elevation information feature extraction sub-network to improve the attention calculation and window attention calculation. Among them:

[0170] Regarding the improvement of the patch embedding part, on the one hand, in order to integrate the temporal information into the image and reduce the number of parameters in image calculation, after cropping and partitioning the remote sensing image, the partition stage uses 3D (Division, D) patch learning, and the patch size is (1,4,4), which are the dimensions of time, height and width respectively. Here, the patch size is set to 1 in the time dimension, mainly because the temporal features of the image need to be learned separately. It can be obtained 3D tokens, learn embedding mapping for each token to a token embedding of dimension D. Represent X through T frame temporal coding T-fea The obtained T×24-dimensional vector is represented by a T×D-dimensional vector based on time embedding learning. On the other hand, since the time of each image is the same value, the patch time value of each image is set to the same embedding, and the time dimension can be obtained by increasing the dimension. Vector. For example, the temporal stride is set to 2, and the four-dimensional vector of the time embedding is patched with the image to obtain The vectors are summed and superimposed, and non-overlapping summation is performed according to the stride channel of the time frame to obtain Feature vectors are used to reduce the amount of time series parameters and are used for subsequent wasteland identification and elevation information feature learning.

[0171] On the other hand, extracting T 1 , T n The time embeddings of the two time periods are additively fused with the image token embeddings of the corresponding two time periods for subsequent feature learning of object segmentation.

[0172] When improving the Attention calculation and window attention calculation based on the abandonment recognition and elevation information feature extraction sub-network, on the one hand, the window calculation range is improved, see Figure 7 As shown in Figure 1, based on the large-scale characteristics of remote sensing images, transformation parameters are introduced to learn 3D windows of variable size and position, including the spatial offset relative to the reference window, the spatial scale scaling factor, and the time offset. In this way, the time offset can be defined. Used to learn the offset value of the window in the time dimension; define the spatial offset factor U s and the scale factor O s Used to learn the offset and scaling of the window in the spatial dimension. During the network training process, the three parameters are calculated using the following formula:

[0173] θs =R·f(Linear(LeakyReLU(GAP(X s )))),

[0174] U s ,O s =Linear(LeakyReLU(GAP(X s ))),

[0175] Among them, X s is the feature of the current window input, f(x) is the Gumbel-max function, and the one-hot representation is obtained. R is the time step vector, R = (0, 1, ..., T). It can be seen that in the process of learning the time offset parameter, since the value is a discrete bias value, using the argmax function to directly obtain the corresponding discrete value is likely to cause the model to be non-differentiable in back propagation. For this reason, the reparameterization idea gumbel-max technique is introduced to obtain discrete time bias values ​​and ensure that the network can be trained in back propagation.

[0176] After learning the above parameters, the model transforms the randomly initialized window. The window size and position calculations before and after the transformation are as follows:

[0177]

[0178]

[0179] Among them, x l ,y l ,x r ,y r To initialize the coordinates of the upper left and lower right corners of the window, T l ,T r is the coordinate corresponding to the l and r time periods; c ,y c ,T c The center point and center period of the current window; is the relative position offset from the center, x l ′ ,y l ′ ,T l ′ ,x ′ r ,y ′ r ,T r ′ is the transformed window coordinate, The elements in are the predicted bias parameters and scale factors. The elements in are the predicted time offset parameters. The converted window range is obtained, and the corresponding Q, K, and V vectors are obtained based on the window for subsequent attention calculations.

[0180] On the other hand, it shows that adding temporal feature learning improves the attention calculation method, see Figure 8 As shown in Figure 2. In the attention calculation stage, we can learn time and space information separately instead of using the initialization vector directly. Specifically, for the spatial relative position parameter The weight vector is learned by initializing the parameters, where A is the time series size in a 3D window, and F is the window height and width; based on the time position parameters Based on the T×D time dimension vector obtained in the input conversion stage, 2×1×1conv is used to obtain Size vector, the vector can be shared by all tokens in the current time sequence. In the overall module learning process, the following formula is used for weighted calculation when calculating the attention of each window of the improved W-MSA and improved VSA submodules.

[0181]

[0182] Among them, Q, K, V are query, key, and value matrices; d is the dimension of query and key; α and β are spatiotemporal weights.

[0183] The feature vectors obtained based on the improved W-MSA and improved VSA modules are input into the subsequent network to implement a system architecture diagram for model encoder encoding feature learning. Fig. 9 shown.

[0184] Step 2.3: Data decoding processing.

[0185] Among them, the data decoding (Decoder) part mainly proposes a lightweight 3D decoding method with multi-task collaboration. The network mainly decodes the features learned by the Encoder, which includes two decoding tasks: one is the ground object classification results of the first frame and the last frame in the pseudo-time series frame, and the other is the abandonment detection results and elevation detection results based on time series data. In the process of decoding the change detection task, the network merges the feature vectors learned from the two-frame segmentation task to strengthen the influence of the two-frame time series information on the abandonment detection results, achieve the effect of multi-task collaboration, and help the network better learn the change results.

[0186] A system architecture diagram of the data decoding (Decoder) module can be referred to Fig.10 shown.

[0187] Step 2.4: Output the results.

[0188] in, Fig.10 After the architecture shown outputs the ground object detection results, the abandonment detection results, and the elevation detection results, the network loss value can be calculated using the following calculation formula. The network loss function is recorded as: Loss = αLoss change +γLoss seg +δLoss h , where α, γ, and δ are the network loss weights corresponding to each part of the network. change For the abandoned change detection loss, the corresponding network loss function can be written as: In the formula, y c is the real abandonment change label, y c-p To predict the abandonment change label, N is the amount of multi-view pseudo video data. seg is the object classification loss, and the corresponding network loss function can be recorded as: In the formula, y seg For real objects

[0189] Classification label, y seg-p To predict the classification label of the ground object, M is the amount of image data of the first and last frames of the multi-view pseudo video, M = 2N. Loss h is the elevation change loss, and the corresponding network loss function can be recorded as: In the formula, y dsm is the actual elevation dsm change result, y dtm is the actual elevation dtm change result, To predict the elevation detection results, N is the amount of multi-view pseudo video data and μ is the loss weight.

[0190] According to the calculated loss value Loss, the parameters in the aforementioned model are adjusted until a trained abandonment recognition network model is obtained.

[0191] Based on the above embodiment, after the wasteland identification network model is trained, the application of the wasteland identification network model can refer to Fig.11 As shown, the following steps are included:

[0192] Step 3.1: Preprocess the image to obtain the pseudo time-series frame to be identified.

[0193] Among them, a set of images in the actual scene that needs to be predicted and recognized is obtained, and after the collected images are pre-processed by atmospheric correction, coordinate conversion, etc., key frames are extracted based on software to form pseudo-time series frames to be recognized.

[0194] Step 3.2: Obtain the temporal one-hot encoding based on the file name or TIFF annotation information of the pseudo time series frame to be identified.

[0195] Step 3.3: Input the time-series frames of the pseudo video to be identified and the corresponding time one-hot encoding into the abandonment recognition network model to obtain multiple sub-image segmentation results, abandonment detection results, elevation detection results, and ground object classification results of the pseudo video to be identified.

[0196] Step 3.4: Use the Gaussian convolution method to stitch multiple sub-image segmentation results to obtain the overall detection result.

[0197] Among them, the overall detection results at least include abandonment detection results, ground and surface change detection results, namely elevation detection results, and object classification results of the first and last frames of multi-view data.

[0198] Based on the model output results, due to the image cropping and patching operations, the returned results need to be spliced. However, the direct splicing results are prone to produce salt-and-pepper effects at the edges. Therefore, the idea of ​​Gaussian filtering can be used for splicing. Specifically, Gaussian weighting can be used for splicing, and the best model can be found based on the evaluation of the full image detection results and the true value after splicing. The splicing objects include the stitching of cropped images, the classification results of the objects corresponding to each cropped sub-image, the splicing of the abandonment detection results and the elevation detection results.

[0199] For the ground object classification results and abandonment detection results, weighted convolution is used for splicing, which can greatly alleviate the salt and pepper effect.

[0200] For the elevation detection results, normalized weighted convolution splicing is used. Since the change result obtained after model decoding is a continuous value, no softmax operation is required, so the merged result needs to be weighted and summed by weight normalization. Specifically, a 2D Gaussian kernel is used for splicing. When superimposing the continuous values ​​of the sub-images, the multiple weights of the Gaussian kernel obtained need to be normalized according to the normalized ratio for the superimposed location area to obtain the weight value of 1, and then the current elevation detection result value is weighted and summed to obtain the final change result, thereby obtaining a high-precision elevation detection result of the full map splicing.

[0201] Step 3.5: Output the overall detection results.

[0202] In this way, remote sensing image abandonment and abandonment type detection can be realized in real scenes, and image object classification and automatic detection of ground surface changes can be realized without inputting elevation information, helping users to make better decisions in practical applications. It can also support viewing the remote sensing image object classification results before and after the change. Based on the multi-task learning mechanism, the remote sensing image object classification results can be compared and analyzed based on the change results.

[0203] In this way, through the idea of ​​video modeling, a spatiotemporal context model of pseudo-time series frame remote sensing images is constructed, and the four-dimensional information of time, space and spectrum is connected to realize end-to-end change detection of long-time continuous spectral information. In this way, based on the changes in continuous time series data, the surface topography, seasonal climate, long-term changes in crops and other conditions can be effectively learned, and whether the cultivated land is fallow, abandoned or lost can be effectively distinguished; and based on deep learning technologies such as change detection and image semantic segmentation, a multi-task high-precision abandonment detection algorithm is used based on ground object classification, refined identification of abandonment types, and elevation change detection. The three subtasks can help to achieve refined detection of abandoned land types, which can be widely used in agricultural scenarios. There is no need to split the images into pairs for detection and post-processing, which greatly reduces the amount of calculation and improves the accuracy of the model. Furthermore, based on multimodal fusion technology, the time series frame step length of remote sensing image data, ground surface elevation and other information are displayed or implicitly fused into the model to help the model improve its accuracy. Finally, based on the Gaussian filtering idea, an image splicing technology suitable for classification and fitting tasks is proposed to help the model achieve high-precision and high-reliability reasoning methods and alleviate the salt and pepper effect generated during the splicing process. In addition, during model training, a data generation and annotation method suitable for continuous time-spectrum remote sensing images is proposed to help the algorithm realize high-precision prediction and model prediction, making it robust and accurate, and realizing refined detection of abandoned land types.

[0204] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.

[0205] The embodiment of the present application provides an information identification method, which obtains a set of ground analysis remote sensing images corresponding to a region to be identified within a preset time period through an information identification device, sorts at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous set of images to be processed, and determines the image acquisition time of each key image in the image set to be processed to obtain at least two target acquisition times, and finally uses a trained abandonment identification network model to perform abandonment identification on the image set to be processed and at least two target acquisition times to obtain a target identification result of the region to be identified. In this way, key frames are extracted from remote sensing images collected within a period of time, sorted in chronological order to obtain a continuous set of images to be processed, and abandonment identification is performed on the image set to be processed and the corresponding target acquisition time through a trained abandonment identification network model to obtain a target identification result of the region to be identified, which solves the problem of low accuracy of abandoned land detection results of remote sensing image change detection, and proposes a method for realizing abandoned land detection using remote sensing images with continuous time series, which fully considers the long-term change trend of ground objects and improves the accuracy of abandoned land detection results.

[0206] Based on the above embodiments, the embodiments of the present application provide an information recognition device, which can be applied to Figures 1-2 In the information identification method provided in the corresponding embodiment, refer to Fig.12 As shown, the information identification device 4 may include: a processor 41, a memory 42 and a communication bus 43, wherein:

[0207] A memory 42, for storing executable instructions;

[0208] A communication bus 43, used to realize the communication connection between the processor 41 and the memory 42;

[0209] The processor 41 is used to execute the information recognition program stored in the memory 42 to implement the following steps:

[0210] Acquire a set of remote sensing images to be analyzed corresponding to the area to be identified within a preset time period;

[0211] At least two key images selected from the remote sensing image set to be analyzed are sorted in chronological order to obtain a continuous image set to be processed;

[0212] Determine the image acquisition time of each key image in the image set to be processed, and obtain at least two target acquisition times;

[0213] Through the trained abandonment recognition network model, abandonment recognition is performed on the image set to be processed and at least two target acquisition times to obtain the target recognition results of the area to be identified.

[0214] In other embodiments of the present application, the processor executes the steps of performing abandonment identification on the image set to be processed and at least two target acquisition times through the trained abandonment identification network model, and before obtaining the target identification result of the area to be identified, it is also used to perform the following steps:

[0215] Determining a first type of sample data set;

[0216] Determine a second type of sample data set; wherein data in the first type of sample data set and the second type of sample data set are different;

[0217] Determine the network model to be trained;

[0218] The first type of sample data set and the second type of sample data set are used to perform model training on the network model to be trained, so as to obtain a trained abandonment recognition network model.

[0219] In other embodiments of the present application, when the processor executes the step of determining the first type of sample data set, it can be implemented by the following steps:

[0220] Acquire a first remote sensing image set collected for a photographed object within a first time period from a first data source; wherein each element in the first remote sensing image set includes at least: image elements of at least two frames of images for the same photographed object, a regional feature label corresponding to each frame of images in each image element regarding a regional feature, an image collection time of each frame of images in each image element, and elevation data information in each frame of images;

[0221] Based on the first remote sensing image set, determining a first sub-image set and a second sub-image set; wherein each image element in the first sub-image set includes: at least two frames of key images corresponding to the same photographed object within a preset time length, and the second sub-image set includes the collective elements in the first remote sensing image set except the first sub-image set;

[0222] Based on the first sub-image set, a first type of sample data set is determined.

[0223] In other embodiments of the present application, when the processor executes the step of determining the first type of sample data set based on the first sub-image set, it can be implemented by the following steps:

[0224] Using preset identification information to identify the non-cultivated land area in each image element in the first sub-image set to obtain a third sub-image set;

[0225] Determine the cultivated land change in each image element in the third sub-image set, and obtain the cultivated land change label of each image element;

[0226] Based on the elevation data information of each frame of the image, determining the elevation change label of each image element in the third sub-image set;

[0227] determining a normalized vegetation index for each image element in the third sub-image set;

[0228] The third sub-image set, the normalized vegetation index corresponding to each image element in the third sub-image set, the elevation change label corresponding to each image element in the third sub-image set, the image acquisition time of each image element in the third sub-image set, and the cultivated land change label corresponding to each image element in the third sub-image set are used as the fourth sub-image set;

[0229] According to the preset size, image segmentation processing is performed on each frame of image corresponding to each image element in the fourth sub-image set to obtain a first type of sample data set.

[0230] In other embodiments of the present application, when the processor executes the step of determining the second type of sample data set, it can be implemented by the following steps:

[0231] Obtaining, from a second data source, a second remote sensing image set collected for the photographed object within a second time period; wherein each element in the second remote sensing image set includes at least: an image, an image collection time of each image, and a ground object segmentation label;

[0232] Determine the ground object segmentation label corresponding to each image element in the second sub-image set;

[0233] fusing the second remote sensing image set, the second sub-image set, and the elements in the ground object segmentation labels corresponding to the second sub-image set to obtain a fifth sub-image set;

[0234] According to a preset size, image segmentation processing is performed on the image elements in the fifth sub-image set to obtain a second type of sample data set.

[0235] In other embodiments of the present application, when the processor executes the step of using the first type of sample data set and the second type of sample data set to perform model training on the network model to be trained and obtain the trained abandonment recognition network model, it can be implemented by the following steps:

[0236] Acquire first training sample data from a first type of sample data set;

[0237] Acquire second training sample data from a second type of sample data set;

[0238] Inputting the first training sample data and the second training sample data into the network model to be trained, and obtaining a first abandoned land recognition result, a first elevation recognition result, and a first land feature segmentation recognition result;

[0239] Determining a target loss value based on the first abandonment identification result, the first elevation identification result, and the first ground feature segmentation identification result;

[0240] If the target loss value is less than or equal to the preset threshold, the abandonment identification network model is determined as the network model to be trained;

[0241] If the target loss value is greater than a preset threshold, the parameters in the network model to be trained are updated based on the target loss value to obtain a reference network model;

[0242] After updating the model to be trained to be the reference network model, repeat the step of “obtaining the first training sample data from the first type of sample data set” until the abandonment identification network model is obtained.

[0243] In other embodiments of the present application, when the processor executes the step of inputting the first training sample data and the second training sample data into the network model to be trained, and obtains the first abandoned land recognition result, the first elevation recognition result and the first land feature segmentation recognition result, it can be achieved by the following steps:

[0244] The first training sample data and the second training sample data are respectively processed by integrating time series information through a patch embedding layer in an encoding module of the network model to be trained to obtain corresponding first fused sample data and second fused sample data;

[0245] Extracting the abandonment feature information and the elevation feature information in the first fused sample data through a first network including a multi-head self-attention mechanism and a rotational self-attention mechanism in the encoding module for performing attention calculation; wherein the multi-head self-attention mechanism at least includes a mechanism for learning three-dimensional window information of variable size and position, a spatial offset relative to a reference window, a spatial scale scaling factor, and a temporal offset, and the rotational self-attention mechanism is used to implement a temporal feature learning mechanism for learning temporal information and spatial information separately;

[0246] Extracting ground object segmentation feature information in the second fused sample data through the second network in the encoding module;

[0247] The abandonment feature information, the elevation feature information and the land object segmentation feature information are decoded by a decoding module to obtain a first abandonment recognition result, a first elevation recognition result and a first land object segmentation recognition result.

[0248] In other embodiments of the present application, when the processor executes the step of determining the target loss value based on the first abandonment recognition result, the first elevation recognition result, and the first ground feature segmentation recognition result, it can be implemented by the following steps:

[0249] Determine an abandonment loss value based on the first abandonment identification result and the abandonment identification label in the first training sample data;

[0250] Determining an elevation loss value based on the first elevation recognition result and the elevation change label in the first training sample data;

[0251] Determining a ground object loss value based on the first ground object segmentation recognition result and the ground object segmentation label in the second training sample data;

[0252] The target loss value is obtained by performing weighted summation on the abandonment loss value, elevation loss value and ground feature loss value.

[0253] In other embodiments of the present application, when the processor executes the step of performing abandonment recognition on the image set to be processed and at least two target acquisition times through the trained abandonment recognition network model to obtain the target recognition result of the area to be recognized, it can be achieved by the following steps:

[0254] Performing image segmentation processing on the images in the to-be-processed image set according to a preset size to obtain an image set to be input; wherein the image set to be input includes a preset number of sub-image groups;

[0255] The abandonment recognition network model is used to recognize and process the input image set and at least two target acquisition times, and the ground object prediction results, elevation prediction results and abandonment prediction results corresponding to a preset number of sub-images are obtained;

[0256] The ground feature prediction results, elevation prediction results and land abandonment prediction results corresponding to a preset number of sub-images are spliced ​​and fused to obtain the target recognition result.

[0257] In other embodiments of the present application, when the processor executes the step of splicing the ground feature prediction results, elevation prediction results, and abandonment prediction results corresponding to a preset number of sub-images to obtain a target recognition result, it can be achieved by the following steps:

[0258] The ground object prediction results and the wasteland prediction results corresponding to the preset number of sub-images are spliced ​​respectively by using the weighted convolution method to obtain the ground object detection results and the wasteland detection results of the area to be identified; wherein the target recognition result includes the ground object detection results and the wasteland detection results of the area to be identified;

[0259] The normalized weighted convolution method is used to splice the elevation prediction results corresponding to a preset number of sub-images to obtain the elevation detection result; wherein, the target recognition result includes the elevation detection result.

[0260] It should be noted that the specific implementation process of the steps performed by the information identification device in this embodiment can refer to Figures 1-2 The implementation process of the information identification method provided in the corresponding embodiment will not be repeated here.

[0261] The embodiment of the present application provides an information identification device, which obtains a set of ground analysis remote sensing images corresponding to a region to be identified within a preset time period through the information identification device, sorts at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous set of images to be processed, and determines the image acquisition time of each key image in the image set to be processed to obtain at least two target acquisition times, and finally uses a trained abandonment identification network model to perform abandonment identification on the image set to be processed and at least two target acquisition times to obtain a target identification result of the region to be identified. In this way, key frames are extracted from remote sensing images collected within a period of time, sorted in chronological order to obtain a continuous set of images to be processed, and abandonment identification is performed on the image set to be processed and the corresponding target acquisition time through a trained abandonment identification network model to obtain a target identification result of the region to be identified, which solves the problem of low accuracy of abandoned land detection results of remote sensing image change detection, and proposes a method for realizing abandoned land detection using remote sensing images with continuous time series, which fully considers the long-term change trend of ground objects and improves the accuracy of abandoned land detection results.

[0262] Based on the foregoing embodiments, the embodiments of the present application provide a computer-readable storage medium, referred to as a storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement reference Figures 1-2 The implementation process of the information identification method provided in the corresponding embodiment will not be repeated here.

[0263] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0264] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0265] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0266] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0267] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.

Claims

1. An information identification method, It is characterized in that The method comprises: Acquire a set of remote sensing images to be analyzed corresponding to the area to be identified within a preset time period; Sorting at least two key images selected from the remote sensing image set to be analyzed in chronological order to obtain a continuous image set to be processed; Determine the image acquisition time of each key image in the image set to be processed to obtain at least two target acquisition times; Through the trained abandonment identification network model, abandonment identification is performed on the image set to be processed and at least two target acquisition times to obtain the target identification result of the area to be identified.

2. The method according to claim 1, It is characterized in that Before performing abandonment identification on the image set to be processed and at least two target acquisition times by using the trained abandonment identification network model to obtain the target identification result of the area to be identified, the method further includes: Determining a first type of sample data set; Determine a second type of sample data set; wherein the data in the first type of sample data set and the data in the second type of sample data set are different; Determine the network model to be trained; The first type of sample data set and the second type of sample data set are used to perform model training on the network model to be trained, so as to obtain a trained abandonment recognition network model.

3. The method according to claim 2, It is characterized in that The determining of the first type of sample data set comprises: Acquire a first remote sensing image set collected for a photographed object within a first time period from a first data source; wherein each element in the first remote sensing image set at least includes: image elements of at least two frames of images for the same photographed object, a regional feature label corresponding to each frame of images in each image element regarding a regional feature, an image collection time of each frame of images in each image element, and elevation data information in each frame of images; Based on the first remote sensing image set, determining a first sub-image set and a second sub-image set; wherein each image element in the first sub-image set includes: at least two frames of key images corresponding to the same photographed object within a preset time length, and the second sub-image set includes the collective elements in the first remote sensing image set except the first sub-image set; Based on the first sub-image set, the first type sample data set is determined.

4. The method according to claim 3, It is characterized in that The determining, based on the first sub-image set, the first type of sample data set comprises: Using preset identification information to identify the non-cultivated land area in each image element in the first sub-image set to obtain a third sub-image set; Determine the cultivated land change in each image element in the third sub-image set to obtain a cultivated land change label for each image element; Determining an elevation change label of each image element in the third sub-image set based on the elevation data information of each image element in the third sub-image set; Determining a normalized vegetation index for each image element in the third sub-image set; The third sub-image set, the normalized vegetation index corresponding to each image element in the third sub-image set, the elevation change label corresponding to each image element in the third sub-image set, the image acquisition time of each image element in the third sub-image set, and the cultivated land change label corresponding to each image element in the third sub-image set are used as the fourth sub-image set; According to a preset size, image segmentation processing is performed on each frame of image corresponding to each image element in the fourth sub-image set to obtain the first type of sample data set.

5. The method according to claim 3, It is characterized in that The determining of the second type of sample data set comprises: Obtaining, from a second data source, a second remote sensing image set collected for the photographed object within a second time period; wherein each element in the second remote sensing image set includes at least: an image, an image collection time of each image, and a ground object segmentation label; Determine a ground object segmentation label corresponding to each image element in the second sub-image set; fusing the second remote sensing image set, the second sub-image set, and the elements in the ground object segmentation label corresponding to the second sub-image set to obtain a fifth sub-image set; Image segmentation processing is performed on the image elements in the fifth sub-image set according to a preset size to obtain the second type sample data set.

6. The method according to any one of claims 2 to 5, It is characterized in that The method of using the first type sample data set and the second type sample data set to perform model training on the network model to be trained to obtain a trained abandonment recognition network model includes: Acquire first training sample data from the first type of sample data set; Acquire second training sample data from the second type sample data set; Inputting the first training sample data and the second training sample data into the network model to be trained to obtain a first abandoned land recognition result, a first elevation recognition result and a first land feature segmentation recognition result; Determining a target loss value based on the first abandonment identification result, the first elevation identification result, and the first ground feature segmentation identification result; If the target loss value is less than or equal to a preset threshold, determining that the abandonment identification network model is the network model to be trained; If the target loss value is greater than the preset threshold, updating the parameters in the network model to be trained based on the target loss value to obtain a reference network model; After updating the model to be trained to the reference network model, the step of "obtaining first training sample data from the first type of sample data set" is repeated until the abandonment identification network model is obtained.

7. The method according to claim 6, It is characterized in that The step of inputting the first training sample data and the second training sample data into the network model to be trained to obtain a first abandoned land recognition result, a first elevation recognition result and a first land feature segmentation recognition result comprises: Performing time series information integration processing on the first training sample data and the second training sample data respectively through the patch embedding layer in the encoding module of the network model to be trained, so as to obtain corresponding first fused sample data and second fused sample data; Extracting the abandonment feature information and elevation feature information in the first fused sample data through a first network including a multi-head self-attention mechanism and a rotational self-attention mechanism in the encoding module for performing attention calculation; wherein the multi-head self-attention mechanism at least includes a mechanism for learning three-dimensional window information of variable size and position, a spatial offset relative to a reference window, a spatial scale scaling factor, and a temporal offset, and the rotational self-attention mechanism is used to implement a temporal feature learning mechanism for learning temporal information and spatial information separately; Extracting ground object segmentation feature information in the second fused sample data through a second network in the encoding module; The abandonment feature information, the elevation feature information and the land object segmentation feature information are decoded by a decoding module to obtain the first abandonment identification result, the first elevation identification result and the first land object segmentation identification result.

8. The method according to claim 6, It is characterized in that The determining of a target loss value based on the first abandonment identification result, the first elevation identification result and the first ground feature segmentation identification result comprises: Determining an abandonment loss value based on the first abandonment identification result and the abandonment identification label in the first training sample data; Determining an elevation loss value based on the first elevation recognition result and the elevation change label in the first training sample data; Determining a ground object loss value based on the first ground object segmentation and recognition result and the ground object segmentation label in the second training sample data; The abandonment loss value, the elevation loss value and the ground feature loss value are weightedly summed to obtain the target loss value.

9. The method according to claim 1, It is characterized in that The trained abandonment identification network model is used to perform abandonment identification on the image set to be processed and at least two target acquisition times to obtain target identification results of the area to be identified, including: Performing image segmentation processing on the images in the image set to be processed according to a preset size to obtain an image set to be input; wherein the image set to be input includes a preset number of sub-image groups; The abandonment identification network model is used to identify the input image set and at least two target acquisition times to obtain ground object prediction results, elevation prediction results and abandonment prediction results corresponding to the preset number of sub-images; The ground feature prediction results, elevation prediction results and abandonment prediction results corresponding to the preset number of sub-images are spliced ​​and fused to obtain the target recognition result.

10. The method according to claim 9, It is characterized in that The step of splicing the ground object prediction results, the elevation prediction results and the wasteland prediction results corresponding to the preset number of sub-images to obtain the target recognition result includes: The ground object prediction results and the abandonment prediction results corresponding to the preset number of sub-images are spliced ​​respectively by weighted convolution to obtain the ground object detection results and the abandonment detection results of the area to be identified; wherein the target recognition result includes the ground object detection results and the abandonment detection results of the area to be identified; The elevation prediction results corresponding to the preset number of sub-images are spliced ​​using a normalized weighted convolution method to obtain an elevation detection result; wherein the target recognition result includes the elevation detection result.

11. An information recognition device, It is characterized in that The device comprises at least: a memory, a processor and a communication bus; wherein: The memory is used to store executable instructions; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the information identification program stored in the memory to implement the steps of the information identification method according to any one of claims 1 to 10.

12. A storage medium, It is characterized in that The storage medium stores an information identification program, and when the information identification program is executed by the processor, the steps of the information identification method according to any one of claims 1 to 10 are implemented.