Deep network model self-supervised training and land use change detection method and device

By constructing positive and negative sample pairs from multi-angle remote sensing images and training a deep network model, the problems of high manual annotation costs and neglect of remote sensing observation mechanisms in existing technologies are solved, achieving efficient land use change detection and improving automation and applicability.

CN118968223BActive Publication Date: 2026-07-21深圳市规划和自然资源数据管理中心(深圳市空间地理信息中心) +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
深圳市规划和自然资源数据管理中心(深圳市空间地理信息中心)
Filing Date
2024-08-12
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing deep learning-based land use monitoring methods rely on a large number of manually labeled samples, which is costly. Furthermore, existing self-supervised learning algorithms neglect the mechanisms of multi-angle, multi-spectral, and multi-scale remote sensing observations, resulting in models that cannot fully describe land surface attributes.

Method used

Positive and negative sample pairs are constructed using multi-angle remote sensing images. A deep network model is trained using a multi-angle contrast loss function. Self-supervised learning is performed using unlabeled data, incorporating the characteristics of multi-angle remote sensing observations.

Benefits of technology

It alleviates the reliance on manually labeled samples, improves the efficiency of land use change detection, promotes the personalized use of deep learning in the field of remote sensing, and has a high degree of automation and wide applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968223B_ABST
    Figure CN118968223B_ABST
Patent Text Reader

Abstract

The application discloses a deep network model self-supervised training method and device and a land use change detection method and device, and belongs to the field of remote sensing image processing. The training method comprises the following steps: acquiring a first time phase training image and a second time phase training image; inputting the first time phase training image and the second time phase training image into a deep network model to acquire a first time phase image feature and a second time phase image feature; constructing a positive sample pair from image features of the same scene and different angles in the first time phase image feature and the second time phase image feature, and constructing a negative sample pair from image features of different scenes; constructing a multi-angle contrast loss function; and performing self-supervised training on the deep network model according to the positive sample pair, the negative sample pair and the multi-angle contrast loss function. The application constructs a self-supervised learning paradigm from multi-view images, takes land use change detection as a downstream task, and realizes land use monitoring without relying on or only relying on a small amount of manually labeled samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image processing technology, and relates to a method and device for self-supervised training of deep network models and detection of land use change. Background Technology

[0002] Monitoring urban land use provides data support for urban planning, construction, and management, and is a crucial guarantee for achieving sustainable urban development. Remote sensing technology, with its advantages of wide coverage, fine granularity, and high revisit frequency, has been widely applied in land use monitoring. Deep learning is an important means of extracting land use information from remote sensing images; however, current deep learning-based land use monitoring relies on a large number of manually labeled samples, which is costly and difficult to adapt to the speed of remote sensing image acquisition and the speed of land use change. Self-supervised learning is a new machine learning paradigm that has gradually developed in recent years. It mainly utilizes the inherent characteristics and structure of data to learn general feature representations from large amounts of unlabeled data for use in downstream tasks, thereby reducing reliance on manual annotation. Self-supervised learning has made initial progress in the field of remote sensing; however, existing algorithms typically only consider the visual features of images, neglecting the multi-angle, multi-spectral, and multi-scale remote sensing observation mechanisms, making the models unable to fully describe surface attributes. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and apparatus for self-supervised training of deep network models and detection of land use change, thereby improving the efficiency of land use change detection.

[0004] To achieve the above objectives, the present invention is implemented using the following technical solution:

[0005] In a first aspect, the present invention provides a self-supervised training method for deep network models, comprising:

[0006] Acquire a first-phase training image and a second-phase training image, wherein the first-phase training image and the second-phase training image include multi-angle images of the same scene;

[0007] The first and second phase training images are input into a deep network model to obtain the features of the first and second phase images.

[0008] Positive sample pairs are constructed from image features of the same scene but different angles in the first and second phase image features, and negative sample pairs are constructed from image features of different scenes in the first and second phase image features.

[0009] Construct a multi-angle contrast loss function;

[0010] The deep network model is trained under self-supervised conditions based on the positive sample pairs, negative sample pairs, and the multi-angle contrast loss function.

[0011] Furthermore, the first and second phase training images are remote sensing images.

[0012] Furthermore, each scene in the first and second phase training images includes images from three perspectives: a front view, a frontal view, and a rear view.

[0013] Furthermore, based on positive sample pairs, negative sample pairs, and a multi-angle contrast loss function, the deep network model undergoes self-supervised training, including:

[0014] The first phase training images are grouped by scene;

[0015] Taking each angle image of each scene as the main viewpoint, the contrast loss is calculated by using the multi-angle contrast loss function to obtain multiple angle loss values.

[0016] The loss values ​​from multiple angles are summed to obtain the multi-angle contrast loss value;

[0017] The deep network model is trained under self-supervised conditions by comparing loss values ​​from multiple perspectives.

[0018] Furthermore, using the first phase training image in the first phase... The formula for calculating the contrast loss of a scene's front view from the main perspective is as follows:

[0019] ,

[0020] in, This represents the front view. This represents the front view. Represents the rear view; Indicates the first One scenario; In the first phase training image, the first... A front view of the scene; In the second phase training image, the first... The scene is Images from the observation angle; This indicates the number of scenes in the second-phase training images; To distinguish the functions, the calculation formula is as follows:

[0021] ,

[0022] in, To make the image Image features obtained from the input deep network model; The L2 norm of the eigenvectors; These are the training parameters for the deep network model; These are the hyperparameters of the deep network model;

[0023] The first phase training image The formula for calculating the multi-angle contrast loss value for a given scenario is:

[0024] .

[0025] Secondly, the present invention also provides a self-supervised training device for deep network models, the device comprising:

[0026] The training image acquisition module is used to acquire a first phase training image and a second phase training image, wherein the first phase training image and the second phase training image include multi-angle images of the same scene;

[0027] The image feature acquisition module is used to input the first phase training image and the second phase training image into the deep network model to acquire the first phase image features and the second phase image features.

[0028] The sample pair construction module is used to construct positive sample pairs from image features of the same scene but different angles in the first and second time phase image features, and to construct negative sample pairs from image features of different scenes in the first and second time phase image features.

[0029] The loss function construction module is used to construct multi-angle contrast loss functions;

[0030] The deep network model self-supervised training module is used to perform self-supervised training on the deep network model based on the positive sample pairs, negative sample pairs, and multi-angle contrast loss function.

[0031] Thirdly, the present invention also provides a method for detecting land use change, comprising:

[0032] Acquire remote sensing images of the same scene at different times to be detected;

[0033] The remote sensing image to be detected is input into a pre-established land use change detection model to obtain the land use change in the remote sensing image;

[0034] The pre-training method for the land use change detection model includes the aforementioned self-supervised training method for deep network models.

[0035] Fourthly, the present invention also provides a land use change detection device, comprising:

[0036] The remote sensing image acquisition module is used to acquire remote sensing images of the same scene at different times to be detected;

[0037] The land use change detection module is used to input the remote sensing image to be detected into a pre-established land use change detection model to obtain the land use change in the remote sensing image.

[0038] Fifthly, the present invention also provides a computer device, comprising:

[0039] Memory, used to store computer programs;

[0040] A processor is used to execute the computer program to implement the steps of the above-described deep network model self-supervised training method or the steps of the above-described land use change detection method.

[0041] In a sixth aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the steps of the above-described deep network model self-supervised training method or the steps of the above-described land use change detection method.

[0042] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0043] This invention constructs multi-angle positive and negative sample pairs from multi-view images across two time phases, inputs them into a deep network for feature extraction, and builds a multi-angle self-supervised contrastive loss function to train the deep network. After training, the network is applied to downstream tasks of land use change detection, realizing self-supervised learning and feature representation of the deep network. It fully utilizes the abundant unlabeled data available, effectively alleviating the over-reliance on manually labeled samples in current deep learning-based land use change detection methods. It incorporates the characteristics of multi-angle remote sensing observation, mining the unique content and structural characteristics of remote sensing data, and promoting the personalized use of deep learning in the field of remote sensing. The model requires fewer manually adjustable parameters, which are relatively easy to set, and has a high degree of automation, making it widely applicable to large-scale urban land use monitoring. Attached Figure Description

[0044] Figure 1 A flowchart illustrating a self-supervised training method for a deep network model provided in an embodiment of the present invention;

[0045] Figure 2 This is a schematic diagram of the framework for self-supervised training of a deep network model in an embodiment of the present invention;

[0046] Figure 3 This is a schematic diagram of the structure of a self-supervised training device for a deep network model provided in an embodiment of the present invention;

[0047] Figure 4A schematic flowchart of a land use change detection method provided in an embodiment of the present invention;

[0048] Figure 5 This is a schematic diagram of the framework of the land use change detection model in an embodiment of the present invention;

[0049] Figure 6 This is a schematic diagram of the structure of a land use change detection device provided in an embodiment of the present invention;

[0050] Figure 7 An internal structural diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0051] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The same reference numerals in the drawings indicate the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. The embodiments and specific features within the embodiments of this application are detailed descriptions of the technical solution of this application, and not limitations thereof. Where there is no conflict, the embodiments and technical features within the embodiments of this application can be combined with each other.

[0052] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0053] Example 1:

[0054] like Figure 1 and Figure 2 As shown, this embodiment of the invention provides a self-supervised training method for deep network models. Figure 1 This is a flowchart illustrating the self-supervised training method for the deep network model. This flowchart only shows the logical order of the method described in this embodiment. In other possible embodiments of the invention, different methods may be used, provided they do not conflict with each other. Figure 1 Complete the steps shown or described in the order indicated.

[0055] The self-supervised training method for deep network models provided in this embodiment can be applied to a terminal and can be executed by a self-supervised training device for deep network models. This device can be implemented in software and / or hardware and can be integrated into the terminal.

[0056] See Figure 1 The method of this embodiment specifically includes steps 101 to 105. Wherein:

[0057] Step 101: Obtain the first phase training image and the second phase training image.

[0058] Here, the first time phase and the second time phase refer to images acquired at two different time points. In this embodiment of the invention, the first time phase training image and the second time phase training image include multi-angle remote sensing images of the same scene, and each scene image includes three angles, namely a front view, a frontal view, and a rear view.

[0059] Furthermore, this embodiment of the invention also performs geometric registration on the acquired first-temporal training image and the second-temporal training image. Geometric registration is used to align image data acquired at different times, in different bands, or from different sensors in terms of spatial location and orientation, so that the same feature points in the images can be accurately matched. This process is particularly important in the field of remote sensing because it involves integrating image data from different sources to ensure the accuracy and consistency of spatial information. Geometric registration of images can be achieved using remote sensing image processing software such as ENVI, ERDAS, and ArcGIS. The original image is then cropped into several fixed-size image blocks to meet the input requirements of the deep network model; typically, the cropped image block size is 256×256 or 512×512.

[0060] Step 102: Input the first time-phase training image and the second time-phase training image into the deep network model to obtain the features of the first time-phase image and the features of the second time-phase image.

[0061] The embodiments of the present invention are based on Figure 2 Taking the model structure shown as an example for self-supervised training, the model consists of two parts: an encoder and a mapping unit. The encoder is used to extract semantic features from the input image and can use deep network structures commonly used in image interpretation, such as ResNet-50 based on convolutional layers or Swin Transformer based on the Transformer structure. The mapping unit takes the encoder output features after global spatial pooling as input and projects the features. It is usually a multi-layer perceptron (MLP) composed of fully connected layers.

[0062] Step 103: Construct positive sample pairs from image features of the same scene but different angles in the first and second time-phase image features, and construct negative sample pairs from image features of different scenes in the first and second time-phase image features.

[0063] Step 104: Construct a multi-angle contrast loss function.

[0064] The goal of the multi-angle comparison loss function is to narrow the distance between positive samples and widen the distance between negative samples.

[0065] Step 105: Perform self-supervised training on the deep network model based on the positive sample pairs, negative sample pairs, and multi-angle contrast loss function.

[0066] Step 105 specifically includes: grouping the first phase training images by scene; taking each angle image of each scene as the main viewpoint, calculating the contrast loss through a multi-angle contrast loss function for the positive sample pairs and negative sample pairs formed by the image and the second phase training images, and obtaining multiple angle loss values; adding the multiple angle loss values ​​to obtain the multi-angle contrast loss value; and performing self-supervised training on the deep network model based on the multi-angle contrast loss value.

[0067] The training image in the first phase Taking a specific scenario as an example, the process of calculating its contrast loss includes:

[0068] First, calculate the first time-phase training image. The contrast loss value of the front view of each scene is calculated using the following formula:

[0069] ,

[0070] in, This represents the front view. This represents the front view. Represents the rear view; Indicates the first One scenario; In the first phase training image, the first... A front view of the scene; In the second phase training image, the first... The scene is Images from the observation angle; This indicates the number of scenes in the second-phase training images; To distinguish the functions, the calculation formula is as follows:

[0071] ,

[0072] in, To make the image Input image features obtained from the deep network model; The L2 norm of the eigenvectors; These are the training parameters for the deep network model; These are the hyperparameters of the deep network model.

[0073] Then, in the same manner described above, the first time-phase training image is calculated... The contrast loss value between the front and back views of a scene.

[0074] Finally, the first phase training image obtained is calculated. The sum of the contrast loss values ​​of the front view, front view, and back view of a scene is calculated using the following formula:

[0075] ,

[0076] Scene in the first phase training image Multi-angle comparison of loss values.

[0077] Based on the constructed loss function, the deep network model is trained using gradient descent until the network converges. Commonly used optimizers for network training include SGD, RMSprop, and Adam, which can be directly called in popular deep learning frameworks such as Tensorflow and PyTorch.

[0078] The self-supervised training method for deep network models in this invention constructs a multi-angle self-supervised contrastive loss function to train the network with the goal of narrowing the distance between positive sample pairs and widening the distance between negative sample pairs. This allows the deep network to learn the remote sensing observation characteristics of the same scene from different perspectives, achieving viewpoint-independent deep semantic feature representation. In land use change detection applications, this helps suppress spurious changes caused by differences in observation angles between two time phases. After multi-angle self-supervised pre-training, the network learns rich feature representations of land features from a large amount of unlabeled data. Its encoder can be applied to downstream tasks of land use change detection. By using unsupervised change detection methods or fine-tuning the network with a small number of samples, relatively ideal land use change detection results can be obtained.

[0079] Example 2:

[0080] Based on the same inventive concept as Embodiment 1, this embodiment of the invention also provides a deep network model self-supervised training device for implementing the above-described deep network model self-supervised training method. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in the following embodiments of the detection model training device can be found in the limitations of the deep network model self-supervised training method described above, and will not be repeated here.

[0081] like Figure 3 As shown, this embodiment of the invention provides a self-supervised training device for deep network models, comprising:

[0082] The training image acquisition module is used to acquire a first phase training image and a second phase training image, wherein the first phase training image and the second phase training image include multi-angle images of the same scene;

[0083] The image feature acquisition module is used to input the first phase training image and the second phase training image into the deep network model to acquire the first phase image features and the second phase image features.

[0084] The sample pair construction module is used to construct positive sample pairs from image features of the same scene but different angles in the first and second time phase image features, and to construct negative sample pairs from image features of different scenes in the first and second time phase image features.

[0085] The loss function construction module is used to construct multi-angle contrast loss functions;

[0086] The deep network model self-supervised training module is used to perform self-supervised training on the deep network model based on the positive sample pairs, negative sample pairs, and multi-angle contrast loss function.

[0087] Example 3:

[0088] like Figure 4 and Figure 5 As shown in the figure, this embodiment of the invention provides a method for detecting land use change. Figure 4 This is a flowchart illustrating the land use change detection method. This flowchart only shows the logical sequence of the method described in this embodiment. Provided there are no conflicts, different methods may be used in other possible embodiments of the invention. Figure 4 Complete the steps shown or described in the order indicated.

[0089] The land use change detection method provided in this embodiment can be applied to a terminal and can be executed by a land use change detection device. This device can be implemented by software and / or hardware and can be integrated into the terminal.

[0090] See Figure 4 The method in this embodiment of the invention specifically includes:

[0091] Step 301: Acquire remote sensing images of the same scene at different time phases to be detected.

[0092] Step 302: Input the remote sensing image to be detected into the pre-established land use change detection model to obtain the land use change in the remote sensing image.

[0093] Among them, such as Figure 5 As shown, the pre-training method of the land use change detection model of the present invention includes the deep network model self-supervised training method in the aforementioned embodiments, and applies the trained detection model to the downstream task of land use change detection to obtain the change detection result.

[0094] Land use change can be detected using the following two methods:

[0095] (1) For the two input remote sensing images at different times, the image features of the two output images at different times are directly differentially analyzed to obtain the change intensity map. :

[0096] ,

[0097] in, and Given two temporal phases of remote sensing images, For a well-trained network encoder, The number of features extracted by the encoder. The change intensity map. After normalization to the range of [0,1], threshold segmentation is performed using the Otsu method to obtain the change detection results.

[0098] (2) In addition to the encoder, a change detection discriminant network (e.g., BiT) is introduced, and the overall network is fine-tuned using a small number of land use change detection samples (i.e., the network is further trained using the cross-entropy loss function based on the self-supervised pre-trained encoder). Then, remote sensing images from two time phases are input into the fine-tuned network to obtain the change detection results.

[0099] Example 4:

[0100] Based on the same inventive concept as Embodiment 3, this embodiment of the invention also provides a land use change detection device for implementing the above-described land use change detection method. The solution provided by this device is similar to the implementation described in the above-described method; therefore, the specific limitations in the embodiments of the land use change detection device provided below can be found in the limitations of the land use change detection method described above, and will not be repeated here.

[0101] like Figure 6 As shown, an embodiment of the present invention provides a land use change detection device, comprising:

[0102] The remote sensing image acquisition module is used to acquire remote sensing images of the same scene at different times to be detected;

[0103] The land use change detection module is used to input the remote sensing image to be detected into a pre-established land use change detection model to obtain the land use change in the remote sensing image.

[0104] Example 5:

[0105] This invention also provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements the deep network model self-supervised training method or the land use change detection method described in the foregoing embodiments.

[0106] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0107] Example 6:

[0108] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned deep network model self-supervised training method or land use change detection method.

[0109] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0110] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0111] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0113] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A self-supervised training method for deep network models, characterized in that, include: Acquire a first-phase training image and a second-phase training image, wherein the first-phase training image and the second-phase training image include multi-angle images of the same scene; The first and second phase training images are input into a deep network model to obtain the features of the first and second phase images. Positive sample pairs are constructed from image features of the same scene in the first and second time-phase image features, and negative sample pairs are constructed from image features of different scenes in the first and second time-phase image features. The image features of the same scene include: (a) image features of the same scene at different times and at the same imaging angle; and (b) image features of the same scene at different times and at different imaging angles. Construct a multi-angle contrast loss function; The deep network model is trained under self-supervised conditions based on the positive sample pairs, negative sample pairs, and multi-angle contrast loss function. The first and second phase training images are remote sensing images; The method further includes: performing geometric registration on the first temporal training image and the second temporal training image; Each scene in the first and second phase training images includes three perspective images: a front view, a frontal view, and a rear view. The deep network model is self-supervised trained based on positive sample pairs, negative sample pairs, and a multi-angle contrast loss function, including: The first phase training images are grouped by scene; Taking each angle image of each scene as the main viewpoint, the contrast loss is calculated by using the multi-angle contrast loss function to obtain multiple angle loss values. The loss values ​​from multiple angles are summed to obtain the multi-angle contrast loss value; Self-supervised training of deep network models is performed by comparing loss values ​​from multiple perspectives. Among them, the first phase training image is the first The formula for calculating the contrast loss of a scene's front view from the main perspective is as follows: , in, This represents the front view. This represents the front view. Represents the rear view; Indicates the first One scenario; Indicates the first One scenario; In the first phase training image, the first... A front view of the scene; In the second phase training image, the first... The scene is Images from the observation angle; This indicates the number of scenes in the second-phase training images; To distinguish the functions, the calculation formula is as follows: , in, To make the image Image features obtained from the input deep network model; The L2 norm of the eigenvectors; These are the training parameters for the deep network model; These are the hyperparameters of the deep network model; Scene in the first phase training image The formula for calculating the multi-angle contrast loss value is: 。 2. A self-supervised training device for deep network models, characterized in that, include: The training image acquisition module is used to acquire a first phase training image and a second phase training image, wherein the first phase training image and the second phase training image include multi-angle images of the same scene; The image feature acquisition module is used to input the first phase training image and the second phase training image into the deep network model to acquire the first phase image features and the second phase image features. The sample pair construction module is used to construct positive sample pairs from image features of the same scene in the first time phase image features and the second time phase image features, and to construct negative sample pairs from image features of different scenes in the first time phase image features and the second time phase image features. The image features of the same scene include: (a) image features of the same scene at different times and at the same imaging angle; and (b) image features of the same scene at different times and at different imaging angles. The loss function construction module is used to construct multi-angle contrast loss functions; The deep network model self-supervised training module is used to perform self-supervised training on the deep network model based on the positive sample pairs, negative sample pairs, and multi-angle contrast loss function. The first and second phase training images are remote sensing images; The device is also capable of: performing geometric registration on the first temporal training image and the second temporal training image; Each scene in the first and second phase training images includes three perspective images: a front view, a frontal view, and a rear view. The deep network model is self-supervised trained based on positive sample pairs, negative sample pairs, and a multi-angle contrast loss function, including: The first phase training images are grouped by scene; Taking each angle image of each scene as the main viewpoint, the contrast loss is calculated by using the multi-angle contrast loss function to obtain multiple angle loss values. The loss values ​​from multiple angles are summed to obtain the multi-angle contrast loss value; Self-supervised training of deep network models is performed by comparing loss values ​​from multiple perspectives. Among them, the first phase training image is the first The formula for calculating the contrast loss of a scene's front view from the main perspective is as follows: , in, This represents the front view. This represents the front view. Represents the rear view; Indicates the first One scenario; Indicates the first One scenario; In the first phase training image, the first... A front view of the scene; In the second phase training image, the first... The scene is Images from the observation angle; This indicates the number of scenes in the second-phase training images; To distinguish the functions, the calculation formula is as follows: , in, To make the image Image features obtained from the input deep network model; The L2 norm of the eigenvectors; These are the training parameters for the deep network model; These are the hyperparameters of the deep network model; Scene in the first phase training image The formula for calculating the multi-angle contrast loss value is: 。 3. A method for detecting land use change, characterized in that, include: Acquire remote sensing images of the same scene at different times to be detected; The remote sensing image to be detected is input into a pre-established land use change detection model to obtain the land use change in the remote sensing image; The pre-training method for the land use change detection model includes the self-supervised training method for deep network models as described in claim 1.

4. A land use change detection device, characterized in that, include: The remote sensing image acquisition module is used to acquire remote sensing images of the same scene at different times to be detected; The land use change detection module is used to input the remote sensing image to be detected into a pre-established land use change detection model to obtain the land use change in the remote sensing image; The pre-training method for the land use change detection model includes the self-supervised training method for deep network models as described in claim 1.

5. A computer device, characterized in that, include: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the self-supervised training method for deep network models as described in claim 1 or the steps of the land use change detection method as described in claim 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps of the self-supervised training method for deep network models as described in claim 1 or the steps of the land use change detection method as described in claim 3.