BERT and self-supervised learning-based small sample city scene image analysis method

By introducing BERT and self-supervised learning methods in small sample urban scene image analysis, combined with cross-modal information fusion and layer-by-layer thawing strategies, the problems of high labeling costs and low recognition accuracy of small sample data are solved, and efficient and accurate urban scene image analysis is achieved.

CN120145112APending Publication Date: 2025-06-13NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510233699.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has problems in the analysis of small sample urban scene images with high data annotation costs, difficulty in obtaining small samples, and complex and diverse image features, resulting in a decrease in the recognition accuracy of the model in complex urban scenes.

Method used

The small sample city market scene image analysis method based on BERT and self-supervised learning is adopted, and feature extraction and cross-modal alignment training is performed through the self-supervised learning framework, combining the cross-modal information fusion module and layer-by-layer thawing fine-tuning strategy to improve the generalization ability of the model in small sample scenarios.

Benefits of technology

It significantly improves the accuracy and applicability of market scene image analysis in small sample cities, improves the generalization ability of the model and the efficiency of cross-modal information utilization, and reduces the dependence on labeled data and the demand for hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145112A_ABST
    Figure CN120145112A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and natural language processing, and discloses a small sample city scene image analysis method based on BERT and self-supervised learning. Potential information of unlabeled urban environment image data is fully mined in combination with self-supervised learning, and multi-modal features of urban environment images and description texts are integrated by using a cross-modal semantic enhancement mechanism, so that accurate diagnosis of small sample urban environment images is realized. According to the method, the generalization ability of the model in a small sample scene is improved, the diagnosis efficiency and accuracy of an existing method in a complex analysis scene are remarkably improved, and the defects that in the prior art, data dependence is high, the requirement for hardware resources is high, and cross-modal information utilization is insufficient are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and natural language processing, and in particular to a small sample urban scene image analysis method based on BERT and self-supervised learning. Background Art

[0002] Urban scene image analysis has a wide range of applications in smart city construction, especially in traffic management, urban planning, security monitoring and other fields. With the rapid development of artificial intelligence (AI) technology, urban scene image analysis methods based on deep learning have become a hot topic of research. These methods promote the transformation from traditional analysis methods that rely on manual experience to intelligent computer-aided analysis (CAA) systems through automated analysis of urban scene images. However, urban scene image data has the following characteristics: high data annotation costs, small samples are difficult to obtain, and image features are complex and diverse. Existing deep learning models usually rely on large-scale annotated data to improve the accuracy of the model. Due to the lack of sufficient annotated data, how to improve the analysis efficiency and accuracy in small sample urban scene image analysis has become a difficult problem that needs to be solved urgently in this field.

[0003] Current research focuses on the following aspects: (1) urban scene image feature extraction and classification technology based on convolutional neural network (CNN); (2) using data enhancement, transfer learning and other methods to improve the generalization ability of the model in small sample scenarios; (3) combining hardware design to optimize computing resources to realize the practical application of urban scene analysis methods; (4) scene segmentation based on traditional machine learning algorithms;

[0004] The following are the relevant research and applications of two existing technologies:

[0005] CN202310580934.6 proposes a semantic segmentation method and system for high-resolution remote sensing urban images, called RS-SwinUnet. It uses Swin Transformer to construct the encoder and CNN to construct the decoder. In the encoding stage, it extracts global features of the image. In the decoding stage, it uses a feature fusion module to fuse the low-level features from the encoder and the high-level features from the decoder, and uses an upsampling dilation layer to restore the detailed position information. At the same time, skip connections from the encoder are added during decoding to assist in restoring the position information. The overall structure is still a U-shaped structure to achieve semantic segmentation of urban scene remote sensing images. However, when dealing with complex urban scene images, this method may have limitations in processing low-resolution or high-noise images, and in the feature fusion process, the combination of low-dimensional and high-dimensional features may not be sufficient to effectively capture multi-modal information. In addition, although skip connections are used to assist in restoring detailed information, in some highly complex urban scenes, it may still not be able to fully restore details, resulting in a decrease in recognition accuracy.

[0006] CN202311687735.1 is a method for classifying urban area scenes in SAR images based on multi-scale wavelet convolution. It includes constructing a multi-scale convolution module with an attention mechanism to extract and fuse features of different scales from the original SAR image, assigning attention weights to the fused multi-scale features to obtain a fused feature map with spatial attention weights; constructing a discrete wavelet transform convolution module to decompose and obtain a multi-scale wavelet feature map with spatial domain information and frequency domain information; constructing a discriminator network to identify the probability of the multi-scale wavelet feature map belonging to the urban area scene category; using the standard cross-entropy loss function to train the overall network constructed by S1-S3 to obtain a pre-trained overall network; using the pre-trained overall network to test the SAR image to be tested to obtain the urban area scene classification result. However, when dealing with complex urban area scenes, this method may have certain limitations, especially in the classification accuracy of low-quality or noisy SAR images. Although the multi-scale convolution module can extract features of different scales, in the allocation of spatial attention weights, it may not be able to fully handle the detailed features of high-complexity urban areas, resulting in insufficient accuracy of the fused feature map. In addition, although the discrete wavelet transform convolution module can extract spatial domain and frequency domain information, in diverse urban areas, this method may not be able to fully capture the high heterogeneity in the scene, affecting the accuracy of the classification result.

[0007] Although the two methods respectively provide innovative technical solutions in the semantic segmentation of remote sensing images and the scene classification of urban areas in SAR images, there are still certain deficiencies in their technical solutions. These deficiencies limit their promotion and further improvement of performance in practical applications. Generally speaking, the main problems of these two technologies lie in the inaccurate modeling of complex scenes and the insufficient processing of small-sample data. Especially when dealing with high-noise or low-resolution images, the existing methods have limited ability to fuse and recover features, and fail to effectively improve the utilization efficiency of multi-modal information. At the same time, these technologies rely too much on hardware acceleration or traditional data augmentation methods, and do not fully combine self-supervised learning and cross-modal information fusion to explore the potential of unlabeled data. The present invention can improve the generalization ability of the model under small-sample conditions by introducing self-supervised learning and cross-modal information enhancement mechanisms, and effectively improve the problem of insufficient cross-modal information fusion between images and texts in the existing technology, thereby significantly improving the analysis accuracy and applicability of complex urban scenes. Summary of the Invention

[0008] The object of the present invention is to propose a small-sample urban scene image analysis method based on BERT and self-supervised learning. By combining self-supervised learning, the potential information of unlabeled medical image data is fully explored. At the same time, a cross-modal semantic enhancement mechanism is used to integrate the multi-modal features of urban environment images and descriptive texts, so as to realize the accurate analysis of small-sample urban environment images. The present invention not only improves the generalization ability of the model in small-sample scenarios, but also significantly improves the recognition efficiency and accuracy of the existing methods in complex urban scenes, and overcomes the defects of strong data dependence, high requirements for hardware resources and insufficient utilization of cross-modal information in the existing technology.

[0009] The technical solution of the present invention is as follows: A small-sample urban scene image analysis method based on BERT and self-supervised learning, comprising the following steps:

[0010] Step 1: Data preprocessing, standardizing the urban scene images and descriptive texts;

[0011] Step 2: Self-supervised learning, extracting features and performing cross-modal alignment training on the preprocessed urban scene images through a designed self-supervised learning framework; in the cross-modal alignment training, a cross-modal information fusion module is added, and a cross-modal attention mechanism is used to semantically fuse the urban scene image features and descriptive text features to generate cross-modal fusion features;

[0012] Step 3: Small-sample optimization, using a fine-tuning strategy of layer-by-layer thawing to optimize the parameters of the self-supervised learning framework and the cross-modal information fusion module, and obtaining an optimized analysis model;

[0013] Step 4: Inference Output. Apply the optimized analysis model to the joint input of the urban scene image and the descriptive text, and output the urban classification result and the scene analysis report.

[0014] The data preprocessing includes normalizing the urban scene image and the descriptive text.

[0015] The specific data preprocessing of the urban scene image is as follows: All urban scene images are adjusted to a unified size, and the size adjustment is completed using the bilinear interpolation algorithm while retaining the detailed features of the image; the adjusted images are subjected to data augmentation, including random horizontal flipping, random rotation, random cropping, and scaling; the pixel values of all images after data augmentation are normalized to the interval [0, 1]; the normalization formula is as follows, where I represents the pixel value of the original image, I min and I max are the minimum and maximum values of the image pixel values respectively:

[0016]

[0017] The specific data preprocessing of the descriptive text is as follows: Use the pre-trained BERT tokenizer to tokenize the descriptive text; the BERT tokenizer decomposes each descriptive text into lexical units and converts them into corresponding word embedding vectors; the formula is as follows:

[0018] T = BERT-Tokenizer(S)

[0019] where S represents the original descriptive text, T = [t 1 , t 2 ..t 3 represents the tokenization result, and each t i is a Token in the BERT tokenizer vocabulary; for texts longer than 128, truncation operations are performed; texts shorter than 128 are padded with the special padding symbol [PAD]. Subsequently, the tokenization result is converted into a 768-dimensional embedding vector through the BERT model:

[0020] E = BERT-Encoder(T)

[0021] Construct the preprocessed urban scene images and the corresponding descriptive texts into paired data; the paired data will be input into the self-supervised learning framework for feature extraction and cross-modal alignment training; at the same time, the dataset of the paired data includes both positive samples and negative samples. The positive samples are formed by pairing the urban scene images with their corresponding descriptive texts, and the negative samples are constructed by randomly shuffling the pairing relationship between the urban scene images and the descriptive texts to form mismatched image-text combinations.

[0022] The specific pairing method of the paired data is as follows:

[0023] The positive samples are formed by pairing urban scene images with their corresponding descriptive texts; the preprocessed image data I 1 , I 2 ,...,, I n is mapped to the normalized image set I norm,1 , I norm,2 ,...,, I norm,n , and the descriptive text data S 1 , S 2 ,...,, S n is mapped to the embedding vector set E 1 , E 2 ,...,, E n ; the paired data set composed of positive samples is expressed as:

[0024] D positive = {(I norm,1 , E 1 ), (I norm,2 , E 2 ), …, (I norm,n , E n )}

[0025] The negative samples are constructed by randomly shuffling the pairing relationship between urban scene images and descriptive texts, forming unmatched image-text combinations; the generation rule of negative samples is: keeping the original image data set I norm,1 , I norm,2 ,...,, I nrom,n unchanged, randomly shuffling the order of the embedding vector set to form a new unmatched combination E rand,1 , E rand,2 ,...,, E rand,n , and finally the paired data set composed of negative samples is expressed as:

[0026] D negative

[0027] = {(I norm,1 , E rand,1 ), (I norm,2 , E rand,2 ), …, (I norm,n , E rand,n )}

[0028] The negative sample ratio is set to 50%, and the final data set is composed of a mixture of positive and negative samples: D = D positive ∪ D negative .

[0029] The self-supervised learning framework includes an image feature extraction task and a cross-modal alignment task; the cross-modal information fusion module, as part of the self-supervised learning framework, is used to fuse urban scene image features and descriptive text features to generate cross-modal fusion features;

[0030] The image feature extraction task includes a jigsaw puzzle reconstruction task and a contrastive learning task; both the jigsaw puzzle reconstruction task and the contrastive learning task are used to train the same pre-trained ResNet-50 model; the loss function during training includes the cross-entropy loss of the jigsaw puzzle reconstruction task and the contrastive loss based on cosine similarity of the contrastive learning task, and the two loss functions jointly optimize the parameters of ResNet-50 with equal weights;

[0031] The core idea of the jigsaw puzzle reconstruction task is to randomly divide the input urban scene image into nine small images and randomly shuffle the order; the pre-trained ResNet-50 model needs to predict the original position of each piece of the urban scene image in this task to restore the overall spatial structure of the image, and the loss function is designed as the cross-entropy loss, and its formula is as follows:

[0032]

[0033] where N represents the number of training samples, M represents the number of possible positions of each piece of the urban scene image, y ij is the distribution of the true label, is the probability distribution predicted by the ResNet-50 model;

[0034] The contrastive learning task is to generate two versions of the augmented images as positive sample pairs by performing different data augmentation operations on the same urban scene image, and at the same time randomly select other images to generate negative sample pairs; the two augmented images are respectively encoded as high-dimensional feature vectors, and the distance between the positive sample pairs is calculated by cosine similarity to optimize the feature distribution of the pre-trained ResNet-50 model between positive and negative samples; the loss function of contrastive learning adopts the contrastive learning loss based on the temperature parameter, and its formula is:

[0035]

[0036] where sim(z i , z j ) represents the cosine similarity between two feature vectors, τ is the temperature parameter, set to 0.5; the feature dimension of contrastive learning is 128, and the batch size is 64.

[0037] In the cross-modal alignment task, combining the image-text matching task and the text generation task, the semantic alignment of the urban scene image and the descriptive text is realized;

[0038] The image-text matching task determines whether the input image and the descriptive text match through a designed binary classification model. First, the image features are extracted by a pre-trained ResNet-50 model to obtain a 512-dimensional embedding vector. The text features are encoded by a BERT model to generate a 768-dimensional embedding vector. The image features and text features are fused through a cross-modal information fusion module to generate cross-modal fusion features. The cross-modal information fusion module dynamically assigns weights of the descriptive text to the image features through a cross-modal attention mechanism, fully integrating the spatial information of the image and the semantic information of the text to improve the accuracy of matching. Finally, the probability distribution of the match is output through an output layer composed of two fully connected networks and a Sigmoid. The loss function uses cross-entropy loss, and the formula is:

[0039]

[0040] where y i represents the true label of the sample, is the probability predicted by the binary classification model;

[0041] The text generation task generates a corresponding diagnostic report from the cross-modal fusion feature F fusion ;

[0042] The Seq2Seq structure includes an encoder and a decoder. In the encoder, a double-layer bidirectional LSTM is used to further encode the cross-modal fusion feature F fusion . The hidden layer dimension of the BiLSTM is 256, and the input of each layer is a 1024-dimensional fusion feature, outputting the context representation and the global hidden state, which are used to initialize the hidden state of the decoder;

[0043] In the decoder, a double-layer unidirectional LSTM structure is adopted, and the hidden layer dimension of each layer is 256. The input of the decoder is the word generated at the previous time step, and the initial input is the [START] token. Combining the context representation of the encoder, the context weight distribution at each time step is dynamically calculated through the attention mechanism, generating the semantic representation at the current time step and predicting the probability distribution of the current word. The word embedding dimension is 300, and the Beam Search strategy is used during the decoding process. The training loss function of the Seq2Seq structure is the sequence cross-entropy loss, and its formula is:

[0044]

[0045] where T is the length of the generated sequence, w t is the generated word, and P(w t |w 1 , w 2 ,...w t-1 ) is the conditional probability of generating the next word.

[0046] The cross-modal attention mechanism organically integrates the semantic information of urban scene images and descriptive texts by dynamically allocating the weights of descriptive texts to image features. The core formula of the cross-modal attention mechanism is as follows:

[0047]

[0048] where Q, K, and V are the query matrix, key matrix, and value matrix respectively, is the normalization factor of the feature dimension, and softmax is used to calculate the attention weights. The image features serve as the query matrix Q, and the text features serve as the key matrix K and value matrix V, thereby calculating the weight distribution of the descriptive texts to the image features. After obtaining the attention weights, the image features are fused with the weighted text features, and the fusion formula is as follows:

[0049] F fusion = W f · [F img + Attention(Q, K, V)]

[0050] where F fusion represents the cross-modal fusion features, F img are the original image features, Attention(Q, K, V) are the weighted text features output by the cross-modal attention mechanism, and W f is the learnable linear transformation matrix of the fusion features. The dimension of the cross-modal fusion features is 1024-dimensional, and the cross-modal fusion features are reduced to the target dimension through a fully connected layer. The dimension reduction formula is:

[0051] F reduced = W r · F fusion + b r

[0052] where W r and b r are the weights and biases of the fully connected layer respectively.

[0053] The few-shot optimization is specifically as follows: Design a fine-tuning strategy of thawing layer by layer, and transfer the image feature extractor and text feature extractor trained in the self-supervised learning framework to a specific urban scene image analysis task by fine-tuning the pre-trained model parameters.

[0054] The few-shot optimization specifically includes the following steps:

[0055] (1) Initialize the model: Load the pre-trained model parameters trained through the self-supervised learning framework, including the ResNet-50 model and the BERT model;

[0056] (2) Freeze model parameters: At the initial stage of fine-tuning, most layers of the ResNet-50 model and the BERT model are frozen, and only the high-level parameters are unfrozen; the specific settings are as follows:

[0057] ResNet-50: Freeze the first convolutional layer to the third convolutional layer, and only unfreeze the fourth convolutional layer and the fully connected layer;

[0058] BERT: Freeze the first 8 layers of the Transformer encoder, and only unfreeze the last 4 layers;

[0059] (3) Unfreeze layer by layer: During the fine-tuning process, after a set number of training epochs, gradually unfreeze the low-level parameters of the ResNet-50 model and the BERT model, release the frozen weights layer by layer, and gradually adjust from high-level features to global features; for the ResNet-50 model, freeze the first convolutional layer to the third convolutional layer at the initial stage, and only unfreeze the fourth convolutional layer and the fully connected layer; when unfreezing layer by layer, unfreeze the low-level convolutional parameters from the third convolutional layer to the first convolutional layer in sequence; for the BERT model, freeze the first 8 layers of the Transformer encoder at the initial stage, and only unfreeze the last 4 layers; when unfreezing layer by layer, gradually release the weights starting from the 8th layer of the Transformer encoder, and finally unfreeze all encoding layers;

[0060] (4) Adjust the learning rate: During the fine-tuning process, set different learning rates for different layers, and the learning rate of the low-level parameters is lower than that of the high-level parameters;

[0061] (5) Optimize the task loss: Optimize the cross-entropy loss of the classification task on the small-sample data, and at the same time introduce regularization constraints.

[0062] The inference output converts the input urban scene image and its related descriptive text data into clear and accurate analysis results by using the feature extraction model, cross-modal information fusion module, and optimized binary classification model trained in the previous steps; the specific implementation process of the inference output module is as follows:

[0063] (1) Urban scene classification: In the inference stage, the image features and text features obtained through the previous steps of training are input into the cross-modal information fusion module to generate the final fusion feature F fusion; This fused feature contains the spatial information of the image and the semantic information of the text, which is the key input for the urban scene classification task. The binary classification model uses a two-layer fully connected network (FCN) followed by Sigmoid for urban scene classification. The first layer reduces the dimension of the fused feature from 1024 to 512. The second layer maps the 512-dimensional feature to the number of target classes, and the classifier outputs the probability distribution of each class. Through the above formula, the probability distribution of each urban scene class is output, and the class with the highest probability is taken as the final classification result. The classifier formula is as follows:

[0064]

[0065] where C is the number of classes, W c and b c are the weight and bias parameters of class c respectively;

[0066] (2) Scene analysis report generation: In the scene analysis report generation task, the fused feature F fusion is further input into a sequence-to-sequence generation model for generating the corresponding natural language scene analysis report. The generation model consists of an encoder-decoder structure: The encoder uses the fused feature of the image and text as input to generate the context semantic representation. The decoder: Through the attention mechanism, it dynamically focuses on different parts of the fused feature when generating words at each step. The decoder starts with a token and recursively generates the scene analysis report by generating the probability distribution of the next word. The maximum generation length is set to 64 words.

[0067] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing a multi-task training framework based on self-supervised learning, the present invention fully utilizes the potential information of unlabeled urban scene image data without the need for a large amount of labeled data. In the small-sample scenario, the method of the present invention improves the accuracy by more than 20%, while significantly reducing the dependence on labeled data, and solves the problems of high cost and insufficient samples in the annotation of urban scene image data. This technology not only improves the generalization ability of the model but also reduces the implementation cost of diagnosis, and has extremely high practical value.

[0068] The present invention adopts a cross-modal semantic enhancement mechanism, which organically combines urban scene images with descriptive texts to achieve semantic alignment between images and texts. Compared with the existing analysis methods that only rely on image features, the present invention significantly improves the classification accuracy in the analysis of complex urban scene images. Especially in the fusion analysis of multi-modal information, the comprehensive judgment ability of the model is significantly improved, and the accuracy is increased by 15%-25%. This mechanism helps the model to more comprehensively understand the characteristics of urban scenes, compensates for the analysis deviation that may be caused by single-modal information, and thus effectively reduces the wrong analysis or omission caused by incomplete information.

[0069] The present invention realizes efficient model migration in small-sample scenarios through a fine-tuning strategy of a small amount of labeled data. Compared with traditional data augmentation techniques, the present invention adopts a transfer learning strategy of layer-by-layer thawing to quickly transfer the feature capabilities of the pre-trained model to specific tasks. When only 10%-20% of the data is labeled using traditional methods, the analysis accuracy of the model almost reaches the level of using the method of large-scale labeled data. This strategy significantly reduces the cost of data acquisition and processing, and provides a more efficient and economical solution for urban scene image analysis tasks.

[0070] The present invention reduces the dependence on high-performance hardware by improving the algorithm design and has strong applicability. Compared with the existing technologies optimized based on FPGA hardware, the present invention can achieve the same or even higher analysis performance without relying on special hardware. In resource-constrained urban scenes, the algorithm of the present invention achieves the effects of 30% reduction in energy consumption and 50% increase in inference speed on general GPU devices, significantly improving the analysis efficiency while reducing the hardware investment and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is a flowchart of the implementation of a small-sample urban scene image analysis method based on BERT and self-supervised learning according to the present invention;

[0072] Figure 2 is a flowchart of the implementation process of the small-sample optimization proposed by a small-sample urban scene image analysis method based on BERT and self-supervised learning according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0073] Figure 1 is the main flowchart of the technical solution of the present invention. As Figure 1As shown in the figure, the present invention proposes a few-shot urban scene image analysis method based on BERT and self-supervised learning. Its overall technical solution significantly improves the accuracy and generalization ability of urban scene image analysis by designing a self-supervised learning framework, a cross-modal information fusion module, and a few-shot optimization strategy. The main technical core of the present invention lies in: fully exploiting the potential information of unlabeled urban scene image data, integrating the multi-modal features of urban scene images and descriptive texts, and achieving efficient and accurate analysis in scenarios with scarce data. This technical solution includes five main steps: data preprocessing, self-supervised learning, cross-modal information fusion, few-shot optimization, and inference output.

[0074] Step 1: Data preprocessing

[0075] In the data preprocessing module, the present invention standardizes urban scene images and texts.

[0076] (1) Urban scene image data preprocessing: First, to adapt to the input size of the model, all urban scene images are resized to a unified size of 224×224. The resizing is completed using the bilinear interpolation algorithm while retaining the detailed features of the images. Since the samples of urban scene images are usually few, to alleviate the overfitting problem and enhance the adaptability of the model to different data distributions, the present invention designs a series of data augmentation operations, including random horizontal flipping, random rotation, random cropping, and scaling. To ensure the uniformity of the input data, the pixel values of all images are normalized to the interval [0, 1]. The normalization formula is as follows, where I represents the pixel value of the original image, I min and I max are the minimum and maximum values of the image pixel values respectively:

[0077]

[0078] (2) Text data preprocessing: The pre-trained BERT tokenizer is used to tokenize the descriptive text. The tokenizer decomposes each text into lexical units (Tokens) and converts them into corresponding word embedding vectors. The formula is as follows:

[0079] T = BERT-Tokenizer(S)

[0080] where S represents the original text, T represents the tokenization result, and each t i is a Token in the BERT vocabulary. For texts with a length exceeding 128, truncation operations are performed; texts with a length less than 128 are padded with the special padding symbol [PAD]. Subsequently, the tokenization result is converted into a 768-dimensional embedding vector through the BERT model:

[0081] E = BERT-Encoder(T)

[0082] (3) Data pairing: The preprocessed urban environment images and their corresponding descriptive texts are constructed into paired data for the training and inference of cross-modal tasks. During the construction process, all image-text paired data are divided into two categories: positive samples - the images match the corresponding diagnostic reports; negative samples - the pairing relationship is randomly shuffled to construct unmatched image-text combinations. The negative sample ratio is set to 50% to balance the number of positive and negative samples and facilitate the training of subsequent image-text matching tasks.

[0083] The image data {I 1 , I 2 ,..., I n} is mapped to the normalized image set {I norm,1 , I norm,2 ,..., I norm,n}, and the text data {S 1 , S 2 ,..., S n} is mapped to the embedded vector set {E 1 , E 2 ,..., E n}. Finally, the paired data set after step one is D = {(I norm,1 , E 1 ),..., (I norm,n , E n )}. These preprocessing steps provide a clear and unified input data form for subsequent model training, significantly improving the model's ability to extract image and text features and the alignment effect.

[0084] Step Two: Self-Supervised Learning

[0085] In the self-supervised learning module, the present invention designs two types of tasks: image feature extraction tasks and cross-modal alignment tasks. The image feature extraction tasks include jigsaw puzzle reconstruction tasks and contrastive learning tasks, aiming to utilize a large number of unlabeled medical images to mine deep feature information, thereby improving the model's learning ability for small sample data.

[0086] The core idea of the jigsaw puzzle reconstruction task is to randomly divide the input image into nine small images of 3×3 and randomly shuffle the order. The model needs to predict the original positions of each small image. This task extracts and encodes the features of the image patches through a convolutional neural network and outputs the prediction results of the original positions of each small image through a fully connected classifier. The loss function uses cross-entropy loss to optimize the accuracy of position prediction, and its formula is:

[0087]

[0088] where N represents the number of training samples, M represents the number of possible positions of each urban scene image (here it is 9), and y ij is the distribution of the true labels, is the probability distribution predicted by the ResNet-50 model. Through this task, the model can learn the spatial structure characteristics of the image and the context relationship of local features. During the training process, the learning rate is set to 0.001, and the batch size is set to 32 to ensure the convergence speed and stability.

[0089] The contrastive learning task constructs positive and negative sample pairs by generating different enhanced versions of the image, further strengthening the robustness of the model to image features. In the specific implementation, two enhanced images are encoded as high-dimensional feature vectors, the cosine similarity is used to calculate the distance of the positive sample pair, and a contrastive learning loss function (NT-Xent Loss) is designed to optimize the feature distribution between positive and negative samples. The formula of the loss function is:

[0090]

[0091] where sim(z i , z j ) represents the cosine similarity between two feature vectors, τ is the temperature parameter, set to 0.5. The feature dimension of contrastive learning is 128, and the batch size is 64. Through this task, the model can learn the impact of different enhancement methods on the essential features of the image and improve the discrimination ability for urban environment images.

[0092] In the cross-modal alignment task, the present invention combines the matching task and the text generation task of image features and text features to achieve semantic alignment between urban environment images and descriptive texts. The image-text matching task judges whether the input image and text match by designing a binary classification model. Image features are extracted by the pre-trained ResNet-50 to obtain 512-dimensional embedding vectors; text features are encoded by the BERT model to generate 768-dimensional embedding vectors. The two features are fused through a cross-modal attention mechanism, and the model outputs the probability distribution of the match. The loss function uses cross-entropy loss, and the formula is:

[0093]

[0094] where, y i represents the true label of the sample, is the probability predicted by the binary classification model. Through this task, the model can effectively learn the semantic associations between images and texts, improving its ability to understand multimodal data. The text generation task generates corresponding diagnostic reports from medical image features through a Seq2Seq structure, which specifically includes an encoder module and a decoder module. The encoder encodes the image features, and the decoder dynamically generates each word of the diagnostic report through an attention mechanism. The maximum generation length is set to 64 words to adapt to the length distribution of actual diagnostic reports. The loss function for the generation task is the sequence cross-entropy loss, and its formula is:

[0095]

[0096] where T is the length of the generated sequence, w t is the generated word, and P(w t |w 1 , w 2 ,... w t-1 ) is the conditional probability of generating the next word. Through this task, the model can generate descriptive reports with rich content and reasonable structures, which highly correspond to the image features and descriptive information. The self-supervised learning module significantly improves the utilization rate of unlabeled data by jointly training the above tasks, providing a powerful representation ability and cross-modal alignment ability for subsequent tasks.

[0097] The cross-modal information fusion module of the present invention realizes the feature alignment and deep fusion between urban environmental images and descriptive texts through a cross-modal attention mechanism, further improving the model's understanding ability in multimodal tasks. The core of this module is to use the attention mechanism to dynamically allocate the weights of texts to image features, organically integrating the semantic information of images and texts. The fused features are processed by a fully connected network for dimensionality reduction, and finally a unified feature vector is generated for use in downstream classification or generation tasks.

[0098] In the cross-modal information fusion module, the features of images and texts are first encoded separately. The urban environmental image features are extracted through a pre-trained ResNet-50 model, outputting a 512-dimensional feature representation; the text features are generated into 768-dimensional embedding vectors through a BERT encoder. To achieve the semantic alignment of image and text features, the present invention designs a cross-modal attention mechanism, and its core formula is as follows:

[0099]

[0100] where Q, K, and V are the query, key, and value matrices respectively, is the normalization factor for the feature dimension, and softmax is used to calculate the attention weights. In the cross-modal attention mechanism of the present invention, the image features serve as the query matrix Q, and the text features serve as the key matrix K and the value matrix V, thereby calculating the weight distribution of the text for the image features. Through the attention mechanism, the text features can dynamically focus on the relevant regional features in the image according to their semantic content. For example, when the description text mentions "city square", the attention mechanism will automatically enhance the image features related to the square area. After obtaining the attention weights, the image features are fused with the weighted text features. The fusion formula is as follows:

[0101] F fusion = W f · [F img + Attention(Q, K, V)]

[0102] where F fusion represents the fused features, F img is the original image features, Attention(Q, K, V) is the weighted text features output by the cross-modal attention mechanism, and W f is the learnable linear transformation matrix for the fused features. The dimension of the fused features is 1024-dimensional. To adapt to the final classification task, the fused features are reduced in dimension to the target dimension through a fully connected layer. The dimension reduction formula is:

[0103] F reduced = W r · F fusion + b r

[0104] where W r and b rThey are the weights and biases of the fully connected layer respectively. Through this cross-modal fusion mechanism, the present invention can effectively integrate multi-modal information of images and texts, and achieve in-depth understanding of complex urban scenes. Especially in the case of incomplete or noisy image information, text features provide important context semantic supplements, thus enhancing the model's ability to extract and discriminate urban environmental features. Taking the multi-modal analysis scenario as an example, when processing a mixed image containing features of "city center" and "commercial area", the information in the descriptive text can guide the attention mechanism to focus on the image areas related to the "city center" or "commercial area", significantly enhancing the accuracy of the analysis. In addition, the cross-modal information fusion module of the present invention has demonstrated superior performance and versatility in practice. Experiments show that compared with traditional analysis methods based only on image features, after adding the cross-modal information fusion module, the analysis accuracy has increased by 15%-20%, especially in scenarios with scarce data and complex multi-modal information. By introducing this module, the present invention has achieved efficient fusion of image and text information in urban scene image analysis, providing a more comprehensive and accurate solution for urban scene analysis tasks.

[0105] Step 4: Small sample optimization

[0106] In the small sample scenario, the present invention designs a fine-tuning strategy of layer-by-layer thawing, which combines the generalization ability of the pre-trained model and the effective utilization of a small amount of labeled data, significantly improving the analysis performance and avoiding the overfitting problem at the same time. The small sample optimization module mainly migrates the image and text feature extractors trained in the self-supervised learning module to specific urban environmental image analysis tasks by fine-tuning the pre-trained model parameters, so as to achieve efficient urban environmental image analysis under the condition of limited labeled data.

[0107] The core idea of small sample optimization is to use a small amount of labeled data to gradually thaw the model parameters for fine-tuning on the basis of freezing most of the pre-trained parameters, so as to ensure the training stability of the model on small sample data. During the fine-tuning process, the model gradually optimizes from general feature extraction to specific task features, and further avoids overfitting by adjusting the learning rate and regularization strategy.

[0108] In the present invention, the implementation process of small sample optimization is as Figure 2 shown, and can be specifically summarized as the following steps:

[0109] (1) Initialize the model: Load the pre-trained model parameters trained by self-supervised learning.

[0110] (2) Freeze model parameters: At the beginning of fine-tuning, most layers of the image encoder and text encoder are frozen, and only the last few layers are kept trainable. This operation ensures the stability of the model at the initial stage of fine-tuning and avoids large fluctuations in model parameters caused by small sample data. The settings for the frozen layers are as follows: ResNet-50: Freeze convolutional layers 1 to 3, and only unfreeze convolutional layer 4 and the fully connected layer. BERT: Freeze the first 8 layers of the Transformer encoder and only unfreeze the last 4 layers.

[0111] (3) Thaw layer by layer: Every 3 training epochs, gradually thaw the lower-level parameters of the model, adjusting from high-level features to global features, enabling the model to gradually adapt to the requirements of specific tasks while maintaining stability. This strategy balances the generalization ability of the model and the specific optimization needs in the small sample scenario by gradually expanding the training scope.

[0112] (4) Adjust the learning rate: During fine-tuning, different learning rates are set for different layers. The learning rate of the lower-level parameters is set lower to protect the pre-trained features, and the learning rate of the higher-level parameters is set higher to quickly adapt to new tasks.

[0113] (5) Optimize the task loss: Optimize the cross-entropy loss of the classification task on small sample data, and at the same time introduce regularization constraints to avoid overfitting.

[0114] (6) Validation and testing: Evaluate the model performance on the validation set, adjust the learning rate and regularization parameters, and finally verify the optimization effect on the test set.

[0115] Step Five: Inference Output

[0116] Inference output: Apply the optimized analysis model to the joint input of urban scene images and descriptive texts, and output the urban classification results and scene analysis reports. The specific implementation process of the inference output module is as follows:

[0117] (1) Urban scene classification: During the inference stage, the image features and text features obtained through the previous training steps are input into the cross-modal information fusion module to generate the final fusion feature F fusion ; This fusion feature contains the spatial information of the image and the semantic information of the text, which is the key input for the urban scene classification task; The binary classification model uses a two-layer fully connected network (FCN) followed by Sigmoid for urban scene classification; The first layer reduces the dimension of the fusion feature from 1024 to 512; The second layer maps the 512-dimensional feature to the number of target categories, and the classifier outputs the probability distribution of each category; Through the above formula, the probability distribution of each urban scene category is output, and the category with the highest probability is taken as the final classification result; The formula for this classifier is as follows:

[0118]

[0119] Among them, C is the number of categories, and W c and b c are the weight and bias parameters of category c, respectively;

[0120] (2) Scene analysis report generation: In the task of generating a scene analysis report, the fused feature F fusion is further input into a sequence-to-sequence generation model for generating the corresponding natural language scene analysis report. The generation model consists of an encoder-decoder structure: where the encoder uses the fused feature of the image and text as input to generate a context semantic representation; the decoder: through the attention mechanism, dynamically focuses on different parts of the fused feature at each step of generating a word; the decoder starts with a token and recursively generates the scene analysis report by generating the probability distribution of the next word; the maximum generation length is set to 64 words.

[0121] The present invention designs a multi-task training framework based on self-supervised learning, uses unlabeled urban scene image data, sets tasks such as image puzzle reconstruction and contrast learning, significantly improves the feature extraction ability and the generalization ability of the model in small-sample scenarios, thereby overcoming the problem of strong dependence on labeled data in the prior art.

[0122] The present invention adopts a cross-modal semantic enhancement mechanism, jointly models urban scene images and descriptive texts, encodes the descriptive text features using the BERT model, and realizes the alignment of image and text semantics through the cross-modal attention mechanism, enhancing the model's understanding ability of complex urban scenes and multi-modal data.

[0123] The present invention designs an optimization strategy suitable for small-sample scenarios, adopts a transfer learning method of layer-by-layer thawing and fine-tuning with a small amount of labeled data, efficiently transfers the learning ability of the pre-trained model to specific urban scene recognition tasks, and further improves the accuracy and efficiency of recognition.

[0124] The present invention reduces the dependence on high-performance hardware by improving the algorithm structure, and with the design of the optimized neural network model, can also achieve efficient urban image recognition in scenarios with limited hardware resources, having high practicality and promotion value.

Claims

1. A small sample urban scene image analysis method based on BERT and self-supervised learning, characterized in that: The following steps are involved: Step 1: Data preprocessing, standardizing urban scene images and description texts; Step 2: Self-supervised learning: feature extraction and cross-modal alignment training are performed on the pre-processed urban scene images through the designed self-supervised learning framework; In the cross-modal alignment training, a cross-modal information fusion module is added to use the cross-modal attention mechanism to semantically fuse the urban scene image features and description text features to generate cross-modal fusion features; Step 3: Small sample optimization: Use a layer-by-layer fine-tuning strategy to optimize the parameters of the self-supervised learning framework and the cross-modal information fusion module to obtain the optimized analysis model. Step 4: Inference output: Apply the optimized analysis model to the joint input of urban scene images and description texts, and output urban classification results and scene analysis reports.

2. The small sample urban scene image analysis method based on BERT and self-supervised learning according to claim 1 is characterized in that: The data preprocessing includes standardizing the city scene images and description texts; The data preprocessing includes standardizing the city scene images and description texts; The data preprocessing of the urban scene images is specifically as follows: all urban scene images are adjusted to a uniform size, and the size adjustment is completed by using a bilinear interpolation algorithm while retaining the detailed features of the image; The resized images are subjected to data augmentation, including random horizontal flipping, random rotation, random cropping and scaling; all image pixel values ​​after data augmentation are normalized to the interval [0, 1]; The normalization formula is as follows, where I represents the pixel value of the original image, I min and I max They are the minimum and maximum values ​​of the image pixel values: The data preprocessing of the description text is specifically as follows: using a pre-trained BERT word segmenter to segment the description text; the BERT word segmenter decomposes each description text into vocabulary units and converts them into corresponding word embedding vectors; the formula is as follows: T=BERT-Tokenizer(S) Among them, S represents the original description text, T = [t1, t2..t3] represents the word segmentation result, and each t i It is a token in the BERT tokenizer vocabulary. For texts longer than 128, truncation is performed. For texts less than 128, a special padding symbol [PAD] is used to fill the gap. Then, the tokenization result is converted into a 768-dimensional embedding vector through the BERT model: E=BERT-Encoder(T) The preprocessed urban scene images and the corresponding description texts are constructed as paired data; the paired data are input into a self-supervised learning framework for feature extraction and cross-modal alignment training; at the same time, the paired data set includes both positive samples and negative samples, the positive samples are formed by pairing the urban scene images with their corresponding description texts, and the negative samples are constructed by randomly disrupting the pairing relationship between the urban scene images and the description texts to form mismatched image-text combinations.

3. The small sample urban scene image analysis method based on BERT and self-supervised learning according to claim 2 is characterized in that: The specific pairing method of the pairing data is as follows: The positive samples are paired with city scene images and their corresponding description texts; the preprocessed image data I1, I2, ..., I n Mapped to the normalized image set I norm,1 ,I norm,2 ,...,I norm,n , describing text data S1,S2,...,S n is mapped into an embedding vector set E1, E2, ..., E n ; The paired data set consisting of positive samples is expressed as: D positive ={(I norm,1 ,E1),(I norm,2 ,E2),…,(I norm,n ,E n )} Negative samples are constructed by randomly disrupting the pairing relationship between urban scene images and description texts to form mismatched image-text combinations; the generation rule of negative samples is: keep the original image dataset I norm,1 ,I norm,2 ,...,I norm,n The order of the embedding vector set is randomly disrupted to form a new mismatched combination E rand,1 ,E rand,2 ,...,E rand,n , the paired data set composed of the final negative samples is expressed as: D negative ={(I norm,1 ,E rand,1 ),(I norm,2 ,E rand,2 ),…,(I norm,n ,E rand,n )} The negative sample ratio is set to 50%, and the final dataset is a mixture of positive and negative samples. Composition: D = D positive ∪D negative .

4. The small sample urban scene image analysis method based on BERT and self-supervised learning according to claim 1, characterized in that: The self-supervised learning framework includes an image feature extraction task and a cross-modal alignment task; the cross-modal information fusion module, as a part of the self-supervised learning framework, is used to fuse urban scene image features with description text features to generate cross-modal fusion features; The image feature extraction task includes a jigsaw reconstruction task and a contrastive learning task; the jigsaw reconstruction task and the contrastive learning task are both used to train the same pre-trained ResNet-50 model; the loss function in the training process includes a cross entropy loss for the jigsaw reconstruction task and a contrastive loss based on cosine similarity for the contrastive learning task, and the two loss functions jointly optimize the parameters of ResNet-50 with equal weights; The core idea of ​​the jigsaw puzzle reconstruction task is to randomly divide the input city scene image into nine small images and randomly shuffle the order; the pre-trained ResNet-50 model needs to predict the original position of each city scene image in this task to restore the overall spatial structure of the image. The loss function is designed to be a cross entropy loss, and its formula is as follows: Where N represents the number of training samples, M represents the number of possible locations of each urban scene image, and y ij is the distribution of true labels, Probability distribution predicted by the ResNet-50 model; The contrastive learning task is to generate two versions of enhanced images as positive sample pairs by performing different data enhancement operations on the same urban scene image, and randomly select other images to generate negative sample pairs; the two enhanced images are respectively encoded into high-dimensional feature vectors, and the distance of the positive sample pairs is calculated by cosine similarity to optimize the feature distribution of the pre-trained ResNet-50 model between positive and negative samples; the loss function of contrastive learning adopts contrastive learning loss based on temperature parameters, and its formula is: Among them, sim(z i ,z j ) represents the cosine similarity between two feature vectors, τ is the temperature parameter, which is set to 0.5; the feature dimension of contrastive learning is 128 and the batch size is 64.

5. The small sample urban scene image analysis method based on BERT and self-supervised learning according to claim 4 is characterized in that: In the cross-modal alignment task, the image-text matching task and the text generation task are combined to achieve semantic alignment of the urban scene image and the description text; The image-text matching task uses a designed binary classification model to determine whether the input image and the description text match. First, the image features are extracted by the pre-trained ResNet-50 model to obtain a 512-dimensional embedding vector; the text features are encoded by the BERT model to generate a 768-dimensional embedding vector. Image features and text features are fused through a cross-modal information fusion module to generate cross-modal fusion features; The cross-modal information fusion module dynamically allocates the weight of the description text to the image features through the cross-modal attention mechanism, fully integrates the spatial information of the image and the semantic information of the text, so as to improve the matching accuracy; finally, the matching probability distribution is output through the output layer composed of two layers of fully connected networks and Sigmoid, and the loss function adopts the cross entropy loss, and the formula is: Among them, y i represents the true label of the sample, The probability predicted by the binary classification model; The text generation task uses the Seq2Seq structure to fuse the cross-modal features F fusion Generate a corresponding diagnostic report; The Seq2Seq structure includes an encoder and a decoder. In the encoder, a two-layer bidirectional LSTM is used to fusion the cross-modal feature F fusion Further encoding is performed; the hidden layer dimension of BiLSTM is 256, each layer input is a 1024-dimensional fusion feature, and the output context representation and global hidden state are used to initialize the hidden state of the decoder; In the decoder, a two-layer unidirectional LSTM structure is used, and the dimension of each hidden layer is 256. The input of the decoder is the word generated in the previous time step. The initial input is the [START] tag. Combined with the context representation of the encoder, the context weight distribution of each time step is dynamically calculated through the attention mechanism to generate the semantic representation of the current time step and predict the probability distribution of the current word. The word embedding dimension is 300, and the Beam Search strategy is used in the decoding process. The training loss function of the Seq2Seq structure is the sequence cross entropy loss, and its formula is: Among them, T is the length of the generated sequence, w t is the generated word, P(w t |w1,w2,...w t-1 ) is the conditional probability of generating the next word.

6. The small sample urban scene image analysis method based on BERT and self-supervised learning according to claim 5, characterized in that: The cross-modal attention mechanism organically integrates the semantic information of the urban scene image and the description text by dynamically assigning the weight of the description text to the image features. The core formula of the cross-modal attention mechanism is as follows: Among them, Q, K, and V are query matrix, key matrix, and value matrix respectively. is the normalization factor of the feature dimension, and softmax is used to calculate the attention weight; the image features are used as the query matrix Q, and the text features are used as the key matrix K and the value matrix V, so as to calculate the weight distribution of the description text on the image features; after obtaining the attention weight, the image features are fused with the weighted text features, and the fusion formula is as follows: F fusion =W f ·[F img +Attention(Q,K,V)] Among them, F fusion represents the cross-modal fusion feature, F img is the original image feature, Attention(Q,K,V) is the weighted text feature output by the cross-modal attention mechanism, W f is the learnable linear transformation matrix of the fusion feature; the dimension of the cross-modal fusion feature is 1024 dimensions, and the cross-modal fusion feature is reduced to the target dimension through the fully connected layer. The dimension reduction formula is: F reduced =W r ·F fusion +b r Among them, W r and b r are the weight and bias of the fully connected layer respectively.

7. The small sample urban scene image analysis method based on BERT and self-supervised learning according to claim 1, characterized in that: The small sample optimization is specifically as follows: a layer-by-layer unfreezing fine-tuning strategy is designed to migrate the image feature extractor and text feature extractor trained in the self-supervised learning framework to a specific urban scene image analysis task by fine-tuning the pre-trained model parameters.

8. The small sample urban scene image analysis method based on BERT and self-supervised learning according to claim 7, characterized in that: The small sample optimization is specifically the following steps: (1) Initialize the model: load the pre-trained model parameters trained by the self-supervised learning framework, including the ResNet-50 model and the BERT model; (2) Freeze model parameters: In the early stage of fine-tuning, freeze most layers of the ResNet-50 model and the BERT model, and only unfreeze the high-level parameters; the specific settings are as follows: ResNet-50: freeze the first to third convolutional layers, and only unfreeze the fourth convolutional layer and the fully connected layer; BERT: Freeze the first 8 layers of Transformer encoder and only unfreeze the last 4 layers; (3) Layer-by-layer unfreezing: During fine-tuning, after a set number of training rounds, the low-level parameters of the ResNet-50 model and the BERT model are gradually unfrozen, and the frozen weights are released layer by layer, gradually adjusting from high-level features to global features. For the ResNet-50 model, the first to third convolutional layers are frozen in the initial stage, and only the fourth convolutional layer and the fully connected layer are unfrozen; when unfreezing layer by layer, the low-level convolutional parameters from the third convolutional layer to the first convolutional layer are unfrozen in turn; for the BERT model, the first 8 layers of Transformer encoders are frozen in the initial stage, and only the last 4 layers are unfrozen; when unfreezing layer by layer, the weights are gradually released starting from the 8th layer of Transformer encoder, and finally all encoding layers are unfrozen; (4) Adjusting the learning rate: During fine-tuning, the learning rates of different layers are set to different values, and the learning rate of low-level parameters is lower than that of high-level parameters; (5) Optimizing task loss: Optimizing the cross entropy loss of the classification task on small sample data while introducing regularization constraints.

Citation Information

Patent Citations

  • High-resolution remote sensing city image semantic segmentation method and system

    CN116310916A

  • SAR image urban area scene classification method based on multi-scale wavelet convolution

    CN117636052A

Cited By

  • Glioma postoperative radiotherapy prediction method based on iconomics of multi-modal MRI (Magnetic Resonance Imaging)

    CN120413056A

  • Deep learning prediction method for rock freeze-thaw damage based on text embedding

    CN120470950A

  • Disease prediction and auxiliary diagnosis system construction method and system based on multi-modal large model

    CN120565123A

  • A disease prediction and auxiliary diagnosis system construction method and system based on a multi-modal large model

    CN120565123B

  • Visual data rapid labeling method based on small samples

    CN121366327A