Image feature extraction method based on multi-scale time and space contrast loss optimization
Through the multi-scale time and space comparison loss optimization image feature extraction method, the problem of difficult to distinguish dynamic and static factors in the prior art is solved, and accurate feature extraction and distinction of complex environmental images is achieved.
Patent Information
- Application Number
- CN202510106616.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to accurately and efficiently distinguish dynamic and static factors, especially when dealing with complex urban environment images.
The image feature extraction method optimized by multi-scale time and spatial contrast loss is adopted. The image data is processed through feature encoder and momentum encoder, and combined with time interval weights and distance weights, multi-scale time contrast loss and spatial consistency contrast loss are constructed to achieve the distinction between dynamic and static factors.
It realizes accurate and efficient distinction between dynamic and static factors in the image, improves the effect of image feature extraction, and is suitable for downstream tasks in complex environments.
Smart Images

Figure CN119992115A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an image feature extraction method with multi-scale time and space contrast loss optimization. Background Art
[0002] Current unsupervised learning focuses on encoding as much semantic and structural information as possible in the scene, which is not applicable to downstream tasks in all fields, such as street-view-based urban environment understanding. Taking the place recognition task as an example, features need to focus on location-invariant information, such as buildings and roads, and filter out dynamic information such as lighting conditions, pedestrians, vehicles, and vegetation. In tasks related to human perception of places, these dynamic factors are very important. In addition, tasks such as socioeconomic forecasting emphasize spatially consistent expressions between adjacent scenes.
[0003] In image representation learning, selectively encoding dynamic and static information in urban environments and the atmosphere they create is crucial but inherently challenging. Accurate encoding of this information usually requires labeling dynamic and static factors separately and adopting specific training strategies, such as masking dynamic factors when encoding static factors. However, both the labeling and training processes are fraught with difficulties. Factors such as lighting conditions, vegetation appearance, and ground debris are difficult to label objectively and consistently, making it almost impossible to accurately represent these complex environmental factors using existing datasets and methods, whether supervised or unsupervised. That is, the existing technology cannot accurately and efficiently distinguish between dynamic and static factors. Summary of the invention
[0004] In response to the problems existing in the prior art, the present invention provides an image feature extraction method with multi-scale temporal and spatial contrast loss optimization, which at least partially solves the problem in the prior art that dynamic factors and static factors cannot be accurately and efficiently distinguished.
[0005] The embodiment of the present invention provides an image feature extraction method with multi-scale temporal and spatial contrast loss optimization, comprising: The image data obtained according to the set requirements is processed by the corresponding feature encoder to obtain a positive sample; The image data obtained according to the set requirements is processed by the momentum encoder to obtain negative samples; The positive samples and negative samples are input into the corresponding loss function to obtain the constructed sample pairs.
[0006] Optionally, the setting requirements include: Set requirements according to time relationships and set requirements according to space relationships.
[0007] Optionally, the step of obtaining a positive sample by processing the image data acquired according to the set requirements with a corresponding feature encoder includes: The image data acquired according to the time relationship setting requirements is processed by the time encoder to obtain positive samples; The image data obtained according to the spatial relationship setting requirements is processed by the spatial encoder to obtain positive samples.
[0008] Optionally, the positive samples and negative samples are input into the corresponding loss function, and time interval weights are introduced so that the weights of sample pairs with smaller time intervals are higher, while the weights of sample pairs with larger time intervals are lower.
[0009] Optionally, the time interval weight , in, is the time interval, is a constant.
[0010] Optionally, the image data required to be acquired according to the time relationship setting includes: Divide time sampling into two scales: short-term and long-term; Select image pairs separated by a short time to construct short-term positive sample pairs and exclude short-term dynamic factors; Select image pairs with a larger time interval to construct long-term positive sample pairs and capture more stable structural changes.
[0011] Optionally, the step of inputting the positive samples and the negative samples into corresponding loss functions includes: The short-term and long-term temporal contrastive losses are combined to define the multi-scale temporal contrastive loss.
[0012] Optionally, the short-term and long-term temporal contrast losses are combined to define a multi-scale temporal contrast loss, including: , in: It is the short-term time contrast loss, is the long-term time contrast loss, and are the weight parameters respectively.
[0013] Optionally, the positive samples and negative samples are input into corresponding loss functions, distance weights are introduced, and image pairs at adjacent positions are selected in a weighted manner.
[0014] Optionally, the distance weight , in, Represents the spacing, which is a constant in the spacing weight.
[0015] The multi-scale time and space contrast loss optimized image feature extraction method provided by the present invention obtains images according to set requirements, and processes the images by corresponding encoders and momentum encoders to obtain positive and negative samples respectively. After the positive and negative samples are processed by the loss function, static factors are retained and dynamic factors are excluded. Images can be processed in batches, thereby achieving the purpose of accurately and efficiently distinguishing dynamic factors from static factors. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present invention.
[0017] Figure 1 This is a principle framework diagram of the image feature extraction method with multi-scale temporal and spatial contrast loss optimization disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0019] It should be clear that the following embodiments of the present invention are described by specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without making creative work belong to the scope of protection of the present invention.
[0020] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. It should be understood by those skilled in the art based on the invention that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.
[0021] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. The illustrations only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0022] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.
[0023] MoCov3, short for Momentum Contrast v3, is an improved method in the field of self-supervised learning, which aims to achieve efficient image feature representation learning through contrastive learning. MoCo v3 generates stable feature representations through momentum encoders, and combines the visual Transformer (ViT) architecture to further enhance the model's ability to capture image semantics and contextual information.
[0024] In MoCo v3, the query and key mechanism is still used to generate positive and negative sample pairs, where the positive sample pairs are different views of the same image generated by applying different data enhancements (such as cropping, color perturbation, etc.), and the negative sample pairs come from other images. MoCov3 uses an improved version of the InfoNCE (Information Noise-Contrastive Estimation) loss function to optimize feature representation by minimizing the distance of positive sample pairs and maximizing the distance of negative sample pairs.
[0025] A significant improvement of MoCo v3 is that it abandons the dynamic negative sample queue mechanism used in the previous two versions and instead uses samples in small batches to generate negative sample pairs, which simplifies the training process and reduces computational overhead. At the same time, the momentum encoder in MoCo v3 uses exponential moving average (EMA) to update parameters, maintaining synchronization stability with the online encoder to ensure a more robust training process. These improvements enable MoCo v3 to perform well in downstream tasks such as image classification and target detection, and is particularly suitable for feature learning tasks on large-scale unlabeled datasets.
[0026] For ease of understanding, Figure 1 As shown, this embodiment discloses an image feature extraction method with multi-scale time and space contrast loss optimization, including: The image data obtained according to the set requirements is processed by the corresponding feature encoder to obtain a positive sample; The image data obtained according to the set requirements is processed by the momentum encoder to obtain negative samples; The positive samples and negative samples are input into the corresponding loss function to obtain the constructed sample pairs.
[0027] Optionally, the setting requirements include: Set requirements according to time relationships and set requirements according to space relationships.
[0028] Optionally, the step of obtaining a positive sample by processing the image data acquired according to the set requirements with a corresponding feature encoder includes: The image data acquired according to the time relationship setting requirements is processed by the time encoder to obtain positive samples; The image data obtained according to the spatial relationship setting requirements is processed by the spatial encoder to obtain positive samples.
[0029] The real world is constantly changing in both time and space, but these changes show a certain continuity. In the time dimension, it is important to capture the constant features of a place that evolve over time, called time-invariant features; while in the space dimension, the focus is on maintaining the consistency of the overall atmosphere within a specific spatial range, called space-invariant neighborhood atmosphere. These temporal and spatial invariances are crucial to improving the performance of various downstream tasks.
[0030] Contrastive learning aims to learn feature representations from unlabeled data by comparing positive and negative samples. Its main goal is to minimize the distance between positive samples in the feature space while maximizing the distance of negative samples. Positive sample pairs are constructed by enhancing street view images. By optimizing the InfoNCE loss function, the model learns to reduce the distance between positive sample pairs in the feature space and increase the distance with negative samples, thereby improving the effect of feature representation learning. Given a query representation and a set of positive and negative sample keys , InfoNCE (Information Noise-Contrastive Estimation). The loss is defined as:
[0031] (1), here, represents the feature representation of the query, represents the feature representation of the positive sample, Represents the feature representation of negative samples. Temperature parameter Used to control the scaling of similarity. The goal is to maximize the query and positive sample keys The similarity between the query and the negative sample key is minimized The similarities between.
[0032] Based on the contrastive learning framework, temporal and spatial contrasts are introduced to construct positive sample pairs from street view images.
[0033] like Figure 1 As shown in Figure 2, street view images are used for spatial and temporal contrastive learning. The contrastive learning task uses street view images taken at the same location over time to learn the time-invariant features of the built environment; using spatially similar street view images from the same period, the learning task aims to learn the spatially invariant community atmosphere, such as the socioeconomic atmosphere.
[0034] Comparison based on multi-scale time: Street view images taken at different times at the same location are different from video frames because the interval between their shooting is not fixed. Unlike remote sensing images, street view images taken at different times will not be completely aligned in position. Since images taken at different times are usually offset in space and angle, certain restrictions are imposed on the conditions of positive time sample pairs: they must be taken within a range of no more than 5 meters from each other and at the same shooting angle. The historical street view image set for each location can be represented as ,in represents images taken at different times. Since the number of images at each location is different, the value of n is also different. The purpose of time comparison is to capture the constant features of the same location over time. This means that even if the images are taken at different times, the model should learn to recognize the consistent features of the scene. To this end, we define a positive sample pair that satisfies the above time conditions .
[0035] Time interval weight: In the temporal contrast loss, we can consider introducing the weight of the time interval so that the sample pairs that are closer in time have higher weights, while the sample pairs with a larger time interval have lower weights.
[0036] Specifically, and are images taken at the same location at different times, and the shooting times are and . Define the time interval and use , as the weight in time comparison, is a small constant used to avoid division by zero.
[0037] (2), is the feature representation of the image taken at the same time (positive sample); are image features from other irrelevant times (negative samples); is the temperature parameter.
[0038] When constructing positive sample pairs with a larger time interval, the position, quantity, and shape of some short-term dynamic factors (such as pedestrians or vehicles) may vary greatly due to the large time span. However, the constant objects in the scene (such as buildings, trees, etc.) remain consistent. In the process of learning this invariance, the model will pay more attention to the characteristics of these constant objects and gradually ignore the dynamic factors with large changes.
[0039] Multi-scale temporal contrast loss: Since the temporal sampling at different locations may be sparse, a multi-scale temporal contrast loss is designed, that is, positive sample pairs are selected at different time intervals, so that the model can learn more robust time-invariant features.
[0040] Specifically, time sampling can be divided into two scales: short-term and long-term: Short-term temporal comparison: Select image pairs that are separated by a short period of time (such as a few months or a year) to construct short-term positive sample pairs. This helps the model ignore short-term dynamic factors (such as vehicles and pedestrians).
[0041] Long-term temporal comparison: Select image pairs with a large time interval (such as comparisons between several years) to construct long-term positive sample pairs to capture more stable structural changes (such as changes in architectural style or modifications to infrastructure).
[0042] The multi-scale temporal contrast loss is defined by combining the short-term and long-term temporal contrast losses. It is defined as follows:
[0043] (3), in: is the short-term temporal contrast loss, calculated using Equation 2 using image pairs with a short time interval.
[0044] is the long-term temporal contrast loss, calculated using Equation 2 using image pairs with a longer time interval.
[0045] and is a weight parameter used to balance the impact of short-term and long-term time comparisons.
[0046] The multi-scale temporal based contrastive loss has the following characteristics: Short-term comparison can weaken the impact of short-term dynamic factors In short-term comparisons (such as time intervals of several months or a year), the model assigns high weights to sample pairs within a short period of time. These sample pairs may contain some slightly changing dynamic factors, such as occasional vehicles or pedestrians. Due to the short time interval, these dynamic factors change less in a short period of time, and their influence is not persistent relative to constant objects such as buildings. Therefore, the model gradually learns to ignore these slightly changing dynamic factors during the short-term comparison process and focus on those core features that remain unchanged (such as streets and buildings). The high weight mechanism of short-term comparison effectively guides the model to focus on those constant factors that are consistent over a short period of time, weakening the impact of short-term dynamic factors.
[0047] Long-term comparison can reduce the interference of dynamic factors on core features In long-term comparisons (such as intervals of several years or even longer), the time span is large, and the appearance of dynamic factors (such as pedestrians, vehicles, billboards, etc.) is highly random. Dynamic factors in different time periods are unlikely to appear in the same location or in the same form. As the time interval increases, the similarity between these dynamic factors is greatly reduced, and the model gradually ignores their inconsistencies. Instead, the model will focus more on scene features that remain stable over large time spans, such as building structures, street layouts, and other unchanging physical environments. Long-term comparisons help the model automatically filter out dynamic factors with higher randomness in the time span, gradually capture the core and long-term stable features of the scene, and further reduce the interference of dynamic factors.
[0048] The time interval weighting can balance the attention of short-term and long-term comparisons and gradually filter dynamic factors. Time interval weight Higher weights are given to short-term sample pairs, while relatively lower weights are given to long-term sample pairs. The effect of this design is that when the time span is small, the model will pay more attention to short-term consistency features, and as the time span increases, the model's attention to invariant features gradually increases. The time interval weight enables the model to first adapt to slight changes in a short period of time in short-term comparisons, and gradually amplify the influence of constant features in long-term comparisons, so that the model can gradually focus on invariant features on both short-term and long-term scales, and weaken the influence of dynamic factors. The time interval weight balances the effects of short-term and long-term comparisons, allowing the model to selectively learn short-term consistency and long-term stability, and gradually filter out dynamic factors in time changes.
[0049] Gravity-weighted spatial consistency comparison: In order to accurately represent the urban physical environment, it is crucial to capture the spatial consistency within an urban area. Spatial consistency refers to the ability to identify the same underlying physical features at different locations in the same area. To achieve this goal, all street view images taken in the same area are considered to represent the same environment, even if they are taken from different angles or slightly different locations. This approach allows the model to account for changes in location while maintaining the overall atmosphere of the area. A set of street view images of a given urban area can be represented as , where each denotes images taken in the area. Together, these images provide a comprehensive spatial representation of the urban environment. Randomly select two images from the set S This encourages the model to learn that these images belong to the same spatial context even if there are slight changes in shooting angle or position.
[0050] In spatial comparison, a distance weight is introduced to enhance the learning of spatial proximity. In this way, image pairs at adjacent positions can be selected weightedly, and the closer the image pairs are, the higher the weight is, simulating the effect of the "gravity model", thereby enhancing the model's learning of spatial consistency.
[0051] Given a collection of images of the same area , assuming that each image The shooting location is . Defines the distance between two locations .
[0052] The gravity-weighted spatial consistency contrast loss is defined as follows: (4), is the distance weight, simulating the attenuation characteristics of gravity. is a small constant used to avoid division by zero.
[0053] q is the feature representation of the query image.
[0054] It is the feature representation of the image taken in the same area (positive sample).
[0055] are image features from other irrelevant regions (negative samples).
[0056] is the temperature parameter.
[0057] The gravity-weighted spatial consistency contrast loss has the following characteristics: Proximity preference and local consistency capture: Distance weighting based on gravity model This makes the model more inclined to focus on the similarities of adjacent image pairs during the learning process. In other words, the model assigns higher weights to image pairs that are closer, and lower weights to image pairs that are farther away. The model is more inclined to focus on the similarities of image pairs in nearby positions and capture common features in the same spatial region. The higher the weight, the more the model tends to identify stable features at these positions, while distant image pairs have less impact on model learning due to their lower weights, thereby effectively suppressing the inconsistency caused by long distances.
[0058] Extension of consistency from local to global: Through multiple sets of overlapping image pairs of adjacent locations, the model can gradually capture the overall consistency of the entire area. In real environments, adjacent locations often have similar physical features (such as architectural style and street layout), and when the model learns these overlapping image pairs, it can gradually build a complete feature representation of the area, expanding from local consistency to overall consistency.
[0059] Negative sample discrimination in contrastive learning: Introducing negative samples (images from other regions) into the loss function further enhances the model’s sensitivity to the features of the current region. By comparing positive samples at nearby locations with negative samples at a distance, the model can more accurately capture the spatial consistency features of the current region and learn how to distinguish the differences between different regions.
[0060] Through the above optimization strategy, the gravity-weighted spatial consistency contrast loss can effectively help the model learn spatial consistency features, thereby accurately characterizing the spatial atmosphere of urban areas.
[0061] 1. A framework based on spatiotemporal contrastive learning is proposed to learn temporal and spatial invariant features from street view images.
[0062] 2. By performing temporal contrastive learning between images taken at different times at the same location, the model is able to capture the time-invariant characteristics of the built environment and is suitable for tasks that rely on temporal stability.
[0063] 3. By performing spatial contrast learning between images taken at different locations at the same time, the model is able to capture the spatial consistency of urban areas and is suitable for tasks that require spatial consistency, such as socioeconomic forecasting.
[0064] Using multi-scale temporal contrast loss, short-term contrast can help the model adapt to slight changes in a short period of time, while long-term contrast can further remove occasional dynamic factors in the scene.
[0065] This embodiment achieves the ability to learn temporal and spatial invariant features from street view images by introducing a spatiotemporal contrast learning framework. Compared with the prior art, the advantage of the present invention is that it can automatically extract temporal and spatial consistency features in the architectural environment, thereby improving the performance of downstream tasks. For example, the temporal invariant representation can effectively retain the encoding of static factors such as buildings and streets, and automatically filter out the information of dynamic factors; the spatial invariant representation can encode the overall neighborhood atmosphere within a specific spatial range without focusing on any specific factors. This enables the model to have stronger generalization and robustness when dealing with tasks such as visual place recognition and socioeconomic forecasting.
[0066] That is, in a scene, only fixed buildings need to be considered, without considering pedestrians and vehicles around the buildings. Fixed buildings are static factors, while people and vehicles are dynamic factors.
[0067] The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details disclosed above are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.
[0068] In the present invention, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Words such as "including", "comprising", "having", etc. are open words, meaning "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.
[0069] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0070] It should also be noted that in the system and method of the present invention, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present invention.
[0071] Various changes, substitutions, and modifications of the techniques described herein may be made without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of the present invention is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and actions described above. Currently existing or later to be developed processes, machines, manufactures, compositions of events, means, methods, or actions that perform substantially the same functions or achieve substantially the same results as the corresponding aspects described herein may be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or actions within their scope.
[0072] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present invention. Therefore, the present invention is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0073] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present invention to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A multi-scale temporal and spatial contrast loss optimized image feature extraction method, characterized in that: include: The image data obtained according to the set requirements is processed by the corresponding feature encoder to obtain a positive sample; The image data obtained according to the set requirements is processed by the momentum encoder to obtain negative samples; The positive samples and negative samples are input into the corresponding loss function to obtain the constructed sample pairs.
2. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 1, characterized in that: The setting requirements include: Set requirements according to time relationships and set requirements according to space relationships.
3. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 2 is characterized in that: The image data acquired according to the set requirements is processed by a corresponding feature encoder to obtain a positive sample, including: The image data acquired according to the time relationship setting requirements is processed by the time encoder to obtain positive samples; The image data obtained according to the spatial relationship setting requirements is processed by the spatial encoder to obtain positive samples.
4. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 3 is characterized in that: The positive samples and negative samples are input into the corresponding loss function, and the time interval weight is introduced, so that the sample pairs with smaller time intervals have higher weights, while the sample pairs with larger time intervals have lower weights.
5. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 4 is characterized in that: The time interval weight , in, is the time interval, is a constant.
6. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 3, characterized in that: The image data acquired according to the time relationship setting requirement includes Divide time sampling into two scales: short-term and long-term; Select image pairs separated by a short time to construct short-term positive sample pairs and exclude short-term dynamic factors; Select image pairs with a larger time interval to construct long-term positive sample pairs and capture more stable structural changes.
7. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 6, characterized in that: The inputting of positive samples and negative samples into corresponding loss functions includes: The short-term and long-term temporal contrastive losses are combined to define the multi-scale temporal contrastive loss.
8. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 7, characterized in that: The multi-scale time contrast loss is defined by combining short-term and long-term time contrast losses, including: , in: It is the short-term time contrast loss, is the long-term time contrast loss, and are the weight parameters respectively.
9. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 3, characterized in that: The positive samples and negative samples are input into the corresponding loss function, distance weights are introduced, and image pairs at adjacent positions are selected in a weighted manner.
10. The image feature extraction method with multi-scale temporal and spatial contrast loss optimization according to claim 9, characterized in that: The distance weight , in, is the spacing, which is a constant.