Feature extractor and concept for training and applying feature extractors

Fine-tuning feature extractors by adjusting them based on similarity matrices improves their adaptability to dynamic scenes, enhancing performance and generalizability in applications like autonomous driving and augmented reality.

GB2701549APending Publication Date: 2026-05-06CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH
Filing Date
2024-10-25
Publication Date
2026-05-06

AI Technical Summary

Technical Problem

Pre-trained neural feature encoders lack sensitivity and capacity to accurately model spatially and temporally evolving scenes, leading to subpar performance and limited generalizability in applications like autonomous driving and augmented reality.

Method used

A method for fine-tuning feature extractors by determining a connection matrix between images using a feature matching algorithm, transferring it to a feature space, and adjusting the extractor based on similarity matrices to reduce deviation, ensuring space and time consistency.

Benefits of technology

Enhances the ability of feature extractors to adapt to new data distributions, improving generalization and performance in dynamic environments, particularly in safety-critical applications such as autonomous driving and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for fine-tuning a machine learning encoder (feature extractor) 100, comprising: obtaining at least two images of a scene 110; using a feature matching algorithm to determine a relationship betw
Need to check novelty before this filing date? Find Prior Art

Description

Embodiments of the present disclosure relate to a feature extractor and a concept for training and applying feature extractor. In particular, the present disclosure relates to an approach for fine-tuning feature extractors and a respective training method. Feature extractors, also referred to herein as “feature encoders” or “encoders”, are versatile tools used in various fields of technology. In 3D reconstruction, they identify and extract important features from 2D images or point clouds to build accurate 3D models of objects or environments, which are essential in virtual reality, gaming, and cultural heritage preservation. In automotive applications, feature extractors may be integral to advanced driver-assistance systems (ADAS) and autonomous vehicles, helping detect and recognize objects such as pedestrians, vehicles, and road signs from sensor data, which is critical for navigation, collision avoidance, and ensuring safety. For dynamic scene reconstruction, feature extractors capture and model changes in a scene over time, which is particularly important in augmented reality, where the system needs to understand and integrate virtual objects into a real-world environment seamlessly, tracking the movement and transformation of objects within the scene to maintain a coherent and realistic experience. These applications highlight the importance of feature extractors in processing and interpreting complex data, enabling advanced functionalities in various technological domains. In the field of 3D reconstruction and dynamic scene reconstruction pre-trained neural feature encoders may be used. These encoders, which have been trained on large datasets for various tasks, may be employed without a thorough examination of their ability to handle scenes that change over time and space. This practice is widespread because it saves time and resources, allowing for quicker development cycles. However, this approach has significant drawbacks. Pre-trained neural feature encoders are designed to extract features from data, but their effectiveness in capturing the nuances of spatially and temporally evolving scenes is not always guaranteed. These scenes are complex and dynamic, requiring models that can adapt to changes in the environment over time. When pre-trained encoders are used without modification, they may lack the sensitivity and capacity to accurately model these changes. This can lead to a mismatch between the capabilities of the encoder and the requirements of the task at hand. The reliance on these off-the-shelf encoders means that new models built on top of them are expected to perform well in final tasks, such as reconstructing 3D scenes or understanding dynamic environments. However, if the foundational encoders are not well-suited to these tasks, the overall performance of the new models can suffer. This is because the encoders may not generalize well to the specific characteristics of the scenes being modeled. Generalizability is crucial in machine learning and computer vision, as it ensures that models can perform well across a variety of different scenarios and datasets. Ignoring the need for generalizability in the backbone networks can lead to several issues. First, the models may produce inferior results, failing to accurately capture the details and dynamics of the scenes. This can be particularly problematic in applications where precision is critical, such as autonomous driving, augmented reality, and robotics. Second, the lack of generalizability can limit the applicability of the models to new and unseen environments, reducing their usefulness in real-world scenarios. In summary, while using pre-trained neural feature encoders can expedite the development process, it is crucial to ensure that these encoders are capable of modeling the spatial and temporal dynamics of the scenes they are applied to. Failing to do so can result in models that perform poorly in end tasks, limiting their effectiveness and generalizability. By paying attention to the sensitivity and capacity of the backbone networks, researchers can develop more robust and accurate models for 3D reconstruction and dynamic scene reconstruction, ultimately leading to better performance and broader applicability in real-world applications. Hence, there may be a demand for an improved concept of feature extraction. This demand may be satisfied by the subject-matter of the appended independent claims. Optional embodiments are disclosed by the appended dependent claims. Embodiments of the present disclosure provide a method for fine-tuning a machine-learning-based feature extractor. The method comprises obtaining at least two images of a scene, determining a connection matrix between the images using a feature matching algorithm, transferring the connection matrix to a feature space of the feature extractor, determining a similarity matrix indicative of a similarity of features extracted by the feature extractor applied to the images, and adjusting the machine-learning-based feature extractor based on a comparison of the matrices such that a deviation of the matrices is reduced for the adjusted feature extractor. In this way, a space and time consistency of the feature extractor may be improved. So, the proposed approach provides a fine-tunning strategy for generating strong feature extractors which can be used for any visual perception task. In some embodiments, the images overlap at least partly with each other. In this way, comparability of the images may be provided. Optionally, obtaining the images comprises selecting the images from frames of a video or of an image series. Optionally, the images are selected with respect to a time period between recording the images such that the time period does not exceed a predefined maximum time period. The predefined maximum time period, e.g., is set such that the images at least partly overlap in view of a motion of a recording device for recording the images. For example, the time period is shorter if the recording device moves faster and vice versa. In this way, it may be ensured that the images at least partly overlap for a comparability of the images. In some embodiments, the method further comprises normalizing the connection matrix transferred to the feature space, and adjusting the machine-learning-based feature extractor based on a comparison of the matrices comprises adjusting the machine-learning-based feature extractor based on a comparison of the similarity matrix and the normalized connection matrix. In this way, comparability of the matrices may be provided. Optionally, the feature extractor is configured for a visual perception task. A skilled person will appreciate, that the feature extractor may be configured for any visual perception task such as object detection, (3D) scene reconstruction, and / or various other visual perception tasks (computer vision tasks). In particular, the feature extractor may be configured for a safety-critical visual perception task. In this case, the proposed training method may provide an improved feature extractor for a higher level of safety. In practice, the feature extractor, e.g., may be applied in safety-critical automotive applications like assisted driving, semi-autonomous driving, or autonomous driving. Further embodiments provide a machine-learning-based feature extractor which is obtainable by the training method proposed herein. As laid out above, such feature extractor may be applicable in different applications. Further embodiments provide a method for a visual perception task. The method comprises providing an embodiment of the proposed machine-learning-based feature extractor obtainable by an embodiment of the fine-tuning method proposed herein and applying the machine-learning-based feature extractor for a visual perception task. Optionally, applying the machine-learning-based feature extractor for a visual perception task comprises applying the machine-learning-based feature extractor for a visual perception task for assisted, semi-autonomous, or autonomous driving. Further embodiments provide a computer program which comprises instructions which, when the computer program is executed by a computer, cause the computer to carry out any one of the methods of the concept proposed herein and / or to provide a machine-learning-based feature extractor proposed herein. Still further embodiments provide an apparatus which comprises one or more interfaces for communication and a data processing circuit configured to carry out any one of the methods proposed herein and / or provide a machine-learning-based feature extractor proposed herein. Further, embodiments are now described with reference to the attached drawings. It should be noted that the embodiments illustrated by the referenced drawings show merely optional embodiments as an example and that the scope of the present disclosure is by no means limited to the embodiments presented: Brief description of the drawings Fig. 1 shows a flow chart schematically illustrating an embodiment of a method for fine-tuning a machine-learning-based feature extractor; Fig. 2 shows a flow chart schematically illustrating an embodiment of a method for computer vision; and Fig. 3 shows a block diagram schematically illustrating an apparatus according to the proposed approach. In the context of 3D reconstruction and dynamic scene reconstruction, many existing approaches rely on pre-trained neural feature encoders without evaluating their ability to handle spatially and temporally evolving scenes. This practice can be problematic because these encoders might not be sensitive or capable enough to accurately model such changes. Consequently, when new models are built on these pre-trained encoders, they are expected to perform well in final tasks. However, this expectation is flawed if the foundational encoders are not adequately generalized. This methodological flaw can result in subpar performance and outcomes. The present disclosures provides a solution addressing the above changes, as outlined in more detail below with reference to the appended drawings and exemplary embodiments of the proposed concept. Fig. 1 shows a flow chart schematically illustrating an embodiment of a method 100 for fine-tuning a machine-learning-based feature extractor. In context of the present disclosure, fine-tuning refers to the process of taking a pre-trained model (here: the feature extractor) and making (slight) adjustments to its parameters to adapt it to a specific task. This may involve training the model on a new, often smaller dataset that is relevant to the desired application. The pre-trained model, which has already learned general features from a large dataset, serves as a starting point. Fine-tuning allows the model to learn task-specific features without starting from scratch, making the process more efficient and leading to better performance. As can be seen from the flowchart, method 100 comprises obtaining 110 at least two images of a scene. In practice, e.g., a pair of images of the scene may be obtained. For the sake of simplicity, and examples described herein, the proposed approach may be applied to only two images. However, it is noted that the proposed approach may be analogously applied to an arbitrary number of images. The images may be specific for a desired application of the feature extractor. As mentioned above, the feature extractor may be supposed to be used for automotive applications. Accordingly, the images may be representative of a traffic scene, e.g., including one or more traffic participants (e.g., vehicles, cyclists, pedestrians, and / or the like) and / or infrastructure objects. For other applications, the images respectively may represent scenes specific for respective applications, for medical applications, e.g., images (e.g., X-ray images) of the human body. The images may be recorded using a camera or video camera. So, in practice, the images may be obtained from a sequence or plurality of images of the scene. The images, e.g., may be selected from frames of a video or of an image series. Further, method 100 comprises determining 120 a connection matrix between the images using a feature matching algorithm. In the context of the present disclosure, a connection matrix may be a mathematical representation indicating relationships between features detected in the images. Each element in the matrix indicates whether a pair of features, one from each image, is considered a match based on certain criteria such as similarity in appearance, spatial proximity, or other matching metrics. The matrix may be binary, where a value of 1 signifies a match and 0 signifies no match. The connection matrix may be a two-dimensional grid of binary values. Each row corresponds to a feature from a first image the images, and each column corresponds to a feature from a second image of the images. For a binary form of the connection matrix, elements of the matrix are either 0 or 1, where 1 indicates a match between the corresponding features and 0 indicates no match. The matrix may be sparse, meaning that most of the elements are 0, reflecting that only a few features from one image match features in the other. A skilled person will appreciate that various feature matching algorithms may be used for determining 120 the connection matrix. Examples of such feature matching algorithm comprise LightGlue and SuperGlue. Method 100 further comprises transferring 130 the connection matrix to a feature space of the feature extractor. In the context of the present disclosure, the term "feature space" refers to the multi-dimensional space where each dimension represents a distinct feature extracted from the images. When images are processed by a feature extractor, it transforms the raw input into a set of features, which may be numerical representations capturing essential characteristics of the images. These features are then plotted in the feature space, where each point corresponds to a data instance with its coordinates determined by the values of its features. The structure and distribution of points in this space can reveal patterns, clusters, or relationships within the data, aiding in tasks such as classification, clustering, and regression. Transferring 130 the connection matrix to the feature space may involve mapping the (binary relationships) indicated by the connection matrix to actual feature vectors. The connection matrix is, e.g., obtained on the full resolution of the image. The feature extractor usually downscales the image, if, e.g., a vision transformer architecture is used for the feature extractor. In practice, for example the extracted features are of size H / 14 (i.e., the height is divided by 14), W / 14 (i.e., the width is divided by 14). So, one can define a grid for the connection matrix where every patch, e.g., of 14x14, in the original connection matrix is mapped to the feature space. So, transferring the connection matrix to the feature space, e.g., comprises downscaling the connection matrix to a resolution of the feature extractor. In doing so, e.g., an average value is determined for connection matrix entries of the patches and the average value is mapped to a respective patch in the feature space. Further, method 100 comprises determining 140 a similarity matrix (may be also referred to as “predicted connection matrix”) indicative of a similarity of features extracted by the feature extractor when it is applied to the images. The similarity matrix, e.g., includes pairwise similarities between extracted features of the images. The pairwise similarities may be calculated between feature vectors generated by the feature extractor when applied to the images. Columns of the similarity matrix may correspond to feature vectors of a first image of the images and lines of similarity matrix may correspond to feature vectors of a second image of the images. For representing the similarity in the similarity matrix, various similarity measures may be applied. Exemplary similarity measures include cosine similarity, Euclidean distance, or Pearson correlation. For instance, cosine similarity measures the cosine of the angle between two feature vectors. In some embodiments, the similarities may assume a value between -1 and 1, where 1 indicates identical orientation. The similarities may be organized into a matrix to obtain similarity matrix. Each element of this matrix may represent the similarity between a pair of feature vectors, e.g., with the diagonal elements being 1, as they represent the similarity of each vector with itself. This matrix can then be used for various tasks such as clustering, classification, or visualization. Further, method 100 comprises adjusting 150 the machine-learning-based feature extractor based on a comparison of the matrices such that a deviation of the matrices is reduced for the adjusted feature extractor. In doing so, the transferred connection matrix may be used as ground truth for the similarity matrix in the fine-tuning. In practice, the feature extractor may be adjusted based on a deviation of the matrices. In embodiments, e.g., a loss function or loss term may be used to determine a deviation. For adjusting 150 the feature extractor, e.g., parameters of the feature extractor may be adapted. To this end, the gradient descent may be applied to minimize or at least reduce the deviation. By backpropagating the error through the network, the parameters of the feature extractor are updated to reduce the loss term / loss function L. For determining the deviation, the transferred connection matrix may be compared to the similarity matrix, the deviation, e.g., is indicative of a loss term, e.g., loss term L = -p ■ logp as a loss term to optimize the network with proper temperature scaling where p is an average prediction value of the ground truth (connection matrix), and p is an average prediction value of the feature extractor applied to the image. For multiple patches, the deviation may be summed over all (matched or mapped) patches (wherein the patches refer to embeddings returned by the feature extractor). This may be applied in an iterative process which may continue until the feature extractor converges to an improved or ideally optimal set of parameters, enhancing its ability to extract relevant features for a desired task. In doing so, the proposed fine-tuning allows the feature extractor to adapt to new data distributions and improve its generalization capabilities. The proposed approach, thus, provides a stronger feature extractor then other fine-tuning approaches not considering features similarities and feature matches as proposed herein. In practice, the images may overlap at least partly with each other for a better comparability of the images and, consequently, a better comparability of the matrices. That is, the images at least partly represent the same part of the scene, e.g., similar objects and / or a similar area of the scene. In other words, afield-of-view, space, and / or area represented by the images may at least partly overlap or even match. The images, e.g., may be recorded (approximately) simultaneously using a multi-camera system. Alternatively, they may be recorded at different points in time, e.g., through continuous shooting or using a video camera. In this case, considering a motion of a recording device (camera, video camera), images recorded within a predefined (short) time period may be selected to ensure that they (sufficiently) overlap with each other (e.g., in terms of area or space captured by the images). For this, the images may be selected with respect to a time period between recording the images such that the time period does not exceed a predefined maximum time period. In doing so, the maximum time period may depend on a relative motion between the recording device and the environment, e.g., on how fast the recording device is moving and / or rotating or expected to move and / or rotate. For this, the motion may be measured or expected motion may be predefined. Accordingly, the time period may be shorter if the recording device is moving or expected to move faster and vice versa. The time period may be specified as a duration or (for a video camera or continuous shooting) by a number of frames between the images. To further improve the comparability of the matrices, the method may further comprise normalizing the connection matrix transferred to the feature space. For example, for some loss terms such as the one specified above, values of the connection matrix may need to be greater than 0. So, the method may include a normalization to make sure that the connection matrix values are greater than 0. For this, e.g., a uniform normalization between 0 and the maximum possible value of the similarity matrix may be applied. Accordingly, the machine-learning-based feature extractor may be adjusted based on a comparison of the similarity matrix and the normalized connection matrix. As the skilled person will appreciate, the proposed approach may be applied for fine-tuning the feature extractor for different applications. In particular, the feature extractor may be fine-tuned for automotive applications such as assisted or (semi-) autonomous driving where the feature extractor may be applied for perception of an environment of a vehicle. Accordingly, embodiments of the present disclosure provide a machine-learning-based feature extractor obtainable by proposed fine-tuning method. As the skilled person will appreciate, the proposed fine-tuning method may consider different points in time and / or different perspectives of a scene during the fine-tuning. As a result, predictions of the fine-tuned feature extractor may be more consistent over time and space. In other words, the predictions may fluctuate less than for other fine-tuning approaches which do not consider changes over time and space during training. In practice, e.g., a classification result may fluctuate less between different object classes. Further embodiments provide a method for computer vision, as outlined in more detail below with reference to Fig. 2. Fig. 2 shows a flow chart schematically illustrating an embodiment of a method 200 for computer vision. As can be seen, method 200 comprises providing 210 an embodiment of the proposed machine-learning-based feature extractor. Such feature extractor, e.g., may be provided by a respective computer program, as laid out in more detail later. Further, method 200 comprises applying 220 the machine-learning-based feature extractor for a visual perception task. For this, the feature extractor may be executed or ran on a computer or any other programmable hardware, as laid out in more detail later. As mentioned above, the proposed fine-tuned feature extractor provides more consistent predictions in terms of time and space. In practice, this may lead to better results of computer vision tasks utilizing such feature extractor. In particular, this may achieve a higher reliability and stability of computer vision applications. In safety-critical applications, such as assisted driving, (semi-) autonomous driving, or medical applications of computer vision (e.g., for visual perception tasks), the proposed approach particularly may lead to a higher level of safety. The proposed approach may be similarly applied for an apparatus, as laid out in more detail below with reference to Fig. 3. Fig. 3 shows a block diagram schematically illustrating an embodiment of such an apparatus 300. The apparatus comprises one or more interfaces 310 for communication and a data processing circuit 320 configured to execute the proposed method. In embodiments, the one or more interfaces 310 may comprise wired and / or wireless interfaces for transmitting and / or receiving communication signals in connection with the execution of the proposed concept. In practice, the interfaces, e.g., comprise pins, wires, antennas, and / or the like. As well, the interfaces may comprise means for (analog and / or digital) signal or data processing in connection with the communication, e.g., filters, samples, analog-to-digital converters, signal acquisition and / or reconstruction means as well as signal amplifiers, compressors and / or any encryption / decryption means. The data processing circuit 320 may correspond to or comprise any type of programable hardware. So, examples of the data processing circuit 320, e.g., comprise a memory, microcontroller, field programable gate arrays, one or more central, and / or graphical processing units. To execute the proposed method, the data processing circuit 320 may be configured to access or retrieve an appropriate computer program for the execution of the proposed method from a memory of the data processing circuit 320 or a separate memory which is communicatively coupled to the data processing circuit 320. In practice, the proposed apparatus may be installed on a vehicle. So, embodiments may also provide a vehicle comprising the proposed apparatus. In implementations, the apparatus, e.g., is part or a component of an assisted and / or (semi-) autonomous driving system. However, in implementations, computing resources for the vehicle may be outsourced to an external server separate from the vehicle. In such implementations, the proposed approach may be also implemented outside of the vehicle. In the foregoing description, it can be seen that various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed examples require more features than are expressly recited in each claim. Rather, as the following claims reflect, subject matter may lie in less than all features of a single disclosed example. Thus, the following claims are hereby incorporated into the description, where each claim may stand on its own as a separate example. While each claim may stand on its own as a separate example, it is to be noted that, although a dependent claim may refer in the claims to a specific combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of each other dependent claim or a combination of each feature with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a specific combination is 5 not intended. Furthermore, it is intended to include also features of a claim to any other independent claim even if this claim is not directly made dependent to the independent claim. Although specific embodiments have been illustrated and described herein, it will 10 be appreciated by those of ordinary skill in the art that a variety of alternate and / or equivalent implementations may be substituted for the specific embodiments shown and described without departing from the scope of the present embodiments. This application is intended to cover any adaptations or variations of the specific embodiments discussed herein. Therefore, it is intended that the 15 embodiments be limited only by the claims and the equivalents thereof.

Claims

1. A method (100) for fine-tuning a machine-learning-based feature extractor, the method (100) comprising:obtaining (110) at least two images of a scene;determining (120) a connection matrix between the images using a feature matching algorithm;transferring (130) the connection matrix to a feature space of the feature extractor;determining (140) a similarity matrix indicative of a similarity of features extracted by the feature extractor applied to the images; andadjusting (150) the machine-learning-based feature extractor based on a comparison of the matrices such that a deviation of the matrices is reduced for the adjusted feature extractor.

2. The method (100) of claim 1, wherein the images overlap at least partly with each other.

3. The method (100) of claim 1 or 2, wherein obtaining the images comprises selecting the images from frames of a video or of an image series.

4. The method (100) of claim 3, wherein the images are selected with respect to a time period between recording the images such that the time period does not exceed a predefined maximum time period.

5. The method (100) of any one of the preceding claims, wherein the method (100) further comprises normalizing the connection matrix transferred to the feature space, and wherein adjusting the machine-learning-based featureextractor based on a comparison of the matrices comprises adjusting the machine-learning-based feature extractor based on a comparison of the similarity matrix and the normalized connection matrix.

6. The method (100) of any one of the preceding claims, wherein the feature extractor is configured for a visual perception task.

7. A machine-learning-based feature extractor obtainable by a method according to any one of the preceding claims.

8. A method (200) for computer vision, wherein the method comprises:providing (210) a machine-learning-based feature extractor according to claim 7; andapplying (220) the machine-learning-based feature extractor for a visual perception task.

9. The method (200) of claim 8, wherein applying the machine-learning-based feature extractor for a visual perception task comprises applying the machine-learning-based feature extractor for a visual perception task for assisted, semi-autonomous, or autonomous driving.

10. A computer program comprising instructions which, when the computer program is executed by a computer, cause the computer to carry out any one of the methods (100, 200) of any one of the claims 1 to 6, 8, and 9 and / or provide a machine-learning-based feature extractor according to claim 7.

11. An apparatus (300) comprising:one or more interfaces (310) for communication; anda data processing circuit (320) configured to carry out any one of the methods (100, 200) of any one of the claims 1 to 6, 8, and 9 and / or provide a machine-learning-based feature extractor according to claim 7.

Citation Information

Patent Citations

  • Near space remote sensing image registration method based on self-supervision

    CN115082533A

  • Panoramic image stitching method and system

    CN117094895A

  • Localization based on neural networks

    WO2024099593A1