SYSTEM AND METHOD FOR TRAINING MODELS USING LOCALIZED TEXT MANAGEMENT - Patent application
By combining the pre-training method of self-supervised comparative losses and supervised losses, the problem of expensive annotation data and training time required for pre-training of computer vision models is solved, and effective performance improvement and training cost reduction for Localization tasks are achieved.
Patent Information
- Application Number
- JP2022041745
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-18
- Filing Date
- 2022-03-16
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-16
AI Technical Summary
The prior art requires expensive annotation training data when pre-training the model of computer vision tasks, and the self-supervised pre-training method requires long-term training schedules and is not effective for Localization tasks.
The model is trained by using a comparative pre-training framework through a combination of self-supervised comparison loss function and supervised loss function. The specific steps include: using the self-supervised contrast loss function to calculate the contrast loss based on the image feature map and the subtitle feature vector, and calculating the positioning loss based on the comparison between the image subtitle attention map and the visual identifier, thereby adjusting the weight of the visual and text backbone network.
Through this method, it is possible to effectively pre-train the model without the need for expensive annotation data, improve the performance of Localization tasks, and reduce training costs and time.
Smart Images

Figure 0007673673000009 
Figure 0007673673000010 
Figure 0007673673000011
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 161,686, entitled “LocTex: Learning Data-Efficient Visual Representations from Localized Text Management,” filed March 16, 2021, the entire contents of which are incorporated herein by reference.
[0002] The subject matter described herein relates generally to systems and methods for training models, and more particularly, to systems and methods for pre-training models for use in computer vision tasks. [Background technology]
[0003] The background art description provided generally presents the context of the disclosure. The inventors' work to the extent that it can be described in this background art section, and aspects of the description that may not qualify as prior art at the time of filing, are not admitted, expressly or impliedly, to be prior art against the present technology.
[0004] Neural networks, such as convolutional neural networks (CNNs), have been utilized to perform computer vision tasks such as object detection and semantic / instance segmentation. These neural networks must first be trained to successfully complete the computer vision task. Training these neural networks may involve pre-training and fine-tuning the neural network to reduce the need for costly annotations. Furthermore, in one example, the CNN backbone must first be pre-trained to perform a specific task. The learned features can then be transferred to other downstream tasks by fine-tuning the neural network using the target dataset.
[0005] However, pre-training still requires annotated training data that can be very costly to obtain, and pre-training for classification tasks can be ineffective for tests that are more sensitive to localization than classification. Efforts to solve these problems include pre-training neural networks with coarse, freely available labels such as metadata or hashtags, or self-supervised pre-training, which learns visual representations from unlabeled images. However, these solutions also have drawbacks. For example, pre-training with coarse labels remains ineffective for those tasks that are more sensitive to localization than classification. As for self-supervised pre-training, these methods require very long schedules to realize their potential. Summary of the Invention
[0006] This section summarizes the disclosure as a whole, but is not an exhaustive description of its entire scope or all of its features.
[0007] In one embodiment, a system for training a model includes a processor and a memory in communication with the processor having a training module. The training module includes instructions that, when executed by the processor, cause the processor to determine a contrastive loss using a self-supervised contrastive loss function based on a feature map describing the visual content of an image having an object and a feature vector describing the meaning of a phrase in a caption describing the object in the image. Then, based on the contrastive loss, the training module causes the processor to adjust model weights of the visual backbone that generated the feature map and / or the textual backbone that generated the feature vector.
[0008] The training module further includes instructions that, when executed by the processor, cause the processor to determine a localization loss using a supervised loss function that compares the image caption attention map to a visual identifier and adjust model weights of the visual backbone and / or the textual backbone based on the localization loss. The visual identifier identifies a location of an object in the image and is associated with a portion of the caption that describes the object, and can be the shape of a mouse trace.
[0009] In another embodiment, a method for training a model includes determining a contrastive loss using a self-supervised contrastive loss function based on a feature map describing the visual content of an image having an object and a feature vector describing the meaning of words in a caption describing the object in the image, and adjusting model weights of a visual backbone that generated the feature map and / or a textual backbone that generated the feature vector based on the contrastive loss.
[0010] The method further includes determining a localization loss using a supervised loss function that compares the image caption attention map to the visual identifier, and adjusting model weights of the visual backbone and / or the textual backbone based on the localization loss. As before, the visual identifier identifies a location of an object in the image and is associated with a portion of the caption describing the object, and can be the shape of a mouse trace.
[0011] In yet another embodiment, a non-transitory computer-readable medium has instructions that, when executed by a processor, cause the processor to determine a contrastive loss using a self-supervised contrastive loss function based on a feature map describing visual content of an image having an object and a feature vector describing meanings of words in a caption describing the object in the image. Then, based on the contrastive loss, the instructions cause the processor to adjust model weights of the visual backbone that generated the feature map and / or the textual backbone that generated the feature vector.
[0012] The non-transitory computer readable medium further includes instructions that, when executed by a processor, cause the processor to determine a localization loss using a supervised loss function that compares the image caption attention map to the visual identifier and adjust model weights of the visual backbone and / or the textual backbone based on the localization loss. Again, the visual identifier identifies the location of an object in the image and is associated with a portion of the caption that describes the object, and can be the shape of a mouse trace.
[0013] Further areas of applicability and various ways of enhancing the disclosed technology will become apparent from the description provided. The description and specific examples in this summary are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure. [Brief description of the drawings]
[0014] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate various systems, methods, and other embodiments of the disclosure. It will be appreciated that the boundaries of elements illustrated in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one embodiment of the boundaries. In some embodiments, one element may be designed as multiple elements, or multiple elements may be designed as one element. In some embodiments, an element shown as an internal component of another element may be realized as an external component, and vice versa. Additionally, elements may not be drawn to scale.
[0015] [Figure 1] 1 illustrates a system for training a model, such as a visual backbone model and / or a textual backbone model.
[0016] [Diagram 2] 1 is a flowchart illustrating the extraction of feature maps and feature vectors from images and associated captions performed by the system for training a model.
[0017] [Diagram 3] 1 is a flowchart illustrating the determination of a contrastive loss using a self-supervised contrastive loss function performed by a system for training a model.
[0018] [Figure 4] 11 is a flowchart illustrating the determination of a localization loss using a supervised loss function that compares an image caption attention map to a visual classifier performed by a system for training a model.
[0019] [Diagram 5] A method is illustrated for training a model such as a visual backbone model and / or a textual backbone model.
[0020] [Figure 6] 1 illustrates a method for computing an image caption attention map.
[0021] [Figure 7] 1 illustrates a method for determining a localization loss using a supervised loss function that compares an image caption attention map to a visual classifier.
[0022] [Figure 8] 10 illustrates a vehicle having an object detection system that utilizes a visual backbone pre-trained using the system of FIG. 1 and / or the method of FIG. 5. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0023] Systems and methods are described for training and / or pre-training models for neural networks such as CNNs. As described in the Background section, training and / or pre-training models typically requires the use of annotated datasets for supervised training or unannotated datasets for self-supervised training. Annotated datasets are difficult and costly to develop, while the use of unannotated datasets typically requires a significant amount of computing resources.
[0024] The system and method described herein utilizes a contrastive pre-training framework to train a model between images and associated captions. In addition, the system and method utilize a supervised training method, where a cross-modal attention map with a represented mouse trace is utilized to provide a coarse localization signal to perform supervised training. Thus, the system and method trains a model in an unsupervised manner using images and associated captions, and in a supervised manner using mouse traces associated with images to provide a coarse localization signal. The two losses from supervised and unsupervised training can be used together to optimize the model weights. This form of annotation can be easily obtained from non-expert personnel, resulting in lower cost and better scalability.
[0025] Referring to Figure 1, a model training system 10 for training a model is illustrated. Training a model may be actual training of the model or may be pre-training the model, where pre-training refers to training a model on one task to help form parameters that can be used on other tasks.
[0026] As shown, model training system 10 includes one or more processors 12. Processor 12 may be a single processor or may be multiple processors working in concert. Thus, processor 12 may be part of model training system 10, or model training system 10 may access processor 12 through a data bus or other communication path. In one or more embodiments, processor 12 may be an application specific integrated circuit configured to implement the functions associated with training module 16. Generally, processor 12 is an electronic processor, such as a microprocessor, capable of performing various functions as described herein.
[0027] In one embodiment, model training system 10 includes memory 14 that stores a training module 16. Memory 14 may be random access memory (RAM), read only memory (ROM), a hard disk drive, flash memory, or other suitable memory for storing training module 16. Training module 16 may be, for example, computer readable instructions that, when executed by processor 12, cause processor 12 to perform various functions disclosed herein.
[0028] Additionally, in one embodiment, the model training system 10 includes one or more data stores 20, which in one embodiment are electronic data structures such as databases stored in memory 14 or other memories configured with routines executable by processor 12 to analyze stored data, present stored data, organize stored data, generate stored data, etc. Thus, in one embodiment, data store 20 stores data used by training module 16 in performing various functions. In one embodiment, data store 20 includes three different models. The three models may include visual backbone model 22, text backbone model 24, and second neural network 26. Visual backbone model 122, text backbone model 124, and second neural network 26 may be different types of neural networks and may include model weights 23, 25, and 27, respectively. Model weights 23, 25, and / or 27 may be parameters, including trainable and non-trainable, of the models used in the layers of the model. Adjustment of the model weights 23, 25, and 27 affects the performance of the visual backbone model 22, the text backbone model 24, and the second neural network 26, respectively.
[0029] The visual backbone model 22 can be utilized to perform any one of a number of different computer vision tasks, such as object detection and semantic / instance segmentation. In one example, the visual backbone model 22 can be a component that is transferred to other downstream vision tasks. Any CNN can be utilized as the visual backbone model 22. In one example, the visual backbone model 22 can be a standard ResNet-50, which can have certain modifications, such as removing the last linear classification layer and the preceding global average pooling layer, to maintain the spatial dimension. In one example, the visual backbone model 22 can output a feature map having a size of 2048×R×R, where R is the output resolution, which may be 1 / 32 of the input resolution. Again, it should be understood that this type of ResNet-50 is just one example of a type of CNN that can be utilized as the visual backbone model 22.
[0030] The text backbone model 24 can be used to encode the input caption into a feature vector capturing the meaning of the word tokens forming the caption. In one example, the text backbone model 24 can employ a transducer architecture as the text backbone implemented in a 4-layer, 1024-wide model with 16 self-attention heads. The activation function can be a Gaussian error linear unit (GELU) instead of a rectified linear unit (ReLU) to achieve better empirical performance. Before inputting the caption, the caption can first be tokenized into lowercase byte pairs that encode with a vocabulary size of 10K. The input sequence can also be padded with start of sequence tokens and end of sequence tokens to mark boundaries. The output feature vector from the text backbone model 24 can have a size of 1024×L, where L is the length of the caption after tokenization.
[0031] The second neural network 26 may include a multi-dimensional fully connected layer to generate transformed feature vectors and transformed feature maps that can be used to train the visual backbone model 22 and / or the textual backbone model 24.
[0032] The data storage device 20 may also include training data 30 for training the visual backbone model 22, the text backbone model 24, and / or the second neural network 26. The training data 30 typically includes three pairs of data. Additionally, the training data 30 includes an image 32 that is paired with a caption 34 and a visual identifier 36. In this example, the image 32 is an image having a cat 32A lying on a blanket 32B, with several books 32C located in the background behind the cat 32A and the blanket 32B. Of course, it should be understood that the image 32 could be an image of many different objects that are arranged in different ways.
[0033] Caption 34 includes geometric descriptions of tokens 34A-34C. The first token 34A describes "there is a yellow cat". The second token 34B describes "lying on a blanket". The third token 34C describes "there are books behind it". Taken together, tokens 34A-34C of caption 34 collectively describe what is happening in image 32. That is, tokens 34A-34C describe the presence of a yellow cat lying on a blanket and some books behind it. Caption 34 is therefore related to what is happening in image 32.
[0034] In general, the caption 34 is a free-form annotation resulting from an annotator being asked to describe the content of the image 32 using natural language. The information captured in the caption 34 may be dense in meaning, i.e., describing the objects 32A-32C in the image 32, their attributes, and their relative spatial relationships. This underlying rich semantic information potentially provides advantages for a variety of downstream vision tasks. The cost of this form of annotation is very low compared to other dense labeling, as it is a very natural task for humans to perform and does not require the annotator to have extensive training or domain knowledge. The caption 34 can be generated using a two-stage data collection pipeline. In the first stage, the annotator is asked to describe the image 32 verbally and apply either speech recognition or manual transcription to generate the caption 34. From this collection protocol, the start and end timestamps of the tokens 34A-34C that form the caption 34 can be obtained, which can be used to synchronize with the visual identifiers 36, as will be explained later.
[0035] Visual identifier 36 may be in the form of one or more mouse traces representing the location of particular objects within the image. For example, visual identifier 36A roughly identifies the location of cat 32A within image 32. Visual identifier 36B roughly identifies the location of blanket 32B within image 32. Finally, visual identifier 36C roughly identifies the location of book 32C within image 32.
[0036] Compared to drawing a sequence of bounding boxes or instance masks, recording the mouse trace of a subject while describing an image32 is an easier and more natural way for a human annotator to identify the location of an object. It can be obtained almost at will in a caption annotation pipeline, since it only requires the annotator to hover their mouse over the region being described. Although the localization and semantic correspondence are too coarse-grained to directly use these annotations for tasks such as object detection, it does capture a wealth of information about "what is where" at a high level.
[0037] The training module 16 generally includes instructions that function to control the processor 12 to train the visual backbone model 22, the text backbone model 24, and / or the second neural network 26. Further, with reference to FIG. 2, the training module 16 may include instructions that cause the processor 12 to generate a feature map that describes the visual content of an image, such as the image 32 having objects 32A-32C. This may occur by first passing the image 32 through the visual backbone model 22 to generate a visual feature map 42 that describes the image 32 as a whole. As explained above, the visual backbone model 22 may output a feature map having a size of 2048×R×R, where R is the output resolution, which may be 1 / 32 of the input resolution.
[0038] The training module 16 may include instructions that cause the processor 12 to generate a text feature vector 44. This may occur by passing the caption 34 through the text backbone model 24. As explained above, the caption 34 includes tokens 34A-34C that describe objects 32A-32C found in the image 32. The text backbone model 24 may encode the caption 34 into a text feature vector 44 that captures the meaning of the tokens 34A-34C. The text feature vector 44 from the text backbone model 24 may have a size of 1024×L, where L is the length of the caption after tokenization.
[0039] Next, the training module 16 may include instructions to cause the processor 12 to determine a contrastive loss using a self-supervised contrastive loss function based on the visual feature map 42 describing the visual content of the image 32 and the text feature vectors 44 describing the meaning of words in the captions 34 of the objects 32A-32C in the image 32. Referring to FIG. 3, a flow chart 50 is illustrated detailing how the contrastive loss is determined. For a batch of feature pairs {(xV, k ,xT, k )|1≦k≦n}, where n is the batch size, processor 12 can transform feature maps 42A and 42B and feature vectors 44A and 44B, respectively, with global average pooling and a single 1024-dimensional fully connected layer. The resulting visual features 46A and 46B and text features 48A and 48B, both of size 1024, are V,k andy T,k This is expressed as:
[0040] y V,k andy T,kConventional methods of guiding pre-training by matching n in feature space using simple regression losses lead to a broken solution where all features are projected to the same location in feature space. Therefore, training module 16 may include instructions to encourage processor 12 to not only project visual feature maps 42 and text feature vectors 44 that match image caption pairs more closely, but also project features that drive non-matching pairs further apart. More specifically, a total of n 2 Image caption vs. {(y V,i ,y T,j )|1≦k≦n}, among which only the n pairs of i=j correspond to the same data and are therefore positive, and the remaining (n 2 -n) pairs are negative. Therefore, training module 16 causes processor 12 to group positive pairs together and separate negative pairs from each other to guide pre-training.
[0041] The symmetric loss function for determining the symmetric loss can be expressed as follows:
number
[0042] Once the contrastive loss is determined, the training module 16 may include instructions to cause the processor 12 to adjust the model weights 23 and / or 25 of the visual backbone model 22 and / or the textual backbone model 24, respectively, based on the contrastive loss. Applying the contrastive loss to the global visual and textual features (after average pooling) provides the visual backbone model 22 with an overall sense of what objects 32A-32C are in the image 32. However, the visual backbone model 22 cannot address each instance in its spatial location, limiting its effectiveness when transferred to downstream tasks sensitive to localization, such as object detection and / or instance segmentation.
[0043] Therefore, the training module 16 may include instructions to cause the processor 12 to determine a localization loss using a supervised loss function that compares the image caption attention map to the visual identifiers 36. Referring to FIG. 4, a flowchart 60 is illustrated detailing how the localization loss is determined. The training module 16 may include instructions to cause the processor 12 to pass the visual feature map 42 and the text feature vector 44 through a second neural network 26. Further, the second neural network 26 linearly transforms the visual feature map 42 and the text feature vector 44 using 1024-dimensional fully connected layers 62 and 64, respectively. Global average pooling cannot be applied to preserve the spatial dimension for learning the localization. Therefore, the transformed visual feature map 42z V , k has a size of 1024 × R × R. The transformed text feature vector 44z V , k has a size of 1024×L.
[0044] The training module 16 may then transmit to the processor 12 the image caption attention map 68 as a transformed visual feature map 42z. V , k and the converted text feature vector 44zV , k This calculation can be expressed in the following equation:
number
[0045] Assuming the visual identifiers 36A-36C can correspond to the locations of the objects 32A-32C in the image 32 and are synchronized with the tokens 34A-34C in the caption 34, the visual identifiers 36A-36C can be utilized to manage the generation of the image caption attention map 68. As such, a localization loss is generated using a loss function that compares the image caption attention map 68 to the visual identifiers 36. And, the training module 16 can include instructions to cause the processor 12 to adjust the visual backbone model 22, the text backbone model 24, and the model weights 23, 25 and / or 27 of the second neural network 26 based on the localization loss.
[0046] To determine the localization loss, training module 16 may include instructions to cause processor 12 to temporally crop portions of visual identifiers 36 using cropping function 70 to generate cropped visual identifiers corresponding to caption phrases associated with each of the objects in image 32. Training module 16 may then include instructions to cause processor 12 to represent covered regions of image 32 associated with the cropped visual identifiers to generate a binary mask having resolution R.
[0047] Training module 16 then transmits to processor 12 the represented attention 72
number
number
number
number
number
number
[0048] Once the regularized regression loss has been determined, as described above, the training module 16 may include instructions that cause the processor 12 to adjust the model weights 23, 25 and / or 27 of the visual backbone model 22, the text backbone model 24, and the second neural network 26, respectively, based on the localization loss.
[0049] If the visual feature maps from the visual backbone model 22 have a low resolution, a localization loss can be applied to the penultimate visual feature map (which can have double the resolution) to provide control at a finer scale. Losses computed at different resolutions can be added with equal weights.
[0050] Referring to FIG. 5, a method 100 for training a model is shown. The method 100 is described from the perspective of the model training system 10 of FIG. 1 with the support of the flow charts of FIGS. 2-4. However, it should be understood that this is merely one example for implementing the method 100. Although the method 100 is discussed in combination with the model training system 10, it should be appreciated that the method 100 is not limited to being implemented within the model training system 10, which is one example of a system that can implement the method 100. Furthermore, when describing the method 100, it should be understood that the operations performed by the model training system 10, previously described in the paragraphs above, are equally applicable to the method 100 and need not be described again as the previous descriptions are pertinent.
[0051] At step 102, the training module 16 may include instructions to cause the processor 12 to determine a contrastive loss using a self-supervised contrastive loss function based on the visual feature map 42 and the text feature vector 44. As explained above, this can be accomplished using a self-supervised contrastive loss function based on the visual feature map 42 describing the visual content of the image 32 and the text feature vector 44 describing the meaning of words in the captions 34 of the objects 32A-32C in the image 32. In essence, the training module 16 may cause the processor 12 to encourage the visual backbone model 22 and the text backbone model 24 not only to project visual feature maps 42 and text feature vectors 44 that more closely match image-caption pairs, but also to project features that drive non-matching pairs further apart.
[0052] In step 104, the training module 16 may include instructions that cause the processor 12 to adjust the model weights 23 and / or 25 of the visual backbone model 22 and the textual backbone model 24, respectively, based on the contrastive loss.
[0053] At step 106, training module 16 may include instructions to cause processor 12 to generate an image caption attention map 68 based on visual feature map 42 and text feature vector 44. Image caption attention map 68 may identify the location and object type of objects 32A-32C within image 32.
[0054] With respect to generating the image caption attention map 68, reference is made to Figure 6. At step 106A of Figure 6, the training module 16 may include instructions to cause the processor 12 to transform each of the feature maps 42A and 42B, and the feature vectors 44A and 44B, with global average pooling and a single 1024-dimensional fully connected layer. At step 106B, the training module 16 may include instructions to cause the processor 12 to utilize the visual identifiers 36 and calculate the image caption attention map 68 as a normalized product between the transformed visual feature map 42 and the transformed text feature vector 44.
[0055] Returning to Figure 5, at step 108, training module 16 may include instructions to cause processor 12 to compute a localization loss using a loss function that compares image caption attention map 68 to visual identifiers 36. For example, with reference to Figure 7, at step 108A, training module 16 may include instructions to cause processor 12 to temporally crop portions of visual identifiers 36 using cropping function 70 to generate cropped visual identifiers that correspond to words of the caption associated with each of the objects in image 32.
[0056] In step 108B, training module 16 may include instructions to cause processor 12 to render covered regions of image 32 associated with the cropped visual identifier to generate a binary mask having resolution R. In step 108C, training module 16 may include instructions to cause processor 12 to stack the rendered masks of all tokens together to generate rendered attention 72. Finally, in step 108D, training module 16 may include instructions to cause processor 12 to use rendered attention 72 to provide a regularized regression loss for supervision over image caption attention map 68.
[0057] 5, in step 110, training module 16 may include instructions to cause processor 12 to adjust visual backbone model 22, text backbone model 24, and model weights 23, 25, and / or 27 of second neural network 26 based on the localization loss. After performing step 110, method 100 may either end, or may continue again if more training data is available.
[0058] Thus, the model training system 10 and associated method 100 can use low-cost localized text annotations to pre-train models such as the visual backbone model 22, the text backbone model 24, and / or the second neural network 26, to reduce annotation effort. The model training system 10 and associated method 100 essentially bridges contrastive learning between visual and language modalities, manages cross-modal attention maps with represented mouse traces, and provides coarse localization information to improve performance of downstream tasks sensitive to localization.
[0059] Pre-training a model, e.g., the visual backbone model 22, allows the features to be transferred to other downstream tasks by fine-tuning on a target dataset. The type of downstream tasks performed by a model trained by the model training system 10 and / or the associated method 100 may vary depending on the application. For example, the visual backbone model 22 can be utilized to perform object detection, object classification, instance segmentation, and other types of computer-related tasks. Again, models pre-trained by the model training system 10 and / or the associated method 100 can be used in many different applications, not necessarily limited to those specifically listed above.
[0060] One such application relates to object detection, and in particular to object detection performed by one or more systems of a vehicle. Again, the applications of any models pre-trained using the model training system 10 and / or associated method 100 are numerous and are not limited to vehicles. It should be understood that incorporating models trained by the model training system 10 and / or associated method 100 is not limited to vehicles.
[0061] Referring to FIG. 8, an example vehicle 200 is shown using one or more models pre-trained using the model training system 10 and / or associated method 100. As used herein, a "vehicle" is any form of motorized transportation device. In one or more implementations, the vehicle 200 is an automobile. Although the arrangements are described herein with respect to an automobile, it will be understood that the embodiments are not limited to automobiles. In some implementations, the vehicle 200 may be any robotic device or any form of motorized transportation device that benefits from the functionality discussed herein, for example, by including one or more automated or autonomous systems.
[0062] Vehicle 200 also includes various elements. It will be understood that in various embodiments, vehicle 200 need not have all of the elements shown in FIG. 8. In some arrangements, vehicle 200 can be implemented without one or more of the elements shown in FIG. 8. Although various elements are shown in FIG. 8 as being located within vehicle 200, it will be understood that one or more of these elements can be located external to vehicle 200. Additionally, the elements shown can be physically separated by a significant distance and provided as remote devices (e.g., cloud computing services).
[0063] In various embodiments, the automated / autonomous system or combination of systems may vary. For example, in one aspect, an automated system is a system that provides autonomous control of a vehicle according to one or more levels of automation, such as levels (e.g., levels 0-5) defined by the Society of Automotive Engineers (SAE). Thus, an autonomous system may provide semi-autonomous control or fully autonomous control, as discussed in connection with autonomous driving system 260.
[0064] As used herein, an "autonomous vehicle" refers to a vehicle operating in an autonomous mode. An "autonomous mode" refers to using one or more processing systems to control vehicle 200 with minimal or no input from a human driver to navigate and / or steer vehicle 200 along a travel route. In one or more embodiments, vehicle 200 is highly automated or fully automated. In one embodiment, vehicle 200 is configured in one or more semi-autonomous operating modes in which one or more processing systems perform portions of the navigation and / or steering of vehicle 200 along a travel route, and an operator (i.e., driver) of the vehicle provides input to the vehicle to perform portions of the navigation and / or steering of vehicle 200 along a travel route. Such semi-autonomous operation may include supervisory control.
[0065] The vehicle 200 may include one or more processors 210. In one or more arrangements, the processor 210 may be the main processor of the vehicle 200. For example, the processor 210 may be an electronic control unit (ECU). The vehicle 200 may include one or more data storage devices 215 for storing one or more types of data. The data storage devices 215 may include volatile and / or non-volatile memory. Examples of the data storage devices 215 include RAM (random access memory), flash memory, ROM (read only memory), PROM (programmable read only memory), EPROM (erasable programmable read only memory), EEPROM (electrically erasable programmable read only memory), registers, magnetic disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The data storage devices 215 may be components of the processor 210, or the data storage devices 215 may be operatively connected to the processor 210 for use by the processor 210. As used throughout this description, the terms "operably connected" and / or "in communication with" can include direct or indirect connections, including connections without direct physical contact.
[0066] In one or more arrangements, the data storage 215 may include map data 216. The map data 216 may include maps of one or more geographic regions. In some instances, the map data 216 may include information or data about roads, traffic control devices, road signs, structures, features, and / or landmarks in one or more geographic regions. The map data 216 may be in any suitable form. In some instances, the map data 216 may include an aerial photograph of the region. In some instances, the map data 216 may include a ground map of the region, including a 360 degree ground map. The map data 216 may include measurements, dimensions, distances, and / or information for one or more items included in the map data 216 and / or about other items included in the map data 216. The map data 216 may include a digital map having information about road geometry. The map data 216 may be of high quality and / or may be highly detailed.
[0067] In one or more arrangements, the map data 216 may include one or more terrain maps 217. The terrain maps 217 may include information about the ground, terrain, roads, surfaces, and / or other features of one or more geographic regions. The terrain maps 217 may include elevation data for one or more geographic regions. Topographical Map 217 The topographical map 217 may be of high quality and / or may be highly detailed. The topographical map 217 may define one or more terrain surfaces, which may include paved roads, unpaved roads, land, and other things that define the terrain.
[0068] In one or more configurations, the map data 216 may include one or more static obstacle maps 218. The static obstacle map 218 may include information about one or more static obstacles located within one or more geographic regions. A "static obstacle" is a physical object whose location does not change or changes very little over time and / or whose size does not change or changes very little over time. Examples of static obstacles may include trees, buildings, curbs, walls, fences, medians, telephone poles, statues, monuments, signs, benches, furniture, mailboxes, large stones, slopes, etc. A static obstacle may be an object that extends above the ground. One or more static obstacles included in the static obstacle map 218 may have location data, size data, dimension data, material data, and / or other data associated with them. The static obstacle map 218 may include measurements, dimensions, distances, and / or information for one or more static obstacles. The static obstacle map 218 may be of high quality and / or may be highly detailed. The static obstacle map 218 can be updated to reflect changes in the area described by the map.
[0069] The one or more data stores 215 may include sensor data 219. In this context, "sensor data" refers to any information about sensors equipped with the vehicle 200, including functional and other information about such sensors. As described below, the vehicle 200 may include a sensor system 220. The sensor data 219 may relate to one or more sensors of the sensor system 220.
[0070] In some instances, at least a portion of the map data 216 and / or the sensor data 219 may be located in one or more data stores 215 located on the vehicle 200. Alternatively or additionally, at least a portion of the map data 216 and / or the sensor data 219 may be located in one or more data stores 215 located remotely from the vehicle 200.
[0071] As noted above, the vehicle 200 can include a sensor system 220. The sensor system 220 can include one or more sensors. A "sensor" refers to any device, component, and / or system that can detect and / or sense something. The one or more sensors can be configured to detect and / or sense in real time. As used herein, the term "real time" refers to a level of processing response that a user or system perceives as sufficiently immediate for a particular process or decision being made, or that allows a processor to keep up with an external process.
[0072] In an arrangement in which the sensor system 220 includes multiple sensors, the sensors can function independently of one another. Alternatively, two or more of the sensors can function in combination with one another. In such a case, the two or more sensors can form a sensor network. The sensor system 220 and / or one or more sensors can be operatively connected to the processor 210, the data storage device 215, and / or other elements of the vehicle 200 (including any of the elements shown in FIG. 8). The sensor system 220 can acquire data of at least a portion of the environment external to the vehicle 200 (e.g., nearby vehicles).
[0073] The sensor system 220 may include any suitable type of sensor. Various examples of different types of sensors are described herein. However, it will be understood that the embodiments are not limited to the particular sensors described. The sensor system 220 may include one or more vehicle sensors 221. The vehicle sensors 221 may detect, determine, and / or sense information about the vehicle 200 itself. In one or more arrangements, the vehicle sensors 221 may be configured to detect and / or sense changes in the position and orientation of the vehicle 200, such as, for example, based on inertial acceleration. In one or more arrangements, the vehicle sensors 221 may include one or more accelerometers, one or more gyroscopes, an inertial measurement unit (IMU), a dead reckoning system, a global navigation satellite system (GNSS), a global positioning system (GPS), a navigation system 247, and / or other suitable sensors. The vehicle sensors 221 may be configured to detect and / or sense one or more characteristics of the vehicle 200. In one or more arrangements, the vehicle sensors 221 may include a speedometer that determines the current speed of the vehicle 200 .
[0074] Alternatively or additionally, the sensor system 220 may include one or more environmental sensors 222 configured to obtain and / or sense driving environment data. "Driving environment data" includes data or information about or about one or more portions of an external environment in which the autonomous vehicle is located. For example, the one or more environmental sensors 222 may be configured to detect, measure, weigh, and / or sense obstacles in at least a portion of the external environment of the vehicle 200 and / or information / data about such obstacles. Such obstacles may be static objects and / or dynamic objects. The one or more environmental sensors 222 may be configured to detect, measure, weigh, and / or sense other objects in the external environment of the vehicle 200, such as, for example, lane markings, signs, traffic lights, traffic signs, lane lines, pedestrian crossings, curbs close to the vehicle 200, objects off the road, etc.
[0075] Described herein are various examples of sensors of sensor system 220. The example sensors may be part of one or more environmental sensors 222 and / or one or more vehicle sensors 221. However, it will be understood that embodiments are not limited to the particular sensors described.
[0076] By way of example, in one or more arrangements, the sensor system 220 can include one or more radar sensors 223, one or more lidar sensors 224, one or more sonar sensors 225, and / or one or more cameras 226. In one or more arrangements, the one or more cameras 226 can be high dynamic range (HDR) cameras or infrared (IR) cameras.
[0077] Vehicle 200 may include an input system 230. An "input system" includes any device, component, system, element or arrangement, or grouping thereof, that allows information / data to enter the machine. Input system 230 may receive input from an occupant of the vehicle (e.g., a driver or passenger). Vehicle 200 may include an output system 235. An "output system" includes any device, component, or arrangement, or grouping thereof, that allows information / data to be presented to an occupant of the vehicle (e.g., a human, a vehicle passenger, etc.).
[0078] Vehicle 200 may include one or more vehicle systems 240. Various examples of one or more vehicle systems 240 are shown in FIG. 8. However, vehicle 200 may include more, less, or different vehicle systems. Although specific vehicle systems are defined separately, it should be appreciated that each or any of the systems, or portions thereof, may be combined or separated within vehicle 200 via hardware and / or software. Vehicle 200 may include a propulsion system 241, a braking system 242, a steering system 243, a throttle system 244, a transmission system 245, a signal system 246, and / or a navigation system 247. Each of these systems may include one or more devices, components, and / or combinations thereof, now known or later developed.
[0079] Navigation system 247 may include one or more now known or later developed devices, applications, and / or combinations thereof configured to determine a geographic location of vehicle 200 and / or determine a driving route for vehicle 200. Navigation system 247 may include one or more mapping applications that determine a driving route for vehicle 200. Navigation system 247 may include a global positioning system, a local positioning system, or a geolocation system.
[0080] The vehicle 200 may include an object detection system 270 that receives information from the sensor system 220. Using the information received from the sensor system 220, the object detection system 270 may detect the presence of an object using the visual backbone model 22 pre-trained using the model training system 10 and / or the associated method 100, as described above. Again, it should be understood that this is just one example of using a model trained by the model training system 10 and / or the associated method 100. In addition to object detection, there are many other uses for the visual backbone model 22, such as semantic / instance segmentation, object detection, or any other computer vision task. The information generated by the object detection system 270 may be provided to an autonomous driving system 260 that may control the movement of the vehicle 200.
[0081] The processor 210 and / or the autonomous driving system 260 can be operatively connected to communicate with the vehicle system 240 and / or their individual components. The processor 210 and / or the autonomous driving system 260 can communicate to send and / or receive information from the vehicle system 240 to control the movement, speed, steering, path, direction, etc. of the vehicle 200. As discussed above, the object detection system 270 can also communicate with the processor 210 and / or the autonomous driving system 260 to provide information related to object detection. Additionally, the autonomous driving system 260 can provide autonomous operation for the vehicle 200, where little or no driver input is required. However, the autonomous driving system 260 can provide semi-autonomous operation of the vehicle 200, where commands from the driver are still required to move the vehicle 200 from one location to another.
[0082] Processor 210 and / or autonomous driving system 260 may be operable to control navigation and / or steering of vehicle 200 by controlling vehicle systems 240 and / or one or more of their components. For example, when operating in an autonomous mode, processor 210 and / or autonomous driving system 260 may control the direction and / or speed of vehicle 200. Processor 210 and / or autonomous driving system 260 may cause vehicle 200 to accelerate (e.g., by increasing the supply of fuel provided to the engine), slow down (e.g., by decreasing the supply of fuel to the engine and / or by applying the brakes), and / or change direction (e.g., by changing the direction of the front two wheels). As used herein, "cause" or "causing" means, directly or indirectly, to cause, compel, direct, command, instruct, and / or enable an event or action to occur, or to at least cause, compel, direct, command, instruct, and / or enable such event or action to occur.
[0083] Vehicle 200 may include one or more actuators 250. Actuator 250 may be any element or combination of elements operable to modify, adjust, and / or alter vehicle system 240, or one or more of its components, in response to receiving signals or other inputs from processor 210 and / or autonomous driving system 260. Any suitable actuator may be used. For example, one or more actuators 250 may include motors, pneumatic actuators, hydraulic pistons, relays, solenoids, and / or piezoelectric actuators, to name a few possibilities.
[0084] In one or more arrangements, one or more of the modules described herein may include artificial or computational intelligence elements, such as, for example, neural networks, fuzzy logic, or other machine learning algorithms. Further, in one or more arrangements, one or more of the modules may be distributed among multiple modules described herein. In one or more arrangements, two or more of the modules described herein may be combined into a single module.
[0085] Detailed embodiments are disclosed herein. However, it should be understood that the disclosed embodiments are intended to be examples only. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but are intended only as a basis for the claims and as a representative basis to teach one skilled in the art to variously employ the aspects herein in substantially any suitable detailed structure. Furthermore, the terms and phrases used herein are not intended to be limiting, but rather to provide an understandable description of possible implementations. Although various embodiments are shown in Figures 1-8, the embodiments are not limited to the illustrated structures or applications.
[0086] According to various embodiments, the flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may be executed substantially simultaneously, or the blocks may be executed in the reverse order, depending on the functions involved.
[0087] The above-mentioned systems, components, and / or processes can be realized in hardware or a combination of hardware and software, either centralized in one processing system or distributed where different elements are spread across several interconnected processing systems. Any kind of processing system, or other device adapted to perform the methods described herein, is suitable. A typical combination of hardware and software can be a processing system having computer usable program code that, when loaded and executed, controls the processing system such that the processing system performs the methods described herein. The systems, components, and / or processors can also be embedded in a computer readable storage device, such as a computer program product or other data program storage device that is readable by a machine and tangibly contains a program of instructions that are executable by the machine to perform the methods and processes described herein. These elements can also be embedded in an application product that comprises all the features enabling the implementation of the methods described herein and that, when loaded into a processing system, can perform these methods.
[0088] Additionally, the arrangements described herein may take the form of a computer program product embodied in one or more computer readable media containing, e.g., storing, computer readable program code. Any combination of one or more computer readable media may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. The phrase "computer readable storage medium" refers to a non-transitory storage medium. The computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of computer readable storage media include the following (a non-exhaustive list): portable computer diskette, hard disk drive (HDD), solid state drive (SSD), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), portable compact disc read only memory (CD-ROM), digital versatile disk (DVD), optical storage device, magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0089] Generally, a module as used herein includes routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular data types. In a further aspect, a memory generally stores the above-mentioned modules. The memory associated with a module may be a buffer or cache embedded within a processor, RAM, ROM, flash memory, or other suitable electronic storage medium. In a further aspect, a module as contemplated by the present disclosure is implemented as an application specific integrated circuit (ASIC) that is a hardware component of a system on a chip (SoC), as a programmable logic array (PLA), or as other suitable hardware component embedded with a defined configuration set (e.g., instructions) to perform the disclosed functions.
[0090] The program code contained on the computer readable medium may be transmitted using any suitable medium, including, but not limited to, wireless, wireline, fiber optic, cable, RF, etc., or any suitable combination of the above. Computer programs for performing operations for aspects of the present arrangement may be implemented using Java. TMThe program code may be written in any combination of one or more programming languages, including object-oriented programming languages such as, for example, Smalltalk, C++, and the like, and traditional procedural programming languages such as the "C" programming language or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server, as a stand-alone software package. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet Service Provider).
[0091] As used herein, the term "a" is defined as one or more than one. As used herein, the term "multiple" is defined as two or more than two. As used herein, the term "another" is defined as at least a second or more. As used herein, the terms "including" and / or "having" are defined as comprising (i.e., open language). As used herein, the phrase "and at least one of" refers to and is inclusive of any and all possible combinations of one or more of the associated listed items. By way of example, the phrase "at least one of A, B, and C" includes A only, B only, C only, or any combination thereof (e.g., AB, AC, BC, or ABC).
[0092] The aspects herein may be embodied in other forms without departing from the spirit or essential attributes thereof, and reference should accordingly be made to the following claims, rather than the foregoing specification, as indicating their scope.
Claims
1. 1. A system for training a model, comprising: Processor and a memory in communication with the processor and having a training module, the training module having instructions that, when executed by the processor, cause the processor to: determining a contrastive loss using a self-supervised contrastive loss function based on a feature map describing the visual content of an image having an object and a feature vector describing the meaning of words in a caption describing the object in the image; Adjusting model weights of at least one of a visual backbone that generated the feature map and a textual backbone that generated the feature vector based on the contrastive loss; determining a localization loss using a supervised loss function that compares an image caption attention map to visual identifiers that identify the location of the object in the image and that are associated with portions of the caption that describe the object; adjusting the model weights of at least one of the visual backbone and the textual backbone based on the localization loss; generating the image caption attention map based on the feature map and the feature vector, the image caption attention map identifying a location and an object type of the object within the image; determining the localization loss by comparing the location and object type of the object defined by the image caption attention map to the visual identifier; transforming the feature vector and the feature map using a second neural network having a multi-dimensional fully connected layer to generate a transformed feature vector and a transformed feature map; Computing the image caption attention map as a normalized product between the transformed feature vector and the transformed feature map; and adjusting the model weights of the second neural network based on the localization loss.
2. The training module further includes instructions that, when executed by the processor, cause the processor to: temporally cropping portions of the visual identifiers to generate cropped visual identifiers corresponding to the words of the caption associated with each of the objects; representing a covered area of the image associated with the cropped visual identifier to generate a binary mask; Stacking the binary masks together to generate a represented attention; 2. The system of claim 1, wherein the localization loss is determined using the supervised loss function that compares the image caption attention map to the represented attention.
3. 2. The system of claim 1, wherein the visual identifier is a mouse trace indicating a location of an object within the image.
4. 2. The system of claim 1, wherein the training module further includes instructions that, when executed by the processor, cause the processor to use the self-supervised contrastive loss function to pull positive pairs of the feature map and the feature vector closer together and push inconsistent pairs of the feature map and the feature vector apart to determine the contrastive loss.
5. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to: determining a contrastive loss using a self-supervised contrastive loss function based on a feature map describing the visual content of an image having an object and a feature vector describing the meaning of words in a caption describing the object in the image; Adjusting model weights of at least one of a visual backbone that generated the feature map and a textual backbone that generated the feature vector based on the contrastive loss; determining a localization loss using a supervised loss function that compares an image caption attention map to visual identifiers that identify the location of the object in the image and that are associated with portions of the caption that describe the object; adjusting the model weights of at least one of the visual backbone and the textual backbone based on the localization loss; generating the image caption attention map based on the feature map and the feature vector, the image caption attention map identifying a location and an object type of the object within the image; determining the localization loss by comparing the location and object type of the object defined by the image caption attention map to the visual identifier; transforming the feature vector and the feature map using a second neural network having a multi-dimensional fully connected layer to generate a transformed feature vector and a transformed feature map; Computing the image caption attention map as a normalized product between the transformed feature vector and the transformed feature map; 4. The method of claim 3, further comprising: adjusting the model weights of the second neural network based on the localized loss.
6. and instructions that, when executed by a processor, cause the processor to: temporally cropping portions of the visual identifiers to generate cropped visual identifiers corresponding to the words of the caption associated with each of the objects; representing a covered area of the image associated with the cropped visual identifier to generate a binary mask; Stacking the binary masks together to generate a represented attention; 6. The non-transitory computer-readable medium of claim 5, wherein the localization loss is determined using the supervised loss function that compares the image caption attention map to the represented attention.
7. 6. The non-transitory computer-readable medium of claim 5, further comprising instructions that, when executed by a processor, cause the processor to use the self-supervised contrastive loss function to pull positive pairs of the feature map and the feature vector closer together and push non-matching pairs of the feature map and the feature vector apart to determine the contrastive loss.
Citation Information
Patent Citations
Image segmentation method, device and equipment and storage medium
CN112184738A
Text-to-Visual Machine Learning Embedding Techniques
US20200380298A1