Image-based generation of descriptive and perceptual messages of automotive scenarios
By introducing traffic object detection module, attention map highlighting module, image encoder and PLM module into the vehicle detection system, the accuracy of traffic object detection and text description generation in the vehicle environment is solved, and a more efficient vehicle autonomous driving and collision warning system is realized.
Patent Information
- Application Number
- CN202410058993.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-01-15
- Publication Date
- 2025-05-16
AI Technical Summary
Existing vehicle detection systems are difficult to accurately detect traffic objects in the vehicle environment and generate accurate text descriptions, resulting in the limitation of the effectiveness of vehicle autonomous driving and collision warning systems.
Using the traffic object detection module, attention map highlighting module, image encoder and pretrained language model (PLM) module, the traffic map, encoded images and text is generated, and the text is selected and attached to iteratively to create a specific description of the content of the vehicle environment.
Accurate detection and text description generation of traffic objects in the vehicle environment are realized, the effectiveness of the vehicle's autonomous driving and collision warning system is improved, and the perception and decision-making ability of the vehicle environment is enhanced.
Smart Images

Figure CN120014075A_ABST
Abstract
Description
[0001] introduce
[0002] The information provided in this section is for the purpose of generally presenting the context of the present disclosure. The work of the presently named inventors is neither explicitly nor implicitly admitted to be prior art against the present disclosure to the extent it is described in this section and in aspects of the description that might not otherwise be prior art at the time of filing.
[0003] The present disclosure relates to object detection and collision warning systems, and more particularly to perception systems for describing and responding to environmental situations.
[0004] The host vehicle may include an object detection and collision warning system for detecting an approaching object and executing countermeasures and / or taking evasive action to prevent a collision. The vehicle may include various sensors for detecting objects, such as other vehicles, pedestrians, cyclists, etc. The controller determines the position of the object relative to the host vehicle and the trajectory of the object and the host vehicle. If it is determined that the host vehicle is likely to collide with one of the objects, a warning signal may be generated and / or the controller may execute some other countermeasures (e.g., slowing the vehicle, applying brakes, changing the steering angle of the vehicle, etc.) to prevent a collision. Summary of the invention
[0005] A system is disclosed and includes: a traffic object detection module configured to detect traffic objects in a vehicle environment; an attention map highlighting module configured to i) generate an attention map, ii) highlight relevant traffic objects in the traffic objects and at least one of the areas in which the relevant traffic objects in the traffic objects are located in the attention map, and iii) not highlight non-relevant traffic objects and other non-traffic objects in the traffic objects in the attention map; and an image encoder configured to encode an environment image received from an imaging device based on the attention map and generate an image embedding vector. The system further includes a pre-trained language model (PLM) module, which is configured to perform an iterative process, which includes iteratively selecting and appending text to create a text message. The PLM module is configured to select text based on at least one score during each iteration of the iterative process, and the text message is a specific description of the content perceived in the vehicle environment. The system further includes: a text encoder configured to encode a portion of the text message created to date to generate a text embedding vector; one or more modules configured to score the portion of the text message created to date based on the image embedding vector and the text embedding vector to generate at least one score, wherein the PLM module is configured to update the portion of the text message created to date based on the at least one score; and at least one output device configured to output the text message for an occupant of the vehicle when completed.
[0006] In other features, the system further includes a perception and decision module configured to perform at least one of an autonomous driving maneuver and a collision avoidance countermeasure based at least on information in the text message describing the vehicle environment.
[0007] In other features, for each iteration of the iterative process, the one or more modules include a vector comparison module configured to compare the text embedding vector with the image embedding vector and generate a first score. The PLM module is configured for each iteration to update a portion of the text message created so far based on the first score.
[0008] In other features, for each iteration of the iterative process, the one or more modules include a language scoring module configured to analyze the text embedding vector to determine whether the portion of the text message created so far is grammatically correct and generate a second score based on the determination. The PLM module is configured for each iteration to update the portion of the text message created so far based on the second score.
[0009] In other features, for each iteration of the iterative process, the one or more modules include a car vocabulary scoring module configured to analyze the text embedding vector to determine how many words in the portion of the text message created so far are not car words and generate a third score based on the determination. The PLM module is configured for each iteration to update the portion of the text message created so far based on the third score.
[0010] In other features, the one or more modules include a plurality of scoring modules and an overall scoring module. The plurality of scoring modules are configured to generate scores based on the portion of the text messages created so far for each iteration of the iterative process. The overall scoring module is configured to generate an overall score based on the scores. The PLM module is configured to update the portion of the text messages created so far based on the overall score.
[0011] Among other features, the scores include i) a first score indicating how closely the text embedding vector matches the image embedding vector, ii) a second score indicating how grammatically correct the portion of the text message created so far is, and iii) a third score indicating how many non-car words are included in the portion of the text message created so far.
[0012] In other features, the system further includes a car scene captioning module configured to receive the image embedding vector and the car vocabulary and, based on the image embedding vector and the car vocabulary, generate a text message such that the text message is specific to the car.
[0013] In other features, the PLM module is configured to replace one or more words of a portion of a text message created to date based on the at least one score.
[0014] In other features, the PLM module is configured to start with a sentence prefix and iteratively append car words to the sentence prefix to generate a text message.
[0015] In other features, the at least one output device includes a display, a speaker, and a haptic device.
[0016] In other features, a method for providing image captions for an image of a vehicle environment is disclosed. The method includes: detecting traffic objects in the vehicle environment; generating an attention map; highlighting at least one of relevant traffic objects in the traffic objects and the area in which the relevant traffic objects in the traffic objects are located in the attention map; not highlighting non-relevant traffic objects and other non-traffic objects in the traffic objects in the attention map; encoding the image and generating an image embedding vector based on the attention map; performing an iterative process including iteratively selecting and appending text to create a text message; during each iteration of the iterative process, selecting text based on at least one score, the text message being a specific description of what is perceived in the environment of the vehicle; encoding a portion of the text message created so far to generate a text embedding vector; scoring the portion of the text message created so far based on the image embedding vector and the text embedding vector to generate at least one score; updating the portion of the text message created so far based on the at least one score; and outputting the text message for an occupant of the vehicle when completed.
[0017] In other features, the method further comprises performing at least one of an autonomous driving maneuver and a collision avoidance maneuver based at least on the information in the text message describing the vehicle environment.
[0018] In other features, the method further includes, for each iteration of the iterative process: comparing the text embedding vector to the image embedding vector and generating a first score; and updating the portion of the text message created so far based on the first score.
[0019] In other features, the method further includes, for each iteration of the iterative process: analyzing the text embedding vector to determine whether the portion of the text message created so far is grammatically correct and generating a second score based on the determination; and updating the portion of the text message created so far based on the second score.
[0020] In other features, the method further includes, for each iteration of the iterative process: analyzing the text embedding vector to determine how many words in the portion of the text message created so far are not car words and generating a third score based on the determination; and updating the portion of the text message created so far based on the third score.
[0021] In other features, the method further comprises: for each iteration of the iterative process, generating a score based on the portion of the text message created to date; generating an overall score based on the score; and updating the portion of the text message created to date based on the overall score.
[0022] Among other features, the scores include i) a first score indicating how closely the text embedding vector matches the image embedding vector, ii) a second score indicating how grammatically correct the portion of the text message created so far is, and iii) a third score indicating how many non-car words are included in the portion of the text message created so far.
[0023] In other features, the method further includes: receiving the image embedding vector and a car vocabulary; and generating a text message based on the image encoding vector and the car vocabulary, such that the text message is specific to the car.
[0024] In other features, the method further comprises replacing one or more words of the portion of the text message created to date based on the at least one score.
[0025] A system is disclosed, comprising: a traffic object detection module, configured to detect traffic objects in a vehicle environment; an attention map highlighting module, configured to i) generate an attention map, ii) highlight at least one of relevant traffic objects among the traffic objects and an area in which the relevant traffic objects among the traffic objects are located in the attention map, and iii) not highlight non-relevant traffic objects and other non-traffic objects among the traffic objects in the attention map; an image encoder, configured to encode an environment image received from an imaging device based on the attention map and generate an image embedding vector; a pre-trained language model (PLM) module, configured to perform an iterative process, including iteratively selecting and appending The PLM module is configured to select text based on at least one score during each iteration of the iterative process, the text message being a specific description of what is perceived in the vehicle environment; a text encoder configured to encode a portion of the text message created so far to generate a text embedding vector; one or more modules configured to score the portion of the text message created so far based on the image embedding vector and the text embedding vector to generate at least one score, wherein the PLM module is configured to update the portion of the text message created so far based on the at least one score; and at least one output device configured to output the text message for an occupant of the vehicle when completed.
[0026] The system further includes a perception and decision module configured to perform at least one of an autonomous driving maneuver and a collision avoidance countermeasure based at least on information in the text message describing the vehicle environment.
[0027] Wherein, for each iteration of the iterative process: the one or more modules include a vector comparison module configured to compare the text embedding vector with the image embedding vector and generate a first score; and the PLM module is configured to update a portion of the text message created so far based on the first score.
[0028] Wherein, for each iteration of the iterative process: the one or more modules include a language scoring module configured to analyze the text embedding vector to determine whether a portion of the text message created so far is grammatically correct and generate a second score based on the determination; and the PLM module is configured to update the portion of the text message created so far based on the second score.
[0029] Wherein, for each iteration of the iterative process: the one or more modules include a car vocabulary scoring module configured to analyze the text embedding vector to determine how many words in the portion of the text message created so far are not car words and generate a third score based on the determination; and the PLM module is configured to update the portion of the text message created so far based on the third score.
[0030] Among them, the one or more modules include multiple scoring modules and an overall scoring module; the multiple scoring modules are configured to generate multiple scores based on the portion of the text message created so far for each iteration of the iterative process; the overall scoring module is configured to generate an overall score based on the multiple scores; and the PLM module is configured to update the portion of the text message created so far based on the overall score.
[0031] wherein the multiple scores include i) a first score indicating how closely the text embedding vector matches the image embedding vector, ii) a second score indicating how grammatically correct the portion of the text message created so far is, and iii) a third score indicating how many non-car words are included in the portion of the text message created so far.
[0032] The system further includes a car scene captioning module configured to receive the image embedding vector and the car vocabulary, and based on the image embedding vector and the car vocabulary, generate a text message such that the text message is specific to the car.
[0033] Therein, the PLM module is configured to replace one or more words of a portion of a text message created to date based on the at least one score.
[0034] Therein, the PLM module is configured to start with a sentence prefix and iteratively append car words to the sentence prefix to generate a text message.
[0035] Wherein, at least one output device includes a display, a speaker, and a haptic device.
[0036] A method for providing image captions for an image of a vehicle environment is disclosed, the method comprising: detecting traffic objects in the vehicle environment; generating an attention map; highlighting at least one of relevant traffic objects among the traffic objects and an area in which the relevant traffic objects among the traffic objects are located in the attention map; not highlighting non-relevant traffic objects and other non-traffic objects among the traffic objects in the attention map; encoding the image and generating an image embedding vector based on the attention map; performing an iterative process including iteratively selecting and appending text to create a text message; selecting text based on at least one score during each iteration of the iterative process, the text message being a specific description of what is perceived in the environment of the vehicle; encoding a portion of the text message created so far to generate a text embedding vector; scoring the portion of the text message created so far based on the image embedding vector and the text embedding vector to generate at least one score; updating the portion of the text message created so far based on the at least one score; and outputting the text message for an occupant of the vehicle when completed.
[0037] The method further includes executing at least one of a vehicle driving maneuver and a collision avoidance maneuver based at least on information in the text message describing the vehicle environment.
[0038] The method further includes, for each iteration of the iterative process: comparing the text embedding vector to the image embedding vector and generating a first score; and updating the portion of the text message created so far based on the first score.
[0039] The method further includes, for each iteration of the iterative process: analyzing the text embedding vector to determine whether the portion of the text message created so far is grammatically correct and generating a second score based on the determination; and updating the portion of the text message created so far based on the second score.
[0040] The method further includes, for each iteration of the iterative process: analyzing the text embedding vector to determine how many words in the portion of the text message created so far are not car words and generating a third score based on the determination; and updating the portion of the text message created so far based on the third score.
[0041] The method further includes: for each iteration of the iterative process, generating a plurality of scores based on the portion of the text message created to date; generating an overall score based on the plurality of scores; and updating the portion of the text message created to date based on the overall score.
[0042] wherein the multiple scores include i) a first score indicating how closely the text embedding vector matches the image embedding vector, ii) a second score indicating how grammatically correct the portion of the text message created so far is, and iii) a third score indicating how many non-car words are included in the portion of the text message created so far.
[0043] The method further includes: receiving the image embedding vector and a car vocabulary; and generating a text message based on the image encoding vector and the car vocabulary, such that the text message is specific to a car.
[0044] The method further includes replacing one or more words of a portion of a text message created to date based on the at least one score.
[0045] Further areas of applicability of the present disclosure will become apparent from the detailed description, claims and drawings.The detailed description and specific examples are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present disclosure will be more fully understood based on the detailed description and accompanying drawings, in which:
[0047] Figure 1 is a functional block diagram of a vehicle including an image captioning module according to the present disclosure;
[0048] Figure 2 is a functional block diagram of an image captioning module including an automobile scene captioning module according to the present disclosure;
[0049] Figure 3 is a functional block diagram of a car scene caption module according to the present disclosure; and
[0050] Figure 4A and Figure 4B (collectively, FIG. 4 ) illustrates a perceptual method including an image captioning process according to the present disclosure.
[0051] Among the drawings, reference numerals may be repeated to identify similar and / or identical elements. DETAILED DESCRIPTION
[0052] Generating text descriptions from vehicle images helps communicate the autonomous vehicle's camera-based observations to the vehicle's occupants (e.g., driver and passengers). The text descriptions (or messages) can be displayed on the vehicle's human-machine interface (e.g., a center counselor display). The observations enhance the vehicle's perception and decision-making capabilities. Conventional image capture techniques are limited and are often inaccurate and / or incorrect in describing the vehicle scene (also referred to as the environmental situation).
[0053] The examples set forth herein include a perception and decision system and an image capture module configured to generate descriptive and perceptual messages (referred to herein as "messages"). The messages may include text messages and / or audio messages indicating environmental situations. The image capture module may be implemented as part of the perception and decision system or separate from and in communication with the perception and decision system. The messages are customized for each environmental situation and are specific to each environmental situation. As an example, the message may include a text message that accurately describes the situation, such as a message indicating "a pedestrian intends to cross the road that the host vehicle is currently traveling on" or "the road (or lane) of the host vehicle is merging into a highway with traffic." The text message may be displayed together with a corresponding image of the environment and / or area to which the text message relates.
[0054] The example further includes a visual language model that has been pre-trained and directs attention to traffic objects present in the image scene. As used herein, "traffic object" refers to a vehicle, pedestrian, cyclist, or other object that is in the path of the host vehicle or close to the path of the host vehicle and is relevant to making autonomous vehicle operation decisions. Traffic objects do not generally refer to, for example, trees, buildings, birds, etc. unless they are in the path of the host vehicle. The image is encoded using the visual language model to generate an image embedding vector. A text description is generated using the image embedding vector with a specific emphasis on automotive vocabulary. Automotive vocabulary refers to vocabulary that can be used to describe the vehicle environment.
[0055] The example further includes generating a textual interpretation of the automotive environment of the host vehicle based on images captured by the host vehicle. The host vehicle may be a fully autonomous, partially autonomous, or non-autonomous vehicle. The textual description i) conveys the perception of the host vehicle's environment as determined by the host vehicle to the driver or passenger, ii) enhances the visual language model used for automotive perception tasks, iii) provides input to the controller (or planner) of the host vehicle to aid in decision making, and iv) captures critical situations (e.g., situations with high probability of collision if no evasive action is taken) during real-time vehicle operation to continuously improve autonomous vehicle operations. The textual description of the automotive scene remains of great value for relaying information to automotive occupants via the automotive human-machine interface and for performing perception and decision operations. The textual description is aligned with the knowledge of the decision system and corresponding decision operations.
[0056] Figure 1A vehicle 100 is shown including a perception and decision system 102 having a vehicle control module 104, which, as shown, includes a perception and decision module 105 and an image captioning module 106. The perception and decision module 105 performs perception determination operations, planning operations, and decision operations. The image captioning module 106 generates descriptive and perceptual messages that describe what the perception and decision system 102 perceives as the environment (or scene) of the vehicle 100. Each message describes not only the captured image, but also what the perception and decision system 102 sees and predicts (i.e., perceives) from the captured image. The image captioning module 106 is trained for automotive scenarios. The messages may be displayed, audibly played, and / or indicated via a display 108 and / or one or more speakers 109 of a human-machine interface (HMI) 107 and / or via one or more other output devices. The following is a description of the image captioning module 106. Figure 2 -4 further describes the image captioning module 106. The vehicle control module 104 can perform various operations based on the message, as further described below. The perception and decision module 105 can perform autonomous operations based on the message.
[0057] The vehicle 100 further includes sensors 110, memory 112, accelerator pedal actuator 114, steering system 116, braking system 117, and propulsion system 118. The sensors 110 may include radar and / or lidar sensors 120, image devices (e.g., cameras) 122, vehicle speed sensors 124, acceleration sensors (e.g., longitudinal and lateral acceleration sensors) 126, and other sensors 128. The memory 112 may store sensor data 132, algorithms 134, applications 136, parameters 138, and other data 140 (e.g., automotive data sets). The sensor data 132 may include data collected from the sensors 110 and / or other sensors such as an accelerator position sensor 141 of the accelerator pedal actuator 114. The accelerator pedal actuator 114 and the accelerator position sensor 141 and / or other devices mentioned herein may be connected to the vehicle control module 104 via a controller area network (CAN) or other network bus 143. The algorithms 134 include the algorithms mentioned herein implemented by the image captioning module 106. The applications 136 may include the image captioning module 106 and / or other applications.
[0058] The propulsion system 118 may include one or more torque sources, such as one or more motors and / or one or more engines (eg, internal combustion engines). Figure 1In the example shown in , the vehicle 100 includes an engine 150 and one or more motors 152. The torque sources are independently controlled. The propulsion system 118 includes a motor control system 154, which includes one or more motors 152 and a motor control module 156, which can control the operation of the one or more motors 152 based on signals from the vehicle control module 104.
[0059] The vehicle control module 104 may further include a mode selection module 160 and / or a parameter adjustment module 162. The module 160 may select different operating modes. As an example, the vehicle control module 104 may operate in a fully autonomous or partially autonomous mode and may control the steering system 116, the braking system 117, and the propulsion system 118. In an embodiment, the vehicle control module 104 controls the operation of the systems 116-118 based on the message generated by the image caption module 106 and / or the corresponding information in the message.
[0060] In an embodiment, the perception and decision module 105 generates text describing the environment, which can be combined and / or compared with the text provided by the image caption module 106. If the text generated by the modules 105, 106 does not match or correspond, so that the text describes the same environment, the perception and decision module 105 can take action. For example, the perception and decision module 105 can perform an operation based on the text generated by the image caption module 106. The caption (or generated text) can add information, and the perception and decision module 105 can integrate this information as part of its input and then determine what operation to perform based on this information. Based on the text from the image caption module 106 and the text generated by the perception and decision module 105 and / or other inputs, the perception and decision module 105 can i) perform autonomous operations such as steering, braking, acceleration, etc., and / or ii) display and / or audibly play descriptive and perceptual messages, perform tactile operations via a tactile device 170 (e.g., a seat vibration device), and / or output messages and / or corresponding signals via other output devices.
[0061] As an example, the perception and decision module 105 may generate a boundary box where the pedestrian is located and determine the road boundary, and based on this information may incorrectly determine that the pedestrian does not intend to cross the road. The perception and decision module 105 may determine that the pedestrian is just standing still and does not plan to cross the road. The perception and decision module 105 may then receive a text message from the image caption module 106. The text message from the image caption module 106 provides an understanding of the scene indicating that the pedestrian does intend to cross the road. If the perception and decision module 105 initially determines that the pedestrian will not cross the street and the image caption module 106 determines that the pedestrian is going to cross the street, the perception and decision module 105 determines that there is a mismatch in the determination and cannot accurately perceive the pedestrian's intention. In an embodiment, the perception and decision module 105 and the image caption module 106 recheck the state determination. In another embodiment, the perception and decision module 105 operates out of sufficient caution, by, for example, reducing the speed of the vehicle to avoid a collision, as if the pedestrian is going to cross the street.
[0062] Figure 2 An image captioning module 106 is shown that can receive output from radar and lidar sensors 120 and imaging devices 122, receive a car vocabulary from memory 112, and generate an output image caption (or descriptive message) of the environment visible in one or more images captured by the imaging device 122. The image captioning module 106 includes a traffic object detection module 200, an attention map highlighting module 202, an image encoder 204, a car scene captioning module 206, and a car vocabulary module 208.
[0063] The traffic object detection module 200 can receive the output of the radar and / or lidar sensor 120 and the imaging device 122 and detect traffic objects, such as vehicles, pedestrians, cyclists, crosswalks, traffic lights, traffic signs, and road features (e.g., lane boundaries, lane types including upcoming traffic lanes and incoming traffic lanes). The traffic object detection module 200 can implement a neural network and / or a convolutional network to detect traffic objects. The output from the radar and lidar sensor 120 is represented as signal 201. The output from the imaging device 122 is represented as signal 203. The output of the traffic object detection module 200 is represented as signal 205.
[0064] The attention map highlighting module 202 creates an image similar to the image captured by the imaging device 122, except that the created image is in the form of an attention map that highlights traffic objects. The created image is represented as a signal 207. In an embodiment, the created image highlights the traffic object of interest. The created image includes pixels of the traffic object having a higher intensity than pixels of other objects (non-traffic objects). The intensity of the pixels of the traffic object increases, while the intensity of the pixels of the non-traffic object decreases.
[0065] The image encoder 204 encodes the image received from the imaging device 122. Each encoded image provides an image embedding vector represented as a signal 209, which includes a value indicating the vehicle environment. The image encoder 204 may include a neural network and implement a visual language model, such as a contrastive language image pre-training (CLIP) model. The CLIP model learns overtime. The CLIP model receives an image and encodes the image to generate an image embedding vector. The neural network is trained to encode the image into a feature space to obtain a corresponding image embedding vector, each of which includes the features of the image.
[0066] The visual language model may include multiple layers. Data passes through each layer and is further processed by each of the layers, making the data more and more descriptive of the image. Each layer includes patches of traffic objects that are respectively used for highlighting. In an embodiment, the attention map highlighting module 202 increases the attention weight of the patch or feature vector in the last attention layer of the visual language model of the image encoder 204 via signal 207. In another embodiment, the attention map highlighting module 202 increases the attention weight of the patch or feature vector of two or more layers of the image encoder 204 associated with the traffic object and reduces the weight of the feature vector associated with the patch that does not involve the traffic object. Therefore, the emphasis on traffic objects can occur at the last layer of the visual language model and optionally at one or more layers before the last layer of the visual language model.
[0067] Each patch has an associated feature vector. Each feature vector changes layer by layer so that the feature vector becomes more descriptive of the feature (or traffic object). Feature vectors associated with non-traffic objects contribute less at each subsequent layer. Each layer has the same number of feature vectors and the same size of feature vectors, but each subsequent layer has feature vectors that process more and are more accurately descriptive of the image than the feature vectors of the previous layer. Each layer combines more and more information and learns more complex functions. The feature vectors of the final attention layer are combined to provide an image embedding vector for that image.
[0068] When there is text describing the image, the text can be encoded by a text encoder to provide a text embedding vector. Figure 3 An example text encoder is shown in . The text embedding vector is similar to the image embedding vector of the same image. The text embedding vector of the text describing the image is similar to the image embedding vector of the image. This is true when the text describes the image. This is not true when the text does not describe the image. When the text does not describe the image, then the image embedding vector is different from the text embedding vector.
[0069] The image encoder 204 is pre-trained for traffic objects and generates an image embedding vector associated with the traffic object. The image embedding vector describes the traffic object highlighted by the attention map highlighting module 202. The image embedding vector includes much fewer values than the number of pixels of the corresponding image. For example, an image may include 1 million pixels, and the image embedding vector may include 500 values. The 500 values represent relevant information describing the image. As a result, modules 202, 204, and 206 highlight the information of interest in the image and provide a description of the traffic object of interest.
[0070] In an embodiment, the image encoder 204 is trained prior to use. This may include feeding a text message to a text encoder (e.g., Figure 3 The image encoder 204 is trained until the vectors match and / or are similar.
[0071] The car scene caption module 206 receives the signals 205, 209 and the car vocabulary 211 from the car vocabulary module 208 and generates and outputs the image caption (or text message) 213. The car vocabulary module 208 obtains car terms from the car data set 215 stored in the memory 112. The car data set 215 has text descriptions of different scenes. The car vocabulary module 208 helps car scene capture by limiting the vocabulary to car vocabulary. The car vocabulary module 208 obtains words, phrases and / or captions from previous car scenes obtained when performing learning operations. Words in the stored messages of many car scenes are used as a vocabulary dataset for describing the encountered scenes (or encountered car situations).
[0072] In an embodiment, the car scene captioning module 206 generates a text message 213 such that: i) the text embedding vector of the text message 213 has a small cosine distance from the image embedding vector of the corresponding image; ii) the words in the text message (or sentence) are restricted to the vocabulary obtained from the description of the car scene and thus restricted to the car vocabulary; and iii) the language used in the sentence is grammatically correct. This is a language model restriction. The car scene captioning module 206 generates a sentence that, when implemented by a visual language model Figure 3 When encoded by the text encoder of , a text embedding vector similar to the image embedding vector generated by the image encoder 204 is provided. When the text embedding vector is similar to the image embedding vector, the sentence describes the image. The car scene captioning module 206 can determine the cosine distance between the vectors to determine whether the text embedding vector is similar to the image embedding vector. The smaller the distance, the more the text embedding vector matches the image embedding vector. Other methods such as Manhattan distance (or L1 norm) or Euclidean distance (or L2 norm) methods can be used to determine the difference between the vectors.
[0073] Figure 3 The car scene caption module 206 is shown, which includes a prefix module 300, a pre-trained language model (PLM) module 302, a text encoder 304, a vector comparison module 306, a language scoring module 308, a car vocabulary scoring module 310 and an overall scoring module 312. The prefix module 300 generates a sentence prefix (referred to as a "prefix" herein), based on which a text message is generated. Words are iteratively added to the prefix to eventually generate a text message that will be conveyed to the vehicle occupants. As an example, the prefix may be "An image of a". In an embodiment, the prefix does not describe the current car situation of the main vehicle. The prefix may be selected from the stored predefined prefixes. In an embodiment, the prefix with the highest probability of being suitable for the current car shape is selected. In an embodiment, the prefix is selected based on the image. In another embodiment, the prefix is not selected based on the image.
[0074] The PLM module 302 receives a prefix or a current message and expands the prefix or the current message by one or more words. In an embodiment, the PLM module 302 expands the current message word by word based on the overall score generated by the overall score module 312. The PLM module 302 selects words that satisfy the functions (or restrictions) associated with the modules 306, 308, 310 to construct the output image caption 213.
[0075] The text encoder 304 receives the prefix or current message (i.e., the portion of the entire message created so far) and encodes the message to generate a text embedding vector. The text encoder 304 is a visual language text encoder, such as a CLIP text encoder. The text encoder 304 is trained with a visual language model. When the message is complete, the text encoder 304 encodes the complete message and provides a text embedding vector that matches the output of the image encoder 204 for the corresponding image that has been encoded.
[0076] The vector comparison module 306 compares the output of the text encoder 304 with the output of the image encoder 204 for each portion of the message generated so far. In an embodiment, this includes determining the cosine distance between the text and image embedding vectors and generating a first score based on the cosine distance. The smaller the cosine distance, the better (or higher) the first score.
[0077] The language scoring module 308 evaluates the message generated so far and generates a second score. The language scoring module 308 checks whether the message generated so far is grammatically correct. When selecting words, the visual language model of the PLM module 302 selects one or more words with a high probability of being grammatically correct. In an embodiment, the PLM module 302 selects words based on historical text rather than on corresponding images. The PLM module 302 has many options to choose from when selecting words and selects the words with the highest probability of being grammatically correct. The selection can be based on words that have been pre-selected according to assumptions. The higher the probability that the message is grammatically correct so far, the higher the second score.
[0078] The automobile vocabulary scoring module 310 generates a third score based on whether the words of the message so far are automobile vocabulary words. The more words that are not automobile vocabulary words, the lower the score. The words of the automobile vocabulary include words that can be used to describe the vehicle environment, including words to describe the state of roads, traffic objects, pedestrians, cyclists, etc. The automobile vocabulary may include terms to describe actions being performed or predicted to be performed by nearby traffic objects, such as vehicles, pedestrians, and / or cyclists.
[0079] The overall score module 312 determines an overall score based on the first, second, and third scores. In an embodiment, this includes summing the first, second, and third scores to provide an overall score. In another embodiment, this includes weighting and summing the scores to provide an overall score. In an embodiment, the PLM module 302 selects words to maximize the overall score. In another embodiment, the PLM module 302 selects words that maximize the first, second, and third scores and thus maximize the overall score. The first, second, and third scores may be provided to the PLM module 302. The PLM module 302 may replace one or more words based on the first, second, third, and / or overall scores. In an embodiment, the PLM module 302 iteratively replaces words to improve one or more of the first, second, third, and overall scores. This iterative process may occur for each word and / or set of words selected in combination with previously selected words for the generated current message. The PLM module 302 may implement a neural network having components including keys and probability values. The keys and probability values of each PLM attention layer may be adjusted to maximize the overall score and / or maximize the first, second, and third scores.
[0080] As an example, the PLM module 302 may be configured to communicate with Figure 3 During the associated process, a selection is initially made from 50,000 words or sets of words, and a subset (e.g., 1,000) of the words or sets of words with the highest scores are selected when encoded and scored as described above. Words or sets of words in the subset that are not car words are removed from the subset. The PLM module 302 can then select the best match to the output of the image encoder from the remaining words or sets of words. In an embodiment, the message being generated is Figure 3 The iteration period of the process of is extended in an autoregressive manner (word by word) as described.
[0081] FIG4 shows a perceptual method including an image captioning process. The operations of FIG4 may be performed iteratively. The method is mainly about Figure 1-3 The embodiments are described below.
[0082] At 400, sensor data is collected from at least one image sensor at the host vehicle (eg, one of the imaging devices 122). This may include collecting output from the sensors 120 and the imaging devices 122. The sensor data is received at the traffic and object detection module 200.
[0083] At 402 , the traffic and object detection module 200 detects traffic objects in the host vehicle's environment, such as traffic objects in front of the host vehicle.
[0084] At 404, the attention map highlighting module 202 generates one or more attention maps including highlighted and / or re-weighted traffic objects based on the output from the sensors 120 and the imaging devices 122. Traffic objects are weighted more heavily than other objects. Any traffic object that is of interest or within and / or crossing the upcoming path of the host vehicle may be weighted more heavily than other traffic objects.
[0085] At 406, the image encoder 204 encodes one or more images received from one or more imaging sensors, such as the imaging device 122. The image encoder 204 generates one or more image embedding vectors for the one or more images, respectively.
[0086] At 408, the car vocabulary module 208 accesses the car vocabulary and provides it to the car scene caption module. The operation of the car scene caption module 206 is described with reference to operations 410, 412, 414, 416, 418, 420, 422, 424, 426, 428 and 430. Although the following operations are described with respect to a single image, the operations may be performed for each of a plurality of images. The plurality of images may have different environmental regions. In an embodiment, the text message generated for the initial image or a portion thereof is updated based on a subsequently captured image associated with the same environmental region of the initial image.
[0087] At 410, the prefix module 300 accesses, obtains, and / or generates a message prefix for the image captured at 400, as described above. At 412, the text encoder 304 encodes the prefix or message generated thus far and generates a corresponding text embedding vector.
[0088] At 414, the vector comparison module 306 compares the text embedding vector to the image embedding vector and generates a first score as described above. At 416, the language scoring module 308 evaluates the prefixes or messages generated so far and generates a second score as described above. At 418, the automotive vocabulary scoring module 310 evaluates the prefixes or messages generated so far and generates a third score as described above.
[0089] At 420, the overall rating module generates an overall rating based on the first, second, and third ratings as described above. At 422, the PLM module 302 determines whether one or more of the ratings are satisfactory. In an embodiment, the PLM module 302 determines whether the overall rating is satisfactory. In another embodiment, the PLM module 302 determines whether one or more of the first, second, and third ratings are satisfactory. One or more of the first, second, and third ratings and / or the overall rating may be compared to a predetermined threshold, and if greater than the predetermined threshold, the rating is deemed satisfactory. If one or more of the ratings are deemed satisfactory, operation 426 may be performed, otherwise operation 424 may be performed.
[0090] At 424, the PLM module 302 may replace the previously generated and / or appended text with different text. Operation 412 may be performed after operation 424. At 426, the PLM module 302 may determine whether the message is complete. If not, operation 428 may be performed, otherwise operation 430 may be performed. At 428, the PLM module 302 selects one or more words to append to the current message and / or prefix generated so far, as described above. To reduce processing time, operations 414, 416, 418, 420, and 422 may be performed simultaneously for multiple word options or selectable sets of words. The word or set of words that provides the best score(s) is selected for this iteration of the process.
[0091] At 430 , the car scene caption module 206 and / or the PLM module 302 may display the text message and / or audibly play the text message to the occupants of the host vehicle. A haptic alert may also or alternatively be generated based on the generated text message.
[0092] The above operations are meant to be illustrative examples. Depending on the application, the operations may be performed sequentially, synchronously, simultaneously, continuously, during overlapping time periods, or in a different order. In addition, depending on the implementation, the order of events, and / or some other logic (e.g., based on user experience practices or decisions), any operation may not be performed or skipped.
[0093] The examples disclosed herein provide text descriptions that serve several functions, including i) conveying the perception of an automotive vehicle to the driver and / or passengers of the vehicle, ii) enhancing visual language models for automotive perception tasks, iii) providing input to automotive vehicle planners to aid decision making, and iv) capturing situations during real-time operation of the vehicle to continuously improve automotive vehicle systems. The disclosed system uses a visual language model specific to automotive applications. The text descriptions are logged and can be sent from the host vehicle to the backend for analysis, refinement, and returning the refined text descriptions back to the host vehicle. This helps learn difficult situations, such as situations with a high probability of collision. As an example, an image showing a pedestrian approaching a vehicle can be sent to the backend for analysis. The visual language model can be pre-trained and directs attention to traffic objects present in the scene. The text description is generated using an image embedding vector, with a particular emphasis on vehicle vocabulary.
[0094] The examples disclosed herein include a method for textually depicting an automotive environment as observed via an autonomous driving perception system. The examples are applicable to explaining the perception of the environment as perceived by the perception system of the vehicle to a human driver and / or a passenger of the vehicle. The generated message can increase the confidence of the driver and / or passenger in the situational awareness of the vehicle. The systems and modules disclosed herein can be used to explain the perception of areas that are not visible to the vehicle occupants as seen by the vehicle's sensors, such as blind spots, areas positioned adjacent to the vehicle laterally, areas seen by the vehicle's rear camera, etc. The systems and modules disclosed herein can be used to explain the perception of the vehicle's surroundings as seen by the vehicle's sensors. It may be difficult for a vehicle occupant to simultaneously view areas of the surroundings.
[0095] The disclosed example integrates generated text subtitles into the input of an automotive vehicle decision system. The example includes comparing and / or combining decision scene understanding and perceptual text subtitles generated based on captured images. The example is also applicable to monitoring driver behavior, such as when a fleet manager monitors fleet drivers or when teenage drivers are monitored. Capturing and reporting critical situations, including providing text explanations (or subtitles) of the situations. The example is further applicable to insurance purposes. Online logging of scene descriptions is implemented, which can be used for vehicle insurance purposes in the event of an accident. Text subtitles can be used to refine an automotive visual language model encoder, which is used for online perception tasks, such as pedestrian road crossing intention estimation.
[0096] Data of automotive vehicles may be captured in high risk situations (e.g., collision risk situations), which may be identified from textual captions. The data may be annotated and used for system refinement and continuous learning as well as for diagnostics. Visual language models are used for perception due to their high-level semantics and open vocabulary classification capabilities. However, visual language models may be trained primarily with generic images and not necessarily specifically for automotive environments. The examples disclosed herein may be used to improve a visual language model of an automotive environment by generating many automotive images with corresponding captions.
[0097] In an embodiment, an image of a scene and a corresponding generated text message are displayed to a vehicle occupant. As a few examples, the image captioning module disclosed herein can generate text messages such as "Image shows a truck driver in the middle of a traffic collision," "Image shows a pedestrian walking along a sidewalk in a city center," "Image shows a pedestrian crossing the street as she walks," "Image shows a pedestrian in front of a busy traffic intersection," "Crossing vehicle detected ahead," and "Pedestrians approaching a vehicle will trigger logging of perception data for offline analysis."
[0098] The foregoing description is essentially merely illustrative and is by no means intended to limit the present disclosure, its application or use. The broad teachings of the present disclosure can be implemented in a variety of forms. Therefore, although the present disclosure includes specific examples, the true scope of the present disclosure should not be so limited, because after studying the drawings, the specification and the appended claims, other modifications will become clear. It should be understood that one or more steps in the method can be performed in different orders (or simultaneously) without changing the principles of the present disclosure. In addition, although each embodiment is described above as having certain features, any one or more of those features described with respect to any embodiment of the present disclosure can be implemented in the features of any other embodiment and / or implemented in combination with the features of any other embodiment, even if the combination is not explicitly described. In other words, the described embodiments are not mutually exclusive, and the replacement of one or more embodiments with each other is still within the scope of the present disclosure.
[0099] Spatial and functional relationships between elements (e.g., between modules, circuit elements, semiconductor layers, etc.) are described using various terms including: "connected," "engaged," "coupled," "adjacent," "near," "on top," "above," "below," and "disposed." Unless explicitly described as "direct," when describing a relationship between a first and a second element in the above disclosure, the relationship may be a direct relationship in the absence of other intervening elements between the first and second elements, but may also be an indirect relationship in the presence of one or more intervening elements (spatially or functionally) between the first and second elements. As used herein, the phrase at least one of A, B, and C should be interpreted to mean a logical (A or B or C), using a non-exclusive logical or, and should not be interpreted to mean "at least one of A, at least one of B, and at least one of C."
[0100] In a diagram, the direction of the arrow, as indicated by the arrow, generally indicates the flow of information (such as data or instructions) of interest to the diagram. For example, when element A and element B exchange a variety of information, but the information transmitted from element A to element B is relevant to the diagram, the arrow may point from element A to element B. This unidirectional arrow does not mean that no other information is transmitted from element B to element A. In addition, for information sent from element A to element B, element B may send a request or receipt confirmation to element A for the information.
[0101] In this application, including the definitions below, the term "module" or the term "controller" may be replaced with the term "circuit". The term "module" may refer to, be part of, or include: an application specific integrated circuit (ASIC); a digital, analog, or mixed analog / digital discrete circuit; a digital, analog, or mixed analog / digital integrated circuit; a combinational logic circuit; a field programmable gate array (FPGA); a processor circuit (shared, dedicated, or group) that executes code; a memory circuit (shared, dedicated, or group) that stores code executed by the processor circuit; other suitable hardware components that provide the described functionality; or a combination of some or all of the above, such as in a system on a chip.
[0102] The module may include one or more interface circuits. In some examples, the interface circuit may include a wired or wireless interface connected to a local area network (LAN), the Internet, a wide area network (WAN), or a combination thereof. The functionality of any given module of the present disclosure may be distributed in multiple modules connected via the interface circuit. For example, multiple modules may allow load balancing. In another example, a server (also referred to as a remote or cloud) module may complete some functionality on behalf of a client module.
[0103] The term code as used above may include software, firmware, and / or microcode, and may refer to programs, routines, functions, classes, data structures, and / or objects. The term shared processor circuit includes a single processor circuit that executes some or all code from multiple modules. The term group processor circuit includes a processor circuit that, in combination with additional processor circuits, executes some or all code from one or more modules. References to multiple processor circuits include multiple processor circuits on discrete dies, multiple processor circuits on a single die, multiple cores of a single processor circuit, multiple threads of a single processor circuit, or a combination of the above. The term shared memory circuit includes a single memory circuit that stores some or all code from multiple modules. The term group memory circuit includes a memory circuit that is combined with additional memory to store some or all code from one or more modules.
[0104] The term memory circuit is a subset of the term computer-readable medium. The term computer-readable medium as used herein does not include transient electrical or electromagnetic signals propagated through a medium (such as on a carrier wave); therefore, the term computer-readable medium may be considered to be tangible and non-transitory. Non-limiting examples of non-transitory tangible computer-readable media are non-volatile memory circuits (such as flash memory circuits, erasable programmable read-only memory circuits, or mask read-only memory circuits), volatile memory circuits (such as static random access memory circuits or dynamic random access memory circuits), magnetic storage media (such as analog or digital magnetic tape or hard disk drives), and optical storage media (such as CDs, DVDs, or Blu-ray discs).
[0105] The apparatus and methods described in this application may be implemented in part or in whole by a special-purpose computer, which is created by configuring a general-purpose computer to perform one or more specific functions embodied in a computer program. The functional blocks, flow chart components and other elements described above serve as software specifications, which can be translated into computer programs by routine work of a skilled technician or programmer.
[0106] The computer program includes processor executable instructions stored on at least one non-transitory tangible computer readable medium. The computer program may also include or rely on stored data. The computer program may include a basic input / output system (BIOS) that interacts with the hardware of the special-purpose computer, device drivers that interact with specific devices of the special-purpose computer, one or more operating systems, user applications, background services, background applications, etc.
[0107] A computer program may include: (i) descriptive text to be parsed, such as HTML (Hypertext Markup Language), XML (Extensible Markup Language), or JSON (JavaScript Object Notation) (ii) assembly code, (iii) object code generated by a compiler from source code, (iv) source code executed by an interpreter, (v) source code compiled and executed by a just-in-time compiler, etc. By way of example only, source code may be written using syntax from a language including: C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, Fortran, Perl, Pascal, Curl, OCaml, HTML5 (Hypertext Markup Language Revision 5), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Visual Lua, MATLAB, SIMULINK, and
Claims
1. A system comprising: a traffic object detection module configured to detect traffic objects in a vehicle environment; an attention map highlighting module configured to i) generate an attention map, ii) highlight at least one of a relevant traffic object among the traffic objects and a region in which the relevant traffic object among the traffic objects is located in the attention map, and iii) not highlight non-relevant traffic objects and other non-traffic objects among the traffic objects in the attention map; an image encoder configured to encode an environment image received from an imaging device and generate an image embedding vector based on the attention map; a pre-trained language model (PLM) module configured to perform an iterative process including iteratively selecting and appending text to create a text message, the PLM module being configured to select text based on at least one score during each iteration of the iterative process, the text message being a specific description of what is sensed in the vehicle environment; a text encoder configured to encode a portion of the text message created so far to generate a text embedding vector; one or more modules configured to score portions of the text message created so far based on the image embedding vector and the text embedding vector to generate at least one score, wherein the PLM module is configured to update a portion of the text message created to date based on the at least one rating; and At least one output device is configured to output a text message to an occupant of the vehicle when completed.
2. The system of claim 1 , further comprising a perception and decision module configured to perform at least one of an autonomous driving operation and a collision avoidance countermeasure based at least on information in the text message describing the vehicle environment.
3. The system according to claim 1, wherein: For each iteration of the iterative process: The one or more modules include a vector comparison module configured to compare the text embedding vector to the image embedding vector and generate a first score; and The PLM module is configured to update a portion of the text message created to date based on the first score.
4. The system according to claim 3, wherein: For each iteration of the iterative process: The one or more modules include a language scoring module configured to analyze the text embedding vector to determine whether a portion of the text message created thus far is grammatically correct and generate a second score based on the determination; and The PLM module is configured to update the portion of the text message created to date based on the second score.
5. The system according to claim 4, wherein: For each iteration of the iterative process: The one or more modules include a car vocabulary scoring module configured to analyze the text embedding vector to determine how many words in the portion of the text message created thus far are not car words and generate a third score based on the determination; and The PLM module is configured to update a portion of the text message created to date based on the third score.
6. The system of claim 1, wherein: The one or more modules include a plurality of scoring modules and an overall scoring module; The plurality of scoring modules are configured, for each iteration of the iterative process, to generate a plurality of scores based on portions of the text message created so far; The overall rating module is configured to generate an overall rating based on the plurality of ratings; and The PLM module is configured to update a portion of the text message created to date based on the overall rating.
7. The system according to claim 6, wherein: The multiple scores include i) a first score indicating how closely the text embedding vector matches the image embedding vector, ii) a second score indicating how grammatically correct the portion of the text message created so far is, and iii) a third score indicating how many non-car words are included in the portion of the text message created so far.
8. The system of claim 1 , further comprising a car scene captioning module configured to receive the image embedding vector and the car vocabulary, and based on the image embedding vector and the car vocabulary, generate a text message such that the text message is specific to the car.
9. The system according to claim 1, wherein: The PLM module is configured to replace one or more words of a portion of a text message created to date based on the at least one score.
10. The system according to claim 1, wherein: The PLM module is configured to start with a sentence prefix and iteratively append car words to the sentence prefix to generate a text message.
Citation Information
Cited By
Techniques for controlling autonomous vehicles using vision-language models
US20250333079A1