Object detection verification for vehicle perception system

By introducing the architecture of semantic extraction, semantic generation, consistency evaluation and diagnosis modules into the perception system, the problem of accuracy verification when the perception system detects objects in the area around the vehicle is solved, and the accuracy verification of object determination is realized, improving the safety and reliability of vehicle operations.

CN120182955APending Publication Date: 2025-06-20GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410196693.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-02-22
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When existing perception systems detect objects in areas around the vehicle, it is difficult to verify the accuracy of object determination, which may lead to inaccurate object determination affecting vehicle operation.

Method used

A architecture is adopted, including a semantic extraction module, a semantic generation module, a consistency evaluation module and a diagnostic module, to verify the accuracy of object determination by extracting visual features from image frames, generating text language sentences, and calculating consistency scores for image and text descriptions.

Benefits of technology

Effectively verify the accuracy of the sensing system detecting objects in the area around the vehicle, reduce the impact of inaccurate object determination on vehicle operation, and improve the safety and reliability of vehicle operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182955A_ABST
    Figure CN120182955A_ABST
Patent Text Reader

Abstract

The invention relates to object detection verification for a vehicle perception system. A system for verifying the accuracy of object determination with a perceptual system for objects detected within a surrounding area of a vehicle. The system may include a semantic extraction module configured to generate semantic information for the object; a semantic generation module configured to generate a plurality of semantic subtitles based on the semantic information; the consistency evaluation module is configured to be used for generating a consistency score for the semantic subtitles; and a diagnostic module configured to verify the accuracy of the object determination based on the consistency score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to verifying the accuracy of object determination using a perception system, such as but not limited to verifying the accuracy of object determination using a perception system for objects detected in a surrounding area of a vehicle. Background Art

[0002] A perception system can be used to detect objects in a surrounding area of a vehicle for purposes of assisting navigation or otherwise influencing the operation of the vehicle. The perception system can be used to detect a wide range of objects, such as other vehicles, pedestrians, road signs, traffic signs, buildings, landmarks, etc. The perception system can generate object determinations to describe the movement, location, size, shape, color, and / or other characteristics detected for the object. The object determinations can be used with other systems on the vehicle to facilitate various dependent processes, which can have varying degrees of influence on the operation of the vehicle. Due to the complexity and variability associated with object determination by the perception system, it may be desirable to verify the accuracy of the object determination to avoid inaccurate object determinations being used to influence the operation of the vehicle in an undesirable manner. Summary of the Invention

[0003] One aspect of the present disclosure relates to an architecture that is operable to verify the accuracy of object determination using a perception system, such as but not limited to verifying the accuracy of object determination using a perception system configured to detect objects of a type in a surrounding area of a vehicle. The contemplated accuracy verification can be based on a language-image model and includes: extracting visual features, such as components, scene graphs, etc., from a scene, generating text language sentences based on visual and non-visual information to provide a description of the objects in the scene based on different logics, and generating a consistency score for verifying the accuracy of an accompanying object determination based on the image and the generated description of the object.

[0004] One aspect of the present disclosure relates to a system for verifying the accuracy of object determination based on detecting objects in a surrounding area of a vehicle using a perception system. The system can include: a semantic extraction module configured to generate semantic information for the object; a semantic generation module configured to generate a plurality of semantic captions based on the semantic information; a consistency evaluation module configured to generate a consistency score for the semantic captions; and a diagnostic module configured to verify the accuracy of the object determination based on the consistency score.

[0005] The semantic extraction module can be configured to generate the semantic information for the object determination based on data included within an image frame processed by the perception system.

[0006] The semantic generation module may be configured to generate the semantic captions to explicitly include text language that describes the scene associated with the object.

[0007] The consistency assessment module may include a language-image model operable to generate the consistency score.

[0008] The consistency assessment module may be configured to verify the accuracy of the object determination based on a relative comparison of the consistency scores.

[0009] The perception system is operable to detect objects across multiple object types, the semantic generation module is configured to generate at least one semantic caption for each object type, and the diagnostic module is configured to verify the accuracy of the object determination when the object determination matches the semantic caption associated with the consistency score having the maximum relative rating within the relative comparison.

[0010] The perception system is operable to detect objects across multiple object types, the semantic generation module is configured to generate at least one semantic caption for each object type, optionally, wherein each semantic caption includes the relative size and location of the object type associated therewith. The diagnostic module may be configured to verify the accuracy of the object determination when the object determination matches the semantic caption associated with the consistency score having the maximum relative rating within the relative comparison.

[0011] The semantic generation module may be configured to generate an object caption as one of the semantic captions, optionally, wherein the object caption is based on an object identifier selected by the perception system for the object. The semantic generation module may be configured to determine a plurality of component identifiers of the object from a set of component identifiers listed in a mapping module of the object identifier.

[0012] The semantic generation module may be configured to generate a plurality of component captions as part of the semantic captions, optionally, wherein each component caption identifies a different one of the component identifiers or a combination of more than one component identifier. The diagnostic module may be configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the component captions.

[0013] The semantic generation module may be configured to generate a plurality of component captions as part of the semantic captions, optionally, wherein each component caption identifies a different one of the component identifiers or a combination of more than one component identifier and includes the relative size and location of the component associated with its component identifier. The diagnostic module may be configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the component captions.

[0014] The semantic generation module may be configured to generate a plurality of detailed captions as part of the semantic captions. Optionally, each detailed caption identifies detailed information of the object, including at least one associated characteristic of text, material, and color. The diagnostic module may be configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the detailed captions.

[0015] The semantic generation module may be configured to generate a plurality of common sense captions as part of the semantic captions. Optionally, each common sense caption identifies common sense information of the object, including at least one associated characteristic of use, scenario, and relationship with neighboring objects. The diagnostic module may be configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the common sense captions.

[0016] The semantic generation module may be configured to generate a plurality of combined captions as part of the semantic captions. Optionally, each combined caption combines at least one of a plurality of component captions, detailed captions, and common sense captions. The diagnostic module may be configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the combined captions.

[0017] The diagnostic module may be configured to determine the object as verified or unverified based on the consistency score.

[0018] The system may include: a mitigation module configured to direct a trajectory planner system on the vehicle when the object is determined to be unverified.

[0019] One aspect of the present disclosure relates to a method for verifying the accuracy of an object determination made using a perception system based on detecting an object in a surrounding area of a vehicle. The method may include: generating semantic information for the object for the object determination based on data included in an image frame processed by the perception system; generating a plurality of semantic captions based on the semantic information, the semantic captions including a text language describing a scenario associated with the object; generating a consistency score for the semantic captions; and determining the object determination as either verified or unverified based on the consistency score.

[0020] The method may include: generating an object caption as one of the semantic captions, where the object caption is based on an object identifier selected for the object by the perception system; determining a plurality of component identifiers of the object from a set of component identifiers listed in a mapping module of the object identifier; generating a plurality of component captions as part of the semantic captions, optionally, where each component caption identifies a different one of the component identifiers or a combination of more than one component identifier; and verifying the accuracy of the object determination when a consistency score of the object caption is less than at least one consistency score of the component captions.

[0021] One aspect of the present disclosure relates to a vehicle that includes: a propulsion system configured to propel the vehicle; a perception system configured to detect objects in a surrounding area of the vehicle; a trajectory planner system configured to guide an operation of the propulsion system based on an object determination of the objects by a perception controller; and a verification system. The verification system may be configured to: generate semantic information for the object for the object determination based on data included in an image frame processed by the perception system; generate a plurality of semantic captions based on the semantic information, the semantic captions including text language describing a scene associated with the object; generate a consistency score for the semantic captions; and determine the object determination as one of verified or unverified based on the consistency score.

[0022] The trajectory planner system may be configured to guide the operation of the propulsion system based on a verification notification provided from the verification system indicating whether the object determination is verified or unverified.

[0023] One aspect of the present disclosure relates to a system for verifying the accuracy of an object determination made using a perception system, the perception system making the object determination based on detecting an object in a surrounding area of a vehicle, the system including: a semantic extraction module configured to generate semantic information for the object; a semantic generation module configured to generate a plurality of semantic captions based on the semantic information; a consistency evaluation module configured to generate a consistency score for the semantic captions; and a diagnostic module configured to verify the accuracy of the object determination based on the consistency score.

[0024] The semantic extraction module is configured to generate the semantic information for the object determination based on data included in an image frame processed by the perception system.

[0025] The semantic generation module is configured to generate the semantic captions to explicitly include text language describing a scene associated with the object.

[0026] The consistency assessment module includes a language-image model operable to generate the consistency score.

[0027] The consistency assessment module is configured to verify the accuracy of the object determination based on a relative comparison of the consistency scores.

[0028] The perception system is operable to detect objects across multiple object types; the semantic generation module is configured to generate at least one semantic caption for each object type; and the diagnostic module is configured to verify the accuracy of the object determination when the object determination matches the semantic caption associated with the consistency score having the maximum relative rating within the relative comparison.

[0029] The perception system is operable to detect objects across multiple object types; the semantic generation module is configured to generate at least one semantic caption for each object type, where each semantic caption includes the relative size and location of the object type associated therewith; and the diagnostic module is configured to verify the accuracy of the object determination when the object determination matches the semantic caption associated with the consistency score having the maximum relative rating within the relative comparison.

[0030] The semantic generation module is configured to generate an object caption as one of the semantic captions, where the object caption is based on an object identifier selected by the perception system for the object.

[0031] The semantic generation module is configured to determine a plurality of component identifiers of the object from a set of component identifiers listed in the mapping module of the object identifier.

[0032] The semantic generation module is configured to generate a plurality of component captions as part of the semantic captions, where each component caption identifies a different one of the component identifiers or a combination of more than one component identifier; and the diagnostic module is configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the component captions.

[0033] The semantic generation module is configured to generate a plurality of component captions as part of the semantic captions, where each component caption identifies a different one of the component identifiers or a combination of more than one component identifier and includes the relative size and location of the component associated with the component identifier; and the diagnostic module is configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the component captions.

[0034] The semantic generation module is configured to generate a plurality of detailed captions as part of the semantic captions, where each detailed caption identifies detailed information of the object, including at least one associated characteristic of text, material, and color; and the diagnostic module is configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one consistency score of the detailed captions.

[0035] The semantic generation module is configured to generate a plurality of common sense captions as part of the semantic captions, where each common sense caption identifies common sense information of the object, including at least one associated characteristic of use, scenario, and relationship with neighboring objects; and the diagnostic module is configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one consistency score of the common sense captions.

[0036] The semantic generation module is configured to generate a plurality of combined captions as part of the semantic captions, where each combined caption combines at least one of a plurality of component captions, detailed captions, and common sense captions; and the diagnostic module is configured to verify the accuracy of the object determination when the consistency score of the object caption is less than at least one consistency score of the combined captions.

[0037] The diagnostic module is configured to determine the object as verified or unverified based on the consistency score.

[0038] The system further includes: a mitigation module configured to direct a trajectory planner system on the vehicle when the object is determined to be unverified.

[0039] One aspect of the present disclosure relates to a method for verifying the accuracy of an object determination made using a perception system that makes the object determination based on detecting an object in a surrounding area of a vehicle. The method includes: generating semantic information for the object for the object determination based on data included in an image frame processed by the perception system; generating a plurality of semantic captions based on the semantic information, the semantic captions including text language describing a scenario associated with the object; generating a consistency score for the semantic captions; and determining the object determination as one of verified or unverified based on the consistency score.

[0040] The method further includes: generating an object caption as one of the semantic captions, where the object caption is based on an object identifier selected for the object by the perception system; determining a plurality of component identifiers of the object from a set of component identifiers listed in a mapping module of the object identifier; generating a plurality of component captions as part of the semantic captions, where each component caption identifies a different one of the component identifiers or a combination of more than one component identifier; and verifying the accuracy of the object determination when the consistency score of the object caption is less than at least one of the consistency scores of the component captions.

[0041] One aspect of the present disclosure relates to a vehicle that includes: a propulsion system configured to propel the vehicle; a perception system configured to detect objects in a surrounding area of the vehicle; a trajectory planner system configured to guide an operation of the propulsion system based on an object determination of the objects by a perception controller; and a verification system configured to: generate semantic information for the objects for the object determination based on data included within an image frame processed by the perception system; generate a plurality of semantic captions based on the semantic information, the semantic captions including a text language that describes a scene associated with the objects; generate a consistency score for the semantic captions; and determine the object determination as one of verified or unverified based on the consistency score.

[0042] The trajectory planner system is configured to guide the operation of the propulsion system based on a verification notice provided from the verification system indicating whether the object determination is verified or unverified.

[0043] These and other features and advantages of the present teachings may be readily apparent with the following detailed description of modes for carrying out the present teachings when taken in conjunction with the accompanying drawings. It should be understood that even though the following drawings and embodiments may be described separately, their individual features may be combined into additional embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings, which may be incorporated in and form a part of this specification, illustrate implementations of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0045] Figure 1 Illustrates a vehicle according to a non-limiting aspect of the present disclosure.

[0046] Figure 2 Illustrates a schematic diagram of a verification system according to a non-limiting aspect of the present disclosure.

[0047] Figure 3 Illustrates a flowchart of a method for verifying the accuracy of an object determination according to a non-limiting aspect of the present disclosure. DETAILED IMPLEMENTATION MANNER

[0048] As needed, detailed embodiments of the present disclosure may be disclosed herein; however, it is understood that the disclosed embodiments may only illustrate the present disclosure that can be embodied in various alternative forms. The drawings may not necessarily be to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, the specific structural and functional details disclosed herein may not need to be construed as restrictive, but only as a representative basis for teaching those skilled in the art to adopt the present disclosure in various ways.

[0049] Figure 1 FIG. 12 illustrates a vehicle 12 according to a non-limiting aspect of the present disclosure. The vehicle 12 may be interchangeably referred to as an electric or hybrid vehicle 12 and may include a traction motor 14 that is operable to convert electrical power into mechanical power for work purposes, such as mechanically driving a driveline 16 to propel the vehicle. Since the powertrain 16 optionally includes an internal combustion engine (ICE) 18 for generating mechanical power, the vehicle 12 is illustrated as a hybrid type. The vehicle 12 may alternatively omit the electric motor 14 and instead be propelled only using the ICE 18. The powertrain 16 may include components to facilitate the conveyance of rotational force from the traction motor 14 and / or the ICE 18 to one or more of the wheels 20, 22, 24, 26. The vehicle 12 may include a rechargeable energy storage system (RESS) 30 that is configured to store electrical power and supply the electrical power to the traction motor 12 and / or other components, systems, etc. on the vehicle 12 via a first bus 34 and a second bus 36, for example. The vehicle 12 may include a vehicle controller 38 to facilitate monitoring, control, measurement, and otherwise guiding the operation, execution, etc. on the vehicle 12, which may include performing measurements, reading readings, or otherwise collecting data to facilitate diagnosing binding events and accordingly managing the RESS 30 to mitigate its impact while keeping the operation of the RESS 30 within defined operating boundaries.

[0050] Vehicle 12 may include a perception system 40 configured to detect an object 42 in or near a surrounding area of vehicle 12 based on information collected using a sensor system 44. Although a single object 42 is shown, perception system 40 may be configured to detect multiple objects 42 simultaneously, including detecting objects 42 that are located in other regions relative to vehicle 12 or otherwise have a different spatial relationship to vehicle 12 compared to the illustrated object 42. Perception system 40 may be operable with a trajectory planner 46 and / or additional systems (not shown) on and / or outside of vehicle 12 for purposes of assisting navigation or otherwise influencing the operation of vehicle 12. Perception system 40 may be used to detect a wide variety of objects 42, such as other vehicles, pedestrians, road signs, traffic signs, buildings, landmarks, etc. Perception system 40 may generate an object determination to describe detected movement, location, size, shape, color, and / or other characteristics of object 42. The object determination may be used with trajectory planner 46 or other systems associated with vehicle 12 to facilitate various dependent processes, which may have varying degrees of impact on the operation of vehicle 12. Due to the complexity and variability associated with object determinations made by perception system 40, vehicle 12 may include a verification system 48 operable to verify the accuracy of the object determination to avoid inaccurate object determinations from being used to influence the operation of vehicle 12 in an undesirable manner. In other embodiments, verification system 48 may be implemented in the background and perform verification tasks offline. In such a use case, the output of verification system 48 may be used to support the analysis and development of perception module 40.

[0051] Figure 2 FIG. shows a schematic diagram of a verification system 48 in accordance with a non-limiting aspect of the present disclosure. For illustrative purposes, verification system 48 is shown as being operable with perception system 40, sensor system 44, and trajectory planner 46; however, the present disclosure fully contemplates that verification system 48 may be operable with a wide variety of other devices and systems. For illustrative purposes, verification system 48 is also described as being included on vehicle 12 for perception system 40, as the present disclosure fully contemplates that verification system 48 may operate in other environments, including non-vehicle-based environments where perception system 40 may be included or implemented as part of a robot, machine, autonomous operating device / equipment, etc. Sensor system 44 may include components for facilitating the detection of object 42, including various sensors configured to sense the surrounding environment of the vehicle. The sensors may include, for example, cameras and other sensors (e.g., radar, lidar (LIDAR), sonar, etc.) disposed at various locations inside and outside of the vehicle. The system may be configured to generate or capture images / frames, metrics, information, and other sensor data for use by perception system 40.

[0052] Using a deep neural network (DNN) or other suitable infrastructure, the perception system 40 can be implemented such that for a given image derived from sensor data, the perception system 40 can output object information for each object 42 detected in the image. For example, the object information can be included in a perception table similar to the table shown below.

[0053] Object ID Category Bounding Box Probability Score 1 Car <![CDATA[x 1,1 ,y 1,1 ,x 2,1 ,y 2,1 > 0.846 2 Ship <![CDATA[x 1,2 ,y 1,2 ,x 2,2 ,y 2,2 > 0.815 3 Car <![CDATA[x 1,3 ,y 1,3 ,x 2,3 ,y 2,3 > 0.975

[0054] The perception system 40 can provide a representation of the object 42 within a bounding box, where the position coordinates of the object indicate the position of the object relative to the vehicle. The perception system 40 can also classify the object 42 (e.g., provide an indication of the type or category of the object, whether the object is a vehicle, traffic sign, landmark, pedestrian, etc.). The perception system 40 can then assign a probability to the detected object 42. This probability can be used to indicate the confidence level of the perception system 40 (e.g., the DNN) in detecting the object. For example, if a car is somewhat blurry in the image but is still captured by the neural network, it may have a low probability. Conversely, if the car is clear in the image, it may have a high probability. For each object detected in the frame, corresponding object information can be generated by the perception system 40, and the object information can include an object ID (e.g., a number assigned to the object in the frame), an object category (e.g., car, truck, boat, etc.), bounding box data, and a probability score. Thus, the perception system 40 can output multiple detection results as shown in the table above, where for each identified object, the table includes an ID, a category, bounding box coordinates, and a probability score.

[0055] The perception table and / or other information generated by the perception system 40 based on the scene data can be compiled in the described manner and / or according to other processes to facilitate the generation of object determination of one or more objects identified in the corresponding scene. The perception system 40 can provide the object determination to the trajectory planner system 46 for facilitating the operation of the vehicle. For example, the trajectory planner system 46 can include guidance techniques or other systems, whereby the vehicle can be controlled to be identified in the object determination, which can include autonomously or semi-autonomously controlling the vehicle to take actions. One aspect of the present disclosure relates to the verification system 48 that can operate with the trajectory planner system 46 to facilitate verifying the accuracy of the object determination performed by the perception system 40. The verification system 48 can be configured to provide the verification to the trajectory planner system 46 for the purpose of indicating whether the object determination provided by the perception system 40 has been verified or not. Consistent with the object determination, the verification can be provided to the trajectory planner system 46 so that the trajectory planner system 46 can compare the verification with the object determination to determine whether the object determination has been verified or not, and based on this, facilitate the relevant control of the vehicle. For other systems on and / or outside the vehicle, similar processes can be implemented later, that is, the object determination generated by the perception system 40 and thus the verification generated by the verification system 48 can be similarly provided to those systems.

[0056] The verification system 48 can include: a semantic extraction module 54, configured to generate semantic information for the object 42; a semantic generation module 56, configured to generate a plurality of semantic captions based on the semantic information; a consistency evaluation module 58, configured to generate a consistency score for the semantic captions; and a diagnostic module 60, configured to verify the accuracy of the object determination based on the consistency score. The verification system 48 can optionally include a cropping module 64, which is configured to crop individual objects from the images from the scene data to associate with the relevant object information generated using the perception system 40. The modules 54, 56, 58, 60, 64, 66 can be: one or more integrated and / or independent software and / or hardware constructs, which can be operable according to the corresponding one or more processors executing a plurality of associated non-transitory instructions stored on one or more relevant computer-readable storage media. For non-limiting purposes, the modules 54, 56, 58, 60, 64, 66 are shown separately from each other to functionally highlight different aspects of the present disclosure associated with the process of verifying the accuracy of the object determination for the object 42 detected by the perception system 40.

[0057] The semantic extraction module 54 may be configured to generate semantic information based on data included within an image frame, the image frame being included as part of the scene data and processed by the perception system 40 for object determination. The semantic extraction module 54 may be based on a variety of computer vision and object detection constructs, such as but not limited to UperNet, DETR (Detection Transformer), scene graph, etc. UperNet may include a network architecture designed for semantic segmentation of tasks in computer vision, where semantic segmentation may involve labeling each pixel in an image with a corresponding class label, allowing for a detailed understanding of the scene, which may include a semantic segmentation model architecture that unifies the partitioning and prediction within the network to provide a semantic segmentation benchmark. DETR may include a specific object detection model that formulates object detection as a set prediction problem using a Transformer architecture, which may be based on using a Transformer encoder-decoder architecture for object detection. A scene graph may include a representation of a scene in computer vision that captures the relationships between objects, such as to provide a structured representation of the objects present in an image and their interactions or spatial relationships, representing the visual scene by modeling the objects and their relationships, using nodes in the scene graph to represent objects, edges to represent the relationships between objects, and otherwise enabling complex image understanding and multimodal tasks.

[0058] The semantic generation module 56 may be configured to generate semantic captions to explicitly include text language describing the scene associated with the object. The semantic generation module 56 may be configured to generate object captions as part of the semantic captions, optionally, where the object captions are based on the object identifiers selected for the object by the perception system 40, e.g., based on the object IDs included in the associated perception table. The semantic generation module 56 may be configured to determine multiple component identifiers of the object from a set of component identifiers listed in the mapping module of the object identifier. For example, the semantic generation module 56 may be configured to generate templates, such as but not limited to using ConceptNet to generate templates. ConceptNet may relate to a knowledge graph that connects words and phrases through common sense relationships according to large-scale multilingual resources representing general knowledge about the world. The information in ConceptNet may be manually curated and collected from various sources, including books, websites, and other texts. ConceptNet may be able to organize knowledge into a network of nodes (concepts or terms) connected by edges (relationships), optionally, where each edge represents a relationship between two concepts. ConceptNet can be used in natural language processing and artificial intelligence applications to provide a broader understanding of language and context. It helps the system infer the meaning and relationships between words beyond what is explicitly stated. A template may refer to a predefined structure or pattern that can be filled with specific content to generate text, which may allow the template to serve as a framework for constructing sentences or larger text units. In other embodiments, the semantic generation module 56 may be implemented based on a large language model or an artificial intelligence (AI) system. In other embodiments, the semantic generation module 56 may be implemented as a combination of the above methods.

[0059] The consistency assessment module 58 may include a language-image model operable to generate a consistency score. The consistency assessment module 58 may be configured to verify the accuracy of object determination based on a relative comparison of the consistency scores. The consistency assessment module 58 may receive the output and other information generated by the cropping module and / or the semantic generation module to generate a consistency score. The consistency assessment module may be based on, for example, CLIP (Contrastive Language-Image Pretraining) and / or ViT (Vision Transformer). In this way, the consistency assessment module may operate as an artificial intelligence model designed to understand and generate both text and visual information, optionally where the goal is to bridge the gap between natural language understanding and computer vision, such that applications can understand and generate content across both modalities. The diagnostic module 60 may be configured to determine whether an object is verified or unverified based on the consistency score. The mitigation module 66 may be configured to provide the corresponding verification to the trajectory planner system 46, that is, to notify the trajectory planner system 46 whether the accompanying object determination is verified or unverified. Optionally, for example, to improve the processing requirements for the trajectory planner system 46, the verification notification may be limited to determinations that are identified as unverified. The trajectory planner system 46 may accordingly be configured to accept the object determination made by the perception system 40 in the absence of a verification notification indicating that the object determination is unverified.

[0060] Figure 3 FIG. 70 is a flow chart of a method for verifying the accuracy of object determination using the perception system 40 according to a non-limiting aspect of the present disclosure. Block 72 relates to a semantic extraction process, whereby the semantic extraction module 54 may generate semantic information for an object based on data included within an image frame processed by the perception system 40 for a corresponding object determination. Block 74 relates to a semantic description process, whereby the semantic generation module 56 may generate a plurality of semantic captions based on the semantic information. Block 76 relates to an evaluation process, whereby the consistency assessment module 58 may generate a consistency score for the semantic captions. Block 78 relates to a diagnostic process, whereby the diagnostic module 60 may verify the accuracy of the object determination based on the consistency score generated as part of the evaluation process. Block 80 relates to a mitigation process, whereby the mitigation module 66 may generate instructions for guiding the trajectory planner system 46 for verified and / or unverified object determinations. The present disclosure contemplates that the described processes include a variety of options for generating semantic captions and verifying the accuracy of the accompanying object determination based on the consistency scores obtained therefor. Accordingly, the description provided herein is not necessarily intended to limit the scope of the contemplated present disclosure, however, corresponding examples are provided to demonstrate the beneficial ability of the dictation system to assist in limiting the likelihood that the trajectory planner system 46 depends on inaccurate object determinations in an undesirable manner.

[0061] For example, semantic captions generated as part of the semantic description process may include text language describing the scene associated with object 42, optionally including at least one semantic caption for each object type of the multiple object types associated with the object. The semantic captions may also include the relative size and location of the object type with which they are associated. The semantic captions may include object captions based on the object identifier selected for object 42 by the perception system 40. The object identifier may be used by the semantic generation module 56 to determine multiple component identifiers for object 42 from a set of component identifiers listed in the mapping module of the object identifier. Additionally or alternatively, the semantic description process may include generating multiple component captions as part of the semantic captions, optionally where each component caption identifies a different one of the component identifiers or a combination of more than one component identifier, and the component caption may also include the relative size and location of the component associated with its component identifier. The semantic description process may include generating multiple detailed captions as part of the semantic captions, optionally where each detailed caption identifies detailed information about object 42, including at least one associated characteristic of text, material, and color. The semantic description process may include multiple common sense captions as part of the semantic captions, optionally where each common sense caption identifies common sense information about object 42, such as by including at least one associated characteristic for use, scene, and relationship with neighboring objects. The semantic description process may include generating multiple combined captions as part of the semantic captions, where each combined caption combines at least one of the multiple component captions, detailed captions, and common sense captions. The semantic description process may include generating multiple component captions as part of the semantic captions, optionally where each component caption identifies one of the multiple possible permutations of the component identifier, and for each possible permutation, includes one of the component captions. The semantic description process may include generating each component caption to identify one of the multiple possible permutations of the component identifier, and for each possible permutation, includes one of the component captions, optionally where each permutation includes the relative size and location of the component associated with its component identifier. As an example, if the object is a car detected using the perception system 40, the semantic generation module 56 may generate the following captions: a picture of a car; a picture of a car on the road; a picture of a car with wheels; a picture of a car with a headlight on the right side; a picture of a car with a windshield made of glass; a picture of a car with a selected length and / or a combination of these descriptive captions.

[0062] For example, a diagnostic process for determining whether an object is verified or unverified based on a consistency score generated for semantic captions may include various verification methods. One verification method may include: verifying the accuracy of the object determination when the consistency score of the object caption is less than each consistency score of each component caption. Another verification method may include: verifying the accuracy of the object determination when the consistency score of the object caption is less than the consistency score of the component caption. Another verification method may include: verifying the accuracy of the object determination when the consistency score of the object caption is less than each consistency score of each component caption. When the consistency score of the object caption is less than at least one consistency score of the component caption. Another verification method may include: verifying the accuracy of the object determination when the consistency score of the object caption is less than at least one consistency score of the component caption. Another verification method may include: verifying the accuracy of the object determination when the consistency score of the object caption is less than at least one consistency score of the detailed caption. Another verification method may include: verifying the accuracy of the object determination when the consistency score of the object caption is less than at least one consistency score of the common sense caption. Another verification method may include: verifying the accuracy of the object determination when the consistency score of the object caption is less than at least one consistency score of the combined caption.

[0063] The following exemplary table may represent the semantic captions generated by the semantic generation module 56 and the associated consistency score analysis generated by the consistency assessment module 58 and the diagnostic module 60. The template rows may correspond to the semantic captions, and the analysis rows may correspond to the consistency score analysis for determining the accuracy of the corresponding object determination.

[0064]

[0065]

[0066]

[0067] Although various embodiments have been described, the description is intended to be exemplary rather than restrictive, and many more embodiments and implementations within the scope of the embodiments will be apparent to those of ordinary skill in the art. Any feature of any embodiment may be used in combination with or substituted for any other feature or element in any other embodiment, unless specifically constrained. Thus, the embodiments should not be restricted except in view of the appended claims and their equivalents. In addition, various modifications and changes may be made within the scope of the appended claims. Although several modes for carrying out many aspects of the present teachings have been described in detail, those skilled in the art familiar with the field to which these teachings pertain will recognize various alternative aspects of the present teachings for implementation within the scope of the appended claims. It is intended that all content included in the above description or shown in the accompanying drawings should be construed as illustrative and exemplifying the entire scope of alternative embodiments, and those of ordinary skill in the art will recognize that the alternative embodiments are implied by the content included, are structurally and / or functionally equivalent to the content included or are otherwise made obvious based on the content included, rather than being limited only to those embodiments explicitly depicted and / or described.

Claims

1. A system for verifying the accuracy of object determination made by a perception system based on detecting an object in a surrounding area of ​​a vehicle, the system comprising: a semantic extraction module configured to generate semantic information for the object; A semantic generation module, configured to generate a plurality of semantic subtitles based on the semantic information; a consistency evaluation module configured to generate a consistency score for the semantic subtitles; and A diagnostic module is configured to verify the accuracy of the object determination based on the consistency score.

2. The system of claim 1, wherein: The semantic extraction module is configured to generate the semantic information based on data included in the image frame processed by the perception system to perform the object determination.

3. The system of claim 2, wherein: The semantic generation module is configured to generate the semantic subtitles to explicitly include textual language describing a scene associated with the object.

4. The system of claim 3, wherein: The consistency assessment module includes a language-image model operable to generate the consistency score.

5. The system of claim 4, wherein: The consistency assessment module is configured to verify the accuracy of the object determination based on a relative comparison of the consistency scores.

6. The system of claim 5, wherein: The perception system is operable to detect objects across a plurality of object types; The semantic generation module is configured to generate at least one semantic subtitle for each object type; and The diagnostic module is configured to verify the accuracy of the object determination when the object determination matches a semantic caption associated with a consistency score having a maximum relative rating within the relative comparison.

7. The system of claim 5, wherein: The perception system is operable to detect objects across a plurality of object types; The semantic generation module is configured to generate at least one semantic subtitle for each object type, wherein each semantic subtitle includes a relative size and location of the object type associated therewith; and The diagnostic module is configured to verify the accuracy of the object determination when the object determination matches a semantic caption associated with a consistency score having a maximum relative rating within the relative comparison.

8. The system of claim 5, wherein: The semantic generation module is configured to generate an object caption as one of the semantic captions, wherein the object caption is based on an object identifier selected by the perception system for the object.

9. The system of claim 8, wherein: The semantic generation module is configured to determine a plurality of part identifiers of the object from a set of part identifiers listed in the object identifier mapping module.

10. The system of claim 9, wherein: The semantic generation module is configured to generate a plurality of component subtitles as part of the semantic subtitles, wherein each component subtitle identifies a different one component identifier or a combination of more than one component identifiers; and The diagnostic module is configured to verify the accuracy of the object determination when the consistency score of the object subtitle is less than at least one consistency score of the component subtitle.