Systems and methods for training and utilizing object-centric occupancy estimation models

By receiving two-dimensional images and comparing the loss of the occupancy estimation model with the ground real-life comparison, the parameters are updated to improve the model accuracy, and the problem of insufficient identification of the occupancy estimation model in the vehicle autonomous system is solved, and a higher accuracy of three-dimensional object detection is achieved.

CN120388338APending Publication Date: 2025-07-29GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410580916.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2024-05-11
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the existing vehicle autonomous or semi-autonomous systems, the occupation estimation model lacks accuracy when identifying objects in the surrounding environment of the vehicle, making it difficult to effectively combine multiple sensor data for high-precision three-dimensional object detection.

Method used

By receiving two-dimensional images, the occupancy estimation model is used to generate voxels, and the loss value is determined by comparing with the occupied ground reality and the object ground reality, and the model parameters are updated in combination with the weights to improve the accuracy of the model.

Benefits of technology

The accuracy of the occupancy estimation model when identifying objects in the environment around the vehicle is improved, and the objects in the three-dimensional space around the vehicle can be more accurately identified, enhancing the safety and reliability of autonomous vehicle operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388338A_ABST
    Figure CN120388338A_ABST
Patent Text Reader

Abstract

A method of updating parameters of an occupancy estimation model. The method includes receiving a two-dimensional image. The occupancy estimation model is used to generate voxels based on the image. Occupancy loss is determined by comparing voxels to floor occupancy truth corresponding to the voxels. Object loss is determined by comparing voxels to object ground live. The object loss is combined with the occupancy loss to determine a total loss of voxels. Parameters of the occupancy estimation model are updated to reduce the determined total loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a system and method for training and utilizing an estimation model, and particularly to a system and method for training and utilizing an object - centered occupancy estimation model. Background Art

[0002] Vehicles are a major part of daily life. Specialized cameras, microcontrollers, laser technology, and sensors can be used in many different applications in vehicles. These devices can be used to provide an improved view of the area around the vehicle to the vehicle's operator, or to provide fully or semi - autonomous functions for the vehicle. A variety of methods have been developed to incorporate autonomous or semi - autonomous functions into vehicles, such as by using artificial intelligence (AI). Summary of the Invention

[0003] Disclosed herein is a method for updating parameters of an occupancy estimation model. The method includes receiving a two - dimensional image. The occupancy estimation model is used to generate voxels based on the image. An occupancy loss is determined by comparing the voxels with the occupancy ground truth corresponding to the voxels. An object loss is determined by comparing the voxels with the object ground truth. The object loss and the occupancy loss are combined to determine the total loss of the voxels. The parameters of the occupancy estimation model are updated to reduce the determined total loss.

[0004] Another aspect of the present disclosure may include applying a weight to the object loss when combined with the occupancy loss to determine the total loss.

[0005] Another aspect of the present disclosure may be to perform the comparison of the voxels with the occupancy ground truth on the basis of voxel - wise cross - entropy.

[0006] Another aspect of the present disclosure may be to multiply the object loss by a weight to determine the total loss before combining it with the occupancy loss.

[0007] Another aspect of the present invention may be that the object loss includes an object detection loss and an object detection model is used to identify the detected objects in the voxels.

[0008] Another aspect of the present invention may be that the object detection loss includes comparing the object ground truth of the voxels with the detected objects.

[0009] Another aspect of the present disclosure may be that comparing the object ground truth of the voxels with the detected objects includes at least one of cross - entropy or shape regression.

[0010] Another aspect of the present disclosure may be that the object loss includes an object fullness loss, and determining the object fullness loss includes comparing a set of occupancy voxels defined within a bounding box from the object ground truth with the occupancy probability estimated by the occupancy estimation model.

[0011] Disclosed herein is a non - transitory computer - readable storage medium comprising programming instructions that are operable, when executed by a processor, to perform a method. The method includes receiving a two - dimensional image. An occupancy estimation model is used to generate voxels based on the image. An occupancy loss is determined by comparing the voxels with occupancy ground truth corresponding to the voxels. An object loss is determined by comparing the voxels with object ground truth. The object loss and the occupancy loss are combined to determine a total loss for the voxels. The parameters of the occupancy estimation model are updated to reduce the determined total loss.

[0012] Disclosed herein is a vehicle system. The vehicle system includes at least one optical sensor and a controller in data communication with the at least one sensor, wherein the controller is configured to receive a two - dimensional optical image and perform occupancy estimation on the image using an occupancy estimation model to generate voxels. The controller is further configured to perform three - dimensional object detection on the image to identify at least one region associated with an object of interest. The controller is further configured to change a probability threshold for identifying an object in the voxels based on the three - dimensional object detection. The probability threshold for identifying an object is reduced for a portion of the voxels corresponding to at least one region associated with the object of interest when compared to the probability threshold for identifying an object outside of at least one region associated with the object of interest.

[0013] Another aspect of the present disclosure may be that the occupancy estimation model is trained by receiving a plurality of training images, wherein the plurality of training images are two - dimensional and the occupancy estimation model is used to generate voxels based on the images. An occupancy loss is determined by comparing the voxels with occupancy ground truth corresponding to the voxels, and an object loss is determined by comparing the voxels with object ground truth to further train the occupancy estimation model. The object loss and the occupancy loss are combined to determine a total loss for the voxels, which is used to update the parameters of the occupancy estimation model to reduce the determined total loss.

[0014] Another aspect of the present disclosure may be that the object loss includes an object detection loss, and an object detection model is used to identify detected objects in the voxels, wherein the object detection loss includes comparing the object ground truth of the voxels with the detected objects. Additionally, comparing the object ground truth of the voxels with the detected objects includes at least one of cross - entropy or shape regression.

[0015] Another aspect of the present disclosure may be that the object loss includes an object fullness loss, and determining the object fullness loss includes comparing a set of occupied voxels defined within a bounding box from the object ground truth with an occupancy probability estimated by the occupancy estimation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1An exemplary vehicle system is shown that includes a drive system having a plurality of sensors communicating with a controller.

[0017] Figure 2 A flowchart of a method for training an occupancy estimation model is shown.

[0018] Figure 3 A flowchart of a method for determining occupancy using an adaptive probability threshold is shown.

[0019] Some embodiments of the present disclosure are now described by way of example only and with reference to the drawings. In all the drawings, the same reference numerals denote the same elements or elements of the same type. Detailed Description

[0020] The present disclosure is susceptible to many different forms of embodiments. Representative examples of the present disclosure are shown in the drawings and are described herein in detail as non-limiting examples of the disclosed principles. For this reason, elements and limitations described in the abstract, background, summary, and detailed description sections but not explicitly recited in the claims should not be incorporated into the claims by implication, inference, or otherwise, singly or collectively.

[0021] For the purposes of this description, unless specifically disclaimed, the use of the singular includes the plural and vice versa, and the terms "and" and "or" shall be conjunctive and disjunctive. The words "comprising," "containing," "including," "having," etc. shall mean "including but not limited to." Additionally, approximating words such as "about," "almost," "substantially," "generally," "approximately," etc. may be used herein in the sense of "at, near, or approaching," or "within 0% to 5% of," or "within acceptable manufacturing tolerances," or a logical combination thereof. As used herein, a component "configured to" perform a specified function can perform the specified function without change, rather than only having the potential to perform the specified function after further modification. In other words, when the described hardware is explicitly configured to perform the specified function, the described hardware is specifically selected, created, implemented, utilized, programmed, and / or designed to perform the specified function.

[0022] According to an exemplary embodiment, Figure 1Vehicle 20 is shown that can operate in an autonomous mode or an automated mode. Vehicle 20 can be a fully autonomous vehicle or a semi-autonomous vehicle. Vehicle 20 includes a driving system 22 that controls the autonomous operation of the vehicle. Driving system 22 includes a sensor system 24 for obtaining information about the surroundings or environment of vehicle 20 and a controller 26 for calculating possible actions of the autonomous vehicle based on the obtained information and for implementing one or more of the possible actions, and a human-machine interface 28 for communicating with the occupants of the vehicle, such as a driver or a passenger. Sensor system 24 can include at least one optical sensor 30 (such as at least one camera), at least one distance sensor 32 (such as a depth camera (RGB-D) or lidar).

[0023] Controller 26 can include a processing circuit, which can include an application-specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or grouped) that executes one or more software or firmware programs, and a memory, combinational logic circuits, and / or other suitable components that provide the described functionality. Controller 26 can include a non-transitory computer-readable medium storing instructions that, when processed by one or more processors of controller 26, implement the methods disclosed below.

[0024] Alternatively, the methods disclosed below can be executed on a remote computer system 50 for training at least one of a feature extraction neural network or an object detection neural network, as will be described in more detail below. Although for simplicity of illustration, Figure 1 computer system 50 is depicted as a single computer module, computer system 50 can be physically embodied as one or more processing nodes having a non-transitory computer-readable storage medium 54 (i.e., sufficient memory) and associated hardware and software (such as, but not limited to, a high-speed clock, a timer, input / output circuitry, buffer circuitry, etc.). Computer-readable storage medium 54 can include sufficient read-only memory, such as magnetic or optical memory. The computer-readable code or instructions embodying the methods described below can be executed during the operation of computer system 50. To this end, computer system 50 can include one or more processors 52, such as logic circuits, application-specific integrated circuits (ASICs), central processing units, microprocessors, and / or other necessary hardware required to provide the programming functionality described herein.

[0025] Figure 2FIG. 0 shows a flowchart of a method 100 for generating an updated occupancy estimation model for improving the accuracy of the model to identify objects in the area around vehicle 20, such as cars, people, bicycles, etc. Method 100 finds update parameters for the occupancy estimation model by determining different types of losses, such as occupancy loss, object detection loss, or object fullness loss, as will be discussed in more detail below. Method 100 can be executed iteratively until the accuracy of the model or the loss generated by the model is below a predetermined threshold.

[0026] The occupancy estimation model is used to identify objects in the three-dimensional (3D) space around vehicle 20. Method 100 begins by collecting two-dimensional (2D) images at block 102. In the example shown, 2D images are collected from the area around vehicle 20. In one example, the 2D images are collected using one or more optical sensors 30 on vehicle 20. The images from one or more optical sensors 30 can include fields of view that at least partially overlap with another image. Using the 2D images, method 100 proceeds to block 104 to perform an initial occupancy estimation using the occupancy estimation model based on the images from block 102.

[0027] In the example shown, the occupancy estimation model provides the probability that a point in the 3D space is occupied by an object. In one example, a point in the 3D space is represented by a voxel. The voxel defines the probability that an object occupies each point in the 3D space around vehicle 20 in a 3D voxel grid based on the images from block 102. The occupancy estimation model uses a neural network to extract features from the 2D images of the area around vehicle 20. The occupancy estimation model then uses the extracted features to generate a 3D voxel grid that describes the scene in the image and can be used to identify the type of detected object. A point cloud can be extracted from the voxels by applying a surface processing algorithm to the voxels to create a visual 3D representation of the scene. Using the voxels from block 104, method 100 advances to block 106 to determine the occupancy loss of the voxels.

[0028] In the example shown, the occupancy loss is determined by comparing the voxels generated at block 104 with a data set, such as occupancy ground truth from block 108. The occupancy ground truth represents the actual scene from the 2D images to determine how closely the occupancy estimation model captures the actual scene represented in the ground truth. In one example, the occupancy ground truth can be generated by using at least one of the distance sensors 32 of the sensor system 24 to identify objects in the area around vehicle 20 to determine which objects the occupancy estimation model should find. For comparison purposes, the occupancy ground truth can assign ground truth values to each space in the 3D voxel grid to indicate whether the space is occupied or empty.

[0029] Voxels from frame 104 can be passed through an occupancy network to determine whether each space in the 3D grid of voxels is occupied or empty before comparing with occupancy ground truth. In one example, an occupancy loss is determined by cross-entropy, such as per-voxel cross-entropy calculation. Then the loss determined at block 106 is used to determine the total object loss at block 114 or the total object occupancy loss at block 140, as will be further discussed below.

[0030] To determine the total object loss, method 100 takes the voxels from frame 104 and proceeds to block 116 to perform 3D object detection on the voxels. The object detection performed at block 116 can add bounding boxes or another type of object identification to the voxels or the 3D point cloud representation of the voxels to identify regions with detected objects. In one example, the object detection performed at block 116 can be done by an object detection model utilizing a convolutional neural network. Method 100 then proceeds to block 118 to perform an object detection loss.

[0031] At block 118, method 100 compares a dataset (e.g., object ground truth from block 118) with the results of the 3D object detection performed at block 114. The comparison can be performed by anchor-wise cross-entropy with or without shape regression. In the example shown, the object ground truth can be generated by manually labeling the objects in the dataset corresponding to the scene represented by the image from frame 102 with bounding boxes to identify the objects actually in the scene. The object detection loss determines how closely the objects detected from block 116 match the ground truth or the actual objects in the scene.

[0032] The object detection loss performed at block 118 outputs a loss that is multiplied at block 112 by a weight (w) from block 122. The weight (w) can address the imbalance between positive and negative identifications between the object ground truth from block 120 and the results of the 3D object detection from block 116. Then the weighted loss is added to the occupancy loss from block 106 at block 110 to determine the total object loss at block 114.

[0033] The total object loss from box 114 is backpropagated to box 104 to update the parameters of the occupancy estimation model to improve the occupancy estimation performed at box 104 and to reduce the total loss provided by the occupancy estimation model when compared to the ground truth. Then, method 100 can be repeated an additional number of times on the images from box 102 to extract additional training values from the images and their associated occupancy ground truth and object ground truth. Then, method 100 can be repeated on an additional set of images received from box 102 with the corresponding ground truth to further refine the occupancy estimation model in box 104 using the total object loss from box 114. Method 100 can be repeated until the total object loss from box 114 is within a predetermined loss range or below a threshold loss value.

[0034] The flowchart of method 100 can also be used to determine the total object occupancy loss at box 140 in addition to the total object loss from box 114 to update the parameters of the occupancy estimation model in box 104 to improve its accuracy.

[0035] To determine the object occupancy loss at box 130, method 100 compares the object ground truth from box 132 with the voxels from box 104 using Equation 1 below. The object ground truth represents the actual objects in the scene.

[0036]

[0037] In Equation 1 above, GT(B) includes a set of occupancy voxels inside box B for the occupancy ground truth, and P(x) is the occupancy probability of voxel x estimated by the occupancy estimation model. The lower the value output by Equation 1, the closer the voxels from box 104 are to the object ground truth from box 132, and thus, the closer the voxels are to representing the actual scene. Each box B corresponds to a 3D bounding box of an object of interest in the previously detected or manually annotated scene.

[0038] The object occupancy loss determined at box 130 is output and multiplied by a weight (w) from box 134 at box 136. The weight (w) can address the imbalance from positive and negative identifications. Then, the weighted loss is added to the occupancy loss from box 106 at box 138 to determine the total object occupancy loss at box 140.

[0039] The total object occupancy loss from block 140 is backpropagated to block 104 to update the parameters of the occupancy estimation model, thereby improving the accuracy of the occupancy estimation model at block 104 and reducing the total loss it generates. Then, method 100 can be repeated an additional number of times for the images from block 102 to extract the maximum training values from the images and their associated occupancy and object ground truth. Then, method 100 can be repeated for an additional set of images with corresponding ground truth in block 104 to further refine the occupancy estimation model. Method 100 can be repeated until the total loss determined by method 100 is within a predetermined range or loss or below a threshold loss value.

[0040] Figure 3 Method 200 for applying an adaptive threshold to 3D occupancy estimation is shown. Method 200 begins at block 202 by receiving a plurality of images, such as 2D images captured by an optical sensor 30 on a vehicle 20. Method 200 then proceeds to perform 3D object detection on the plurality of images from block 202 at block 206. In one example, the object detection performed at block 206 locates potential objects in bounding boxes in the images from block 202.

[0041] Method 200 also performs occupancy estimation at block 204 using the images from block 202. In the example shown, the occupancy estimation can be performed by the occupancy estimation model trained by method 100 described above or by another occupancy estimation model. As described above, the occupancy estimation model outputs an occupancy probability P(x). Then, method 200 proceeds to block 208 to utilize an adaptive threshold.

[0042] In block 210, method 200 uses the adaptive threshold from block 208 to determine 3D occupancy. In particular, the adaptive threshold uses a higher threshold to determine the occupancy of regions outside the bounding boxes identified by block 206 and a lower threshold to determine the occupancy of regions within the bounding boxes from block 206. This adaptive method for determining 3D occupancy can be applied in real time to vehicle 20 to identify objects in the area around vehicle 20.

[0043] The terms "a" and "an" do not denote a limitation of quantity, but rather denote the presence of at least one of the referenced item. Unless the context clearly dictates otherwise, the term "or" means "and / or". References to "aspect" throughout the specification mean that the particular elements (e.g., features, structures, steps, or characteristics) described in connection with that aspect are included in at least one aspect described herein and may or may not be present in other aspects. Additionally, it should be understood that the described elements can be combined in a suitable manner in the various aspects.

[0044] When an element such as a layer, film, region, or substrate is referred to as being “on” another element, it can be directly on the other element or intervening elements may also be present. In contrast, when an element is referred to as “directly on” another element, intervening elements are absent.

[0045] Unless otherwise stated herein, the test standards are the latest standards in effect as of the filing date of this application, or, if priority is claimed, the filing date of the earliest priority application in which the test standards appear.

[0046] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0047] Although the foregoing disclosure has been described with reference to exemplary embodiments, those of ordinary skill in the art will understand that various changes may be made and equivalents may be substituted for its elements without departing from its scope. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the disclosure without departing from its scope. Accordingly, it is intended that the disclosure not be limited to the particular embodiments disclosed, but that it will include embodiments falling within its scope.

Claims

1. A method for updating parameters of an occupancy estimation model, the method comprising: Receiving a plurality of images, wherein the plurality of images are two-dimensional; Generating voxels based on the plurality of images using the occupancy estimation model; Determining an occupancy loss by comparing the voxels with occupancy ground truth corresponding to the voxels; Determining an object loss by comparing the voxels with object ground truth; Combining the object loss and the occupancy loss to determine a total loss of the voxels; And Updating the parameters of the occupancy estimation model to reduce the determined total loss.

2. The method according to claim 1, comprising applying a weight to the object loss when combined with the occupancy loss to determine the total loss.

3. The method according to claim 1, wherein comparing the voxels with the occupancy ground truth is performed on the basis of per-voxel cross-entropy.

4. The method according to claim 1, wherein the object loss is multiplied by a weight before being combined with the occupancy loss to determine the total loss.

5. The method according to claim 1, wherein the object loss includes an object detection loss and using an object detection model to identify detected objects in the voxels.

6. The method according to claim 5, wherein the object detection loss includes comparing the object ground truth of the voxels with the detected objects.

7. The method according to claim 6, wherein comparing the object ground truth of the voxels with the detected objects includes at least one of cross-entropy or shape regression.

8. The method according to claim 1, wherein the object loss includes an object fullness loss, and determining the object fullness loss includes comparing a set of occupied voxels defined within a bounding box from the object ground truth with an occupancy probability estimated by the occupancy estimation model.

9. A non-transitory computer-readable storage medium comprising programming instructions that are operable, when executed by a processor, to perform a method, the method comprising: Receiving a plurality of images, wherein the plurality of images are two-dimensional; Generating voxels based on the plurality of images using an occupancy estimation model; Determining an occupancy loss by comparing the voxels with occupancy ground truth corresponding to the voxels; Determining an object loss by comparing the voxels with object ground truth; Combining the object loss and the occupancy loss to determine a total loss of the voxels; And Updating the parameters of the occupancy estimation model to reduce the determined total loss.

10. The computer-readable storage medium according to claim 9, wherein the method includes applying a weight to the object loss when combined with the occupancy loss to determine the total loss.

Citation Information

Cited By

  • Occupancy prediction model training method, related method, device, equipment and medium

    CN122049887A

  • Method for training occupancy prediction model and related method, device, equipment and medium

    CN122049887B