System and method for occlusion reconstruction in surrounding views using temporal information
Patent Information
- Application Number
- CN202211346450.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-08
- Filing Date
- 2022-10-31
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-10-31
Smart Images

Figure CN117237520B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to a system and method for occlusion reconstruction in a surrounding view using temporal information. Background Technology
[0002] Autonomous navigation systems, semi-autonomous navigation systems, and driver-assistance / driver-warning systems interpret sensor data and determine whether the path is clear for navigation and whether there are any objects in the operating environment that may obstruct that clear path. Summary of the Invention
[0003] A system for occlusion reconstruction in an surrounding view using temporal information is provided. The system includes: an active camera device for generating image data describing a first view of an operating environment; and a computerized vision data controller. The computerized vision data controller includes programming for: analyzing the image data to generate a three-dimensional computerized representation of the operating environment; synthesizing a virtual camera view of the operating environment from a desired viewpoint using the image data and the three-dimensional computerized representation of the operating environment; and identifying occlusions in the virtual camera view. The computerized vision data controller further includes programming for: utilizing historical iterations of the image data and identifying fixed objects within the operating environment; and utilizing historical iterations of the image data to estimate occlusion filling information. The computerized vision data controller further includes programming for: reconstructing the occlusion in the three-dimensional computerized representation using pixel data from the fixed objects with filling information; and providing navigation guidance within the operating environment using the three-dimensional computerized representation.
[0004] In some embodiments, the programming for analyzing image data includes programming for performing deep interpretation and semantic segmentation on the image data.
[0005] In some embodiments, the programming for identifying occlusion includes programming for identifying potential moving objects in the operating environment and identifying occlusions in the virtual camera view caused by such moving objects.
[0006] In some embodiments, the system further includes a plurality of active camera devices that generate image data.
[0007] In some embodiments, the programming for synthesizing a virtual camera view using image data and a three-dimensional computerized representation includes programming for synthesizing a virtual camera view using epipolar reprojection.
[0008] In some embodiments, the system further includes a sensor device selected from a light detection and ranging (LIDAR) device or a radar device. The computerized vision data controller further includes programming for improving the three-dimensional computerized representation using data from the sensor device.
[0009] According to an alternative embodiment, a system is provided for occlusion reconstruction in an ambient view using temporal information. The system includes an active camera device that generates image data describing a first view of an operating environment, and a computerized vision data controller. The computerized vision data controller includes programming for: analyzing the image data to generate a three-dimensional computerized representation of the operating environment; synthesizing a virtual camera view of the operating environment from a desired viewpoint using the image data and the three-dimensional computerized representation of the operating environment; and identifying occlusions in the virtual camera view. The computerized vision data controller further includes programming for: utilizing historical iterations of the image data and identifying fixed objects within the operating environment; and utilizing historical iterations of the image data to estimate filling information for occlusion. The computerized vision data controller further includes programming for: reconstructing occlusions in the three-dimensional computerized representation using pixel data from the fixed objects with the filling information; and providing navigation guidance within the operating environment using the three-dimensional computerized representation.
[0010] In some embodiments, the apparatus includes a vehicle.
[0011] In some embodiments, the programming for analyzing image data includes programming for performing deep interpretation and semantic segmentation on the image data.
[0012] In some embodiments, the programming for identifying occlusion includes programming for identifying potential moving objects in the operating environment and identifying occlusion in the virtual camera view caused by such potential moving objects.
[0013] In some embodiments, the system further includes a plurality of active camera devices that generate image data.
[0014] In some embodiments, the programming for synthesizing a virtual camera view using image data and a three-dimensional computerized representation includes programming for synthesizing a virtual camera view using epipolar reprojection.
[0015] In some embodiments, the system further includes a sensor device selected from a light detection and ranging (LIDAR) device or a radar device. The computerized vision data controller further includes programming for improving the three-dimensional computerized representation using data from the sensor device.
[0016] According to an alternative embodiment, a method for occlusion reconstruction in an surrounding view using temporal information is provided. The method includes operating an active camera device to collect image data describing a first view of an operating environment, and within a computerized processor: receiving the image data from the active camera device; and analyzing the image data to generate a three-dimensional computerized representation of the operating environment. The method further includes, within the computerized processor: synthesizing a virtual camera view of the operating environment from a desired perspective using the image data and the three-dimensional computerized representation of the operating environment; identifying occlusions in the virtual camera view; and using historical iterations of the image data to identify fixed objects within the operating environment. The method further includes, within the computerized processor: estimating padding information for occlusion using historical iterations of the image data; reconstructing the occlusion in the three-dimensional computerized representation using the padding information with pixel data from the fixed objects; and providing navigation guidance within the operating environment using the three-dimensional computerized representation.
[0017] In some embodiments, analyzing image data includes performing deep interpretation and semantic segmentation on the image data.
[0018] In some embodiments, the computerized processor is located within the vehicle.
[0019] In some embodiments, identifying occlusion includes identifying potential moving objects in the operating environment and identifying occlusion in the virtual camera view caused by potential moving objects.
[0020] In some embodiments, the method further includes operating a plurality of active camera devices to collect image data.
[0021] In some embodiments, synthesizing a virtual camera view using historical iterations of image data and a three-dimensional computerized representation includes synthesizing a virtual camera view using epipolar reprojection.
[0022] In some embodiments, the method further includes: operating a sensor device selected from a light detection and ranging (LIDAR) device or a radar device; and utilizing data from the sensor device to improve the three-dimensional computerized representation.
[0023] In addition, the present invention also includes the following technical solutions.
[0024] Option 1. A system for occlusion reconstruction in a surrounding view using temporal information, the system comprising: An active camera device that generates image data describing a first view of the operating environment; and A computerized vision data controller, comprising programming for the following situations: Analyze the image data to generate a three-dimensional computerized representation of the operating environment; Using the image data and the three-dimensional computerized representation of the operating environment, a virtual camera view of the operating environment is synthesized from a desired perspective; Identify occlusions in the virtual camera view; The image data is iterated historically to identify fixed objects within the operating environment. The historical iterations of the image data are used to estimate the padding information for the occlusion; The occlusion is reconstructed in the 3D computerized representation using pixel data from the fixed object and the filling information; and The three-dimensional computerized representation is used to provide navigation guidance within the operating environment.
[0025] Option 2. The system according to Option 1, wherein the programming for analyzing the image data includes programming for performing depth interpretation and semantic segmentation on the image data.
[0026] Option 3. The system according to Option 1, wherein the programming for identifying the occlusion includes programming for the following situations: Identify potential moving objects in the operating environment; and Identify the occlusion in the virtual camera view caused by the potential moving object.
[0027] Option 4. The system according to Option 1 further includes a plurality of active camera devices for generating the image data.
[0028] Option 5. The system according to Option 1, wherein the programming for synthesizing the virtual camera view using the image data and the three-dimensional computerized representation includes programming for synthesizing the virtual camera view using epipolar reprojection.
[0029] Option 6. The system according to Option 1 further includes a sensor device selected from a light detection and ranging (LIDAR) device or a radar device; and The computerized vision data controller further includes programming for improving the three-dimensional computerized representation using data from the sensor device.
[0030] Option 7. A system for occlusion reconstruction in a surrounding view using temporal information, the system comprising: The device includes: An active camera device that generates image data describing a first view of the operating environment; and A computerized vision data controller, comprising programming for the following situations: Analyze the image data to generate a three-dimensional computerized representation of the operating environment; Using the image data and the three-dimensional computerized representation of the operating environment, a virtual camera view of the operating environment is synthesized from a desired perspective; and Identify occlusions in the virtual camera view; The image data is iterated historically to identify fixed objects within the operating environment. The historical iterations of the image data are used to estimate the padding information for the occlusion; The occlusion is reconstructed in the 3D computerized representation using pixel data from the fixed object and the filling information; and The three-dimensional computerized representation is used to provide navigation guidance within the operating environment.
[0031] Option 8. The system according to Option 7, wherein the device includes a vehicle.
[0032] Option 9. The system according to Option 7, wherein the programming for analyzing the image data includes programming for performing deep interpretation and semantic segmentation on the image data.
[0033] Option 10. The system according to Option 7, wherein the programming for identifying the occlusion includes programming for the following situations: Identify potential moving objects in the operating environment; and Identify the occlusion in the virtual camera view caused by the potential moving object.
[0034] Option 11. The system according to Option 7 further includes a plurality of active camera devices for generating the image data.
[0035] Option 12. The system according to Option 7, wherein the programming for synthesizing the virtual camera view using the image data and the three-dimensional computerized representation includes programming for synthesizing the virtual camera view using epipolar reprojection.
[0036] Option 13. The system according to Option 7 further includes a sensor device selected from a light detection and ranging (LIDAR) device or a radar device; and The computerized vision data controller further includes programming for improving the three-dimensional computerized representation using data from the sensor device.
[0037] Option 14. A method for occlusion reconstruction in a surrounding view using time information, the method comprising: Operate an active camera device to collect image data describing a first view of the operating environment; and Within a computerized processor: Receive the image data from the active camera device; Analyze the image data to generate a three-dimensional computerized representation of the operating environment; Using the image data and the three-dimensional computerized representation of the operating environment, a virtual camera view of the operating environment is synthesized from a desired perspective; Identify occlusions in the virtual camera view; The image data is iterated historically to identify fixed objects within the operating environment. The historical iterations of the image data are used to estimate the padding information for the occlusion; The occlusion is reconstructed in the 3D computerized representation using pixel data from the fixed object and the filling information; and The three-dimensional computerized representation is used to provide navigation guidance within the operating environment.
[0038] Option 15. The method according to Option 14, wherein analyzing the image data includes performing depth interpretation and semantic segmentation on the image data.
[0039] Option 16. The method according to Option 14, wherein the computerized processor is located within a vehicle.
[0040] Option 17. The method according to Option 14, wherein identifying the occlusion includes: Identify potential moving objects in the operating environment; and Identify the occlusion in the virtual camera view caused by the potential moving object.
[0041] Option 18. The method according to Option 14 further includes operating a plurality of active camera devices to collect the image data.
[0042] Option 19. The method according to Option 14, wherein synthesizing the virtual camera view using the historical iterations of the image data and the three-dimensional computerized representation comprises: synthesizing the virtual camera view using epipolar reprojection.
[0043] Option 20. The method according to Option 14 further includes: The operation is selected from sensor devices of optical detection and ranging (LIDAR) devices or radar devices; and The data from the sensor device is used to improve the three-dimensional computerized representation. Attached Figure Description
[0044] The above-described features and advantages, as well as other features and advantages of this disclosure, will become apparent from the following detailed description of the best mode for carrying out this disclosure, taken in conjunction with the accompanying drawings. Wherein: Figure 1 An exemplary first view of an operating environment generated by an active camera device according to the present disclosure and an exemplary synthetic virtual camera view of the operating environment are illustrated, wherein the synthetic virtual camera view can be used to reconstruct occlusion in sensor data. Figure 2 The illustration schematically shows sensor data generated by an active camera device as viewed from a synthetic virtual camera view according to the present disclosure, showing occlusion in the sensor data; Figure 3 The data stream according to this disclosure is schematically illustrated and can be used to perform depth interpretation and semantic segmentation of sensor data including an input image; Figure 4 It is a flowchart that schematically illustrates an exemplary method for occlusion reconstruction in a surrounding view based on usage time information of this disclosure; Figure 5 An exemplary vehicle is schematically shown that operates the disclosed system to interpret an operating environment according to the present disclosure, including a synthetic virtual camera view used to reconstruct occlusions in sensor data collected by the vehicle. Figure 6 An exemplary computerized vision data controller configured to operate the disclosed method is illustrated schematically. Figure 7 An exemplary device including a vehicle according to the present disclosure is schematically shown, the vehicle including... Figure 6 A computerized vision data controller and multiple sensors that collect data about the vehicle's operating environment; and Figure 8 It is a flowchart illustrating the flow of sensor data according to this disclosure and the corresponding operations and determinations that can be used to reconstruct occlusions identified in the sensor data. Detailed Implementation
[0045] Autonomous navigation systems, semi-autonomous navigation systems, and driver-assistance / driver-warning systems utilize sensor data to interpret the operating environment. Such an operating environment may include driving surfaces on the road and may include complex lane geometry, lane markings, unexpected features (such as building barriers and lane closures), and moving or potentially moving objects (such as other vehicles and pedestrians). Analysis of the sensor data enables computerized controllers to identify a defined path of travel within the operating environment. However, sensor data may be imperfect or incomplete; for example, objects in the operating environment may obscure details or create occlusions in data relevant to the operating environment.
[0046] Human thought develops in youth the concept of object permanence. Once this concept of object permanence develops, a person realizes that an object still exists simply because he or she can no longer see it. Similar strategies can be used for the computerized analysis of the operational environment. Through image analysis techniques, including semantic separation, pixels and groups of pixels in an image, or scenes captured in an image, can be identified as representing certain objects or surfaces in the environment. For example, a road surface has certain visual characteristics that distinguish it from a patch of grass or a telephone pole next to it. Objects and surfaces can be analyzed and categorized in the view into fixed features and potentially moving features. Identifying fixed features can help construct a three-dimensional computerized representation of the operational environment because it can be assumed that fixed features remain constant throughout the operational environment over a period of time. The three-dimensional computerized representation can be described as a static structure matrix. Other analyses can be performed to estimate the depth of objects and surfaces in the operational environment. For example, lane markings on a road can be physically equidistant from the road surface, and the distances between pixels representing lane markings can be used to estimate depths in the scene, such as the distance between a specific portion of a lane marking and the main vehicle in the analyzed scene. Once fixed features are identified and their depth in the scene is estimated, these fixed features can be assumed or fixed in a computerized model or 3D computerized representation of the operating environment. This computerized model or 3D computerized representation of the operating environment can be used to define and update the explicit path of the vehicle as it travels within the operating environment.
[0047] A system and method are provided for occlusion reconstruction in an ambient view using temporal information. This system and method can be used to correct occlusion caused by changes in view angle in natural ambient vision (NSV) images generated by epipolar reprojection. The system and method utilize temporal information from past frames, including pixel data from ambient cameras, inferred depth and segmentation data, and vehicle pose data (e.g., ranging, GPS, inertial navigation, etc.). The disclosed system and method can be used in vehicles. The disclosed system can be used by autonomous robots operating, for example, in manufacturing environments. The disclosed system can be used in a variety of applications where autonomous navigation is desired in the operating environment.
[0048] The disclosed systems and methods utilize temporal information for occlusion correction. The system may include a computerized process or engine that can be used to handle occlusion for either static or dynamic master vehicles. The disclosed method uses information from surrounding cameras to create virtual images (view angle changes). In one embodiment, the system selects a virtual camera position indicating the desired viewpoint and fills in an image captured from that virtual camera position pixel-by-pixel using historical camera and / or sensor data to estimate what each pixel would look like. To avoid ambient visual artifacts (e.g., object removal, distortion, ghosting, etc.), the disclosed method uses inferred depth to adapt to the three-dimensional structure of the scene. Virtual images can be synthesized from multiple alternative viewpoints; however, angle changes inevitably introduce areas not covered by surrounding cameras, which can be described as occluded areas or occlusions. Angle changes may include observer movement and / or movement of objects in the observer's operating environment.
[0049] Using depth and segmentation inference, combined with epipolar reprojection (ER) between the virtual view location and one or more physical cameras, relevant information about the virtual cameras (pixels + depth + segmented objects + mask) can be calculated at each time step. Historical or temporal information can be used to fill in occlusions caused by dynamic objects in the scene or by the movement of the master vehicle.
[0050] A single timestamp (TS) image has occlusion; however, in time... t = n Some occluded areas at time t = n- 1 , n-2 , n-3 The image data is occluded by surrounding cameras. For moving vehicles, using odometer data, the positions of previous surrounding cameras relative to the desired viewpoint can be estimated, and historical data can be incorporated into the ER. For static vehicles, static background pixels are occluded and exposed by moving objects in the scene. Analysis of the image data may include surrounding scene depth / segmentation inference. Segmentation can be used to locate global static objects (roads, buildings, trees). In one embodiment, occlusion can be reconstructed using pixels determined to be associated with static or fixed objects or features, thereby avoiding motion artifacts.
[0051] According to an alternative embodiment, a system for occlusion reconstruction in an ambient view using temporal information is provided. The system includes an active camera device that generates image data describing a first view angle of an operating environment and a computerized vision data controller. The computerized vision data controller includes: programming for: analyzing the image data to generate a three-dimensional computerized representation of objects in the operating environment; and image pixel classification / image segmentation to generate a virtual camera view angle from a desired viewpoint; and identifying one or more occlusions in the virtual camera view. The computerized vision data controller further includes programming for: analyzing the virtual camera view angle using the image data, the three-dimensional computerized representation of the operating environment, and a historical iteration of image pixel classification / image segmentation to estimate filling information for occlusion. The computerized vision data controller further includes programming for reconstructing the occlusion in the three-dimensional computerized representation using the filling information. The computerized vision data controller further includes programming for directing movement within the operating environment using the three-dimensional computerized representation.
[0052] Referring now to the accompanying drawings, the same reference numerals refer to the same features throughout the various views. Figure 1 An exemplary first view 10 of an operating environment 5 generated by an active camera device and an exemplary synthetic virtual camera view 20 of the operating environment 5 are schematically illustrated, wherein the synthetic virtual camera view 20 can be used to reconstruct occlusion 50 in sensor data. The first view 10 includes a first field of view 12. The synthetic virtual camera view 20 includes a second field of view 22. A first object 30 and a second object 40 are shown in the operating environment 5. The first object 30 can be defined as a moving or potentially moving object, and the second object 40 can be defined as a stationary object.
[0053] The synthetic virtual camera view 20 can be generated through various methods. In one embodiment, the synthetic virtual camera view 20 may include data collected sometime previously by the same sensor device that collects sensor data currently defined as the first view 10. Thus, for example, a forward-facing sensor device implemented as a camera device may be mounted in front of a host vehicle moving on the road. Current or recent images captured by the camera device can define or provide data about the first view 10. An image or a series of images captured by the same camera device two seconds ago may be used to generate the synthetic virtual camera view 20. In one instance, the first view 10 may be clear when there is an open road in the scene of the image and the entire road surface can be estimated. In such instances, the synthetic virtual camera view 20 may be irrelevant because there is no occlusion in the estimated three-dimensional computerized representation of the road surface or the host vehicle's operating environment. In another instance, a second vehicle may pass to the left of the host vehicle. The left side of the road, including lane markings on one side, may have been visible two seconds ago; however, since the second vehicle is in the lane to the left of the host vehicle, the left side of the road and its markings are no longer visible. By using the available details in the synthesized virtual camera view 20, it can be assumed that data about fixed features in the scene (including the geometry on the left side of the road and the markings on the left side of the road) still exists, despite the fact that they are not currently visible to the camera device that generates the first view 10.
[0054] exist Figure 1 In the process, although the occlusion 50, which includes details of the second object 40, is hidden from the view of the sensor device that generates the first view 10 by the first object 30, the analysis of the synthetic virtual camera view 20 enables the computerized system to reconstruct the occlusion 50 and estimate details of the second object 40 within the occlusion 50.
[0055] Figure 2 The illustration schematically shows sensor data generated by an active camera device from scene 100 as viewed from a synthetic virtual camera view, showing occlusion 150 in the sensor data. Scene 100 shows sensor data that can be used to synthesize a virtual camera view, which may include, for example, data collected by the camera device at some time prior to the current time. Objects 130, including pedestrians, are shown in scene 100. A background 140 showing a fixed object or feature in scene 100 is shown. The first view of scene 100 is shown from... Figure 2 The points shown in the synthesized virtual camera view are generated on the left and below. The first view includes occlusion 150, wherein the data generated in the first view on the left and below the shown synthesized virtual camera view does not include data or details about the area of background 140 of the fixed object or feature represented by occlusion 150, because object 130 blocks these details from the first view.
[0056] Figure 3 A data stream 200, schematically illustrated, is used to perform depth interpretation and semantic segmentation of sensor data including an input image. An input image 210, comprising a two-dimensional matrix of colorized pixels, is provided as input to a depth interpretation and semantic segmentation programming module 220. The depth interpretation and semantic segmentation programming module 220 represents a program that can operate within a computerized device or computerized controller, including a processor operable to execute programmed code. The depth interpretation and semantic segmentation programming module 220 provides a semantic segmentation output 230 that groups the pixels of the input image 210 according to identified objects, features, or classifications, for example, that can be performed using computerized image recognition techniques. In one embodiment, the depth interpretation and semantic segmentation programming module 220 can compare pixel groups with a stored image library and identify pixel groups as representing objects or features based on the comparison results. The depth interpretation and semantic segmentation programming module 220 can further use objects and features identified from the input image 210 to estimate the depth within the input image 210, for example, estimating the distance between each of the identified objects and features and the camera device that collected the data. The estimated depth within the input image 210 can be provided as a depth estimation output 240.
[0057] Figure 4 This is a flowchart illustrating method 300 for occlusion reconstruction in an surrounding view using temporal information. Method 300 begins at step 310. In step 320, the computerized system acquires one or more images captured by one or more camera devices representing the current or first view, wherein semantic segmentation and depth estimation are performed based on data from the one or more provided images, enabling estimation of the pose or self-pose of the main vehicle. The data collected and generated in step 320 can be described as occurring in time. t = n The data collected and generated in step 320 can be utilized in two ways. The data collected and generated in step 320 can be provided to step 360, where time can be manipulated. t = n The analysis of current data at that time. Furthermore, the data collected and generated in step 320 can be stored in a computerized or digital memory in step 330. In step 340, a three-dimensional computerized representation of the operating environment can be generated or updated based on the data collected and generated in step 320. In step 350, references and utilization from previous iterations ( t < n The data (i.e., historical or temporal data) is used to estimate historical camera poses that can be used to generate a synthetic virtual camera view. In step 360, the temporal data can be manipulated.t = n Analysis of the current data at that time can identify occlusions in the current data, and the occluded / obscured data can be reconstructed or filled using data from the synthetic virtual camera view generated in step 350. Method 300 ends at step 370. Method 300 can iterate, for example, data from the current time is relabeled in step 360. t = n-1 Furthermore, when method 300 returns to step 310, a new set of data is collected. Method 300 is provided as an example, and method 300 may include additional or alternative steps, and this disclosure is not limited to the examples provided herein.
[0058] The three-dimensional computerized representation generated and updated in step 340 can be provided as output or used as a computerized model for navigating vehicles.
[0059] Figure 5 An exemplary primary vehicle 400 is schematically illustrated using the disclosed method to interpret an operating environment 440, including a synthetic virtual camera view 420 used to reconstruct occlusions 442, 444 from sensor data collected by the primary vehicle 400. The primary vehicle 400 is shown, including sensor devices including a camera device arranged in front of the vehicle. The primary vehicle 400 generates a first view 410 at the current time based on data from the camera device, resulting in the first view 410 including a field of view defined at angles to the left and right sides of the longitudinal axis of the primary vehicle 400. The operating environment 440 of the primary vehicle 400 is shown, simplified, as a segmented shape having facets at different angles representing fixed features in the operating environment 440. (Shown at time...) t = n-1 Object 430 at time, and at time t = n The object at time 430'. At time... t = n-1 Object 430 is shown between the main vehicle 400 and the operating environment 440, which is in time t = n-1 Occlusion 442 is created at a certain time. A synthetic virtual camera view 420 is shown, which may include camera data collected by the main vehicle 400 at a certain prior time. t = n-1 At that time, the camera device of the main vehicle 400 is unable to collect data about a portion of the operating environment 440 represented by occlusion 442. However, data provided by the synthetic virtual camera view 420 can be used to reconstruct or fill in occlusion 442. t = n-1At that time, the main vehicle can use data about the operating environment 440 (including reconstructed occlusions 442 provided by the synthetic virtual camera view 420) for navigation.
[0060] exist t = n-1 and t = n During the time span between these points, the main vehicle 410 remained stationary, resulting in a time... t = n First view of time 410 with time t = n-1 The first view 410 is the same or substantially the same at both times because the camera device is in the same position and pose at both times. t = n-1 Object 430 moves during this time span and is shown at its new position as in time. t = n The object 430' at time t = n. The object 430' at time t = n serves as occlusion 444, occluding part of the operating environment 440. Part of the operating environment 440 represented by occlusion 444 is visible or can be estimated from information obtained by the synthesized virtual camera view 420. Additionally or alternatively, the new synthesized virtual camera view can be based on the time... t = n-1 The collected data is defined as coinciding with the first view 410 because of time. t = n-1 At that time, a portion of the operating environment 440, represented by occlusion 444, is visible to the main vehicle 400. Therefore, the time frame can be reconstructed through analysis of the time data. t = n The occlusion at that time was 444.
[0061] Figure 6 An exemplary computerized vision data controller 500 configured to operate the disclosed methods is illustrated schematically. Figure 8 schematically shown Figure 7 A computerized vision data controller 500 is provided. The computerized vision data controller 500 includes a processing unit 510, a communication unit 520, a data input / output unit 530, and a storage unit 540. It should be noted that the computerized vision data controller 500 may include other components, and some components are not shown in some embodiments.
[0062] Processing device 510 may include memory storing processor-executable instructions, such as read-only memory (ROM) and random access memory (RAM), and one or more processors executing the processor-executable instructions. In embodiments where processing device 510 includes two or more processors, the processors may operate in parallel or distributed manner. Processing device 510 may execute an operating system of computerized vision data controller 500. Processing device 510 may include one or more modules that execute programmed code or computerized processes or methods including executable steps. The illustrated modules may include functionality across a single physical device or across multiple physical devices. In exemplary embodiments, processing device 510 also includes an image processing module 512, a 3D computerized representation module 514, and an occlusion reconstruction module 516, which will be described in more detail below.
[0063] Data input / output device 530 is an apparatus operable to receive data collected from sensors and devices throughout the vehicle and process the data into a format readily usable by processing device 510. Data input / output device 530 is further operable to process output from processing device 510 and enable other devices or control modules throughout the vehicle to use that output.
[0064] The communication device 520 may include a communication / data connection to a bus device configured to transmit data to different components of the system, and may include one or more wireless transceivers for performing wireless communication.
[0065] The memory storage device 540 is a means of storing data generated or received by the computerized vision data controller 500. The memory storage device 540 may include, but is not limited to, hard disk drives, optical disk drives, and / or flash drives.
[0066] The image processing module 512 includes programming for processing data collected by the sensor devices of the main vehicle. The image processing module 512 may include... Figure 3 The deep interpretation and semantic segmentation programming module 220 can be programmed, or it can communicate with a separate deep interpretation and semantic segmentation programming module 220 to receive its output. The image processing module 512 includes techniques and methods for defining objects and features in the operating environment of the master vehicle. The image processing module 512 includes techniques and methods for classifying objects and features into either fixed features or potentially moving features.
[0067] The 3D computerized representation module 514 can receive analyzed images from the image processing module 512 and can store information. The 3D computerized representation module 514 may include programming for generating synthetic virtual camera views based on historical data. The 3D computerized representation module 514 can utilize data iteration to generate a 3D computerized representation of the operating environment of the main vehicle. The 3D computerized representation module 514 can determine a defined path within the generated representation. The 3D computerized representation module 514 can provide data from the 3D computerized representation, data from the operating environment, and / or the determined defined path to other vehicle systems such as an autonomous navigation system.
[0068] The occlusion reconstruction module 516 may include programming for identifying occlusions in the 3D computerized representation. The occlusion reconstruction module 516 may also include programming for reconstructing the identified occlusions using historical or temporal data. The reconstructed data may be provided to the 3D computerized representation module 514 to improve or update the 3D computerized representation generated and updated by the 3D computerized representation module 514.
[0069] The computerized vision data controller 500 is provided as an exemplary computerized device capable of executing programmed code to perform the methods and processes described herein. Various different embodiments of the computerized vision data controller 500, devices connected thereto, and modules operable therein are contemplated, and this disclosure is not limited to the examples provided herein.
[0070] Figure 7 The illustration shows including Figure 5 An exemplary device for a main vehicle 400, the vehicle including a system 600, the system including... Figure 6 A computerized vision data controller 500 and multiple sensors that collect data about the operating environment of a host vehicle 400 are shown. The host vehicle 400 includes: a sensor device 610 including a field of view 612, a rear-facing camera device 620 including a field of view 622, and a side mirror camera device 630. The sensor device 610 may include a front-facing camera device. In another embodiment, the sensor device 610 may additionally or alternatively include a LiDAR sensor device and / or a radar sensor device. Two side mirror camera devices 630 may be provided, one on the driver's side and the other on the passenger side of the host vehicle 400. Multiple different sensor configurations and fields of view are contemplated, and this disclosure is not limited to the examples provided. The computerized vision data controller 500 communicates with the camera devices 610, 620, and 630, as well as other vehicle systems.
[0071] Figure 8This is a flowchart illustrating the sensor data flow and the corresponding operations and determinations that help reconstruct the occlusions identified in the sensor data. Method 700 begins at step 702. In step 704, the new input image is analyzed, inferences are made regarding depth and segmentation, and occlusions in the input image are detected. In step 706, each pixel (red-green-blue, segmentation, occlusion status) is reprojected into the virtual view. Steps 708 through 726 illustrate the pixel-by-pixel process of generating and updating the virtual view and generating and updating the static structure matrix. The static structure matrix is one embodiment of a three-dimensional computerized representation of the operating environment. In step 708, it is determined whether the currently examined pixel is occluded. If the pixel is occluded, method 700 proceeds to step 710. If the pixel is not occluded, method 700 proceeds to step 714.
[0072] In step 710, it is determined whether the inspected pixel is found in the static structure matrix of the virtual view. If the pixel is found in the static structure matrix, method 700 proceeds to step 712. In step 712, the pixel value is updated, and the method proceeds to step 726. If the pixel is not found in the static structure matrix, method 700 proceeds to step 726.
[0073] In step 714, it is determined whether the pixel being inspected is segmented into static or fixed features. If the pixel being inspected is not segmented into fixed features, method 700 proceeds to step 722. If the pixel being inspected is segmented into fixed features, method 700 proceeds to step 716. In step 716, the depth of the pixel being inspected is compared with the depth of pixels stored in the static structure matrix to confirm whether the pixel being inspected has a smaller depth value than the stored pixel. If the pixel being inspected has a smaller depth value than the stored pixel, method 700 proceeds to step 718, where the pixel being inspected is used to update the static structure matrix. If the pixel being inspected does not have a smaller depth value than the stored pixel, method 700 proceeds to step 720, where the virtual view can be updated with the stored value from the static structure matrix.
[0074] In step 722, the input image is examined to determine if it is the first image to be examined. If the input image is the first image to be examined, method 700 proceeds to step 726. If the input image is not the first image to be examined, method 700 proceeds to step 724. In step 724, the virtual view is updated with the examined pixels.
[0075] In step 726, the process of steps 708 to 726 is repeated iteratively for each pixel to be inspected. When no more pixels need to be inspected, virtual view data can be provided as output, for example, to be used for occlusion reconstruction as disclosed herein. Method 700 is exemplary and may have additional and / or alternative method steps, and this disclosure is not limited to the examples provided herein.
[0076] Information about the operating environment can be used in a variety of ways. For example, in a user-driven vehicle, information about the operating environment can be used to provide navigation guidance in the form of lane-keeping outputs, collision avoidance outputs, and driving line graphic projections. In another example, in autonomous or semi-autonomous vehicles, this information can be used to provide navigation guidance by navigating the vehicle through its environment.
[0077] While the best mode of implementing this disclosure has been described in detail, those skilled in the art associated with this disclosure will recognize various alternative designs and embodiments that implement this disclosure within the scope of the appended claims.
Claims
1. A system for occlusion reconstruction in an surrounding view using temporal information, the system comprising: An active camera device that generates image data describing a first view of the operating environment; and A computerized vision data controller, comprising programming for the following situations: Analyze the image data to generate a three-dimensional computerized representation of the operating environment; Using the image data and the three-dimensional computerized representation of the operating environment, a virtual camera view of the operating environment is synthesized from a desired perspective; Identify occlusions in the virtual camera view; The image data is iterated historically to identify fixed objects within the operating environment. The historical iterations of the image data are used to estimate the padding information for the occlusion; The occlusion is reconstructed in the 3D computerized representation using pixel data from the fixed object and the filling information; and The three-dimensional computerized representation is used to provide navigation guidance within the operating environment.
2. The system according to claim 1, wherein, The program design for analyzing the image data includes a program design for performing deep interpretation and semantic segmentation on the image data.
3. The system according to claim 1, wherein, The programming for identifying the occlusion includes programming for the following situations: Identify potential moving objects in the operating environment; and Identify the occlusion in the virtual camera view caused by the potential moving object.
4. The system according to claim 1, further comprising a plurality of active camera devices for generating the image data.
5. The system according to claim 1, wherein, The programming for synthesizing the virtual camera view using the image data and the three-dimensional computerized representation includes programming for synthesizing the virtual camera view using epipolar reprojection.
6. The system of claim 1, further comprising a sensor device selected from optical detection and ranging (LIDAR) devices or radar devices; and in, The computerized vision data controller further includes programming for improving the three-dimensional computerized representation using data from the sensor device.
7. A system for occlusion reconstruction in a surrounding view using temporal information, the system comprising: The device includes: An active camera device that generates image data describing a first view of the operating environment; and A computerized vision data controller, comprising programming for the following situations: Analyze the image data to generate a three-dimensional computerized representation of the operating environment; Using the image data and the three-dimensional computerized representation of the operating environment, a virtual camera view of the operating environment is synthesized from a desired perspective; and Identify occlusions in the virtual camera view; The image data is iterated historically to identify fixed objects within the operating environment. The historical iterations of the image data are used to estimate the padding information for the occlusion; The occlusion is reconstructed in the 3D computerized representation using pixel data from the fixed object and the filling information; and The three-dimensional computerized representation is used to provide navigation guidance within the operating environment.
8. The system according to claim 7, wherein, The device includes a vehicle.
9. The system according to claim 7, wherein, The program design for analyzing the image data includes a program design for performing deep interpretation and semantic segmentation on the image data.
10. The system according to claim 7, wherein, The programming for identifying the occlusion includes programming for the following situations: Identify potential moving objects in the operating environment; and Identify the occlusion in the virtual camera view caused by the potential moving object.
11. The system of claim 7, further comprising a plurality of active camera devices for generating the image data.
12. The system according to claim 7, wherein, The programming for synthesizing the virtual camera view using the image data and the three-dimensional computerized representation includes programming for synthesizing the virtual camera view using epipolar reprojection.
13. The system of claim 7, further comprising a sensor device selected from optical detection and ranging (LIDAR) devices or radar devices; and in, The computerized vision data controller further includes programming for improving the three-dimensional computerized representation using data from the sensor device.
14. A method for occlusion reconstruction in a surrounding view using time information, the method comprising: Operate an active camera device to collect image data describing a first view of the operating environment; and Within a computerized processor: Receive the image data from the active camera device; Analyze the image data to generate a three-dimensional computerized representation of the operating environment; Using the image data and the three-dimensional computerized representation of the operating environment, a virtual camera view of the operating environment is synthesized from a desired perspective; Identify occlusions in the virtual camera view; The image data is iterated historically to identify fixed objects within the operating environment. The historical iterations of the image data are used to estimate the padding information for the occlusion; The occlusion is reconstructed in the 3D computerized representation using pixel data from the fixed object and the filling information; and The three-dimensional computerized representation is used to provide navigation guidance within the operating environment.
15. The method according to claim 14, wherein, Analyzing the image data includes performing deep interpretation and semantic segmentation on the image data.
16. The method of claim 14, wherein, The computerized processor is located inside the vehicle.
17. The method of claim 14, wherein, Identifying the occlusion includes: Identify potential moving objects in the operating environment; and Identify the occlusion in the virtual camera view caused by the potential moving object.
18. The method of claim 14, further comprising operating a plurality of active camera devices to collect the image data.
19. The method of claim 14, wherein, Synthesizing the virtual camera view using the historical iterations of the image data and the three-dimensional computerized representation includes: synthesizing the virtual camera view using epipolar reprojection.
20. The method of claim 14, further comprising: The operation is selected from sensor devices of optical detection and ranging (LIDAR) devices or radar devices; and The data from the sensor device is used to improve the three-dimensional computerized representation.
Citation Information
Patent Citations
Image processing device and image processing method
JP2022061397A
Recovering dis-occluded areas using temporal information integration
US20130294710A1