Intelligent vehicle systems and control logic for surround view enhancement
By using natural surround vision technology and depth inference, AI neural networks are used to identify and replace objects in vehicle sensor images to generate virtual view images. This solves the problem of narrow near-field field of view in sensor systems and achieves more accurate target detection and improved virtual scene quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GM GLOBAL TECHNOLOGY OPERATIONS LLC
- Filing Date
- 2022-10-11
- Publication Date
- 2026-04-24
Smart Images

Figure CN116137655B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to control systems for motor vehicles. More specifically, aspects of this disclosure relate to advanced driving systems with image segmentation, depth inference, and object recognition capabilities. Background Technology
[0002] Currently manufactured motor vehicles (such as Hyundai cars) can be equipped with networks of onboard electronics that provide autonomous driving capabilities to help minimize driver effort. For example, one of the most identifiable types of autonomous driving features in automotive applications is cruise control. Cruise control allows the vehicle operator to set a specific vehicle speed and has the onboard vehicle computer system maintain that speed without the driver operating the accelerator or brake pedal. Next-generation adaptive cruise control (ACC) is an autonomous driving feature that regulates vehicle speed while managing headway between the lead vehicle and the leading "target" vehicle. Another type of autonomous driving feature is collision avoidance systems (CAS), which detect impending collisions and warn the driver while also taking preventative measures, such as steering or braking without driver input. Intelligent Parking Assist Systems (IPAS), Lane Monitoring and Automatic Steering ("Automatic Steering") systems, Electronic Stability Control (ESC) systems, and other Advanced Driver Assistance Systems (ADAS) are also available in many modern cars.
[0003] As vehicle processing, communication, and sensing capabilities continue to improve, manufacturers will continue to offer more automated driving capabilities, aiming to produce fully autonomous "self-driving" vehicles capable of operating among heterogeneous vehicle types in urban and rural scenarios. Original equipment manufacturers (OEMs) are moving towards vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) "talking" cars with higher levels of driving automation, employing intelligent control systems to achieve vehicle routing with steering, lane changes, scene planning, and more. Automated path planning systems utilize vehicle state and dynamic sensors, geolocation information, map and road condition data, and path prediction algorithms to provide route derivation with automatic lane centering and lane change prediction. Computer-aided route reselection technology provides alternative route predictions that can be updated, for example, based on real-time and virtual vehicle data.
[0004] Many cars are now equipped with in-vehicle navigation systems that utilize a Global Positioning System (GPS) transceiver in conjunction with navigation software and geolocation mapping services to obtain road terrain, traffic, and speed limit data associated with the vehicle's real-time location. Autonomous driving and advanced driver assistance systems (ADAS) are often able to adapt autonomous driving maneuvers based on road information obtained from the in-vehicle navigation system. For example, ADAS based on self-organizing networks can use GPS and map data combined with multi-hop geocast V2V and V2I data exchange to facilitate autonomous vehicle maneuvering and powertrain control. During assisted and unassisted vehicle operation, the in-vehicle navigation system can identify a recommended route based on the estimated shortest travel time or estimated shortest travel distance between the route's origin and destination for a given trip. This recommended route can then be displayed on a geocoded and annotated map as a map track or navigation driving directions, with optional voice commands output by the in-vehicle audio system.
[0005] Automated and autonomous vehicle systems can employ a variety of sensing components to provide object detection and ranging. For example, Radio Detection and Ranging (RADAR) systems detect the presence, distance, and / or speed of a target object by returning a pulse of high-frequency electromagnetic waves reflected from the object to a suitable radio receiver. Alternatively, vehicles can employ Laser Detection and Ranging (LADAR) backscattering systems, which emit and detect pulsed laser beams for precise distance measurement. Synonymous with LADAR-based detection and often used as a general term for LADAR-based detection is Light Detection and Ranging (LIDAR) technology, which uses various forms of light energy (including invisible, infrared, and near-infrared laser spectra) to determine the distance to a stationary or moving target. Onboard sensor arrays with various digital cameras, ultrasonic sensors, etc., also provide real-time target data. Historically, these object detection and ranging systems have been limited in accuracy and application due to narrow field of view in the vehicle's near field, and inherently limited by constraints on the total sensor count and available packaging locations. Summary of the Invention
[0006] This paper proposes an intelligent vehicle system, methods for manufacturing and using such a system, and a motor vehicle equipped with such a system, the intelligent vehicle system having a networked onboard vehicle camera and accompanying control logic for camera view enhancement using object recognition and model replacement. As an example, a system and method are proposed that employ Natural Surround Vision (NSV) techniques to replace detected objects in a surround camera view with 3D models retrieved from a model set database based on object recognition, image segmentation, and depth inference. Such object classification and replacement can be particularly useful for scenarios where parts of the object are blurred from or invisible to the resident camera system and therefore the object rendering from the desired viewpoint is partial. For model collection, computer vision techniques identify identifiable objects (e.g., vehicles, pedestrians, bicycles, lampposts, etc.) and aggregate them into corresponding object sets for future reference. Image segmentation is implemented to divide two-dimensional (2D) camera images into individual segments; the images are then evaluated on a segment-by-segment basis to group each pixel into a corresponding class. Additionally, depth inference infers the 3D structure in the 2D image by assigning a range (depth) to each image pixel. Using an AI neural network (NN), object recognition analyzes images generated by a 2D camera to classify features from the images and associates the features of a target object with a corresponding set of models. The target object is then removed and replaced with a 3D object model that is generic to that set, for example, in a new "virtual" image.
[0007] At least some of the disclosed concepts include a target acquisition system that more accurately detects and classifies target objects within two-dimensional (2D) camera images, for example, to provide more comprehensive contextual awareness and thus enable more intuitive driver and vehicle responses. Other accompanying benefits may include providing virtual camera sensing capabilities to derive virtual perspectives from assisted vantage points (e.g., bird's-eye view or a following vehicle "chase view" or "dashcam" view) using real-world images captured by a network of stationary cameras. Other benefits may include the vehicle system's ability to generate virtual views with improved virtual scene quality from data captured by physical cameras, while preserving the 3D structure of the captured scene and enabling real-time changes to the virtual scene perspective.
[0008] Various aspects of this disclosure relate to system control logic, intelligent control techniques, and computer-readable media (CRM) for manufacturing and / or operating any of the disclosed vehicle sensor networks. In one example, a method for controlling the operation of a motor vehicle having a sensor array comprising a network of cameras mounted at discrete locations on the vehicle is proposed. This representative method, in any order and in any combination with any of the options and features disclosed above and below, includes, for example, receiving camera data from a sensor array via one or more stationary or remote system controllers, indicating one or more camera images of a target object having one or more camera views from one or more cameras; for example, analyzing the camera images via an object recognition module using a system controller to identify characteristics of the target object and classify the characteristics into a corresponding model collection selected from multiple predefined model collections and associated with the type of the target object; for example, retrieving a three-dimensional (3D) object model common to the corresponding model collection associated with the type of the target object from an object library stored in stationary or remote memory via a controller; generating one or more new images (e.g., virtual images) by replacing the target object with a 3D object model positioned in a different orientation; and for example, sending one or more command signals to one or more stationary vehicle systems via a system controller to perform one or more control operations using the new images.
[0009] A non-transitory CRM for storing instructions is also proposed, which can be executed by one or more processors of a system controller for a motor vehicle including a sensor array having a network of cameras mounted at discrete locations on the vehicle body. When executed by one or more processors, these instructions cause the controller to perform operations including: receiving camera data from the sensor array indicating a target object having a camera image from the viewpoint of one of the cameras; analyzing the camera image to identify characteristics of the target object and classifying the characteristics into a corresponding one of multiple model collections associated with the type of the target object; determining a 3D object model of the corresponding model collection associated with the type of the target object; generating a new image by replacing the target object with a 3D object model positioned in a different orientation; and sending command signals to a parked vehicle system to perform control operations using the new image (e.g., displaying the new image as a virtual image).
[0010] Additional aspects of this disclosure relate to motor vehicles equipped with an intelligent control system employing networked onboard cameras with camera view enhancement capabilities. As used herein, the terms "vehicle" and "motor vehicle" are used interchangeably and synonymously to include any relevant vehicle platform, such as passenger cars (ICE, HEV, FEV, fuel cell, fully and partially autonomous, etc.), commercial vehicles, industrial vehicles, tracked vehicles, off-road and all-terrain vehicles (ATVs), motorcycles, agricultural equipment, boats, aircraft, etc. A prime mover, such as an electric traction motor and / or an internal combustion engine assembly, drives one or more wheels, thereby propelling the vehicle. A sensor array is also mounted to the vehicle body, comprising a network of cameras (e.g., front, rear, port side, and starboard side cameras) mounted at discrete locations on the vehicle body.
[0011] Continuing the discussion of the above example, the vehicle also includes one or more stationary or remote electronic system controllers that communicate with a sensor array to receive camera data indicating one or more camera images of a target object from one or more viewpoints of the camera. Using an object recognition module, the controller analyzes the camera images to identify characteristics of the target object and categorizes these characteristics into a corresponding model collection associated with that type of target object. The controller then accesses a memory-stored object library to retrieve a common 3D object model from a predefined model collection associated with the target object type. A new image is then generated by replacing the target object with the 3D object model, such as a virtual image of a tail vehicle chase view, where the 3D object model is positioned in a new orientation different from the camera orientation. Subsequently, one or more command signals can be sent to one or more stationary vehicle systems to perform one or more control operations using the new image.
[0012] For any disclosed vehicle, system, and method, the system controller may implement an image segmentation module to divide a camera image into multiple distinct segments. The segments are then evaluated to group the pixels contained within them into corresponding predefined classes. In this case, the image segmentation module may be operable to execute a computer vision algorithm that analyzes each distinct segment individually to identify pixels within the segment that share at least one predefined attribute. The segmentation module then assigns bounding boxes to delineate pixels sharing one or more predefined attributes associated with a target object.
[0013] For any publicly disclosed vehicle, system, and method, the system controller can implement depth inference to derive the corresponding depth ray for each image pixel in 3D space. In this case, the depth inference module is operable to receive camera data indicating overlapping camera images of a target object containing multiple camera views from multiple cameras. The depth inference module then processes this camera data via a neural network trained to output depth data and semantic segmentation data using a loss function that combines multiple loss terms, including a semantic segmentation loss term and a panorama loss term. The panorama loss term provides a similarity measure for overlapping blocks of camera data, each overlapping block corresponding to a region of the overlapping field of view of the camera.
[0014] For any disclosed vehicle, system, and method, the system controller can implement a polar reprojection module to generate a virtual image of a target object (e.g., as a new image) from an alternative viewpoint of a virtual camera. In this case, the polar reprojection module can be operable to determine the real-time orientation of the camera capturing the camera image and the desired orientation of the virtual camera used to present the target object from the alternative viewpoint in the virtual image. The reprojection module then defines the epipolar geometry between the real-time orientation of the physical camera and the desired orientation of the virtual camera. The virtual image is generated based on the calculated epipolar relationship between the real-time orientation of the camera and the desired orientation of the virtual camera.
[0015] For any disclosed vehicle, system, and method, the system controller can be programmed to implement an object orientation inference module to estimate the orientation of a target object in 3D space relative to a predefined origin axis of the sensor array. In this case, the estimated orientation of the target object in 3D space is used to determine the orientation of the 3D object model within a new image. Alternatively, the system controller can communicate with the occupant input device of the motor vehicle to receive an avatar usage protocol having a set of rules defining how and / or when the target object is replaced by the 3D object model. The generation of new images can be further based on the avatar usage protocol. The system controller can also derive the size, position, and / or orientation of the target object and use the derived characteristics of the target object to calculate rendering parameters for the 3D object model. These rendering parameters can include the 2D projection of the 3D object model into one or more new images.
[0016] For any disclosed vehicle, system, and method, the parked vehicle system may include a virtual imaging module that displays virtual images and driver cues based on those virtual images. In this case, control operations include simultaneously displaying virtual images from a following vehicle's dashcam view and prompting the driver of the vehicle to take driving actions based on those virtual images. Optionally, the parked vehicle system may include an advanced driver assistance system control module for automatically controlling the motor vehicle. As yet another option, the parked vehicle system may include a vehicle navigation system with an in-vehicle display device. In this case, control operations may include the display device displaying a new image with a 3D object model, for example, as an integrated augmented reality (AR) chase view.
[0017] The foregoing summary is not intended to represent every embodiment or aspect of this disclosure. Rather, it provides only examples of some novel concepts and features set forth herein. The foregoing features and advantages, as well as other features and accompanying advantages, will become apparent from the following detailed description of the illustrated examples and representative modes for carrying out this disclosure when considered in conjunction with the accompanying drawings and appended claims. Furthermore, this disclosure expressly includes any and all combinations and sub-combinations of the elements and features presented above and below. Attached Figure Description
[0018] Figure 1 The diagram illustrates a representative intelligent motor vehicle according to various aspects of this disclosure, the intelligent motor vehicle having a network of onboard controllers, sensing devices, and communication devices for performing surround view enhancement using a recognizable object model.
[0019] Figure 2 This is a flowchart illustrating a representative object recognition and model replacement protocol for operating a vehicle sensor array with a networked camera, which may correspond to memory-stored instructions executable by a resident or remote controller, control logic circuit, programmable control unit, or other integrated circuit (IC) device or network of devices, according to aspects of the disclosed concept.
[0020] This disclosure is adaptable to various modifications and alternatives, and some representative embodiments are illustrated by way of example in the accompanying drawings and will be described in detail herein. However, it should be understood that the novel aspects of the invention are not limited to the specific forms shown in the drawings listed above. Rather, this disclosure will cover all modifications, equivalents, combinations, sub-combinations, substitutions, groupings, and alternatives that fall within the scope of this disclosure, for example, as covered by the appended claims. Detailed Implementation
[0021] This disclosure allows for numerous different forms of embodiments. Representative embodiments of this disclosure are shown in the accompanying drawings and will be described in detail herein. It should be understood that these embodiments are provided as examples of the disclosed principles and not as limitations on the broad aspects of this disclosure. To this extent, elements and limitations described, for example, in the abstract, introduction, summary, and detailed description sections but not expressly set forth in the claims should not be incorporated, individually or collectively, by implication, inference, or otherwise, into the claims.
[0022] For the purposes of this detailed specification, unless specifically denied: the singular includes the plural, and vice versa; the words “and” and “or” should be both connected and separate; the words “any” and “all” should both mean “any and all”; furthermore, approximate words such as “about,” “almost,” “basically,” “usually,” “approximately,” etc., may each be used herein in the sense of, for example, “in,” “nearly,” or “within 0-5%” or “within acceptable manufacturing tolerances,” or any logical combination thereof. Finally, directional adjectives and adverbs, such as forward, aft, inside, outside, starboard, port, vertical, horizontal, up, down, forward, aft, left, right, etc., can be relative to the motor vehicle, such as the forward driving direction of the motor vehicle when the vehicle is operably oriented on a level driving surface.
[0023] Referring now to the accompanying drawings, in which the same reference numerals denote the same features in several views, Figure 1 A representative automobile, generally designated 10, is shown herein and, for the purposes of discussion, is depicted as a sedan-type electric passenger vehicle. The automobile 10 shown (also referred to herein simply as a "motor vehicle" or "vehicle") is merely an exemplary application that can practice the novel aspects of this disclosure. Similarly, incorporating this concept into an all-electric vehicle powertrain system should also be understood as a non-limiting implementation of the disclosed features. Therefore, it should be understood that aspects and features of this disclosure can be applied to other powertrain architectures, can be implemented for any logically related type of vehicle, and can be used for various navigation and automated vehicle operations. Furthermore, only selected components of the motor vehicle and vehicle control system are shown and described in additional detail herein. However, the vehicles and vehicle systems discussed below may include numerous additional and alternative features for performing the various methods and functions of this disclosure, as well as other available peripheral components.
[0024] Figure 1 The representative vehicle 10 is initially equipped with a vehicle telecommunications and information (“telematics”) unit 14, which connects to a remote or “external” cloud computing host service 24, for example, via a cell tower, base station, mobile switching center, satellite service, etc. Wireless communication. As a non-limiting example, Figure 1Other vehicle hardware components 16 shown in the overall diagram include an electronic video display device 18, a microphone 28, an audio speaker 30, and various user input controls 32 (e.g., buttons, knobs, pedals, switches, touchpads, joysticks, touchscreens, etc.). These hardware components 16 partially serve as a human-machine interface (HMI) to enable users to communicate with the telematics unit 14 and other system components within the vehicle 10. The microphone 28 provides vehicle occupants with a means of inputting verbal or other auditory commands; the vehicle 10 may be equipped with an embedded voice processing unit utilizing audio filtering, editing, and analysis modules. Conversely, the speaker 30 provides audible output to vehicle occupants and may be a separate speaker dedicated to use with the telematics unit 14, or it may be part of an audio system 22. The audio system 22 is operatively connected to a network connection interface 34 and an audio bus 20 to receive analog information via one or more speaker components and present it as sound.
[0025] The telematics unit 14 is communicatively coupled to a network interface 34, suitable examples of which include a twisted-pair / fiber Ethernet switch, a parallel / serial communication bus, a local area network (LAN) interface, a controller area network (CAN) interface, a media-oriented system transport (MOST) interface, a local interconnect network (LIN) interface, etc. Other suitable communication interfaces may include communication interfaces conforming to ISO, SAE, and / or IEEE standards and specifications. The network interface 34 enables the vehicle hardware 16 to send and receive signals to and from each other, and to send and receive signals with various systems and subsystems within or "residing" in the vehicle body 12, as well as outside or "away" from the vehicle body 12. This allows the vehicle 10 to perform various vehicle functions, such as modulating powertrain output, controlling the operation of the vehicle's transmission, selectively engaging friction and regenerative braking systems, controlling vehicle steering, regulating the charging and discharging of the vehicle's battery modules, and other autonomous driving functions. For example, the telematics unit 14 sends signals and data to / receives signals and data from the powertrain control module (PCM) 52, the advanced driver assistance system (ADAS) module 54, the electronic battery control module (EBCM) 56, the steering control module (SCM) 58, the brake system control module (BSCM) 60, and various other vehicle ECUs such as the transmission control module (TCM), the engine control module (ECM), the sensor system interface module (SSIM), etc.
[0026] Continue to refer to Figure 1The telematics unit 14 is an onboard computing device that provides hybrid services, both independently and through communication with other networked devices. The telematics unit 14 typically comprises one or more processors 40, each of which may be embodied as a discrete microprocessor, an application-specific integrated circuit (ASIC), or a dedicated control module. The vehicle 10 may be provided with centralized vehicle control via a central processing unit (CPU) 36, which is operatively coupled to a real-time clock (RTC) 42 and one or more electronic memory devices 38, each of which may take the form of a CD-ROM, disk, IC device, flash memory, semiconductor memory (e.g., various types of RAM or ROM), etc.
[0027] Remote vehicle communication capabilities with remote, non-vehicle-connected devices can be provided via one or more of a cellular chipset / component, a navigation and positioning chipset / component (e.g., a Global Positioning System (GPS) transceiver), or a wireless modem, all of which are collectively represented at point 44. Short-range wireless connectivity can be provided via a short-range wireless communication device 46 (e.g., The vehicle 10 may be provided with a unit or near-field communication (NFC) transceiver, a dedicated short-range communication (DSRC) component 48, and / or dual antennas 50. It should be understood that the vehicle 10 may be implemented without one or more of the components listed above, or alternatively, may include additional components and functions required for a particular end use. The various communication devices described above may be configured to exchange data as part of periodic broadcasts in vehicle-to-vehicle (V2V) or vehicle-to-everything (V2X) communication systems (e.g., vehicle-to-infrastructure (V2I), vehicle-to-pedestrian (V2P), vehicle-to-device (V2D), etc.).
[0028] CPU 36 receives sensor data from one or more sensing devices using, for example, optical detection, radar, laser, ultrasonic, optical, infrared, or other suitable technologies, including short-range communication technologies (e.g., DSRC) or ultra-wideband (UWB) radio technologies, for performing autonomous driving operations or vehicle navigation services. According to the illustrated example, vehicle 10 may be equipped with one or more digital cameras 62, one or more range sensors 64, one or more vehicle speed sensors 66, one or more vehicle dynamic sensors 68, and any necessary filtering, classification, fusion, and analysis hardware and software for processing the raw sensor data. The type, placement, quantity, and interoperability of the distributed array of onboard sensors can be individually or collectively adapted to a given vehicle platform to achieve a desired level of autonomous vehicle operation.
[0029] Digital camera 62 can use a complementary metal-oxide-semiconductor (CMOS) sensor or other suitable optical sensor to generate images indicating the field of view of vehicle 10, and can be configured for continuous image generation, for example, at least about 35+ images per second. In comparison, distance sensor 64 can emit and detect reflected radio, infrared, light-based, or other electromagnetic signals (e.g., short-range radar, long-range radar, EM sensing, light detection and ranging (LIDAR), etc.) to detect, for example, the presence, geometry, and / or proximity of a target object. Vehicle speed sensor 66 can take various forms, including wheel speed sensors that measure wheel speed, which are then used to determine real-time vehicle speed. Additionally, vehicle dynamics sensor 68 can have properties such as a single-axis or three-axis accelerometer, angular rate sensor, inclinometer, etc., for detecting longitudinal and lateral acceleration, yaw, roll and / or pitch rates, or other dynamically relevant parameters. Using data from sensing devices 62, 64, 66, and 68, CPU 36 identifies the surrounding driving conditions, determines road characteristics and surface conditions, identifies target objects within the vehicle's detectable range, determines the target objects' attributes such as size, relative position, orientation, distance, approach angle, and relative speed, and performs automatic control maneuvers based on these actions.
[0030] These sensors can be distributed throughout the motor vehicle 10 in an operationally unobstructed position relative to the front or rear, port or starboard sides of the vehicle. Each sensor generates an electrical signal indicating the characteristics or condition of the host vehicle or one or more target objects, typically as an estimate with a corresponding standard deviation. While the operating characteristics of these sensors are generally complementary, some sensors are more reliable than others in estimating certain parameters. Most sensors have different operating ranges and coverage areas and are capable of detecting different parameters within their operating range. For example, radar-based sensors can estimate the distance, rate of change of distance, and azimuth position of an object, but may not be robust in estimating the range of a detected object. On the other hand, cameras with optical processing may be more robust in estimating the shape and azimuth position of an object, but may be less efficient in estimating the range and rate of change of range of a target object. Sensors based on scanning lidar can perform this function effectively and accurately relative to estimating the distance and azimuth position, but may not be able to accurately estimate the rate of change of distance, and therefore may be inaccurate in acquiring / identifying new objects. In contrast, ultrasonic sensors can estimate distance, but generally cannot accurately estimate the rate of change of distance and azimuth position. Furthermore, the performance of many sensor technologies can be affected by different environmental conditions. Therefore, sensors typically exhibit parameter variance, and their operational overlap provides opportunities for sensory fusion.
[0031] To propel the electric-drive vehicle 10, the electric powertrain is operable to generate traction torque and transmit it to one or more of the vehicle's wheels 26. Figure 1 The concept is typically represented by a rechargeable energy storage system (RESS), which may have the characteristics of a chassis-mounted traction battery pack 70 operatively connected to an electric traction motor 78. The traction battery pack 70 typically comprises one or more battery modules 72, each having a stack of battery cells 74, such as pouch, can, or prismatic lithium-ion, lithium-polymer, or nickel-metal hydride battery cells. One or more motors, such as a traction motor / generator (M) unit 78, draw power from the RESS's battery pack 70 and optionally deliver power to the RESS's battery pack 70. A dedicated power inverter module (PIM) 80 electrically connects the battery pack 70 to the motor / generator (M) unit 78 and modulates the current transfer therebetween. The disclosed concept is similarly applicable to HEVs and ICE-based powertrain architectures.
[0032] Battery pack 70 can be configured to integrate module management, cell sensing, and module-to-module or module-to-host communication functions directly into each battery module 72 and to perform them wirelessly via a wireless-enabled cell monitoring unit (CMU) 76. CMU 76 can be a microcontroller-based printed circuit board (PCB) mounted sensor array. Each CMU 76 can have a GPS transceiver and RF capability and can be packaged on or within the battery module housing. Battery module units 74, CMU 76, housing, coolant lines, busbars, etc., collectively define the unit module assembly.
[0033] Next reference Figure 2 The flowchart, according to aspects of this disclosure, describes in general 100 places the method for use with the main vehicle (such as...) Figure 1 An improved method or control strategy for target acquisition, object recognition, and 3D model replacement with enhanced surround view of the vehicle (10). Figure 2 Some or all of the operations shown and further described in detail below may represent algorithms corresponding to processor-executable instructions, such as those stored in main memory, secondary memory, or remote memory (e.g., ...). Figure 1 In the memory device 38), and for example by an electronic controller, processing unit, logic circuit or other module or device or module / device network (e.g., Figure 1 The CPU 36 and / or cloud computing service 24) execute to perform any or all of the functions described above and below associated with the disclosed concepts. It should be understood that the execution order of the illustrated operation boxes can be changed, additional operation boxes can be added, and some operations in the described operations can be modified, combined, or eliminated.
[0034] Figure 2 Method 100 begins at a start terminal box 101, where memory-stored processor-executable instructions are used by a programmable controller or control module, or a similar suitable processor, to call an initialization process for camera view enhancement via a generic object model protocol. This routine can be executed in real-time, near real-time, continuously, systematically, sporadically, and / or at regular intervals (e.g., every 10 or 100 milliseconds) during normal and ongoing operation of the vehicle 10. Alternatively, terminal box 101 can be initialized in response to user command prompts, resident vehicle controller prompts, or broadcast prompt signals received from an “external” centralized vehicle service system (e.g., host cloud computing service 24). Upon completion... Figure 2 When performing control operations, method 100 can proceed to the end terminal box 121 and temporarily terminate, or alternatively, it can loop back to terminal box 101 and run continuously in a loop.
[0035] Method 100 proceeds from terminal box 101 to sensor data input box 103 to acquire image data from one or more available sensors on the vehicle body. For example, Figure 1 The vehicle 10 may initially be equipped with or modified to include a front camera 102 mounted near the front of the vehicle body (e.g., on the front grille), a rear camera 104 mounted near the rear of the vehicle body (e.g., on the tailgate or trunk lid), and a driver-side camera 106 and a passenger-side camera 108 respectively mounted near the respective lateral sides of the vehicle body (e.g., on the starboard and port side mirrors). According to the illustrated example, the front camera 102 captures a real-time forward-facing view of the vehicle (e.g., an outer field of view pointing in front of the front bumper assembly), and the rear camera 104 captures a real-time rearward-facing view of the vehicle (e.g., an outer field of view pointing behind the rear bumper assembly). Similarly, the left-side camera 106 captures a real-time port-side view of the vehicle (e.g., an outer field of view laterally to the driver-side door assembly), and the right-side camera 108 captures a real-time starboard-side view of the vehicle (e.g., an outer field of view laterally to the passenger-side door assembly). Each camera generates and outputs a signal indicating its respective view. These signals can be retrieved directly from the camera or from a memory device tasked with receiving, classifying, and storing such data.
[0036] When aggregating, filtering, and preprocessing image data received from sensor data frame 103, method 100 initiates a Natural Surround Vision (NSV) protocol by executing a depth and segmentation predefined processing frame 105 and a reprojection predefined processing frame 107. Depth inference can generally be represented as the process of inferring the 3D structure of a 2D image by assigning a corresponding range (depth) to each image pixel. For a vehicle-calibrated camera, each pixel may have orientation information, where the pixel position in the image corresponds to a directional ray in 3D space, e.g., defined by pixel coordinates and camera intrinsic parameters. The 3D position of the pixel in 3D space is established by the added range along this ray inferred by a specially trained neural network (e.g., using deep learning monocular depth estimation). Figure 2 In this embodiment, 2D camera images 118 output by vehicle-mounted cameras 102, 104, 106, and 108 can be converted into computer-enhanced grayscale images 118a, wherein the values of pixels associated with a specific target are assigned individual grayscale shadows representing the average light intensity of those pixels. This facilitates the delineation of pixel depths of one target object relative to another. In at least some embodiments, the "grayscale image" is not displayed to the driver or otherwise visible to vehicle occupants. Furthermore, each pixel within a given image segment can be assigned one of several classes into which it is segmented.
[0037] Image segmentation can be represented as a computer vision algorithm that divides a 2D image into discrete segments and evaluates the segments to group each related pixel with similar attributes into a corresponding class (e.g., grouping all pixels of a cyclist into a class). During segmentation evaluation, object detection can construct bounding boxes corresponding to each class in the image. Figure 2 In this process, the 2D camera images 118 output by the vehicle-mounted cameras 102, 104, 106, and 108 can be converted into computer-enhanced grayscale images 118b, in which all pixels associated with a specific target are assigned a single grayscale shading. This helps to distinguish a target object from other target objects contained in the same image.
[0038] For depth inference, dense depth maps can be estimated from image data input by vehicle cameras 102, 104, 106, and / or 108. The dense depth map can be estimated in part by processing one or more camera images using a deep neural network (DNN). Examples of such deep neural networks may include encoder-decoder architectures programmed to generate depth data and semantic segmentation data. The DNN can be trained based on a loss function combining loss terms, including depth loss, depth smoothness loss, semantic segmentation loss, and panorama loss. In this example, the loss function is a single multi-task learning loss function. The depth data output by the trained DNN can be used for various vehicle applications, including stitching or joining vehicle camera images, estimating the distance between the host and detected objects, dense depth prediction, modifying perspective views, and generating surround views.
[0039] Image data from the vehicle's surround-view cameras can be processed by a DNN to enforce consistency in depth estimation. These surround-view cameras can possess high-resolution, wide-lens camera properties (e.g., 180°+1080p+ and 60 / 30fps+ resolution). A panorama loss term can utilize reprojection of images from the surround-view cameras at a common viewpoint as part of the loss function. Specifically, a similarity metric can be employed, comparing overlapping image patches from neighboring surround-view cameras. The DNN employs multi-task learning to jointly learn both depth segmentation and semantic segmentation. As part of evaluating the panorama loss term, a 3D point cloud can be generated using predefined camera extrinsic and intrinsic parameters and the inferred depth. This 3D point cloud is then projected onto a common plane to provide a panoramic image. The panorama loss term evaluates the similarity of overlapping regions in the panoramic images as part of the loss function. The loss function can combine disparity, its smoothness, semantic segmentation, and panorama loss into a single loss function. Additional information relating to depth estimation and image segmentation can be found, for example, in U.S. Patent Application No. 17 / 198,954 entitled “Systems and Methods for Depth Estimation in a Vehicle”, co-owned by Albert Shalumov et al., the entire contents of which are incorporated herein by reference and used for all purposes.
[0040] Using image data output from sensor data frame 103 and image depth and segmentation data output from predefined processing frame 105, the epipolar reprojection predefined processing frame 107 can use epipolar geometry to correlate the projection of 3D points in space with 2D images, so that corresponding points in the multi-view camera images are correlated. In a non-limiting example, the epipolar reprojection module can use a co-located depth sensor to process images captured by one or more of physical cameras 102, 104, 106, 108 and depth information assigned to pixels in these captured images. The physical orientation of each camera in the captured images is determined; epipolar geometry is established between the physical camera and the virtual camera. Generating a virtual image output by the virtual camera may involve resampling the pixel depth information of the captured images in epipolar coordinates, identifying target pixels on a specified epipolar of the physical camera, deriving a disparity map of the output epipolar of the virtual camera, and generating an output image based on one or more of these output epipolars. The system controller can obtain orientation data from the physical camera and / or depth sensor from a pose measurement sensor or a pose estimation module. Additional information relating to the use of epipolar reprojection to generate virtual camera views can be found, for example, in U.S. Patent Application No. 17 / 189,877, co-owned by Michael Slutsky et al., entitled “Using Epipolar Reprojection for Virtual View Perspective Change,” the entire contents of which are incorporated herein by reference and used for all purposes.
[0041] At the virtual image data display box 109, one or more virtual images are generated from the super-reprojection data output from the predefined processing box 107, representing one or more alternative viewpoints from one or more virtual cameras. For example, a first virtual image 116A from a third-person "chasing view" perspective (e.g., as if observing the main vehicle 110 from a following vehicle located directly behind the main vehicle 110) shows a target object 112 depicted by bounding box 114 and adjacent to the driver's side fender of the main vehicle 110. A second virtual image 116b from the same virtual viewpoint can be generated to show the target object 112 removed from bounding box 114. One or both of these images 116A, 116B can be presented to the occupants of the main vehicle (e.g., in...). Figure 1 (on the video display device 18 or the remote information processing unit 14).
[0042] Before, during, or after executing the NSV protocol, method 100 initiates the surround view enhancement protocol by performing object classification and orientation inference predefined processing box 111. Using the captured image 118 of the surround view of the main vehicle, the object classification module identifies one or more target objects (e.g., pedestrians 112) within the image and then assigns each target to a corresponding class from various predefined classes (e.g., motor vehicles, pedestrians, cyclists and other wheeled pedestrians, lampposts and other signs, houses, pets and other animals, etc.). As a non-limiting example, the AI NN analyzes multiple 2D images generated by a single vehicle camera or multiple vehicle cameras to locate the expected target, classifies the characteristics of each expected target within the image, and then systematically associates these characteristics with a corresponding model collection to characterize each expected target accordingly. For example, although presented as 2D images, almost all objects within the captured images are inherently three-dimensional.
[0043] During object classification in predefined processing box 111, the orientation inference module derives the corresponding orientation of each target object in 3D space relative to the host vehicle. Generally, it may be insufficient to simply detect expected target objects without associating them with their corresponding object classes. To appropriately replace the target objects with computer-generated 3D object models, as described further below, the targets are oriented in space, for example, relative to a predefined origin axis of the vehicle sensor array. As a non-limiting example, a convolutional neural network classifies the viewpoints (e.g., front, rear, port, starboard) of vehicles within an image by generating bounding boxes around target vehicles near the host and deriving the orientation or travel of each target vehicle relative to the host vehicle. Using the relative viewpoints of the host vehicle and a large-scale target dataset, the orientation inference module can identify the orientations of these target vehicles relative to the host. The position, size, and orientation of target objects can also be estimated using a single monocular image and inverse perspective mapping to estimate distances to specified portions of the image. Specifically, inertial offset units are used to eliminate camera pitch and roll motion; the corrected camera images are then projected using inverse perspective mapping (e.g., via a bird's-eye view). The convolutional neural network simultaneously detects the vehicle's position, size, and orientation data. The predicted orientation bounding box from the bird's-eye view image is transformed using an inverse projection matrix. Through this process, the projected bird's-eye view image can be aligned parallel and linear to the xy-plane of a universal coordinate system, allowing the target object's orientation to be evaluated against this universal coordinate system's xy-plane.
[0044] Once the target object has been classified at predefined processing box 111, method 100 invokes processor-executable instructions of object extraction predefined processing box 113 to retrieve a computer-generated 3D object model that is generic for a model collection set associated with the classification type of the target object. For example, a vehicle sensor system can “retrieve” a 3D object model from a memory-stored 3D object library by specifying a “key” for the 3D object model; then, a suitable database management system (DBMS) routine retrieves the 3D object assigned to that key from the model collection database. Commonly encountered and identifiable objects (such as vehicles, pedestrians, bicycles, lampposts, etc.) are aggregated into corresponding object sets for subsequent reference from the model collection database. A “generic” 3D object model can then be assigned to each object set within the collection; this 3D object replaces the target object from a camera image in a virtual (AR) image. For example, all target vehicles that have corresponding vehicle characteristics (e.g., wheels, windows, rearview mirrors, etc.) and are within a specified range of sedan vehicle sizes (e.g., approximately 13-16 feet in length) can be associated with the corresponding "sedan vehicle" object set and replaced with a 3D model of a basic sedan motor vehicle. Similarly, all pedestrians that have corresponding human characteristics (e.g., head, feet, arms, torso, legs, etc.) and are within a specified "large adult" human height (e.g., over 5 feet 10 inches) can be associated with the corresponding "large adult" object set and replaced with a 3D model of a basic large adult 112a.
[0045] At the Protocol Input / Output Data Frame 117, the user or occupant of the main vehicle can use any of the aforementioned input devices (e.g., ...). Figure 1 The user-defined avatar usage protocol can be defined using a user input control 32 in vehicle 10 or a personal computing device (e.g., a smartphone, laptop, tablet, etc.). The user-defined avatar usage protocol can contain one or more rule sets that define how a target object is replaced (if any) by a 3D object model. Generally, the driver can be allowed to input a custom rule set that determines how and when a target object in the surrounding camera view will be replaced by a computer graphics image, including the “look and feel” of the replacement image. In this regard, the driver can be allowed to select which 3D object models or characteristics of those 3D object models will be assigned to a given set. Furthermore, the driver can be allowed to disable target object replacement features, restrict their use to only certain objects, limit their use for a specific time, restrict their use to certain driving conditions, restrict their use to certain driving settings, etc.
[0046] Method 100 proceeds from processing box 113 to rendering parameter predefined processing box 115 to calculate rendering parameters for the 3D object model of the target object within the virtual image that will replace the virtual camera view. According to the example shown, the object rendering module determines the estimated size of the target object 112 to be replaced, the estimated position of the target object 112 within the surround view of the main vehicle, and the estimated orientation of the target object 112 relative to the main vehicle 110. Using the estimated size, position, and orientation of the target object, the computer graphics routine renders a 2D projection of the 3D object model 112a with the corresponding size, position, and orientation onto the image and overlays the 3D object model 112a at the appropriate position on the virtual image.
[0047] At virtual image data display box 119, the super-projection data output from predefined processing box 107 and the rendering parameter data output from predefined processing box 117 generate a virtual image from an alternative perspective of the virtual camera. For example, a third virtual image 116c from the perspective of a third-person "dashcam" shows a 3D object model 112a inserted into a virtual image adjacent to the driver's side fender of the main vehicle 110. Alternatively, the ADC module for computer-controlled operation of the motor vehicle can manage vehicle maneuvers based on the virtual image containing the 3D object model. Optionally, the ASAS module for computer-controlled operation of the motor vehicle can use the virtual image to perform vehicle maneuvers. At this point, method 100 can proceed to terminal box 121 and temporarily terminate.
[0048] In some embodiments, aspects of this disclosure may be implemented by a computer-executable program of instructions, such as a program module, which is generally referred to as a software application or application executed by any of the controllers or controller variants described herein. In non-limiting examples, the software may include routines, programs, objects, components, and data structures that perform specific tasks or implement specific data types. The software may form an interface to allow a computer to react to an input source. The software may also cooperate with other code segments to initiate various tasks in response to received data in conjunction with the source of the received data. The software may be stored on any of a variety of storage media, such as CD-ROM, magnetic disk, and semiconductor memory (e.g., various types of RAM or ROM).
[0049] Furthermore, aspects of this disclosure can be practiced with various computer systems and computer network configurations, including multiprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframes, etc. Additionally, aspects of this disclosure can be practiced in distributed computing environments, where tasks are performed by resident and remote processing devices linked via communication networks. In distributed computing environments, program modules can reside in local and remote computer storage media, including memory storage devices. Therefore, aspects of this disclosure can be implemented in combination with various hardware, software, or combinations thereof in computer systems or other processing systems.
[0050] Any method described herein may include machine-readable instructions for execution by: (a) a processor, (b) a controller, and / or (c) any other suitable processing device. Any algorithm, software, control logic, protocol, or method disclosed herein may be embodied as software stored on a tangible medium, such as, for example, flash memory, solid-state drive (SSD) memory, hard disk drive (HDD) memory, CD-ROM, digital versatile disc (DVD), or other storage devices. The entire algorithm, control logic, protocol, or method and / or portions thereof may alternatively be executed by a device other than a controller and / or embodied in firmware or dedicated hardware (e.g., implemented by application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable logic devices (FPLDs), discrete logic, etc.). Furthermore, while a particular algorithm may be described with reference to the flowcharts and / or workflow diagrams depicted herein, many other methods for implementing the example machine-readable instructions may be used alternatively.
[0051] Various aspects of this disclosure have been described in detail with reference to the illustrated embodiments; however, those skilled in the art will recognize that many modifications can be made thereto without departing from the scope of this disclosure. This disclosure is not limited to the precise constructions and compositions disclosed herein; any and all modifications, alterations, and variations apparent from the foregoing description are within the scope of this disclosure as defined by the appended claims. Furthermore, this concept expressly includes any and all combinations and sub-combinations of the foregoing elements and features.
Claims
1. A method for controlling the operation of a motor vehicle, the motor vehicle having a sensor array including a network of cameras mounted at discrete locations on the motor vehicle, the method comprising: Camera data indicating a target object having a camera image from the viewpoint of one of the cameras is received from the sensor array via an electronic system controller; The system controller uses an object recognition module to analyze the camera images to identify the characteristics of the target object and classify the characteristics into one of a set of multiple models associated with the type of the target object. The controller retrieves a general 3D object model from an object library stored in memory, which is associated with the type of the target object and is a collection of corresponding models. The system controller receives an avatar usage protocol from the occupant input device of the motor vehicle, the avatar usage protocol having a set of rules defining how and / or when the target object is replaced by the 3D object model; as well as Based on the avatar usage protocol, a new virtual image is generated from a selected viewpoint using camera data from at least one vehicle camera, and the target object in the virtual image is replaced with the 3D object model positioned in different orientations.
2. The method according to claim 1, further comprising: The system controller uses an image segmentation module to divide the camera image into multiple different segments; as well as The fragment is evaluated to group the pixels contained therein into corresponding classes in a plurality of predefined classes.
3. The method according to claim 2, wherein, The image segmentation module is operable to execute a computer vision algorithm that individually analyzes each of the different segments to identify pixels in the segments that share predefined attributes.
4. The method of claim 2, further comprising deriving the corresponding dense depth for each pixel corresponding to the camera image mapping via the depth inference module using the system controller.
5. The method according to claim 4, wherein, The depth inference module can be operated as follows: Supplementary camera data indicating the overlap between camera images and the target object is received from multiple camera views of multiple cameras in the camera; as well as The supplementary camera data is processed by a neural network trained to output depth data and semantic segmentation data using a loss function that combines multiple loss terms, including a semantic segmentation loss term and a panorama loss term. The panorama loss term includes a similarity measure of overlapping blocks of the camera data, each of which corresponds to a region of the overlapping field of view of the multiple cameras.
6. The method of claim 4, further comprising generating a virtual image from an alternative perspective of the virtual camera via the super-projection module using the system controller, wherein, The new image is the virtual image.
7. The method according to claim 6, wherein, The extremely heavy projection module is operable as follows: The camera's real-time orientation is received when capturing images from the camera; Receive the desired orientation of the virtual camera used to present the target object in the virtual image from the alternative perspective; as well as Define the epipolar geometry between the real-time orientation of the camera and the desired orientation of the virtual camera. The virtual image is generated based on a calculated epipolar relationship between the real-time orientation of the camera and the desired orientation of the virtual camera.
8. The method according to claim 1, further comprising: The system controller uses an object orientation inference module to estimate the orientation of the target object in 3D space relative to a predefined origin axis of the sensor array; as well as The estimated orientation of the target object in 3D space is used to determine the different orientations of the 3D object model within the new image.
9. The method of claim 1, wherein the generation of the new image is further based on the avatar usage protocol.
10. The method according to claim 1, further comprising: Determine the size, position, and orientation of the target object; as well as The size, position, and orientation of the target object are used to calculate rendering parameters for the 3D object model, the rendering parameters including the 2D projection of the 3D object model in the new image, wherein the generation of the new image is also based on the calculated rendering parameters.
Citation Information
Patent Citations
Using epipolar reprojection for virtual view perspective change
US20220284660A1
Systems and methods for depth estimation in a vehicle
US20220292289A1
Depth sensing using an RGB camera
US20150248765A1
Method for classification and segmentation and forming 3D models from images
US20150287211A1
Augmented reality in vehicle platforms
US20180225875A1