Neural radiation field for vehicle

By applying neural radiation field technology in dynamic scenarios, training benchmarks and deformation networks, and combining event camera data for supervision and training, the accuracy problem of dynamic scenario reconstruction is solved, and efficient modeling and automatic actuation functions of the vehicle operating environment are realized.

CN120107453APending Publication Date: 2025-06-06FORD GLOBAL TECH LLC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411753822.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-12-02
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively model dynamic scenes geometrically and visually, especially in vehicle operations, and it is difficult to accurately reconstruct scenes from different perspectives.

Method used

Neural radiation field (NeRF) technology is used to train reference networks and deformation networks to model the geometry and light intensity of the scene, and supervised training is performed using data from event cameras.

Benefits of technology

It realizes a more accurate and efficient reconstruction of dynamic scenes, and can generate synthetic images from different perspectives, supporting automatic actuation and path planning of vehicle components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107453A_ABST
    Figure CN120107453A_ABST
Patent Text Reader

Abstract

The invention provides a neural radiation field for a vehicle. A computer includes a processor and a memory, and the memory stores instructions executable by the processor to train a NeRF network to model a dynamic scene, and during the training, supervise the NeRF network with data from an event camera. The NeRF network is a neural radiation field that models the geometry of the scene and the light intensity of the scene. The NeRF network includes a reference network that models a scene at an initial time and a morphing network that models a change in the scene since the initial time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to neural radiation fields of vehicles. Background Art

[0002] Modern vehicles typically include a variety of sensors. Some sensors detect the outside world, for example, objects and / or features of the vehicle's surroundings, such as other vehicles, road lane markings, traffic lights and / or signs, road users, etc. Types of vehicle sensors include radar sensors, ultrasonic sensors, scanning laser rangefinders, light detection and ranging (lidar) devices, and image processing sensors (such as cameras). Summary of the invention

[0003] The present disclosure describes a technique for geometrically and visually modeling a dynamic scene to operate a vehicle through the scene. The scene is modeled using a neural radiance field (NeRF). A neural radiance field is a neural network trained to implicitly represent a specific scene. In the present disclosure, a NeRF network models both the geometry and light intensity of a scene. The NeRF network can be used to reconstruct a scene from a different perspective compared to the sensor used to train the NeRF network. The NeRF network includes a baseline network that models the scene at an initial time and a deformation network that models the change of the scene since the initial time. The baseline network can be used to initially learn a specific scene, such as the environment around the vehicle, and the deformation network can be used to track changes in the scene, for example, when an object in the environment moves or when the vehicle moves through the environment. A computer (e.g., on a vehicle) is programmed to train the NeRF network to model a dynamic scene, and during training, the NeRF network is supervised with data from an event camera mounted to the vehicle. The event camera can record an asynchronous log of intensity changes at different pixels, rather than recording data from each pixel at each time step as a traditional frame-based camera does. Thus, the event camera can generate image data that is free of motion blur and free of certain lighting issues. Thus, using data from an event camera to supervise a NeRF network can provide a more accurate reconstruction of a scene through the NeRF network than, for example, using a frame-based camera. A computer can operate a vehicle based on the NeRF network (e.g., based on one or more reconstructions generated by the NeRF network). For example, a computer can use a reconstruction from the perspective of a future point on a planned path of the vehicle to determine how to actuate a propulsion system, a braking system, and / or a steering system of the vehicle at the future point.

[0004] A computer includes a processor and a memory, and the memory stores instructions executable by the processor to train a NeRF network to model a dynamic scene, and during the training, the NeRF network is supervised by data from an event camera. The NeRF network is a neural radiance field that models the geometry of the scene and the light intensity of the scene. The NeRF network includes a baseline network that models the scene at an initial time and a deformation network that models changes in the scene since the initial time.

[0005] In one example, the instructions may further include instructions for actuating components of a vehicle including the computer and the event camera based on the NeRF network after the training.

[0006] In one example, the data from the event camera may include multiple events, each event being a change in the light intensity, each event including a pixel position and a time of the change in the light intensity. In another example, each event may indicate that the change in the light intensity at a corresponding pixel position and a corresponding time is greater than a contrast threshold.

[0007] In yet another example, each event may include a polarity indicating the direction of a corresponding change in light intensity.

[0008] In one example, the instructions may also include instructions for performing the following operations: updating the NeRF network based on a loss function. In another example, the data from the event camera may include multiple events, each event being a change in the light intensity, and the loss function may include an event loss based on the event at the pixel location. In yet another example, the event loss may include the difference between the predicted change in the light intensity at the pixel location according to the NeRF network and the sum of the events at the pixel location.

[0009] In another further example, the data from the event camera may include a plurality of events, each event being a change in the light intensity, and the loss function may include a non-event loss based on a period between consecutive events at a pixel location. In yet another further example, the non-event loss may include a difference between a predicted change in the light intensity at the pixel location during the period according to the NeRF network and a preset value.

[0010] In yet another example, the non-event loss for the pixel position may depend on whether an event occurs at a neighboring pixel position of the pixel position during the time period. In yet another example, in response to the event occurring at the neighboring pixel position during the time period, the non-event loss may include a difference between a predicted change in the light intensity at the pixel position during the time period according to the NeRF network and a contrast threshold.

[0011] In yet another example, in response to no event occurring at the neighboring pixel location during the time period, the non-event loss may force a predicted change in the light intensity at the pixel location during the time period according to the NeRF network to be zero.

[0012] In one example, the reference network may receive as input a position and a direction, and the reference network may output light intensity and volume density as seen in the direction from the position.

[0013] In one example, the deformation network may receive a current position and a current time as input, and the deformation network may output a spatial change at the current time from the initial time to a point at the current position at the current time. In another example, a reference network may receive as input a current position adjusted by the spatial change.

[0014] In one example, the baseline network and the deformed network may be multi-layer perceptrons.

[0015] A method includes training a NeRF network to model a dynamic scene, and during the training, supervising the NeRF network with data from an event camera. The NeRF network is a neural radiance field that models the geometry of the scene and the light intensity of the scene. The NeRF network includes a baseline network that models the scene at an initial time and a deformation network that models changes in the scene since the initial time.

[0016] In one example, the method may further include, after training, actuating a component of a vehicle based on the NeRF network, the vehicle including the event camera.

[0017] In one example, the deformation network may receive a current position and a current time as input, the deformation network may output a spatial change of the current time from the initial time to a point at the current position at the current time, and the reference network may receive as input the current position adjusted by the spatial change. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1is a block diagram of an example vehicle.

[0019] Figure 2 is a depiction of exemplary data from a vehicle's event camera over two time periods.

[0020] Figure 3 is an illustration of an exemplary neural radiance field (NeRF) network that receives data from an event camera.

[0021] Figure 4 is a flowchart of an exemplary process for training a NeRF network. DETAILED DESCRIPTION

[0022] Referring to the accompanying drawings, wherein like numbers indicate like parts throughout the several views, a computer 105 includes a processor and a memory, and the memory stores instructions executable by the processor to train a NeRF network 300 to model a dynamic scene, and during the training, supervise the NeRF network 300 with data 200 from an event camera 110. The NeRF network 300 is a neural radiance field that models the geometry of the scene and the light intensity of the scene. The NeRF network 300 includes a baseline network 305 that models the scene at an initial time and a deformation network 310 that models changes in the scene since the initial time.

[0023] refer to Figure 1 , the computer 105 and the event camera 110 may be part of a vehicle 100. The vehicle 100 may be any passenger vehicle or commercial vehicle, such as a car, truck, sport utility vehicle, crossover, van, minivan, taxi, bus, etc. The vehicle 100 may include the computer 105, the event camera 110, a communication network 115, a propulsion system 120, a braking system 125, a steering system 130, and a user interface 135.

[0024] The computer 105 is a microprocessor-based computing device, such as a general-purpose computing device (which includes a processor and a memory, an electronic controller, etc.), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a combination of the foregoing, etc. Typically, hardware description languages ​​such as VHDL (VHSIC (Very High Speed ​​Integrated Circuit) Hardware Description Language) are used in electronic design automation to describe digital and mixed-signal systems such as FPGAs and ASICs. For example, an ASIC is manufactured based on VHDL programming provided before manufacturing, and the logic components inside the FPGA can be configured based on, for example, VHDL programming stored in a memory electrically connected to the FPGA circuit. Therefore, the computer 105 may include a processor, a memory, etc. The memory of the computer 105 may include a medium for storing instructions executable by the processor and for electronically storing data and / or a database, and / or the computer 105 may include a structure such as the foregoing structure that provides programming. The computer 105 may be a plurality of computers coupled together.

[0025] The computer 105 may transmit and receive data via the communication network 115. The communication network 115 may be, for example, a controller area network (CAN) bus, Ethernet, WiFi, a local interconnect network (LIN), an on-board diagnostic connector (OBD-II), and / or any other wired or wireless communication network. The computer 105 may be communicatively coupled to the event camera 110, the propulsion system 120, the braking system 125, the steering system 130, the user interface 135, and other components via the communication network 115.

[0026] Event camera 110 is an imaging sensor that responds to local changes in intensity, sometimes referred to as a neuromorphic camera or a dynamic vision sensor. Each pixel in event camera 110 responds independently to the change in intensity when a change occurs, and each pixel returns no data in the absence of an intensity change. Each pixel stores a reference intensity value and compares the reference intensity value to the current intensity value. In response to the intensity difference exceeding a threshold, the pixel resets the reference intensity value to the current intensity value and outputs data indicating the change in intensity along with a timestamp of when the change occurred. Figure 2 Discuss data 200 generated by event camera 110. Event camera 110 may be of any suitable type, for example, a temporal contrast sensor such as a dynamic vision sensor (DVS) or sensitive DVS (sDVS), a temporal image sensor, a dynamic and active pixel vision sensor (DAVIS), a retinal morphology sensor, and the like.

[0027] The propulsion system 120 of the vehicle 100 generates energy and converts the energy into motion of the vehicle 100. The propulsion system 120 can be a conventional vehicle propulsion subsystem, such as a conventional powertrain system, which includes an internal combustion engine coupled to a transmission that transfers rotational motion to the wheels; an electric powertrain system, which includes a battery, an electric motor, and a transmission that transfers rotational motion to the wheels; a hybrid powertrain system, which includes elements of a conventional powertrain system and an electric powertrain system; or any other type of propulsion device. The propulsion system 120 can include an electronic control unit (ECU) that communicates with and receives input from the computer 105 and / or a human operator. The human operator can control the propulsion system 120 via, for example, an accelerator pedal and / or a gear shift lever.

[0028] The braking system 125 is typically a conventional vehicle braking subsystem and resists the motion of the vehicle 100, thereby slowing and / or stopping the vehicle 100. The braking system 125 may include friction brakes, such as disc brakes, drum brakes, band brakes, etc.; regenerative brakes; any other suitable type of brake; or a combination thereof. The braking system 125 may include an electronic control unit (ECU) or the like that communicates with and receives input from the computer 105 and / or a human operator. The human operator may control the braking system 125 via, for example, a brake pedal.

[0029] The steering system 130 is typically a conventional vehicle steering subsystem and controls the turning of the wheels. The steering system 130 can be a rack and pinion system with electric power steering, a steer-by-wire system (both of which are known), or any other suitable system. The steering system 130 can include an electronic control unit (ECU) or the like that communicates with and receives input from the computer 105 and / or a human operator. The human operator can control the steering system 130 via, for example, a steering wheel.

[0030] The user interface 135 presents information to and receives information from an operator of the vehicle 100. The user interface 135 may be located, for example, on a dashboard in the passenger compartment of the vehicle 100, or anywhere else that the operator can easily see it. The user interface 135 may include a dial, a digital readout, a screen, a speaker, etc., for providing information to the operator, such as, for example, known human-machine interface (HMI) elements. The user interface 135 may include buttons, knobs, keypads, microphones, etc., for receiving information from the operator.

[0031] refer to Figure 2 , the data 200 from the event camera 110 includes a plurality of events. Each event is a change in light intensity. Each event includes the pixel location of the change in light intensity, the time, and possibly the polarity indicating the direction of the corresponding change in light intensity, e.g. i =(ui ,t i ,p i ), where e is the event, i is the index of time, u is the two-dimensional pixel position, t is the timestamp, and p is the polarity. The polarity p can be a binary variable, for example, it can take the value of –1 for a decrease in light intensity, or +1 for an increase in light intensity. Figure 2 205 are events during which the light intensity increased (ie, p=1), and the black dots 210 are events during which the light intensity decreased (ie, p=-1). (For clarity, the Figure 2 Only a few of the points 205, 210 are marked in FIG. 1 . ) The light intensity may be monochromatic, i.e., not decomposed into specific colors, and the light intensity may therefore be a scalar value. Each event indicates that a change in light intensity at a corresponding pixel position and a corresponding time is greater than a contrast threshold; in other words, each event may be generated in response to a change in light intensity at a corresponding pixel position and a corresponding time being greater than a contrast threshold. The change in light intensity may be expressed as a logarithmic intensity difference, for example, as shown in the following expression:

[0032] |log(I(u,t i ))-log(I(u,t i-1 ))|≥C

[0033] Where I is the light intensity at a particular pixel and time, and C is a contrast threshold. The contrast threshold C may be a pre-programmed property or a physical property of the event camera 110.

[0034] refer to Figure 3 In general, a Neural Radiance Field (NeRF) is a neural network trained to implicitly represent a particular scene. Once a NeRF is trained on a scene, it can be used to generate new views of the scene from perspectives not included in the training data. Conventionally, the training data is simply multiple image frames of the scene from a frame-based camera, and the trained NeRF takes a three-dimensional position x = (x, y, z) and a viewing direction d = (x d ,y d ) as input and output the scalar volume density σ at that location and the color c = (r, g, b) emitted at that location towards the viewing direction. NeRF is trained for a specific scene. For conventional NeRF, if a different scene is of interest or if the scene changes, a new NeRF needs to be trained.

[0035] The NeRF network 300 is a neural radiation field that models the geometry of a scene and the light intensity of the scene. Once trained, the NeRF network 300 receives a three-dimensional position, a viewing direction, and a time as input 315, and generates light intensity and volume density as output, i.e., Ψ:(x, d, t)→(I, σ), where ψ is the NeRF network 300, x is a three-dimensional position, d is a viewing direction, t is time, I is light intensity, and σ is a scalar volume density. The volume density and light intensity are as seen in the viewing direction from the position at that time. In other words, the volume density is the predicted volume density at the position and time, and the light intensity is the light intensity emitted at the position toward the viewing direction at that time. The volume density and light intensity can be scalars. The viewing direction d can be a two-dimensional vector, such as an orthogonal horizontal dimension. The NeRF network 300 includes a reference network 305 and a deformation network 310.

[0036] The reference network 305 is at the initial time t 0 Once trained, the reference network 305 receives a three-dimensional position and a viewing direction as input, and generates as output a volume density at that position and the light intensity emitted at that position towards the viewing direction, i.e., Ψ x :(x,d)→(I,σ), where ψ x is the reference network 305. The volume density and light intensity are as seen in the viewing direction from the position at the initial time. In other words, the volume density is the predicted volume density at the position and the initial time, and the light intensity is the light intensity emitted at the position towards the viewing direction at the initial time. The volume density and light intensity may be scalars.

[0037] The deformation network 310 models the change of the scene since the initial time. Once trained, the deformation network 310 receives the current position and the current time as input, and the deformation network 310 outputs the spatial change of the current time from the initial time to the point at the current position at the current time, that is, Ψ t :(x,t)→Δx, where Ψ t is the deformation network 310 and Δx is the spatial change. In other words, the spatial change Δx restores the point at position x at time t to its original position at the initial time t. 0 The spatial variation Δx has the same dimensions as the position, for example, three dimensions.

[0038] The reference network 305 and the deformed network 310 can be multilayer perceptrons (MLPs). MLPs are fully connected feedforward artificial neural networks. Each MLP includes an input layer, at least one hidden layer, and an output layer. Layers are composed of nodes. The nodes in each layer receive as input the outputs of the nodes in the previous layer, which start from the input layer and end at the output layer. Each connection between nodes in adjacent layers has a weight. Each node has an activation function that takes the weighted input of the node as its independent variable. MLPs are fully connected because each node in a layer is connected to each node in an adjacent layer. Training an MLP results in changing the weights via backpropagation.

[0039] The NeRF network 300 includes a deformation network 310 and a reference network 305, which is arranged in series with the reference network 305 following the deformation network 310. The deformation network 310 receives the current position x and the current time t, and outputs the initial time t. 0 =The spatial change Δx to the current position x at the current time t. The reference network 305 receives as input the current position adjusted by the spatial change (i.e., x′=x+Δx) and the viewing direction d, and the reference network 305 outputs the volume density and light intensity as seen in the viewing direction from the position at that time. The reference network 305 and the deformed network 310 are different networks. Therefore, the reference network 305 and the deformed network 310 do not have any interaction terms, i.e., there are no connections from nodes in one of the reference network 305 and the deformed network 310 to nodes in the other of the reference network 305 and the deformed network 310, except that the final output of the deformed network 310 is used as an input to the reference network 305.

[0040] The computer 105 is programmed to render the ray 320 extending from the event camera 110 by executing the NeRF network 300 at sample points along the ray 320. The location of the sample point and the corresponding time may be input to the NeRF network 300 along with the direction of the ray 320. The computer 105 then outputs the expected light intensity of the ray 320 The expected light intensity is the expected light intensity of the corresponding pixel (u, v) of the image that will be returned by the frame-based camera at the origin of the ray 320. For example, the computer 105 can determine the expected light intensity of the ray 320 by summing or integrating the intensity terms of the sample points, the terms being weighted based on the volume density and the distance between the sample points (e.g., based on an exponent of the product of a scalar volume density and the distance between consecutive sample points, e.g., an orthogonal approximation of the volume rendering equation), as shown in the following expression:

[0041]

[0042] in is the expected light intensity of ray 320, k and m are the indices of the sample points, N is the total number of sample points on ray 320, exp() is the exponential function, i.e., the Euler number e raised to the power of its argument, σ() is the volume density at the location of its argument, and b m is the distance from the origin of ray 320 to sample point m, δ m is the distance between sample points m and m+1, x′(b m ,t) is the position of the sample point m adjusted for the spatial variation at time t, and I() is the light intensity at the position and direction of its independent variable. The distance between sample points can be expressed as the difference between the distance from the origin to one sample point and the distance to the next sample point, i.e., δ m =b m+1 –b m The adjusted position x′ is determined by executing the deformation network 310 . The volume density σ() and the light intensity I() are determined by executing the reference network 305 on the independent variables derived from executing the deformation network 310 .

[0043] Computer 105 can be programmed to generate synthetic image 325 from a different viewpoint and a different time than event camera 110 by executing NeRF network 300. For example, computer 105 can receive an input three-dimensional position, an input viewing direction, and an input time; generate a plurality of viewing directions extending within a vertical and horizontal range around the input viewing direction; and calculate a plurality of expected light intensities at the input time in the manner described above based on the input position and the corresponding viewing directions. Each expected light intensity is a pixel in synthetic image 325.

[0044] The computer 105 is programmed to train the NeRF network 300 to model dynamic scenes. During training, the computer 105 supervises the NeRF network 300 with data 200 from the event camera 110. In other words, the computer 105 trains the NeRF network 300 to model dynamic scenes, which are represented by the data 200 from the event camera 110. As a general overview, the computer 105 calculates a loss function that compares the value derived from the NeRF network 300 in training with the ground truth data 200 from the event camera 110. The computer 105 supervises the NeRF network 300 with the data 200 from the event camera 110 using the loss function. The loss function may include an event loss and a non-event loss, as will be described in detail below. The computer 105 updates the NeRF network 300 based on the loss function. For example, as known, back propagation can be used to adjust the weights of the MLP used as the deformation network 310 and the reference network 305 to minimize the value of the loss function.

[0045] To determine event loss, computer 105 can be programmed to sample multiple pairs of events from data 200, which will be referred to as event pairs. For the purposes of this disclosure, an event pair is defined as two events that occur at the same pixel location at different times, which are not necessarily continuous. The sample is designed to capture changes in light intensity so that the NeRF network 300 is trained to reflect these changes. Using non-continuous event pairs in event loss can help prevent errors from accumulating over time. Sampled event pairs can cover a large number of pixel locations and times.

[0046] The event loss is based on the events at the pixel position and is then aggregated across many pixel positions. Specifically, for each event pair in the sample, the event loss includes the difference between the predicted change in the light intensity at the corresponding pixel position according to the NeRF network 300 and the sum of the events at the pixel position. The predicted change in light intensity can be from the earlier event in the event pair to the later event in the event pair. The predicted change in light intensity can be the logarithmic difference of the predicted intensity at the earlier event and the later event. The sum can be the sum of the intermediate events between the earlier event and the later event of the event pair, for example, the sum of the product of the polarity of each event and the contrast threshold. For example, the event loss can be the mean square error of the difference between the predicted change in light intensity at the pixel position and the sum of the events at the pixel position across the sample event pair, as shown in the following expression:

[0047]

[0048] where i is the index of the earlier event in the event pair, j is the index of the later event in the event pair, and u is the pixel position of the event pair i, j.

[0049] To determine non-event losses, the computer 105 may be programmed to sample multiple time periods between consecutive events at corresponding pixel locations, i.e., each time period is between two consecutive events at the same pixel location. The computer 105 may sample events in the data 200 and the corresponding time between the corresponding event and the corresponding next consecutive event at the corresponding pixel location. Each sampled event e i Has a period running from an earlier sampling time to a later sampling time. The earlier sampling time t i It is event i time, and the later sampling time t i' is greater than time t i and is smaller than the next consecutive event e at the same pixel position i+1 Time t i+1 The random time, t i <t i′ <t i+1The samples are designed to capture intervals at pixel locations when no changes in light intensity occur, so that the NeRF network 300 is trained to not include changes when no changes occur.

[0050] The non-event loss is based on the time period between consecutive events at a particular pixel location. Specifically, the non-event loss includes the difference between the predicted change in the light intensity at the pixel location during the time period according to the NeRF network 300 and a preset value. The non-event loss (e.g., a preset value) for the pixel location may depend on whether an event occurs at a neighboring pixel location of the pixel location during the time period. For the purposes of this disclosure, a "neighboring pixel location" is defined as a pixel location within a preset range of the pixel location of interest, e.g., for a pixel of interest (u, v), the neighboring pixels may be within (u±1, v±1). For example, for a sample event e i , in response to the i to i' An event occurs at an adjacent pixel during the period, the preset value may be a contrast threshold C, and in response to the occurrence of an event at an adjacent pixel during the period t i to i' If no event occurs at the adjacent pixel during this period, the preset value can be zero. In other words, for the sample event e i , in response to the i to i' During the period t, an event occurs at a neighboring pixel, and the non-event loss can force the predicted change in light intensity to be less than the contrast threshold C, and in response to the i to i' During a period when no events occurred at neighboring pixels, the non-event loss can force the predicted change in light intensity to be zero. The form of the expression for the non-event loss can also change depending on whether an event occurred at the neighboring pixel location during the period, for example, the rectified linear unit (ReLU) activation function for neighboring events and the mean squared error for the lack of neighboring events, for example, as shown in the following expression:

[0051]

[0052] where relu() is the ReLU activation function and neighborhood(u) is the set of all neighboring pixel locations from pixel location u.

[0053] Figure 4is a flow chart illustrating an exemplary process 400 for training a NeRF network 300 and operating a vehicle 100 based on the NeRF network 300. The memory of the computer 105 stores executable instructions for performing the steps of the process 400, and / or may be programmed to implement a structure such as that mentioned above. As a general overview of the process 400, the computer 105 receives data 200 from the event camera 110 and performs sampling for event losses and non-event losses. For each event pair in the sample of event losses, the computer 105 calculates the event loss and updates the NeRF network 300. For each event in the sample of non-event losses, the computer 105 calculates the non-event loss and updates the NeRF network 300. The computer 105 then actuates components of the vehicle 100 based on the trained NeRF network 300. The process 400 is repeated as long as the vehicle 100 is running.

[0054] Process 400 begins at block 405 where computer 105 receives data 200 from event camera 110 , as described above.

[0055] Next, in block 410 , the computer 105 generates a sample of event losses and non-event losses, as described above.

[0056] Next, the computer 105 advances to the next event pair in the sample of event losses in block 415. For example, event pairs may be assigned index values, and the computer 105 may advance to the next index value in ascending order starting from the smallest index value.

[0057] Next, in box 420, the computer 105 calculates the event loss for the current event pair, as described above.

[0058] Next, in block 425 , the computer 105 updates the NeRF network 300 based on the event losses (eg, using back-propagation), as described above.

[0059] Next, in decision block 430, computer 105 determines whether the current event pair is the last event pair in the sample, e.g., the event pair with the highest index value. If not, process 400 returns to block 415 to continue with the next event pair. If yes, process 400 proceeds to block 435.

[0060] In block 435, the computer 105 advances to the next event and corresponding time period in the sample of non-event losses. For example, events may be assigned index values, and the computer 105 may advance to the next index value in ascending order starting from the smallest index value.

[0061] Next, in box 440, the computer 105 calculates the non-event losses for the time period of the current event, as described above.

[0062] Next, in block 445 , the computer 105 updates the NeRF network 300 based on the non-event losses (eg, using back-propagation), as described above.

[0063] Next, in decision block 450, computer 105 determines whether the current event is the last event in the sample, e.g., the event with the highest index value. If not, process 400 returns to block 435 to continue with the next event. If yes, process 400 proceeds to block 455.

[0064] In box 455, i.e., after training, the computer 105 actuates components, such as components of the vehicle 100, based on the NeRF network 300. For example, the computer 105 may actuate one or more of the propulsion system 120, the braking system 125, the steering system 130, or the user interface 135. For example, the computer 105 may actuate components when executing an advanced driver assistance system (ADAS). ADAS is an electronic technology that assists drivers in achieving driving functions and parking functions. Examples of ADAS include forward approach detection, lane departure detection, blind spot detection, brake actuation, adaptive cruise control, and lane keeping assist systems. The computer 105 may actuate the braking system 125 according to an auxiliary braking algorithm to stop the vehicle 100 before reaching an object in the environment as indicated by the NeRF network 300. The computer 105 may actuate the user interface 135 to output a message to the operator informing them of the object indicated by the NeRF network 300 according to the forward approach detection algorithm. Computer 105 may operate vehicle 100 autonomously, i.e., actuate propulsion system 120, braking system 125, and steering system 130 based on NeRF network 300. Computer 105 may use synthetic images 325 at possible future points on the planned path to execute a path planning algorithm to navigate vehicle 100 around objects in the environment.

[0065] Next, in decision block 460, computer 105 determines whether vehicle 100 is still on. In response to vehicle 100 being still on, process 400 returns to block 405 to retrain NeRF network 300, for example, by updating deformation network 310 for an additional time period in the next iteration. In response to vehicle 100 being off, process 400 ends.

[0066] Generally speaking, the computing systems and / or devices described may employ any of a variety of computer operating systems, including but not limited to the following versions and / or varieties: Ford Application; AppLink / Smart Device Link middleware; Microsoft Operating system; Microsoft Operating system; Unix operating system (for example, distributed by Oracle Corporation of Redwood Shores, California operating systems); AIX UNIX operating system distributed by International Business Machines Corporation of Armonk, New York; Linux operating system; Mac OSX and iOS operating systems distributed by Apple Inc. of Cupertino, California; BlackBerry operating system distributed by BlackBerry Ltd. of Waterloo, Canada; and Android operating system developed by Google Inc. and the Open Handset Alliance; or provided by QNX Software Systems CAR Infotainment Platform. Examples of computing devices include, but are not limited to, an in-vehicle computer, a computer workstation, a server, a desktop, notebook, laptop or handheld computer, or some other computing system and / or device.

[0067] Computing devices typically include computer-executable instructions, wherein the instructions are executable by one or more computing devices such as those listed above. Computer-executable instructions may be compiled or interpreted from a computer program created using a variety of programming languages ​​and / or technologies, including, alone or in combination, but not limited to Java. TM , C, C++, Matlab, Simulink, Stateflow, Visual Basic, Java Script, Python, Perl, HTML, etc. Some of these applications can be compiled and executed on virtual machines such as Java virtual machines, Dalvik virtual machines, etc. Typically, a processor (e.g., a microprocessor) receives instructions from, for example, a memory, a computer-readable medium, etc., and executes these instructions to perform one or more processes, including one or more of the processes described herein. Such instructions and other data can be stored and transmitted using a variety of computer-readable media. Files in a computing device are typically a collection of data stored on a computer-readable medium such as a storage medium, a random access memory, etc.

[0068] Computer-readable media (also referred to as processor-readable media) include any non-transitory (e.g., tangible) media that participate in providing data (e.g., instructions) that can be read by a computer (e.g., by a processor of a computer). Such media can take many forms, including, but not limited to, non-volatile media and volatile media. Instructions can be transmitted via one or more transmission media, including optical fibers, wires, wireless communications, including internals that make up a system bus coupled to a processor of a computer. Common forms of computer-readable media include, for example, RAM, PROM, EPROM, FLASH-EEPROM, any other memory chip or cassette, or any other medium from which a computer can read.

[0069] The database, data repository or other data storage described herein may include various mechanisms for storing, accessing / accessing and retrieving various data, including hierarchical databases, file sets in file systems, application databases in special formats, relational database management systems (RDBMS), non-relational databases (NoSQL), graphic databases (GDB), etc. Each such data storage is typically included in a computing device using a computer operating system such as one of the above-mentioned, and is accessed via a network in any one or more of a variety of ways. The file system can be accessed from the computer operating system and may include files stored in various formats. In addition to the language (such as the above-mentioned PL / SQL language) for creating, storing, editing and executing stored programs, RDBMS typically also uses structured query language (SQL).

[0070] In some examples, system elements may be implemented as computer-readable instructions (e.g., software) on one or more computing devices (e.g., servers, personal computers, etc.) stored on computer-readable media associated therewith (e.g., disks, memories, etc.). A computer program product may include such instructions stored on a computer-readable medium for performing the functions described herein.

[0071] In the accompanying drawings, the same reference numerals indicate the same elements. In addition, some or all of these elements may be changed. With respect to the media, processes, systems, methods, heuristics, etc. described herein, it should be understood that although the steps of such processes, etc. have been described as occurring in a certain ordered sequence, such processes may be practiced by performing the steps in an order different from that described herein. It should also be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. The operations, systems, and methods described herein should always be implemented and / or performed in accordance with applicable owner / user manuals and / or safety guidelines.

[0072] The present disclosure has been described in an illustrative manner, and it should be understood that the terminology that has been used is intended to be in the nature of descriptive words rather than limiting. The use of "in response to," "after determining...", etc. indicates a causal relationship, not just a temporal relationship. In light of the above teachings, many modifications and variations of the present disclosure are possible, and the present disclosure may be practiced in other ways than specifically described.

[0073] According to the present invention, a computer is provided, the computer having a processor and a memory, the memory storing instructions, the instructions executable by the processor to: train a NeRF network to model a dynamic scene, the NeRF network being a neural radiance field that models the geometry of the scene and the light intensity of the scene, the NeRF network comprising a baseline network that models the scene at an initial time and a deformation network that models changes in the scene since the initial time; and during the training, supervise the NeRF network with data from an event camera.

[0074] According to an embodiment, the instructions further comprise instructions for actuating components of a vehicle comprising the computer and the event camera based on the NeRF network after the training.

[0075] According to an embodiment, the data from the event camera comprises a plurality of events, each event being a change in the light intensity, each event comprising a pixel position and a time of the change in the light intensity.

[0076] According to an embodiment, each event indicates that said change in said light intensity at a respective pixel position and a respective time is greater than a contrast threshold.

[0077] According to an embodiment, each event comprises a polarity indicating the direction of the corresponding change in light intensity.

[0078] According to an embodiment, the instructions further comprise instructions for updating the NeRF network based on a loss function.

[0079] According to an embodiment, the data from the event camera comprises a plurality of events, each event being a change in the light intensity, and the loss function comprises an event loss based on the events at a pixel location.

[0080] According to an embodiment, the event loss comprises a difference between a predicted change of the light intensity at the pixel location according to the NeRF network and a sum of the events at the pixel location.

[0081] According to an embodiment, the data from the event camera comprises a plurality of events, each event being a change in the light intensity, and the loss function comprises a non-event loss based on a period between consecutive events at a pixel location.

[0082] According to an embodiment, the non-event loss comprises a difference between a predicted change of the light intensity at the pixel location during the period according to the NeRF network and a preset value.

[0083] According to an embodiment, the non-event loss for the pixel position depends on whether an event occurs at a neighboring pixel position of the pixel position during the time period.

[0084] According to an embodiment, in response to the event occurring at the neighboring pixel location during the time period, the non-event loss comprises a difference between a predicted change in the light intensity at the pixel location during the time period according to the NeRF network and a contrast threshold.

[0085] According to an embodiment, in response to no event occurring at the neighboring pixel location during the time period, the non-event loss forces a predicted change in the light intensity at the pixel location during the time period according to the NeRF network to be zero.

[0086] According to an embodiment, the reference network receives as input a position and a direction, and the reference network outputs light intensity and volume density as seen in said direction from said position.

[0087] According to an embodiment, the deformation network receives as input a current position and a current time, and the deformation network outputs a spatial change in the current time from the initial time to a point at the current position at the current time.

[0088] According to an embodiment, the reference network receives as input the current position adjusted by the spatial variation.

[0089] According to an embodiment, the baseline network and the deformed network are multi-layer perceptrons.

[0090] According to the present invention, a method includes: training a NeRF network to model a dynamic scene, wherein the NeRF network is a neural radiance field that models the geometry of the scene and the light intensity of the scene, wherein the NeRF network includes a baseline network that models the scene at an initial time and a deformation network that models the change of the scene since the initial time; and during the training, supervising the NeRF network with data from an event camera.

[0091] In one aspect of the invention, the method comprises, after the training, actuating a component of a vehicle based on the NeRF network, the vehicle comprising the event camera.

[0092] In one aspect of the invention, the deformation network receives as input a current position and a current time, the deformation network outputs a spatial change in the current time from the initial time to a point at the current position at the current time, and the reference network receives as input the current position adjusted by the spatial change.

Claims

1. A method comprising: training a NeRF network to model a dynamic scene, the NeRF network being a neural radiance field that models the geometry of the scene and the light intensity of the scene, the NeRF network comprising a baseline network that models the scene at an initial time and a deformation network that models changes to the scene since the initial time; as well as During the training, the NeRF network is supervised with data from an event camera.

2. The method of claim 1 further comprising, after the training, actuating components of a vehicle based on the NeRF network, the vehicle including the event camera.

3. The method of claim 1, wherein the data from the event camera comprises a plurality of events, each event being a change in the light intensity, each event comprising a pixel position and time of the change in the light intensity. 4 . The method of claim 3 , wherein each event indicates that the change in the light intensity at a corresponding pixel location and a corresponding time is greater than a contrast threshold.

5. The method of claim 1, further comprising updating the NeRF network based on a loss function.

6. The method of claim 5, wherein the data from the event camera comprises a plurality of events, each event being a change in the light intensity, and the loss function comprises an event loss based on the events at a pixel location.

7. The method of claim 6, wherein the event loss comprises a difference between a predicted change in the light intensity at the pixel location according to the NeRF network and a sum of the events at the pixel location.

8. The method of claim 5, wherein the data from the event camera comprises a plurality of events, each event being a change in the light intensity, and the loss function comprises a non-event loss based on a period between consecutive events at a pixel location.

9. The method of claim 8, wherein the non-event loss comprises a difference between a predicted change in the light intensity at the pixel location during the time period according to the NeRF network and a preset value.

10. The method of claim 8, wherein the non-event penalty for the pixel location depends on whether an event occurs at a neighboring pixel location of the pixel location during the time period.

11. The method of claim 10, wherein in response to the event occurring at the neighboring pixel location during the time period, the non-event loss comprises a difference between a predicted change in the light intensity at the pixel location during the time period according to the NeRF network and a contrast threshold.

12. The method of claim 10, wherein the non-event loss forces a predicted change in the light intensity at the pixel location during the time period according to the NeRF network to be zero in response to no event occurring at the neighboring pixel location during the time period.

13. The method of claim 1, wherein the reference network receives a position and a direction as input, and the reference network outputs light intensity and volume density as seen in the direction from the position.

14. The method of claim 1, wherein the deformation network receives as input a current position and a current time, the deformation network outputs a spatial change in the current time from the initial time to a point at the current position at the current time, and the reference network receives as input the current position adjusted by the spatial change.

15. A computer comprising a processor and a memory, the memory storing instructions executable by the processor to perform the method of one of claims 1 to 14.