Frequency-based feature constraints for neural networks

CN116168209BActive Publication Date: 2026-09-29GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211255260.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-24
Filing Date
2022-10-13
Publication Date
2026-09-29
Estimated Expiration
2042-10-13

Smart Images

  • Figure CN116168209B_ABST
    Figure CN116168209B_ABST
Patent Text Reader

Abstract

A system includes a computer comprising a processor and a memory. The memory includes instructions causing the processor to be programmed to receive frequency filtered spatial domain data at a neural network; compare an output generated by the neural network to a loss function including a frequency-based feature consistency constraint; and update at least one weight of the neural network in accordance with the loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to training neural networks using a loss function that includes frequency-based feature consistency constraints. Background Technology

[0002] Deep neural networks (DNNs) can be used to perform many image understanding tasks, including classification, segmentation, and captioning. Typically, DNNs require a large number of training images (tens of thousands to millions). Additionally, these training images usually need to be annotated, such as labeled, for training and prediction purposes. Summary of the Invention

[0003] A system includes a computer, which includes a processor and a memory. The memory includes instructions that program the processor to: receive frequency-filtered spatial domain data at a neural network; compare the output generated by the neural network with a loss function including frequency-based feature consistency constraints; and update at least one weight of the neural network according to the loss function.

[0004] Among other features, the processor is also programmed to use the Fourier transform process to transform data from the spatial domain to the frequency domain.

[0005] Among other features, the processor is also programmed to filter features from the transformed data based on a predetermined frequency.

[0006] Among other features, the processor is also programmed to transform the filtered transformed data from the frequency domain to the spatial domain to generate frequency-filtered spatial domain data.

[0007] Among other features, the processor is programmed to filter features based on at least one of a high-pass frequency or a low-pass frequency.

[0008] Among other features, the Fourier transform process includes the Fast Fourier Transform process.

[0009] Among other features, the output generated by the neural network includes a latent representation of the spatial domain data after frequency filtering.

[0010] Among other features, neural networks include convolutional neural networks.

[0011] Among other features, the frequency-filtered spatial domain data corresponds to the image captured within the field of view of the vehicle's camera.

[0012] Among other features, the image includes red-green-blue images.

[0013] One method includes: receiving frequency-filtered spatial domain data at a first neural network; comparing the output generated by the neural network with a loss function including frequency-based feature consistency constraints; and updating at least one weight of the neural network according to the loss function.

[0014] Among other features, the method includes using a Fourier transform process to transform the data from the spatial domain to the frequency domain.

[0015] Among other features, the method includes filtering features from the transformed data based on a predetermined frequency.

[0016] Among other features, the method includes transforming the filtered transformed data from the frequency domain to the spatial domain to generate frequency-filtered spatial domain data.

[0017] Among other features, the method includes filtering features based on at least one of a high-pass frequency or a low-pass frequency.

[0018] Among other features, the Fourier transform process includes the Fast Fourier Transform process.

[0019] Among other features, the output generated by the neural network includes a potential representation of the spatial domain data filtered by the frequency.

[0020] Among other features, neural networks include convolutional neural networks.

[0021] Among other features, the frequency-filtered spatial domain data corresponds to the image captured within the field of view of the vehicle's camera.

[0022] Among other features, the image includes red-green-blue images.

[0023] Further areas of application will become apparent from the description provided herein. It should be understood that the descriptions and specific examples are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description

[0024] The accompanying drawings described herein are for illustrative purposes only and are not intended to limit the scope of this disclosure in any way.

[0025] Figure 1 This is a block diagram of an example system including a vehicle;

[0026] Figure 2 This is a block diagram of the sample server within the system;

[0027] Figure 3 This is a diagram illustrating an example neural network;

[0028] Figure 4 This is a block diagram of an example frequency feature extraction system;

[0029] Figure 5 This is a block diagram of an example convolutional neural network;

[0030] Figures 6A to 6C This is a block diagram illustrating an example process for training a neural network;

[0031] Figure 7 This is a block diagram of an example domain adaptive network;

[0032] Figure 8 This is a flowchart illustrating an example process for controlling a vehicle; and

[0033] Figure 9 This is a flowchart illustrating an example process for training a neural network. Detailed Implementation

[0034] The following description is exemplary in nature and is not intended to limit this disclosure, application, or use.

[0035] Typically, standard deep neural networks (DNNs) are pre-trained using labeled training datasets. These DNNs can be validated during testing by comparing the model's output to real-world ground truth. However, obtaining realistic data in real-world testing scenarios can be difficult.

[0036] Domain adaptation refers to generalizing a model from the source domain to the target domain. Typically, the source domain has a large amount of training data, while the target domain may have limited data. For example, the availability of backup camera lane data may be limited due to limitations of camera vendors, rewiring issues, lack of relevant applications, etc. However, there may be many datasets that include forward-looking camera images containing lane information.

[0037] This disclosure discloses systems and methods for training neural networks using a loss function that includes frequency-based feature consistency constraints. The trained neural network can receive data from a source domain and generate data from a target domain. For example, the trained neural network can receive an image including one or more features captured during the day and generate an image including those features such that the image appears to have been captured at night.

[0038] Figure 1 This is a block diagram of an example vehicle system 100. System 100 includes a vehicle 105, which is a land vehicle such as a car, truck, etc. Vehicle 105 includes a computer 110, vehicle sensors 115, actuators 120 for actuating various vehicle components 125, and a vehicle communication module 130. The communication module 130 allows the computer 110 to communicate with a server 145 via a network 135.

[0039] Computer 110 includes a processor and memory. The memory includes one or more forms of computer-readable medium and stores instructions executable by computer 110 for performing various operations, including those disclosed herein.

[0040] Computer 110 can operate vehicle 105 in autonomous, semi-autonomous, or non-autonomous (manual) mode. For the purposes of this disclosure, autonomous mode is defined as a mode in which each of the propulsion, braking, and steering of vehicle 105 can be controlled by computer 110; in semi-autonomous mode, computer 110 controls one or both of the propulsion, braking, and steering of vehicle 105; in non-autonomous mode, a human operator controls each of the propulsion, braking, and steering of vehicle 105.

[0041] Computer 110 may include being programmed to operate one or more of the following functions of vehicle 105: braking, propulsion (e.g., controlling vehicle-related acceleration by controlling one or more of an internal combustion engine, electric motor, hybrid engine, etc.), steering, climate control, interior and / or exterior lights, etc., and to determine whether and when these functions are controlled by computer 110 (rather than a human operator). Additionally, computer 110 may be programmed to determine whether and when these functions are controlled by a human operator.

[0042] Computer 110 may include more than one processor, or, for example, be communicatively coupled to more than one processor (e.g., included in an electronic control unit (ECU) included in vehicle 105) via communication module 130 of vehicle 105 (as further described below) for monitoring and / or controlling various vehicle components 125, such as powertrain controllers, brake controllers, steering controllers, etc. Furthermore, computer 110 may communicate with a navigation system using a Global Positioning System (GPS) via communication module 130 of vehicle 105. As an example, computer 110 may request and receive location data of vehicle 105. The location data may be in a known form, such as geographic coordinates (latitude and longitude coordinates).

[0043] Computer 110 is typically arranged to communicate on communication module 130 of vehicle 105, and also to internal wired and / or wireless networks of vehicle 105 (e.g., buses in vehicle 105 such as controller area network (CAN) and other wired and / or wireless mechanisms).

[0044] Via the communication network of vehicle 105, computer 110 can transmit messages to and / or receive messages from various devices within vehicle 105 (e.g., vehicle sensors 115, actuators 120, vehicle components 125, human-machine interfaces (HMIs), etc.). Alternatively or additionally, where computer 110 actually comprises multiple devices, the communication network of vehicle 105 can be used for communication between devices represented herein as computer 110. Further, as mentioned below, various controllers and / or vehicle sensors 115 can provide data to computer 110. The communication network of vehicle 105 may include one or more gateway modules that provide interoperability between various networks and devices within vehicle 105 (such as protocol converters, impedance matching devices, rate converters, etc.).

[0045] Vehicle sensors 115 may include various devices, such as those known to provide data to computer 110. For example, vehicle sensors 115 may include one or more light detection and ranging (LiDAR) sensors 115 disposed on the top of vehicle 105, behind the windshield of vehicle 105, around vehicle 105, etc., providing information on the relative position, size, and shape of objects and / or the conditions around vehicle 105. As another example, one or more radar sensors 115 fixed to the bumper of vehicle 105 may provide data to provide and measure the velocity, etc., of an object (potentially including a second vehicle 106) relative to the position of vehicle 105. Vehicle sensors 115 may also include one or more camera sensors 115, such as front-view camera sensors, side-view camera sensors, rear-view camera sensors, etc., to provide images of the field of view from inside and / or outside vehicle 105.

[0046] The actuator 120 of vehicle 105 is implemented via circuits, chips, motors, or other electronic and / or mechanical components that can actuate various vehicle subsystems according to known appropriate control signals. The actuator 120 can be used to control components 125, including braking, acceleration, and steering of vehicle 105.

[0047] In the context of this disclosure, vehicle component 125 is one or more hardware components adapted to perform mechanical or electromechanical functions or operations, such as moving vehicle 105, slowing or stopping vehicle 105, steering vehicle 105, etc. Non-limiting examples of component 125 include propulsion components (which include, for example, internal combustion engines and / or electric motors), transmission components, steering components (e.g., which may include one or more of a steering wheel, bogie, etc.), braking components (described below), parking assist components, adaptive cruise control components, adaptive steering components, movable seats, etc.

[0048] Furthermore, computer 110 can be configured to communicate with devices outside vehicle 105 via vehicle-to-vehicle communication module or interface 130, for example, via vehicle-to-vehicle (V2V) or vehicle-to-infrastructure (V2X) wireless communication to another vehicle, or to a remote server 145 (typically via network 135). Module 130 may include one or more mechanisms through which computer 110 can communicate, including any desired combination of wireless (e.g., cellular, wireless, satellite, microwave, and radio frequency) communication mechanisms and any desired combination of network topologies (or topologies when using multiple communication mechanisms). Exemplary communications provided via module 130 include cellular, Bluetooth, IEEE 802.11, Private Short Range Communication (DSRC), and / or Wide Area Network (WAN), including the Internet providing data communication services.

[0049] Network 135 can be one or more of various wired or wireless communication mechanisms, including any desired wired (e.g., cable and fiber optic) and / or wireless (e.g., cellular, wireless, satellite, microwave, and radio frequency) communication mechanisms and any desired network topology (or topology when using multiple communication mechanisms) and any desired combination of these. Exemplary communication networks include wireless communication networks that provide data communication services (e.g., using Bluetooth, Bluetooth Low Energy (BLE), IEEE 802.11, vehicle-to-vehicle (V2V) communication such as Dedicated Short Range Communication (DSRC), etc.), local area networks (LANs), and / or wide area networks (WANs) (including the Internet).

[0050] Computer 110 can receive and analyze data from sensor 115 substantially continuously, periodically, and / or under the instruction of server 145. Furthermore, object classification or identification techniques can be used in computer 110, for example, based on data from lidar sensor 115, camera sensor 115, etc., to identify the type of object (such as vehicle, person, rock, crater, bicycle, motorcycle, etc.) and the physical characteristics of the object.

[0051] Figure 2 This is a block diagram of example server 145. Server 145 includes computer 235 and communication module 240. Computer 235 includes a processor and memory. Memory includes one or more forms of computer-readable medium and stores instructions executable by computer 235 for performing various operations, including those disclosed herein. Communication module 240 allows computer 235 to communicate with other devices, such as vehicle 105.

[0052] Figure 3This is a diagram illustrating an example deep neural network (DNN) 300 that can be used in this paper. The DNN 300 includes multiple nodes 305, and the nodes 305 are arranged such that the DNN 300 includes an input layer, one or more hidden layers, and an output layer. Each layer of the DNN 300 may include multiple nodes 305. Although Figure 3 Three (3) hidden layers are shown, but it should be understood that the DNN 300 may include additional or fewer hidden layers. The input and output layers may also include more than one (1) node 305.

[0053] Nodes 305 are sometimes referred to as artificial neurons 305 because they are designed to mimic biological neurons, such as human neurons. Each neuron 305 receives a set of inputs (indicated by arrows) multiplied by corresponding weights. The weighted inputs are then summed in an input function to provide (possibly adjusted for bias) a net input. The network inputs can then be fed to an activation function, which in turn provides the output to the connected neurons 305. The activation function can be a variety of suitable functions, typically chosen based on empirical analysis. For example... Figure 3 As shown by the arrow, the output of neuron 305 can then be provided as a set of inputs to be included in one or more neurons 305 in the next layer.

[0054] The DNN 300 can be trained to take data as input and generate output based on that input. In one example, the DNN 300 can be trained using real-world data (i.e., data about real-world conditions or states). For example, the DNN 300 can be trained by a processor using real-world data or updated using additional data. For example, the weights can be initialized using a Gaussian distribution, and the bias of each node 305 can be set to zero. Training the DNN 300 can include updating the weights and biases using appropriate techniques, such as optimized backpropagation. Ground-based real-world data can include, but is not limited to, data specifying objects within an image or data specifying physical parameters, such as angle, velocity, distance, color, hue, or the angle of an object relative to another object. For example, ground-based real-world data can be data representing objects and object labels.

[0055] Machine learning services (such as those based on recurrent neural networks (RNNs), convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, or gated recurrent units (GRUs)) can be implemented using the DNN 300 described in this disclosure. In one example, service-related content or other information (such as words, sentences, images, videos, or other such content / information) can be converted into a vector representation.

[0056] Figure 4This is a diagram of an example frequency feature extraction system 400. The frequency feature extraction system 400 can be a software program that can be loaded into memory and executed by a processor, such as in vehicle 105 and / or server 145. As shown, the frequency feature extraction system 400 may include a transform module 405 and an inverse transform module 410. In various embodiments, the transform module 405 and the inverse transform module 410 perform an appropriate Fourier transform process, such as a Fast Fourier Transform, on the received data. For example, the transform module 405 transforms the input data 415 from the spatial domain to the frequency domain. The input data 415 may include image data, audio data, etc.

[0057] Transformation module 405 provides transformed data, i.e., data existing in the frequency domain, to high-pass filter 420 and low-pass filter 425. High-pass filter 420 allows data with frequencies above a predetermined high-pass cutoff frequency to pass through, and low-pass filter 425 allows data with frequencies below a predetermined low-pass cutoff frequency to pass through. High-pass filter 420 provides the filtered data to inverse transform module 410, which converts the filtered data from the frequency domain to the spatial domain, and low-pass filter 425 provides the filtered data to inverse transform module 410, which converts the filtered data from the frequency domain to the spatial domain. For example, inverse transform module 410 can generate a modified image based on the corresponding filtering from high-pass filter 420 or low-pass filter 425.

[0058] As described in more detail below, filtered spatial domain data (such as altered images) can be fed to the DNN 300 to apply feature constraints to the input data based on frequency features. For example, the DNN 300 can be trained using a loss function that includes frequency-based feature consistency constraints to allow it to learn domain-independent features. Thus, the labeled features can be preserved within the image in the source domain.

[0059] Figure 5 This is a block diagram illustrating an example DNN 300. Figure 5 In the embodiment shown, DNN 300 is a convolutional neural network 500. Based on connectivity and weight sharing, the convolutional neural network 500 may include multiple layers of different types. As shown, the convolutional neural network 500 includes convolutional blocks 505A and 505B. Each of the convolutional blocks 505A and 505B may be configured with a convolutional layer (CONV) 510, a normalization layer (LNorm) 515, and a max pooling layer (MAX POOL) 520.

[0060] Convolutional layer 510 may include one or more convolutional filters, which are applied to input data 545 to generate output 540. Although Figure 5Only two convolutional blocks 505A and 505B are shown, but this disclosure may include any number of convolutional blocks 505A and 505B. A normalization layer 515 can normalize the output of the convolutional filter. For example, a normalization layer 515 can provide whitening or lateral suppression. A max-pooling layer 520 can provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.

[0061] The deep convolutional network 500 may also include one or more fully connected layers 525 (FC1 and FC2). The deep convolutional network 500 may also include a logistic regression (LR) layer 530. Weights that can be updated are present between each of the layers 510, 515, 520, 525, and 530 of the deep convolutional network 500. The output of each of the layers (e.g., 510, 515, 520, 525, and 530) can be used as input to subsequent layers in the convolutional neural network 500 (e.g., 510, 515, 520, 525, and 530) to learn features from the input data 540 (e.g., image, audio, video, sensor data, and / or other input data provided at the first location in convolutional block 505A). The output 535 may represent a latent representation based on one or more features of the input data. For example, the output 535 may include latent features derived from a first domain of the input image (such as a red-green-blue (RGB) image captured during the day). Output 535 can be converted by the decoder into a synthetic image in the second domain, such as a synthetic RGB image showing the characteristics of the nighttime input image.

[0062] Figure 6A and Figure 6B Example procedures for training a DNN 300 according to one or more embodiments of this disclosure are shown. Figure 6A As shown, during the initial training phase, the DNN 300 receives a set of labeled training data 605 and training labels 610. (Based on the above reference...) Figure 4 The described process describes a training data 605 that may include frequency-filtered transform spatial domain data. For example, the filtered spatial domain data may include one or more images depicting objects within the field of view (FOV) of a vehicle camera. Training labels 610 may include object labels, object type labels, domain type, and / or the distance of the object relative to the source of the image.

[0063] Following the initial training phase, in the supervised training phase, a set of N training data points 615 are input into the DNN 300. The DNN 300 generates output transformation data for each of the N training data points 615 input. For example, the DNN 300 can generate a synthetic image that includes features from the training data in the second domain. Figure 6BAn example of generating an output based on training data 615 (e.g., unlabeled training images) from N training data 615 is shown. Based on the initial training, the DNN 300 outputs a vector representation 620 of the output data, such as a latent representation of the training data. The vector representation 620 is compared with ground truth data 625. Ground truth data 625 may include frequency-based feature consistency constraints. For example, frequency-based feature consistency constraints may include a portion of a loss function such that features corresponding to low frequencies in the data are consistent across domains, and features corresponding to high frequencies in the data are mitigated or reduced across domains.

[0064] The DNN 300 updates its network parameters based on comparisons with real ground data 625. For example, network parameters, such as the weights associated with neurons, can be updated via backpropagation. The DNN 300 can be trained at server 145 and provided to vehicle 105 via communication network 135. Vehicle 105 can also provide server 145 with data captured by the vehicle 105 system for further training.

[0065] Figure 7 This is an illustration of an example domain adaptive network 700 that can transform data from a source domain into data from a target domain. For example, the domain adaptive network 700 can be a software program that can be loaded into memory and executed by a processor, such as in vehicle 105 and / or server 145. In the example implementation, the domain adaptive network 700 can receive image sequences from a source domain (e.g., daytime) and output image sequences from a target domain (e.g., nighttime).

[0066] As shown in the figure, the domain adaptive network 700 includes an autoencoder, which comprises an encoder 705 and a decoder 710. In an example embodiment, the encoder 705 may include the components described above. Figure 6A and Figure 6B The trained DNN 300 is described above. In various embodiments, the decoder 710 is symmetrical to the encoder 705. The encoder 705 receives input data from the source domain and generates a latent representation 715 of the input data, and the decoder 710 reconstructs the output data in the source domain based on the latent representation 715 of the input data in the target domain.

[0067] Figure 8This is a flowchart of an example process 800 for controlling vehicle 105 based on determined outputs of a neural network trained according to the process described herein. The various blocks of process 800 can be executed by computer 110. Process 800 begins at block 805, where computer 110 determines whether to actuate vehicle 105 based on the determined outputs. Computer 110 may include a lookup table that establishes a relationship between the determined outputs and vehicle actuation actions. For example, based on images captured by one or more sensors 115 of vehicle 105, computer 110 can cause vehicle 105 to perform a specified action, such as initiating a turn, adjusting the direction of vehicle 105, adjusting the speed of vehicle 105, etc. In another example, based on a determined distance between vehicle 105 and an object, computer 110 can cause vehicle 105 to perform a specified action, such as initiating a turn, activating an external alarm, adjusting the speed of vehicle 105, etc.

[0068] If the computer determines that no actuation has occurred, process 800 returns to block 805. Otherwise, at block 810, computer 110 actuates vehicle 105 according to a specified action. For example, computer 110 transmits appropriate control signals to the corresponding actuator 120.

[0069] Figure 9 This is a flowchart of an example process 900 for training a DNN 300. The various blocks of process 900 can be executed by computer 235. Process 900 begins at block 905, where computer 235 trains the DNN 300. For example, as described in more detail above, the DNN 300 can be trained using filtered transform spatial domain data.

[0070] At box 910, computer 235 transmits the trained DNN 300 to vehicle 105. At box 815, computer 235 determines whether additional data has been received. For example, the data could be sensor data already uploaded by computer 110. If no additional sensor data has been uploaded, process 900 returns to box 915. If additional sensor data has been uploaded, process 900 returns to box 905, allowing DNN 300 to be trained based on the uploaded sensor data using filtered transform spatial domain data.

[0071] The description in this disclosure is merely exemplary in nature, and variations thereof without departing from the spirit and scope of this disclosure are intended to be within its scope. Such variations should not be considered as departing from the spirit and scope of this disclosure.

[0072] Generally, the described computing system and / or device may employ any of a variety of computer operating systems, including, but not limited to, Microsoft Automotive Operating System, Microsoft Windows Operating System, Unix operating systems (e.g., Solaris distributed by Oracle Corporation of Redwood Shores, California), AIX UNIX operating system distributed by International Business Machines Corporation of Armonk, New York, Linux operating system, Mac OSX and iOS operating systems distributed by Apple Inc. of Cupertino, California, BlackBerry OS distributed by Blackberry Corporation of Waterloo, Canada, and Android operating system developed by Google and the Open Handset Alliance, or infotainment systems provided by QNX Software Systems. Versions and variants of the CAR platform. Examples of computing devices include, but are not limited to, in-vehicle computers, computer workstations, servers, desktop computers, laptops, handheld computers, or other computing systems and / or devices.

[0073] Computers and computing devices typically include computer-executable instructions, which can be executed by one or more computing devices, such as those listed above. Computer-executable instructions can be compiled or interpreted in computer programs created using various programming languages ​​and / or technologies, including, but not limited to (alone or in combination) Java, C, C++, Matlab, Simulink, Stateflow, Visual Basic, JavaScript, Perl, HTML, etc. Some of these applications can be compiled and executed on virtual machines, such as the Java Virtual Machine, Dalvik Virtual Machine, etc. Generally, a processor (e.g., a microprocessor) receives instructions from, for example, memory, computer-readable media, etc., and executes those instructions to perform one or more processes, including one or more of the processes described herein. These instructions and other data can be stored and transferred using various computer-readable media. A file in a computing device generally refers to a collection of data stored on computer-readable media, such as storage media, random access memory, etc.

[0074] Memory can include computer-readable media (also known as processor-readable media), including any non-transitory (e.g., tangible) medium that contributes to providing data (e.g., instructions) that can be read by a computer (e.g., by the computer's processor). Such media can take many forms, including but not limited to non-volatile and volatile media. Non-volatile media can include, for example, optical discs or magnetic disks, and other permanent storage. Volatile media can include, for example, dynamic random access memory (DRAM), which typically constitutes main memory. These instructions can be transmitted via one or more transmission media, including coaxial cables, copper wires, and optical fibers, including wires on the system bus coupled to the processor of the ECU. Common forms of computer-readable media include, for example, floppy disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs, any other optical media, punched cards, paper tape, any other physical media with a perforated pattern, RAM, PROMs, EPROMs, flash EEPROMs, any other memory chips or cassette tapes, or any other computer-readable medium.

[0075] The databases, data repositories, or other data storage devices described herein can include various mechanisms for storing, accessing, and retrieving various types of data, including hierarchical databases, file sets in file systems, application databases in proprietary formats, relational database management systems (RDBMS), etc. Each such data storage device is typically contained within a computing device employing a computer operating system (as mentioned above) and is accessed via a network in any one or more of various ways. File systems can be accessed from the computer operating system and can include files stored in various formats. In addition to languages ​​used for creating, storing, editing, and executing stored procedures, RDBMS typically use a structured query language (SQL), such as the PL / SQL language mentioned above.

[0076] In some examples, system elements may be implemented as computer-readable instructions (e.g., software) on one or more computing devices (e.g., servers, personal computers, etc.) and stored on an associated computer-readable medium (e.g., disks, storage, etc.). A computer program product may include such instructions stored on a computer-readable medium for performing the functions described herein.

[0077] In this application, including the following definitions, the term "module" or "controller" may be replaced by the term "circuit". The term "module" may refer to or include, or be part of, the following: application-specific integrated circuit (ASIC); digital, analog, or mixed-signal analog / digital discrete circuit; digital, analog, or mixed-signal analog / digital integrated circuit; combinational logic circuit; field-programmable gate array (FPGA); processor circuitry that executes code (shared processor circuitry, dedicated processor circuitry, or group of processor circuitry); memory circuitry that stores code executed by the processor circuitry (shared memory circuitry, dedicated memory circuitry, or group of memory circuitry); other suitable hardware components that provide the aforementioned functionality; or some or all of the above, such as in a system-on-a-chip.

[0078] This module may include one or more interface circuits. In some examples, the interface circuits may include wired or wireless interfaces that connect to a local area network (LAN), the Internet, a wide area network (WAN), or a combination thereof. The functionality of any given module in this disclosure can be distributed among multiple modules connected via the interface circuits. For example, multiple modules can allow for load balancing. In another example, a server (also referred to as a remote or cloud) module may perform some functions on behalf of a client module.

[0079] Regarding the media, processes, systems, methods, and inspirations described herein, it should be understood that although the steps of these processes are described as occurring in a certain order, these processes can be practiced by performing the described steps in an order other than that described herein. It should also be understood that some steps may be performed simultaneously, other steps may be added, or some steps described herein may be omitted. In other words, the process descriptions herein are provided for the purpose of illustrating certain embodiments and should not be construed as limiting the claims in any way.

[0080] Therefore, it should be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications beyond the examples provided will be apparent to those skilled in the art upon reading the above description. The scope of the invention should not be determined by reference to the foregoing description, but rather by reference to the appended claims and the full scope of their equivalents. Future developments are anticipated and intended to occur in the field discussed herein, and the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that modifications and variations are possible with respect to the invention, and it is defined solely by the appended claims.

[0081] Unless otherwise expressly indicated herein, all terms used in the claims are intended to give the simple and common meaning as understood by one of ordinary skill in the art. In particular, unless the claims set forth an explicit limitation to the contrary, the use of singular articles such as “a,” “the,” “the,” etc., should be understood to refer to one or more of the elements indicated in the statement.

Claims

1. A system including a computer, the computer including a processor and a memory, the memory including instructions that program the processor to: The image sequence in the source domain is acquired by the vehicle camera, and the image sequence is transformed from the spatial domain to the frequency domain to filter the image sequence in the frequency domain to obtain filtered frequency domain data. The filtered frequency domain data is converted into spatial domain data to obtain the frequency-filtered spatial domain data. The frequency-filtered spatial domain data is input into a trained neural network to obtain the latent representation of the frequency-filtered spatial domain data. Reconstruct the data in the source domain based on the latent representation of the spatial domain data in the target domain, and output the image sequence in the target domain; Wherein, the source domain and the target domain are respectively either daytime or nighttime; the neural network is trained using a loss function that includes frequency-based feature consistency constraints; the training labels include domain types; and the neural network learns domain-independent features so that the labeled features are maintained in the image within the source domain; and The learning of domain-independent features through the neural network includes: maintaining the consistency of low-frequency features in the spatial domain data across all domains, and mitigating or reducing high-frequency features in the spatial domain data across all domains.

2. The system according to claim 1, wherein, The processor is also programmed to use a Fourier transform process to transform data from the spatial domain to the frequency domain.

3. The system according to claim 2, wherein, The processor is also programmed to filter features from the transformed data based on a predetermined frequency.

4. The system according to claim 3, wherein, The processor is also programmed to transform the filtered transformed data from the frequency domain to the spatial domain to generate frequency-filtered spatial domain data.

5. The system according to claim 2, wherein, The processor is also programmed to filter the features based on at least one of a high-pass frequency or a low-pass frequency.

6. The system according to claim 2, wherein, The Fourier transform process includes the Fast Fourier Transform process.

7. The system according to claim 1, wherein, The neural network includes a convolutional neural network.

8. The system according to claim 1, wherein, The frequency-filtered spatial domain data corresponds to the image captured within the field of view of the vehicle camera.

9. The system according to claim 8, wherein, The image includes red, green, and blue images.

Citation Information

Patent Citations

  • Method for reconstructing sparse MRI (Magnetic resonance imaging) based on convolutional neural network in combination with iterative method

    CN108717717A