Method for training a neural network for detecting an object and method for detecting an object by means of a neural network
Patent Information
- Application Number
- EP2023751280
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-15
- Filing Date
- 2023-07-27
- Publication Date
- 2025-07-23
AI Technical Summary
Current methods for object detection using neural networks in image processing, especially with radar data, face limitations due to dependency on preset parameters and inadequate results under certain conditions, whereas traditional algorithms like CFAR provide suboptimal performance.
A method for training a neural network that combines geometric dimensions from camera images and radar signals to generate mixed spectra, which are then used to create training data, allowing the network to learn object localization and classification independently of image conditions, utilizing convolutional neural networks (CNNs) in an autoencoder structure.
This approach enhances the neural network's ability to accurately detect objects by providing robust classification and localization capabilities, improving the independence from image conditions and parameter settings, and enabling effective object detection in various scenarios.
Smart Images

Figure 1.1
Abstract
Description
[0001] Method for training a neural network for detecting an object and method for detecting an object using a neural network
[0002] Description:
[0003] The invention relates to a method for training a neural network for object detection. Training data is generated and fed to the neural network. The invention also relates to a method for detecting an object using a neural network.
[0004] In image processing for camera systems, the use of neural networks for object detection in images is already common practice and delivers excellent results. The state of the art in the field of neural networks used to locate objects in images is primarily convolutional neural networks (CNNs) with an autoencoder structure. The prerequisite for successful object localization and classification using a neural network is that the neural network has been appropriately trained.
[0005] A convolutional neural network comprises various convolutional layers, which collectively represent the intelligence of the neural network. These include an input layer and an output layer. The layers are linked to each other using mathematical convolution operations. In image processing, an image is fed to the input layer, and a map containing the object positions of the objects in the image is output at the output layer.
[0006] Algorithms such as CFAR (Constant False Alarm Rate) are typically used to detect objects using radar data. These algorithms often produce unsatisfactory results under many conditions. Algorithms such as CFAR follow strict patterns to generate their output and are dependent on preset parameters. If these parameters are poorly selected, the results can be significantly worse than originally expected. The current state of image processing with neural networks shows that they can locate and classify objects more independently of the state of the available images. A method for training a neural network is known from EP 3 690 727 A1.
[0007] A camera and a radar are used together.
[0008] From DE 10 2018 203 684 A1 an evaluation device, a training system and a training method for obtaining a segmentation of a radar image of an environment are known.
[0009] From DE 11 2021 000 135 T5 a system and a method are known which relate to sensor fusion based on machine learning for applications of autonomous machines.
[0010] From DE 10 2019 219 894 A1 a device and a method for generating verified training data for a self-learning system are known.
[0011] EP 3 832 341 A1 describes a neural network for obstacle detection for use in autonomous vehicles. It uses radar sensors.
[0012] The invention is based on the object of developing a method for training a neural network for detecting an object, as well as a method for detecting an object by means of a neural network.
[0013] The problem is solved by a method for training a neural network for detecting an object with the features specified in claim 1. Advantageous embodiments and further developments are the subject of the dependent claims. The problem is also solved by a method for detecting an object using a neural network with the features specified in claim 12. Advantageous embodiments and further developments are the subject of the dependent claims.
[0014] A method for training a neural network for object detection is proposed. Geometric dimensions of a test object from an object class are recorded. Images of the test object are generated by a plurality of cameras over a period of time. Occupancy maps are generated from the recorded geometric dimensions and the generated images. A radar device transmits a radar signal and receives a radar signal reflected from the test object. The transmitted radar signal and the received radar signal are mixed into a complex baseband to form a mixed signal, and a complex four-dimensional mixed spectrum of the mixed signal is calculated. A first complex two-dimensional subspectrum and a second complex two-dimensional subspectrum are calculated from the complex four-dimensional mixed spectrum.The occupancy maps and the partial spectra are fused into training data, and the training data is fed to the neural network.
[0015] The training data contains a sufficient number of mixed spectra containing radar images, linked to the information about the location of a test object and the object class the test object belongs to. This information is referred to as ground truth. When generating the occupancy maps, cells containing the test object are assigned a high value, and cells that are free are assigned a low value. The probability that a test object of a specific object class occupies a location in the radar's field of view can be determined from the respective occupancy maps. The most probable test object can be extracted using a hard decision threshold.
[0016] According to an advantageous embodiment of the invention, the mixed spectrum contains information about the distance, azimuth angle, elevation angle, and radial velocity of the test object. The radial velocity is determined via a frequency shift between the transmitted radar signal and the reflected radar signal. This frequency shift results from the Doppler effect for moving objects.
[0017] According to an advantageous embodiment of the invention, the first partial spectrum contains information about a distance and an azimuth angle of the test object, and the second partial spectrum contains information about a distance and a radial velocity of the test object.
[0018] According to an advantageous embodiment of the invention, the first partial spectrum contains a first radar image with information about the absolute distance and the azimuth angle of the test object. The first partial spectrum also contains a second radar image with information about the phase of the distance and the azimuth angle of the test object. The second partial spectrum contains a third radar image with information about the absolute distance and the radial velocity of the test object. The second partial spectrum also contains a fourth radar image with information about the phase of the distance and the radial velocity of the test object. The radar images are preferably available in polar coordinates.
[0019] According to an advantageous embodiment of the invention, the test object is moved during the time period. Thus, several different images of the test object are generated at different locations and with different orientations, and corresponding mixed spectra are calculated. A higher number of different images with corresponding mixed spectra improves the quality of the training data.
[0020] According to a preferred embodiment of the invention, the occupancy maps are first generated in Cartesian coordinates, and the Cartesian coordinates are subsequently transformed into polar coordinates. The transformation of the Cartesian coordinates into polar coordinates is performed before the occupancy maps and the partial spectra are merged into the training data. This ensures compatibility between the occupancy maps and the partial spectra, whose radar images are also available in polar coordinates.
[0021] According to an advantageous development of the invention, prior to the creation of the images, markers are applied to the test object in such a way that the markers are visible in the generated images. When such markers are applied to a test object, a six-dimensional pose of the test object can be calculated, which includes a position of the test object and an orientation of the test object.
[0022] According to an advantageous embodiment of the invention, the cameras are designed as infrared cameras, and / or the markings are designed as infrared markers. The cameras and the markings are part of a position detection system. In particular, such markings on the test object allow for the precise calculation of a six-dimensional pose of the test object, which includes a position of the test object and an orientation of the test object.
[0023] According to a preferred embodiment of the invention, a pose of the test object is calculated from each of the images, each of which includes a position of the test object and an orientation of the test object. The calculated poses are discretized and integrated into the occupancy maps. According to a preferred development of the invention, the method steps are repeated for at least one additional test object from another object class. The occupancy maps and / or the training data are assigned to the respective object class. Such object classes include, for example, people, forklifts, or autonomous transport vehicles.
[0024] According to a preferred embodiment of the invention, the neural network is designed as a convolutional network, which has an input layer, an output layer, and a plurality of convolutional layers. In particular, the neural network is preferably designed as a CNN (Convolutional Neural Network) with an autoencoder structure. The layers are arranged serially one after the other and linked to one another via mathematical convolution operations. A convolution operation is performed from one layer to the next. The convolution operators used to perform said convolution operations are determined by processing the training data supplied to the neural network.
[0025] A method for detecting an object using a neural network is also proposed, wherein training data was previously fed to the neural network. The training data was fed to the neural network using the inventive method for training a neural network. A radar sensor emits a radar signal, and a radar signal reflected by the object is received. The emitted radar signal and the received radar signal are mixed to form a mixed signal, and a mixed spectrum of the mixed signal is calculated. Input data containing the mixed spectrum is fed to the neural network. The input data is processed in the neural network. The neural network detects the object and its position. The neural network outputs an object class of the detected object and the detected position of the object as output data.
[0026] The more independent dimensions the neural network has available as input data, the more effective the classification of the object using unique features. A unique signature in the mixed spectrum enables robust classification.
[0027] According to a preferred embodiment of the invention, the neural network is designed as a convolutional network comprising an input layer, an output layer, and a plurality of convolutional layers. A convolution operation is performed from one layer to the next. The layers are arranged serially and linked to one another via mathematical convolution operations. A convolution operation is performed from one layer to the next.
[0028] According to a preferred embodiment of the invention, the calculated mixed spectrum comprises at least a distance and an azimuth angle of a radar measurement. First input data containing the distance of the radar measurement is fed to the neural network. Second input data containing the azimuth angle of the radar measurement is fed to the neural network. The first input data and the second input data represent a first complex image composed of complex data. The first complex image thus comprises two simple images containing magnitude and phase.
[0029] According to an advantageous development of the invention, the calculated mixed spectrum comprises at least a distance and a radial velocity of a radar measurement. Third input data containing the distance of the radar measurement is fed to the neural network. Fourth input data containing the radial velocity of the radar measurement is fed to the neural network. The third input data and the fourth input data represent a second complex image composed of complex data. The second complex image thus comprises two simple images containing magnitude and phase.
[0030] The invention is not limited to the combination of features in the claims. Further possible combinations of claims and / or individual claim features and / or features of the description and / or the figures will become apparent to those skilled in the art, particularly from the problem and / or the problem posed by comparison with the prior art.
[0031] The invention will now be explained in more detail with reference to the accompanying drawings. The invention is not limited to the exemplary embodiments shown in the drawings. The drawings only represent the subject matter of the invention schematically. They show:
[0032] Figure 1 : a schematic representation of an arrangement for obtaining training data,
[0033] Figure 2: a schematic representation of a neural network,
[0034] Figure 3: a schematic representation of input data of a neural network and
[0035] Figure 4: a schematic representation of output data of a neural network.
[0036] Figure 1 shows a schematic representation of an arrangement for obtaining training data for a neural network 7. The arrangement has a measuring range 40 and a radar range 42. The measuring range 40 and the radar range 42 largely overlap. A test object (not shown here) is located within the measuring range 40 and within the radar range 42.
[0037] The arrangement includes a radar device 25. The radar device 25 is arranged such that a test object located within the radar range 42 can be detected by the radar device 25. The radar range 42 is designed in the form of a circular sector. The radar device 25 is arranged at the tip of the circular sector.
[0038] The radar device 25 transmits a radar signal and receives a radar signal that is reflected by the test object. The radar device 25 has a multiplier. The transmitted radar signal and the received radar signal are mixed by the multiplier into a complex baseband composite signal. A complex four-dimensional composite spectrum of the composite signal is also calculated in the radar device 25. The composite spectrum is calculated using a discrete Fourier transform from sampled raw data of the composite signal.
[0039] The radar device 25 has a 2D MIMO (Multiple Input Multiple Output) antenna array. The transmitted radar signal has FMCW (Frequency-Modulated Continuous Wave) modulation. The calculated mixed spectrum is therefore four-dimensional and contains information about the distance, azimuth angle, elevation angle, and radial velocity of the test object from which the radar signal is reflected.
[0040] A first complex two-dimensional subspectrum and a second complex two-dimensional subspectrum are calculated from the complex four-dimensional mixed spectrum. The first subspectrum contains information about the distance and azimuth angle of the test object, and the second subspectrum contains information about the distance and radial velocity of the test object.
[0041] The first sub-spectrum contains a first radar image with information about the absolute distance and azimuth angle of the test object. The first sub-spectrum also contains a second radar image with information about the phase of the test object's distance and azimuth angle. The second sub-spectrum contains a third radar image with information about the absolute distance and radial velocity of the test object. The second sub-spectrum also contains a fourth radar image with information about the phase of the test object's distance and radial velocity. The radar images are provided in polar coordinates.
[0042] The arrangement comprises a plurality of cameras 21 for generating images. Six cameras 21 are provided here. The cameras 21 are arranged such that a test object located within the measuring area 40 can be captured by all cameras 21. The measuring area 40 is rectangular. The cameras 21 are arranged at the corners and along the sides of the rectangle. The cameras 21 are designed as infrared cameras and are part of a position detection system.
[0043] The arrangement further comprises a digital computer 32 and a processing unit 34. The cameras 21 are connected to the processing unit 34 and transmit generated images to the processing unit 34. The radar device 25 is also connected to the processing unit 34 and transmits data to the processing unit 34. The processing unit 34 is connected to the digital computer 32 and transmits data to the digital computer 32.
[0044] To obtain the training data for neural network 7, a test object is first selected from an object class. Object classes include, for example, people, forklifts, or autonomous transport vehicles. The selected test object could therefore be, for example, a person, a forklift, or an autonomous transport vehicle.
[0045] First, the geometric dimensions of the test object are recorded. In particular, the length, width, and height of the test object are measured. Furthermore, markings are applied to the test object. These markings are designed as infrared markers. The cameras 21, as already mentioned, are designed as infrared cameras. The markings are applied to the test object in such a way that they are visible in images subsequently generated by the cameras 21.
[0046] The acquisition of training data for the neural network 7 using the selected test object takes place over a previously defined period of time. During this period, the test object is moved within a range that lies within the measurement range 40 and within the radar range 42. If necessary, the test object moves independently within this range during this period.
[0047] During this time, the cameras capture 21 images of the test object. A pose of the test object is calculated from each image. This pose is six-dimensional and includes both the position and orientation of the test object.
[0048] Occupancy maps are generated from the previously recorded geometric dimensions of the test object and the resulting images. The calculated poses are integrated into the occupancy maps. The occupancy maps are assigned to the object class of the selected test object. The occupancy maps are initially generated in Cartesian coordinates, and the Cartesian coordinates are then transformed into polar coordinates.
[0049] During said period, radar device 25 simultaneously transmits a radar signal and receives a radar signal reflected by the test object. The transmitted radar signal and the received radar signal are combined to form a composite signal. A complex four-dimensional composite spectrum of the composite signal is also calculated. From the complex four-dimensional composite spectrum, a first complex two-dimensional subspectrum and a second complex two-dimensional subspectrum are calculated. The subspectra contain radar images in polar coordinates.
[0050] The occupancy maps and the partial spectra are then merged into training data. The training data is assigned to the respective object class of the selected test object. The resulting training data is fed to neural network 7.
[0051] The described procedural steps for obtaining training data for neural network 7 are repeated for additional test objects from additional object classes. Test objects from other object classes are selected. Furthermore, the described procedural steps for obtaining training data for neural network 7 are performed once without a real test object, but with a free space. The occupancy maps and the training data are assigned to the respective object class or free space.
[0052] Figure 2 shows a schematic representation of a neural network 7. The neural network 7 is designed as a convolutional network. The neural network 7 has an input layer 6, a first convolutional layer 11, a second convolutional layer 12, a third convolutional layer 13, a fourth convolutional layer 14, a fifth convolutional layer 15, a sixth convolutional layer 16, a seventh convolutional layer 17, and an output layer 9.
[0053] Input data 1, 2, 3, 4 are fed to the input layer e of the neural network 7. The input layer e, the convolutional layers 11, 12, 13, 14, 15, 16, 17, and the output layer 9 are arranged serially one after the other. A convolution operation is performed from one layer to the next. Output data 51, 52, 53, 54 are output from the output layer 9 of the neural network 7.
[0054] Furthermore, an intermediate connection 8 is provided between the first convolutional layer 11 and the seventh convolutional layer 17. An intermediate connection 8 is also provided between the second convolutional layer 12 and the sixth convolutional layer 16. An intermediate connection 8 is also provided between the third convolutional layer 13 and the fifth convolutional layer 15. The intermediate connections 8 represent direct transfers between two layers, with no convolutional operation being performed via the intermediate connection 8. The intermediate connections 8 are used to accelerate the training phase.
[0055] This is a heuristic.
[0056] Each of the layers represents a three-dimensional matrix of individual pixels. In this case, the input layer 6 has a size of 4x128x128 pixels. The first convolutional layer 11 has a size of 16x64x64 pixels. The second convolutional layer 12 has a size of 32x32x32 pixels. The third convolutional layer 13 has a size of 64x16x16 pixels. The fourth convolutional layer 14 has a size of 128x8x8 pixels. The fifth convolutional layer 15 has a size of 64x16x16 pixels. The sixth convolutional layer 16 has a size of 32x32x32 pixels. The seventh convolutional layer 17 has a size of 16x64x64 pixels. The output layer 9 has a size of 4x64x64 pixels.
[0057] To detect an object using neural network 7, which has previously been fed the training data, a radar measurement is performed. A radar sensor emits a radar signal and receives a radar signal reflected by the object. The emitted radar signal and the received radar signal are combined to form a composite signal. A composite spectrum of the composite signal is calculated.
[0058] The calculated composite spectrum includes a radar measurement range and a radar measurement azimuth angle. The calculated composite spectrum also includes a radar measurement range and a radar measurement radial velocity.
[0059] Input data 1, 2, 3, 4 containing the mixed spectrum are fed to the input layer 6 of the neural network 7. Figure 3 shows a schematic representation of input data 1, 2, 3, 4 of the neural network 7.
[0060] The first input data 1, which contains the distance of the radar measurement, is fed to the input layer 6 of the neural network 7. The first input data 1 represents a two-dimensional matrix of individual pixels. In this case, the first input data
[0061] 1 has a size of 128x128 pixels.
[0062] The input layer 6 of the neural network 7 is fed with the second input data 2, which contains the azimuth angle of the radar measurement. The second input data
[0063] 2 represent a two-dimensional matrix of individual pixels. In this case, the second input data 2 has a size of 128x128 pixels. The third input data 3, which contains the distance of the radar measurement, is fed to the input layer 6 of the neural network 7. The third input data 3 represents a two-dimensional matrix of individual pixels. In this case, the third input data 3 has a size of 128x128 pixels.
[0064] The fourth input data 4, which contains the radial velocity of the radar measurement, is fed to the input layer 6 of the neural network 7. The fourth input data 4 represents a two-dimensional matrix of individual pixels. In this case, the fourth input data 4 has a size of 128x128 pixels.
[0065] In neural network 7, the input data 1, 2, 3, and 4 are processed. A convolution operation is performed from one layer to the next. Through the successive convolution operations, neural network 7 detects the object and its position. Through the successive convolution operations, neural network 7 also detects an object class for the object.
[0066] Output data 51, 52, 53, 54 are output from the output layer 9 of the neural network 7. Figure 4 shows a schematic representation of output data 51, 52, 53, 54 of the neural network 7.
[0067] The first output data 51 is assigned to an object from a first object class, for example, a person. The first output data 51 contains the detected position of the object. The first output data 51 represents a two-dimensional matrix of individual pixels. In this case, the first output data 51 has a size of 64x64 pixels.
[0068] The second output data 52 is assigned to an object from a second object class, for example, a forklift. The second output data 52 contains the detected position of the object. The second output data 52 represents a two-dimensional matrix of individual pixels. In this case, the second output data 52 has a size of 64x64 pixels. The third output data 53 is assigned to an object from a third object class, for example, an autonomously driving transport vehicle. The third output data 53 contains the detected position of the object. The third output data 53 represents a two-dimensional matrix of individual pixels. In this case, the third output data 53 has a size of 64x64 pixels.
[0069] The fourth output data 54 is assigned to an object from a fourth object class, for example, free space. The fourth output data 54 contains the detected position of the object. The fourth output data 54 represents a two-dimensional matrix of individual pixels. In this case, the fourth output data 54 has a size of 64x64.
[0070] pixels.
[0071] List of reference symbols
[0072] 1 first input data
[0073] 2 second input data
[0074] 3 third input data
[0075] 4 fourth input data
[0076] 6 Input layer
[0077] 7 neural network
[0078] 8 Intermediate connection
[0079] 9 Output layer
[0080] 11 first convolutional layer
[0081] 12 second convolutional layer
[0082] 13 third convolutional layer
[0083] 14 fourth convolutional layer
[0084] 15 fifth convolutional layer
[0085] 16 sixth convolutional layer
[0086] 17 seventh convolutional layer
[0087] 21 Camera
[0088] 25 radar device
[0089] 32 digital computers
[0090] 34 processing unit
[0091] 40 measuring range
[0092] 42 radar range
[0093] 51 initial data
[0094] 52 second output data
[0095] 53 third initial data
[0096] 54 fourth output data
Claims
Patent claims:
1. A method for training a neural network (7) for detecting an object, wherein geometric dimensions of a test object from an object class are recorded, and images of the test object are generated by a plurality of cameras (21) over a period of time; occupancy maps are generated from the recorded geometric dimensions and the generated images; a radar device (25) transmits a radar signal and receives a radar signal reflected from the test object; the transmitted radar signal and the received radar signal are mixed in a complex baseband to form a mixed signal; a complex four-dimensional mixed spectrum of the mixed signal is calculated; a first complex two-dimensional partial spectrum and a second complex two-dimensional partial spectrum are calculated from the complex four-dimensional mixed spectrum; the occupancy maps and the partial spectra are fused to form training data;and the training data are fed to the neural network (7); 2. Method according to one of the preceding claims, characterized in that the mixed spectrum contains information about a distance, an azimuth angle, an elevation angle and a radial velocity of the test object.
3. Method according to one of the preceding claims, characterized in that the first partial spectrum contains information about a distance and an azimuth angle of the test object, and that the second partial spectrum contains information about a distance and a radial velocity of the test object.
4. The method according to claim 3, characterized in that the first partial spectrum contains a first radar image with information about an absolute value of the range and the azimuth angle of the test object; and that the first partial spectrum contains a second radar image with information about a phase of the range and the azimuth angle of the test object; and that the second partial spectrum contains a third radar image with information about an absolute value of the range and the radial velocity of the test object; and that the second partial spectrum contains a fourth radar image with information about a phase of the range and the radial velocity of the test object.
5. Method according to one of the preceding claims, characterized in that the test object is moved during the period of time.
6. Method according to one of the preceding claims, characterized in that the occupancy maps are generated in Cartesian coordinates and that the Cartesian coordinates are transformed into polar coordinates; 7. Method according to one of the preceding claims, characterized in that before the images are produced, markings are applied to the test object in such a way that the markings are visible in the images produced.
8. Method according to claim 7, characterized in that the cameras (21) are designed as infrared cameras, and / or that the markings are designed as infrared markers.
9. Method according to one of the preceding claims, characterized in that a pose of the test object is calculated from the images, which pose comprises a position of the test object and an orientation of the test object, and in that the calculated poses are integrated into the occupancy maps.
10. Method according to one of the preceding claims, characterized in that the method steps are repeated for at least one further test object from a further object class, and that the occupancy maps and / or the training data are assigned to the respective object class 11. Method according to one of the preceding claims, characterized in that the neural network (7) is designed as a convolutional network which has an input layer (6), an output layer (9) and a plurality of convolutional layers (11, 12, 13, 14, 15, 16, 17).
12. A method for detecting an object by means of a neural network (7), to which training data have previously been supplied using the method according to one of the preceding claims, wherein a radar signal is emitted by a radar sensor and a radar signal reflected by the object is received; the emitted radar signal and the received radar signal are mixed to form a mixed signal; a mixed spectrum of the mixed signal is calculated; input data (1, 2, 3, 4) containing the mixed spectrum are fed to the neural network (7); the input data (1, 2, 3, 4) are processed in the neural network (7); the object and a position of the object are detected by the neural network (7); an object class of the detected object and the detected position of the object are output by the neural network (7) as output data (51, 52, 53, 54).
13. The method according to claim 12, characterized in that the neural network (7) is designed as a convolutional network which has an input layer (6), an output layer (9) and a plurality of convolutional layers (11, 12, 13, 14, 15, 16, 17), and that from one layer to the subsequent layer a convolution operation is carried out.
14. Method according to one of claims 12 to 13, characterized in that the calculated mixed spectrum comprises at least a distance and an azimuth angle of a radar measurement, and in that first input data (1) containing the distance of the radar measurement are supplied to the neural network (7), and in that second input data (2) containing the azimuth angle of the radar measurement are supplied to the neural network (7).
15. Method according to one of claims 12 to 14, characterized in that the calculated mixed spectrum comprises at least a distance and a radial velocity of a radar measurement, and in that third input data (3) containing the distance of the radar measurement are supplied to the neural network (7), and in that fourth input data (4) containing the radial velocity of the radar measurement are supplied to the neural network (7).