Apparatus and method for training a neural network and for putting it into effect
Patent Information
- Application Number
- CN202010587668.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2020-06-24
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2040-06-24
AI Technical Summary
Existing neural networks have difficulty handling context-free classification problems during training and validation, resulting in misclassification, and are unable to effectively verify the context-free validation of classification during the validation process.
By generating saliency maps and utilizing classification importance values and loss functions, combined with neural network adaptive training methods, the classification process is optimized, context dependence is reduced, and classification accuracy is improved.
Context-free classification training and validation are achieved, which reduces misclassification caused by context and improves classification precision and accuracy.
Smart Images

Figure CN112232364B_ABST
Abstract
Description
Technical Field
[0001] Various embodiments generally relate to an apparatus and method for training a neural network and an apparatus and method for validating a neural network. Background Art
[0002] Different neural networks are used, for example, to classify data. These neural networks can be trained based on ground truth data. Because this ground truth data does not cover all possible scenarios or contexts, unexpected correlations may be learned, which can lead to misclassification. Therefore, it may be necessary to perform classification independently of the data context. Furthermore, within the scope of validation and / or verification processes, it may be necessary to validate or verify the context-free nature of the classification determined by the neural network.
[0003] A method for determining a saliency map (also called a saliency map) is described in K. Simonyan et al., Deep Inside Convolutional Networks: Visualizing Image Classification Models and Saliency Maps, ICLR Workshop, 2014. Summary of the Invention
[0004] The method and device having the features of the first example and the twenty-ninth example enable: training a neural network on context-free classification.
[0005] The method and apparatus having the features of the 40th example and the 54th example enable: enabling a neural network to take effect regarding context-free classification.
[0006] At least a portion of the neural network may be implemented by one or more processors.The features described in this paragraph in combination with the first example form a second example.
[0007] The classification importance value of each piece of input data in the saliency map assigned to a category can indicate the importance of the corresponding piece of input data when assigning the category. In other words, at least one piece of input data to be classified can have one of multiple categories, and a saliency map can be generated for the category based on the classified input data, wherein the saliency map has a classification importance value for each piece of input data, and wherein the classification importance value indicates the importance or significance / relevance of the corresponding piece of input data when the category is assigned to the at least one piece of input data having the category. The features described in this paragraph can be combined with the first or second example to form a third example.
[0008] Generating the saliency map further includes assigning a first classification importance value and a second classification importance value to each of the plurality of classification importance values. The features described in this paragraph, combined with one or more of the first to third examples, form a fourth example.
[0009] Each classification importance value below a threshold can be assigned a first classification importance value, while each classification importance value above the threshold can be assigned a second classification importance value. This has the advantage that each of these input data can be assigned whether the corresponding input data is "unimportant" or "important" in terms of assigning a particular category. In other words, the first classification importance value or the second classification importance value can indicate whether the corresponding input data has an impact on the assignment of the particular category. The features described in this paragraph, combined with the fourth example, form a fifth example.
[0010] Determining the first classification error may include determining a first loss value based on a first loss function for the assigned category of the plurality of categories. The features described in this paragraph are combined with one or more of the first to fifth examples to form a sixth example.
[0011] The first loss function may be a cross entropy loss function.The features described in this paragraph are combined with the sixth example to form a seventh example.
[0012] For each of these input data, the category affiliation of the target segmentation may have one category of multiple categories and one segmentation. A segmentation may have multiple input data of these input data, wherein the multiple input data may be assigned to the same category. In other words, multiple input data of the same category may be assigned to one segmentation. This has the following advantage: if the input data have multiple objects, the input data belonging to the corresponding object of the multiple objects may be assigned to the corresponding segmentation, so that the corresponding object of the multiple objects is different from the other objects of the multiple objects. The features described in this paragraph are combined with one or more of the first to seventh examples to form an eighth example.
[0013] Each of these input data may be assigned to exactly one of the plurality of partitions.The features described in this paragraph are combined with the eighth example to form a ninth example.
[0014] Each of the plurality of segmentations can be different from all other segmentations in the plurality of segmentations. This has the advantage that each of the input data can be assigned to exactly one of the plurality of objects, so that each of the plurality of objects can be clearly distinguished from the other objects in the plurality of objects. The features described in this paragraph can be combined with the eighth example or the ninth example to form a tenth example.
[0015] Comparing the classified input data with the target segmentation may include comparing the classified input data with the categories of the plurality of categories assigned to the corresponding ones of the input data. The features described in this paragraph are combined with one or more of the eighth to tenth examples to form an eleventh example.
[0016] Comparing the saliency map with the target segmentation may include comparing the classification importance values assigned to corresponding input data in the plurality of classification importance values of the saliency map assigned to a category with the segmentation assigned to the corresponding input data in the plurality of segmentations. In other words, the classification importance values of the input data may be compared with the segmentation assigned to the category. The features described in this paragraph, combined with one or more of the eighth to eleventh examples, form a twelfth example.
[0017] Each of the multiple classification importance values of the category in the multiple categories to which the assigned input data has been assigned can be set equal to the value "0". In other words, the classification importance value of each input data of the segmentation assigned to the corresponding category can be set equal to the value "0". This has the following advantage: the multiple input data of the segmentation assigned to the corresponding category can be regarded as "important" for the assignment of the category. In other words, thereby, the multiple input data of the segmentation assigned to the corresponding category can have a great influence on the assignment of the category. In other words, thereby, the multiple input data of the segmentation assigned to the corresponding category can have no influence on the second classification error. The features described in this paragraph are combined with the twelfth example to form the thirteenth example.
[0018] Determining the second classification error may include determining a second loss value for the saliency map assigned to a class. The features described in this paragraph are combined with one or more of the first to thirteenth examples to form a fourteenth example.
[0019] The second loss value assigned to a category may have: the sum of all category importance values among the plurality of category importance values assigned to the saliency map of the category. The features described in this paragraph are combined with the fourteenth example to form a fifteenth example.
[0020] Adaptation of the neural network may include minimizing at least one total loss value based on a total loss function, wherein the total loss value may be based on a first classification error for the assigned class and a second classification error for the saliency map assigned to the class. This has the advantage that not only the error in class assignment, i.e., the first classification error, but also the importance of multiple input data not assigned to the class in assigning the class, i.e., the second classification error, can be reduced. The features described in this paragraph, combined with one or more of the first to fifteenth examples, form the sixteenth example.
[0021] The total loss value of the assigned category may be the sum of the first classification error and the second classification error. That is, the total loss value may be the sum of the first loss value and the second loss value. The features described in this paragraph are combined with the sixteenth example to form the seventeenth example.
[0022] The total loss value assigned to each class can be a weighted sum of the first classification error multiplied by the first weighting factor and the second classification error multiplied by the second weighting factor. This has the advantage of assigning a higher relevance to the first classification error or the second classification error. The features described in this paragraph, combined with the sixteenth or seventeenth example, form the eighteenth example.
[0023] The method for training a neural network can be repeated until the total loss value meets a predefined target criterion. The features described in this paragraph are combined with one or more of the sixteenth to eighteenth examples to form a nineteenth example.
[0024] Determining the first classification error may include determining the first classification error for each of the plurality of categories.The features described in this paragraph are combined with one or more of the first to nineteenth examples to form a twentieth example.
[0025] Generating the saliency map may further include generating a saliency map for each of the plurality of categories. The features described in this paragraph are combined with one or more of the first to twentieth examples to form a twenty-first example.
[0026] Determining the second classification error may further include determining the second classification error for each saliency map in the plurality of saliency maps. The features described in this paragraph are combined with the twenty-first example to form a twenty-second example.
[0027] Adapting the neural network may include determining a total loss value for each of the plurality of classes and minimizing each of the plurality of total loss values. In other words, the neural network may be trained for each of the plurality of classes. The features described in this paragraph may be combined with one or more of the first to twenty-second examples to form a twenty-third example.
[0028] The method for training a neural network may be repeated until each of the plurality of total loss values satisfies a respectively predefined target criterion. The features described in this paragraph are combined with the twenty-third example to form a twenty-fourth example.
[0029] The plurality of total loss values may have a common total loss value and the method for training the neural network may be repeated until the common total loss value satisfies a predefined common target criterion. This has the advantage that each of the plurality of categories can be assigned a different relevance, such as by a weighting factor. The features described in this paragraph are combined with the twenty-third example to form a twenty-fifth example.
[0030] Object segmentation may be provided by at least one additional neural network.The features described in this paragraph are combined with one or more of the first to twenty-fifth examples to form a twenty-ninth example.
[0031] The input data may include digital image data, and each of the input data may include one of a plurality of image points or be formed from such image points. Each image point or plurality of image points may be assigned a color value (e.g., three or four color values, for example, a red value, a green value, and a blue value when using the RGB color space) and / or a brightness value or other values, depending on the color space used. The features described in this paragraph, combined with one or more of the first to twenty-sixth examples, form the twenty-seventh example.
[0032] The input data may include multiple digital image data. Each input data of the allocated image data in the multiple image data may include one of multiple image points or be formed by such an image point. The method for training the neural network can be executed for each of the multiple image data. In other words, the input data may include multiple digital images. Each of the multiple digital images may include multiple image points. The method for training the neural network can be executed for each of the multiple digital images. The features described in this paragraph, combined with the twenty-seventh example, form the twenty-eighth example.
[0033] At least a portion of the neural network may be implemented by one or more processors.The features described in this paragraph are combined with the twenty-ninth example to form a thirtieth example.
[0034] The system may include an apparatus according to the twenty-ninth example or the thirtieth example. The system may include a sensor, such as an imaging sensor, configured to provide these input data. The features described in this paragraph form the thirty-first example.
[0035] The imaging sensor may be a video sensor.The features described in this paragraph are combined with the thirty-first example to form a thirty-second example.
[0036] The input data may include multiple digital image data. Each input data of the allocated image data in the multiple image data may include one of multiple image points or be formed from such image points. The neural network may be configured to process each of the image data. In other words, the input data may include multiple digital images, each of the multiple digital images may include multiple image points, and the neural network may be configured to process each of the multiple digital images. The features described in this paragraph may be combined with the thirty-first example or the thirty-second example to form the thirty-third example.
[0037] The system may also have at least one additional neural network configured to provide object segmentation. The features described in this paragraph may be combined with one or more of the thirty-first to thirty-third examples to form a thirty-fourth example.
[0038] The system may be a medical imaging system.The features described in this paragraph may be combined with one or more of the thirty-first to thirty-fourth examples to form a thirty-fifth example.
[0039] A vehicle may include a driving assistance system. The driving assistance system may include a system according to one or more of the thirty-first to thirty-fifth examples. The features described in this paragraph form a thirty-sixth example.
[0040] The vehicle may have at least one imaging sensor configured to provide digital image data. These digital image data may have multiple image objects. The vehicle may also have a driving assistance system. The driving assistance system may have a neural network trained according to the twenty-seventh example or the twenty-eighth example. The neural network of the driving assistance system may be configured to classify and segment these digital image data. The driving assistance system may be configured to control the vehicle based on the classified or segmented digital image data. That is, the driving assistance system may be configured to process these classified or segmented digital image data and to be able to output at least one control instruction based on these classified or segmented digital image data. This has the following advantage: each image object can be classified independently of the context, that is, independently of the other image objects in the multiple image objects. That is, the number of misclassifications due to context can be reduced (e.g., prevented). In other words, these image objects can be correctly classified with higher accuracy. The features described in this paragraph form the thirty-seventh example.
[0041] The computer program may have program instructions which, when executed by one or more processors, implement the method according to one or more of the first to twenty-eighth examples. The features described in this paragraph form a thirty-eighth example.
[0042] The computer program may be stored in a machine-readable storage medium. The features described in this paragraph are combined with the thirty-eighth example to form a thirty-ninth example.
[0043] At least a portion of the neural network may be implemented by one or more processors.The features described in this paragraph are combined with the 40th example to form a 41st example.
[0044] The neural network may have been trained by the method for training according to one of the first to twenty-eighth examples. The features described in this paragraph may be combined with the fortieth example or the forty-first example to form a forty-second example.
[0045] Each of the plurality of classification importance values, to which the assigned input data has a predefined category affiliation, can be set equal to a value of "0." In other words, each of the classification importance values assigned to a saliency map of a category, to which the assigned input data has a category assigned to the saliency map, can be set equal to a value of "0." The features described in this paragraph, combined with one or more of the 40th to 42nd examples, form a 43rd example.
[0046] The method may further include: if the segmentation error of the saliency map assigned to a class is less than a predefined value, the neural network is activated. This segmentation error can be combined with the fifty-fourth example to illustrate how much influence the input data not assigned to the class has on the assignment of the class. In other words, if the input data not assigned to the class has no influence on the assignment of the class, that is, if the segmentation error is less than a predefined value, the neural network can be activated. The features described in this paragraph, combined with one or more of the fortieth to forty-third examples, form the forty-fourth example.
[0047] Generating the saliency map may further include generating a saliency map for each of the plurality of categories. The features described in this paragraph are combined with one or more of the 40th to 44th examples to form a 45th example.
[0048] The determining of the segmentation error may further include determining the segmentation error for each saliency map of the plurality of saliency maps.The features described in this paragraph are combined with the forty-fifth example to form a forty-sixth example.
[0049] Determining whether the segmentation error is less than a predefined value may include determining whether each segmentation error assigned to the saliency map among the plurality of segmentation errors is less than a corresponding predefined value. The features described in this paragraph, in combination with one or more of the 40th to 46th examples, form a 47th example.
[0050] The segmentation errors in the plurality of segmentation errors may have a common segmentation error, and it may be determined whether the common segmentation error is less than a predefined common value. The features described in this paragraph may be combined with one or more of the 40th to 46th examples to form a 48th example.
[0051] Object segmentation may be provided by at least one additional neural network.The features described in this paragraph are combined with one or more of the 40th to 48th examples to form a 49th example.
[0052] The input data may include digital image data, and each of the input data may include one of a plurality of image points or be formed of such image points. The features described in this paragraph may be combined with one or more of the 40th to 49th examples to form the 50th example.
[0053] The input data may include multiple digital image data. Each input data of the allocated image data in the multiple image data may include one of multiple image points or be formed from such image points. The method for validating the neural network may be executed for each of the multiple image data. The features described in this paragraph, combined with the fiftieth example, form the fifty-first example.
[0054] The computer program may have program instructions which, when executed by one or more processors, implement the method according to one or more of the 40th to 51st examples. The features described in this paragraph form a 52nd example.
[0055] The computer program may be stored in a machine-readable storage medium. The features described in this paragraph are combined with the fifty-second example to form a fifty-third example.
[0056] At least a portion of the neural network may be implemented by one or more processors.The features described in this paragraph are combined with the fifty-fourth example to form a fifty-fifth example. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Embodiments of the present invention are illustrated in the accompanying drawings and described in detail below. In the drawings, like reference numerals generally refer to like parts throughout the several views. The drawings are not necessarily true to scale, with emphasis instead generally being placed on illustrating the principles of the invention.
[0058] Exemplary embodiments of the invention are illustrated in the drawings and are explained in detail in the following description.
[0059] in:
[0060] Figure 1 shows devices according to various embodiments;
[0061] Figure 2 An imaging device according to various embodiments is shown;
[0062] Figure 3 An exemplary digital image is shown;
[0063] Figure 4 A processing system for training a neural network according to various embodiments is shown;
[0064] Figure 5A An exemplary classified digital image is shown;
[0065] Figure 5B An exemplary object segmentation is shown;
[0066] Figure 5C An exemplary saliency map is shown;
[0067] Figure 5D An exemplary processed saliency map is shown;
[0068] Figure 6 Methods for training a neural network according to various embodiments are shown;
[0069] Figure 7 A vehicle according to various embodiments is shown;
[0070] Figure 8 A processing system for validating a neural network according to various embodiments is shown; and
[0071] Figure 9 Methods for implementing a neural network according to various specific embodiments are shown. DETAILED DESCRIPTION
[0072] In one embodiment, a "circuit" may be understood as any type of logical implementation entity, which may be hardware, software, firmware, or a combination thereof. Thus, in one embodiment, a "circuit" may be a hardwired logic circuit or a programmable logic circuit, such as a programmable processor, for example a microprocessor (e.g., a CISC (Complex Instruction Set Computer) or a RISC (Reduced Instruction Set Computer)). A "circuit" may also be software implemented or executed by a processor, such as any type of computer program, for example, a computer program using a virtual machine code, such as Java. According to an alternative embodiment, any other type of implementation of the corresponding functions may be understood as a "circuit," which are described in more detail below.
[0073] Figure 1A system 100 is shown according to various embodiments. System 100 may include one or more sensors 102. Sensor 102 may be configured to provide input data 104. Sensor 102 may be an imaging sensor. Alternatively, sensor 102 may be a LIDAR (laser radar) sensor or a microphone. According to various embodiments, input data 104 may include digital image data (within the scope of this description, detected LIDAR sensor signals are also understood to be image data). Each of these input data 104 may include one of a plurality of image points or be formed from such image points. The sensors of the plurality of sensors may include sensors of the same type or sensors of different types.
[0074] System 100 may also include a storage device 106. Storage device 106 may include memory. This memory may be used, for example, in processing performed by a processor. The memory used in these embodiments may be: volatile memory, such as DRAM (dynamic random access memory); or non-volatile memory, such as PROM (programmable read-only memory), EPROM (erasable PROM), EEPROM (electrically erasable PROM); or flash memory, such as floating gate memory, charge trap memory, MRAM (magnetoresistive random access memory), or PCRAM (phase change random access memory). Storage device 106 may be configured to store input data 104. System 100 may also include at least one processor 108 (e.g., exactly one processor, e.g., two processors, e.g., more than two processors). As described above, the at least one processor 108 may be any type of circuit, that is, any type of logically implemented entity. In various embodiments, the at least one processor 108 is configured to process input data 104.
[0075] In the following, the embodiments are described with reference to digital images as input data. However, it should be pointed out that other (digital) input data can also be used, such as digital audio data or digital video data.
[0076] Figure 2 An imaging system 200 is shown according to various embodiments, wherein the sensor is implemented as an imaging sensor 202. Imaging sensor 202 may be a camera sensor or a video sensor. Imaging sensor 202 may be configured to provide digital image data 204. Digital image data 204 may include at least one digital image, for example, a plurality of digital images 210. Each of the plurality of digital images 210 may be a scene having a plurality of image objects, such as a road, a car, a pedestrian, a cyclist, etc. According to various embodiments, imaging system 200 may include a plurality of imaging sensors.
[0077] Each of the plurality of digital images 210 may have a plurality of image points. A digital image may have one or more image objects. An image object may be assigned with a plurality of image points from the plurality of image points.
[0078] Figure 3 A digital image 300 is shown having a plurality of image points 302. The digital image 300 may have a plurality of image objects, such as a road, a plurality of vehicles, a plurality of cyclists, pedestrians, houses, trees, etc. For example, the digital image 300 may have a first image object 304 and a second image object 306.
[0079] As in Figure 3 As shown in FIG, first image object 304 may be a vehicle, for example, and second image object 306 may be a cyclist, for example. First image object 304 may have a plurality of first object image points 308, and second image object 306 may have a plurality of second image object points 310.
[0080] Figure 4 A processing system 400 for training a neural network according to various embodiments is shown. The processing system 400 may have a storage device 106 for storing digital image data 204, such as digital image 300. The processing system 400 may also have at least one processor 108. The processor 108 implements at least a portion of a neural network 402.
[0081] Neural network 402 is configured to process digital image data 204. Neural network 402 may be configured to process, for example, classify, each pixel in a plurality of pixels of one or more digital images 204. For each pixel in the plurality of pixels 302, the classified digital image 404 may have an assigned class from a plurality of classes. Figure 5A A classified digital image 504 is shown as an example. The classified digital image 504 may be generated by the neural network 402 based on the digital image 300. For each image point in the plurality of image points 302, the classified digital image 504 may have an assigned category X^, where each category X^ of an image point may be different from the categories X^ of other image points in the plurality of image points 302. For example, the first object image point 308 assigned to the first image object 304 may have a common assigned first category from the plurality of categories, while the second object image point 310 assigned to the second image object 306 may have a common assigned second category from the plurality of categories. The common assigned second category may be different from the common assigned first category.
[0082] As in Figure 4As shown in , the storage device 106 is also configured to store an object segmentation 408. The object segmentation 408 can be provided at least in part (e.g., a portion of the object segmentation, such as the entire object segmentation) by an additional neural network. For each of the plurality of image points 302 of the digital image 300, the object segmentation 408 can have an assigned class affiliation. According to various embodiments, for each of the plurality of image points 302, the class affiliation of the object segmentation 408 has an assigned class from the plurality of classes. Furthermore, the object segmentation 408 can assign multiple object image points from the plurality of image points 302 to an image object. The object image points assigned to an image object can have the same assigned class. The class assigned to an object image point from the plurality of classes can be different from all other classes from the plurality of classes. A segmentation can have an object image point assigned to an image object. In other words, the object segmentation 408 can have multiple segmentations, each of which can be different from the other segmentations from the plurality of segmentations. That is, the matrix K (see formula (1)) may have the size of the digital image 300, that is, each matrix element k ij Each matrix element k to which the image point is assigned corresponds to an object image point assigned to an image object. ij For example, the matrix K can have the assigned category, and all other matrix elements k ij There may be one or more other assigned categories. Each matrix element k ij can have natural numbers.
[0083] K∈M(x×y,N) (1).
[0084] The class membership can be provided by a first additional neural network. The segmentation, ie the assignment of object image points to an image object, can be provided by a second additional neural network.
[0085] Figure 5B An object segmentation 508 is shown as an example. The object segmentation 508 may include an object segmentation of the digital image 300. For each image point of the plurality of image points 302, the object segmentation 508 may include a class membership X, that is, an assigned class from the plurality of classes. Figure 5BIn the exemplary object segmentation 508 illustrated in FIG, first object pixel 308 assigned to first image object 304 may have category X1, while second object pixel 310 assigned to second image object 306 may have category X2. Category X1 and category X2 may be different from each other. The pixel in the plurality of pixel 302 that is not assigned to either first image object 304 or second image object 306 may each have category X0 that is different from category X1 and category X2. First object pixel 308 assigned to first image object 304 may be a first segmentation. Second pixel 310 assigned to second image object 306 may be a second segmentation.
[0086] As in Figure 4 As shown in , the neural network 402 can also be set up to generate a saliency map (also called a Saliency Map) 406. The saliency map 406 can be generated for one of the multiple categories. The saliency map 406 can be based on the classified digital image 404. For each of the multiple image points 302, the saliency map 406 can have an assigned classification importance value. The classification importance value of each of the multiple image points 302 of the saliency map 406 assigned to a category can illustrate the importance of the corresponding image point when assigning the category. In other words, the classification importance value can illustrate the importance or significance / relevance of the corresponding image point in the multiple image points 302 when assigning the category. That is, the matrix S (see formula (2)) can have the size of the digital image 300, that is, each matrix element s ij Each of the matrix elements s may be assigned exactly one of the plurality of image points 302. Therefore, the matrix S may have a size of x×y. ij For example, the matrix S can have assigned class importance values. Each matrix element k ij can have real numbers.
[0087] S∈M(x×y,R) (2).
[0088] The processor 108 is configured to assign a first classification importance value and a second classification importance value to each of the plurality of classification importance values. The processor 108 can be configured to assign the first classification importance value to each classification importance value below a threshold and to assign the second classification importance value to each classification importance value above the threshold. According to various embodiments, the first classification importance value is equal to a first binary value of "0" and the second classification importance value is equal to a second binary value of "1." Figure 5CThe saliency map 506A is shown as an example. The saliency map 506A may be generated based on the classified digital image 404. The saliency map 506A may be generated for the class X1 and may have an assigned classification importance value for each of the plurality of image points 402. That is, each of the plurality of classification importance values Each of the plurality of classification importance values may have importance for assigning the class X1 to an image point in the plurality of image points 302. Each classification importance value in the plurality of classification importance values may be different from the other classification importance values in the plurality of classification importance values.
[0089] As in Figure 4 As shown in FIG, the processor 108 can be configured to determine a first classification error 410. The first classification error 410 can be determined by comparing the classified digital image 404 with the target segmentation 408. The comparison of the classified digital image 404 with the target segmentation 408 can include comparing the classified digital image 404 with the target segmentation 408 for each of the plurality of image points 302. The comparison of the classified digital image 404 with the target segmentation 408 can include comparing the class assigned to the respective image point in the plurality of image points 302 of the classified digital image 404 with the class membership assigned to the respective image point in the plurality of image points 302. According to various embodiments, determining the first classification error 410 includes determining a first loss value based on a first loss function. The first loss value can be determined for each assigned class of the plurality of classes. The first loss function can be a cross-entropy loss function.
[0090] The processor 108 may also be configured to determine a second classification error 412. The second classification error 412 may be determined by comparing the saliency map 406 with the target segmentation 408, for example, for each pixel in the plurality of pixels 302. The classification importance values assigned to the respective pixels in the plurality of pixels 302 of the saliency map 406 assigned to a class may be compared with the assigned segmentation of the image object assigned to the class. The processor 108 may be configured to set one or more of the plurality of classification importance values to “0.” According to various embodiments, each classification importance value of the plurality of classification importance values assigned to an object pixel of the class is set to “0.” In other words, each classification importance value of the saliency map assigned to a class whose assigned pixel has the class assigned to it is set to “0.” This can be achieved by having each matrix element k of the matrix K given in formula (1) have a class assigned to the saliency map ij Each matrix element k that does not have an assigned category is set equal to the value "0". ij is set equal to the value "1", so that the matrix K' is obtained. Each matrix element k of the matrix K' ij ' can be multiplied by the matrix element s of the matrix S given in formula (2) ij , so that we get a matrix with elements s ij 'Matrix S'.
[0091] According to different embodiments, the classification importance value assigned to the corresponding image point in the plurality of image points 302 is a first classification importance value or a second classification importance value. According to different embodiments, the determination of the second classification error 412 includes: determining a second loss value. The second loss value can be determined for the saliency map 406 assigned to a category. The second loss value can be determined after each classification importance value of the object image point assigned to the image object of the category in the plurality of classification importance values has been set equal to "0". The second loss value assigned to a category may have: the sum of all classification importance values in the plurality of classification importance values of the saliency map 306 assigned to the category. That is, the second loss value (V2, see formula (3)) can be determined based on the matrix S' and V2 may have all matrix elements s ij 'sum.
[0092] V2=∑ ij s ij ′=∑ ij s ij ×kij ′ (3).
[0093] Figure 5D Processed saliency map 506B is shown as an example. Processed saliency map 506B may be generated based on saliency map 506A. Saliency map 506A may be generated for category X1. Category X1 may be assigned to first object image points 308 assigned to first image object 304. Processed saliency map 506B may have a classification importance value equal to "0" for each first object image point 308 assigned to first image object 304. In other words, the processor 108 may be configured to assign each classification importance value of category X1 to That is, the first object image point 308 or the first segmentation setting is equal to “0”. The second loss value may have all the classification importance values of the plurality of classification importance values assigned to the processed saliency map 506B of the class X1. sum.
[0094] As in Figure 4 As shown in FIG, the processor 108 can also be configured to adapt the neural network 402 (in other words, train the neural network 402), for example, by minimizing at least one total loss value 414. The total loss value 414 is based on a first classification error 410 for the assigned class and a second classification error 412 for the saliency map 406 assigned to the class. The total loss value 414 for the assigned class can be a sum (optionally a weighted sum) of the first classification error 410 (e.g., the first loss value) and the second classification error 412 (e.g., the second loss value).
[0095] According to various embodiments, the processing system 400 further includes a sensor 102 .
[0096] The processing system 400 may also have at least one additional neural network configured to provide at least a portion of the object segmentation 408 (eg, the entire object segmentation).
[0097] The processing system 400 may be a medical imaging system.
[0098] According to one embodiment, a computer-controlled machine, such as a robot, a vehicle, a home appliance, a power tool, a production tool, an intelligent personal assistant, or an access control system, is provided. The computer-controlled machine may have a processing system 400 .
[0099] According to one embodiment, a vehicle is provided, which includes a driving assistance system. The driving assistance system may include a processing system 400 .
[0100] Figure 6 A method 600 for training a neural network according to various embodiments is shown. The method 600 may include classifying input data 104 by means of the neural network 402. The input data 104 may include digital image data 204, such as the digital image 300. The method 600 may include classifying the digital image 300 by means of the neural network 402 (at 602). Classifying the digital image 300 may include assigning each of the plurality of image points 302 to one of a plurality of categories. The method 600 may also include generating a saliency map 406 (at 604). The saliency map 406 may be generated for one of the plurality of categories. The saliency map 406 may be generated based on the classified digital image 404. Generating the saliency map 406 may include assigning a classification importance value to each of the plurality of image points 302. The method 600 may include providing an object segmentation 408 (at 606). For each image point in the plurality of image points 302, the target segmentation 408 may have an assigned class membership. Method 600 may also include determining a first classification error 410 (at 608). The first classification error 410 may be determined by comparing the classified digital image 404 with the target segmentation 408. Method 600 may also include determining a second classification error 412 (at 610). The second classification error 412 may be determined by comparing the saliency map 406 with the target segmentation 408. Method 600 may include adapting the neural network 402 (at 612). The neural network 402 may be adapted based on the first classification error 410 and the second classification error 412. The adaptation of the neural network 402 may include minimizing a total loss value 414, and the total loss value 414 may be based on the first classification error 410 (e.g., a first loss value) and the second classification error 412 (e.g., a second loss value). According to various embodiments, method 600 is repeated until the total loss value 414 meets a predefined target criterion.
[0101] According to various embodiments, generating a saliency map 406 may include generating corresponding saliency maps 406 for a plurality of classes. According to various embodiments, determining a first classification error 410 may include determining a first classification error 410 for a plurality of classes. According to various embodiments, determining a second classification error 412 may include determining a second classification error 412 for a plurality of saliency maps from a plurality of saliency maps. Determining the second classification error 412 may include determining a second classification error 412 for each saliency map 406 from the plurality of saliency maps. According to various embodiments, adapting the neural network 402 may include determining a total loss value 414 for a plurality of classes. Adapting the neural network 402 may include determining a total loss value 414 for each class from the plurality of classes. Adapting the neural network 402 may also include minimizing each total loss value 414 from the plurality of total loss values. According to various embodiments, method 600 is repeated until each total loss value 414 from the plurality of total loss values satisfies a respectively predefined target criterion. The plurality of total loss values may have a common total loss value, and method 600 may be repeated until the common total loss value satisfies a predefined common target criterion.
[0102] Figure 7 A vehicle 700 according to one embodiment is shown. Vehicle 700 may be a vehicle with an internal combustion engine, an electric vehicle, a hybrid vehicle, or a combination thereof. Vehicle 700 may also be a car, a truck (LKW), a ship, an unmanned aerial vehicle, an airplane, or the like.
[0103] The vehicle 700 may have at least one sensor (e.g., imaging sensor) 702 (e.g., sensor 102). The vehicle 700 may have a driver assistance system 704. The driver assistance system 704 may have a storage device 106. The driver assistance system 704 may have a processor 108. The processor 108 may implement a neural network. The neural network of the driver assistance system 704 may be configured to classify and segment the digital image data. According to various embodiments, the neural network is trained according to the method 600 for training a neural network so that the neural network can classify the digital image data without context. Context-free classification prevents the learning of unexpected correlations that could lead to misclassification. In other words, the driver assistance system 704 can better assign multiple image objects to corresponding categories, better segment multiple image objects, and thereby better recognize multiple image objects. Therefore, one aspect is to provide a vehicle that enables improved recognition of image objects in digital image data.
[0104] Driver assistance system 704 can be configured to control vehicle 700 based on context-free classified or segmented digital image data. In other words, driver assistance system 704 can be configured to process this context-free classified or segmented digital image data and to output at least one control command to one or more actuators of vehicle 700 based on this context-free classified or segmented digital image data. In other words, driver assistance system 704 can influence the current driving behavior based on this context-free classified or segmented digital image data, for example, maintaining or changing the current driving behavior. Changes to driving behavior can include, for example, intervention in the driving behavior for safety reasons, such as emergency braking.
[0105] Figure 8 A processing system 800 for implementing a neural network according to various embodiments is shown. The processing system 800 may include a storage device 106 for storing input data 104. The input data 104 may be provided by a sensor 102. The processing system 800 may also include a processor 108. The processor 108 may be configured to implement at least a portion of the neural network 802.
[0106] A neural network 802 may be configured to classify digital image data 204, such as the digital image 300. For each pixel in the plurality of pixels 302, the classified digital image 804 may have an assigned category from a plurality of categories. The neural network 802 may also be configured to generate a saliency map 806. The saliency map 806 may be generated for one of the plurality of categories. The saliency map 806 may be based on the classified digital image 804. For each pixel in the plurality of pixels 302, the saliency map 806 may have an assigned category importance value. The category importance value of each pixel in the plurality of pixels 302 in the saliency map 806 assigned to a category may indicate the importance of the corresponding pixel in assigning the category. The processor 108 may be configured to assign a first category importance value and a second category importance value to each of the plurality of category importance values. The processor 108 may be configured to assign a first classification importance value (eg, "0") to each classification importance value below a threshold and a second classification importance value (eg, "1") to each classification importance value above the threshold.
[0107] The storage device 106 may also be set up to store target partitions 808 .
[0108] The processor 108 may be configured to process the target segmentation 808, for example to determine a segmentation error 812. The segmentation error 812 may be determined by comparing the saliency map 806 with the target segmentation 808, for example by means of a loss function.
[0109] Processor 108 may be configured to determine whether segmentation error 812 is less than a predefined value (at 814). Processor 108 may be configured to validate neural network 802 at 816 if segmentation error 812 is less than the predefined value ("yes" at 814); and processor 108 may be configured not to validate neural network 802 at 818 if segmentation error 812 is greater than or equal to the predefined value ("no" at 814).
[0110] Figure 9 A method 900 for implementing a neural network according to various embodiments is shown. The method 900 may include classifying input data (eg, digital image data) 104 by means of a neural network 802 .
[0111] Method 900 may include classifying digital image 300 using neural network 802 (at 902). Classifying digital image 300 may include assigning one of a plurality of categories to each of the plurality of image points 302. According to various embodiments, neural network 802 is trained or adapted according to method 600 for training a neural network. Method 900 for validating the neural network may also include generating a saliency map 806 (at 904). Saliency map 806 may be generated for one of the plurality of categories. Saliency map 806 may be generated based on the classified digital image 804. Generating saliency map 806 may include assigning a classification importance value to each of the plurality of image points 302. Method 900 may include providing an object segmentation 808 (at 906). Object segmentation 808 may have an assigned category affiliation for each of the plurality of image points 302. Method 900 may include determining a segmentation error 812 (at 908). The segmentation error 812 can be determined by comparing the saliency map 806 with the target segmentation 808. The method 900 may also include determining whether the segmentation error 812 is less than a predefined value (at 910). According to various embodiments, the method 900 includes validating the neural network 802 if the segmentation error 812 is less than the predefined value.
Claims
1. A method for training a neural network, the method being implemented by one or more processors, the method comprising: classifying input data by means of a neural network, wherein the input data comprises digital image data, wherein each input data is assigned one of a plurality of classes; generating a saliency map for a category among the plurality of categories based on the classified input data, wherein each of the input data is assigned a category importance value; providing object segmentation in which each of the input data is assigned a class affiliation, wherein the object segmentation is provided by at least one additional neural network; determining a first classification error by comparing the classified input data to the target segmentation; determining a second classification error by comparing the saliency map with the object segmentation; The neural network is adapted based on the first classification error and the second classification error.
2. The method according to claim 1, The classification importance value of each input data in the input data of the saliency map assigned to a class indicates the importance of the corresponding input data in assigning the class.
3. The method according to claim 1 or 2, The comparing of the classified input data with the target segmentation comprises comparing the classified input data with the target segmentation for each of the input data.
4. The method according to claim 1 or 2, The adaptation of the neural network comprises minimizing at least one total loss value, wherein the total loss value is based on a first classification error of the assigned class and a second classification error of a saliency map assigned to the class.
5. The method according to claim 1 or 2, The determining of the first classification error includes determining the first classification error for each category of the plurality of categories.
6. The method according to claim 1 or 2, The generating of the saliency map further comprises: generating a saliency map for each of the plurality of categories; and the determining of the second classification error further comprises: determining the second classification error for each of the plurality of saliency maps.
7. The method according to claim 1 or 2, wherein the categories of the category affiliations are provided by a first additional neural network; and The segmentation of the classes is provided by a second additional neural network, the second additional neural network being different from the first additional neural network.
8. The method according to claim 1 or 2, Each of the input data has a pixel from a plurality of pixels or is formed by such pixels.
9. The method according to claim 8, The input data comprises a plurality of digital image data; wherein each input data of the assigned image data of the plurality of image data has one of the plurality of image points or is formed by such an image point; and The method for training a neural network is performed for each image data in the plurality of image data. 10 . A device for training a neural network, the device being configured to carry out the method according to claim 1 .
11. A system comprising: The apparatus according to claim 10; and An imaging sensor is configured to provide the input data to the device.
12. A vehicle comprising: at least one imaging sensor configured to provide digital image data; and A driver assistance system having a neural network trained according to claim 8 , wherein the neural network is configured to classify the digital image data, and wherein the driver assistance system is configured to control the vehicle based on the classified digital image data.
13. A method for validating a neural network, the method being implemented by one or more processors, the method comprising: classifying input data by means of a neural network, wherein the input data comprises digital image data, wherein each input data is assigned one of a plurality of classes; generating a saliency map for a category among the plurality of categories based on the classified input data, wherein each of the input data is assigned a category importance value; providing object segmentation in which each of the input data is assigned a class affiliation, wherein the object segmentation is provided by at least one additional neural network; determining a segmentation error by comparing the saliency map with the target segmentation; It is determined whether the segmentation error is smaller than a predefined value.
14. A device for validating a neural network, the device being configured to carry out the method according to claim 13.
Citation Information
Patent Citations
Domain adaptation via class-balanced self-training with spatial priors
CN109724608A
Semantic object proposal generation and validation
US20150170006A1