Evaluating machine learning system for semantic segmentation of video data

By determining the baseline truth consistency and temporal consistency in the semantic segmentation of video data, the false excitation problem caused by temporal inconsistency in the existing technology is solved, and the training effect and application accuracy of the machine learning system are improved.

CN120689701APending Publication Date: 2025-09-23ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340212.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2025-03-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing semantic segmentation methods for video data are inconsistent in time, which may lead to false excitations in machine learning systems during training and evaluation, affecting system performance and practical application effects.

Method used

By determining ground truth consistency and temporal consistency, restricting the evaluation to only pixels or components that give ground truth consistency, fast matrix operations are used to evaluate the performance of machine learning systems, avoiding false activations, and weighting the availability of training examples.

Benefits of technology

It improves the training effect of the machine learning system, ensures that it can more accurately reflect the semantic content of video data in practical applications, reduces the waste of computing resources, and improves the accuracy and reliability of system response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689701A_ABST
    Figure CN120689701A_ABST
Patent Text Reader

Abstract

The invention relates to evaluating a machine learning system for semantic segmentation of video data. The method comprises the following steps: providing video frames X1, X2,..., Xt-1, Xt, segmented frames Y1, Y2,..., Yt-1, Yt and at least one rated segmented frame St; determining a relative motion between the camera and the scene; from the at least one segmented frame Yt-1, determining an expected segmented frame using the relative motion to determine a reference true consistency, which indicates to what extent the actual segmented frame Yt and / or the expected segmented frame is consistent with a predetermined rated segmented frame St; a temporal consistency is determined for those pixels or other components of the actual split frame Yt for which case is such or for the corresponding pixels or other components of the expected split frame, the pixel or other components are consistent with the pixel or other components corresponding to the pixel or other components of the expected segmentation frame or the actual segmentation frame Yt to what degree; and evaluating the sought evaluation of the machine learning system from the temporal consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to semantic segmentation of video data, which can be used, for example, for environmental monitoring of autonomously controlled vehicles and / or robots. Background Art

[0002] At least partially automated driving of vehicles and / or robots on business premises or in public road traffic requires continuous monitoring of the vehicle's and / or robot's environment. To this end, one or more cameras are used, in particular, which provide a sequence of video frames.

[0003] The evaluation of these video frames can, in particular, include semantic segmentation, which assigns classes, such as object types, to pixels or other components of the corresponding frame. Using this semantic segmentation, the scene represented in the video frames can be converted into a machine-readable form that can be used by downstream systems, such as trajectory planners. This allows, for example, the trajectory of a vehicle or robot to be planned so as to avoid collisions with other objects.

[0004] In this case, it is important that the semantic segmentation is consistent over time. Thus, for example, it would not be sensible for the same object visible in two consecutive video frames to be assigned to different classes based on these two video frames. Summary of the Invention

[0005] The present invention provides a computer-implemented method for evaluating a machine learning system for semantic segmentation of video data. The video data comprises video frames X1, X2, ..., X recorded in a time-discrete sequence. N The semantic segmentation extends into the future for a given horizon of t time steps and thus includes the segmented frames Y1, Y2, ..., Y t-1 、Y t , where t≤N. Each of these segmented frames Y1, Y2, ..., Y t-1 、Y t Assign classes from a predefined classification to the corresponding video frames X1, X2, ..., X t-1 、X t pixels or other components of a .

[0006] In the context of the method, video frames X1, X2, ..., X are provided. t-1 、X t and the segmented frames Y1, Y2, ..., Y determined for this purpose by the machine learning system t-1 、Y t In addition, for video frame X t Set at least one rated segmentation frame S t Rated split frame S tWhat should the machine learning system ideally be able to do for video frame X? t The segmented frame Y provided t Therefore, the nominal segmented frame is also called the “ground truth”.

[0007] Determine on one hand the camera used for recording the video data and the video frames X1, X2, ..., X N The relative motion between the scenes represented in [ 1 ] can be composed in any manner of the motion of the camera on the one hand and the motion of the scene on the other hand. For example, without limiting the generality, the relative motion can be expressed by the fact that the position of the camera relative to the stationary scene, i.e., the combination of position and orientation, changes from frame to frame.

[0008] From at least one segmented frame Y t-1 , the expected segmented frame is now determined using the determined relative motion The expected segmented frame It is therefore the segmentation frame that should occur when the video frame changes in time step from t-1 to t only due to the relative motion between the camera and the scene. Determine the expected segmentation frame This may for example include, in particular, segmenting the frame Y according to the determined relative motion. t-1 Deformation (distortion). That is, the perspective of the camera pose at time t in the segmented frame Y t-1 The information contained in the expected segmentation frame In the case of real video sequences, perfect temporal consistency cannot usually be expected for this, since in the expected split frames Not considered:

[0009] The object may be obscured (occluded) by other objects to different degrees from the viewpoint of the two camera poses, and

[0010] • Objects may appear in or disappear from the scene between time points t-1 and t (such as people getting off or getting on a vehicle).

[0011] In addition, the ground truth consistency is determined. The ground truth consistency indicates that the actual segmented frame Y t and / or expected split frames To what extent is the video frame X t The pre-set rated split frame S t consistent.

[0012] Now only for the actual segmentation frame Y tTemporal consistency is determined for those pixels or other components for which the ground truth consistency is given. The temporal consistency describes to what extent these pixels or other components are consistent with the expected segmentation frame. The corresponding pixels or other components are consistent. The actual segmentation frame Y t All pixels or other components for which no ground truth consistency is given are not included in the determination of temporal consistency.

[0013] Alternatively, it is also possible to segment the frame Y from the actual t Starting from the pixels or other components for which the ground truth consistency is given, the temporal consistency evaluation is expected to be performed on the segmented frame The corresponding pixels or other components. So the actual segmentation frame Y t Instead, only the expected segmented frame Extract pixels or other components from a

[0014] Evaluating the sought-after evaluation of machine learning systems from the perspective of temporal consistency.

[0015] It has been found that limiting the determination of temporal consistency to pixels or other components for which a ground truth consistency is also given leads to a more accurate measurement of the performance of the machine learning system. In particular, this limitation avoids false incentives for developing the machine learning system when using the evaluation determined using the method as feedback for training the machine learning system. If only temporal consistency were taken into account during the evaluation, the machine learning system could in extreme cases "cheat" into a good evaluation by simply throwing out the same segmented frame for all video frames, for example in the form of a homogeneous area that fills the entire frame and is assigned to a specific class. In this case, the maximum temporal consistency is always achieved, but the result no longer has any relation to the actual semantic content of the video data.

[0016] Furthermore, it has been recognized that during the training of a machine learning system, the video frames X1, X2, ..., X t-1 、X t Video sequences may differ from one another in terms of their usability and effectiveness. Thus, training examples may, for example, include video sequences recorded during the day and under good visual conditions, so that the content can be well recognized in each entire frame. However, there may also be video sequences in which only individual contents can be recognized, and the majority of the frame cannot be used for further evaluation. By determining temporal consistency only for pixels or other components that can be evaluated correctly in a semantically correct manner, the evaluation determined based on the video sequence can be weighted, for example, by the number of pixels or other components for which a ground truth consistency is given.

[0017] In a particularly advantageous embodiment, the ground truth consistency is determined as the actual segmented frame Y t 、Y t and / or expected split frames of frame with rated split S t The ground truth consistency set of pixels or other components whose corresponding pixels or other components together satisfy a predetermined consistency criterion. In particular, the consistency criterion may for example specify that the actual segmented frame Y t The pixels or other components can be divided into frames S t The corresponding pixels or other components of deviate only by a certain value. Thus, the cardinality of the ground truth consistency set gives information in aggregate about the degree of ground truth consistency for the training examples.

[0018] Therefore, it is particularly advantageous to determine the cardinality of the reference truth consistency set for the video frames X1, X2, ..., X N and rated split frame S t A measure of ground truth consistency of composed training examples.

[0019] In another particularly advantageous embodiment, temporal consistency is determined for pixels or other components of a reference truth consistency set. This allows pixels or other components for which semantic segmentation is not relevant to be excluded from the temporal consistency determination from the outset. Thus, compared to a solution that first calculates temporal consistency for all pixels or other components and then subsequently discards temporal consistency for insignificant pixels or other components, the computational effort required for this purpose can be completely reduced.

[0020] For example, the actual segmentation frame Y t The element-wise product with a binary mask describing the actual segmented frame Y can be fed to the temporal consistency check. t Whether the pixel or other component belongs to the reference truth consistency set can then be checked for temporal consistency using fast matrix operations, which are much more efficient than processing individual pixels or other components in sequence. Nevertheless, unnecessary processing of insignificant pixels or other components can still be avoided.

[0021] In another particularly advantageous embodiment, the temporal consistency is determined as the actual segmented frame Y t Expected segmentation frame The temporal consistency set of pixels or other components whose corresponding pixels or other components satisfy the predetermined consistency criterion. Alternatively, the expected segmented frame can also be determined The actual segmentation frame Y tThe two operation modes provide information on the actual segmentation of the frame Y. t A direct statement of which spatial regions of t exist and which do not. At the same time, the cardinality of the temporal consistency set can be viewed as the temporal correlation for the time step from t-1 to t. An indicator of the degree of.

[0022] Therefore, it is particularly advantageous to evaluate the sought evaluation of the machine learning system according to the cardinality of the temporal consistency set. For example, the evaluation can be calculated as the expected segmentation frame for which the ground truth consistency is given on the one hand The actual segmentation frame Y t A "mean intersection over union" (mIoU) between the two: the intersection between the two corresponds to the temporally consistent set. In the mIoU calculation, the cardinality of the intersection is divided by the cardinality of the union, which in this case is the set of all pixels in the frame. "Average" involves performing this calculation separately for all classes and averaging the results.

[0023] In another particularly advantageous embodiment, the evaluation of the machine learning system is used as feedback to optimize parameters characterizing the behavior of the machine learning system. The evaluation of the machine learning system better corresponds to its actual performance, thereby guiding training towards actual improvement with a greater probability. In particular, as explained above, no false incentives are provided for the further development of the machine learning system.

[0024] In another particularly advantageous embodiment, the evaluation determined by the machine learning system is assigned as a confidence value to the segmented frames Y1, Y2, . . . , Y provided by the machine learning system. t-1 、Y t In this way, when further processing these segmented frames Y1, Y2, ..., Y t-1 、Y t When considering how well the machine learning system is trained overall, you can consider: if, for example, you segment the frames Y1, Y2, ..., Y t-1 、Y t Still aggregated with segmented frames from other sources, the segmented frames can be weighted using the rating determined according to the method proposed here as confidence.

[0025] Alternatively or in combination therewith, in response to the determined evaluation exceeding a predetermined threshold, the machine learning system can be released for use. For example, this can be used, in particular, as a termination criterion for training the machine learning system.

[0026] In another particularly advantageous embodiment, the video frames X1, X2, . . . , X recorded by at least one camera are t-1 、X t The semantic segmentation frames Y1, Y2, ..., Y are then provided to the trained machine learning system. t-1 、Y t The control signal is determined in the video frame X1, X2, ..., X2. The control signal is used to control a vehicle, a driver assistance system, a robot, a system for quality control, a system for monitoring an area, and / or a system for medical imaging. In this regard, the improved training due to the more accurate evaluation of the actual performance of the machine learning system has the following effect: the reaction of the respectively controlled system to the control signal corresponds with a higher probability to the reaction in the video frame X1, X2, ..., X2. t-1 、X t The situation reflected in the sequence corresponds to this.

[0027] In particular, the method can be fully or partially computer-implemented. Therefore, the present invention also relates to a computer program having machine-readable instructions that, when executed on one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the described method. In this sense, control devices for vehicles and embedded systems for technical devices that are also capable of executing machine-readable instructions can also be considered computers. A computing instance can be, for example, a virtual machine, a container, or a serverless execution environment, which can be provided in particular in the cloud.

[0028] The invention also relates to a machine-readable data carrier and / or a download product containing the computer program. A download product is a digital product that can be transmitted via a data network, i.e., can be downloaded by a user of the data network, and can be sold for immediate downloading, for example, in an online store.

[0029] Furthermore, one or more computers and / or computing instances can be equipped with a computer program, a machine-readable data carrier or a download product.

[0030] Further measures for improving the invention are described in more detail below with reference to the figures together with the description of preferred exemplary embodiments of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 An embodiment of a method 100 for evaluating a machine learning system 1 for semantic segmentation of video data is shown;

[0032] Figure 2 Show video frame X t-1 、X t The example processing process is used for evaluation 5;

[0033] Figure 3 Shows the segmented frames Y with different ground truth consistencies included in evaluation 5 t-1 、Y t . DETAILED DESCRIPTION

[0034] Figure 1 is a schematic flow chart of an embodiment of a method 100 for evaluating a machine learning system 1 for semantic segmentation of video data. The video data comprises video frames X1, X2, ..., X N Semantic segmentation involves segmenting frames Y1, Y2, ..., Y t-1 、Y t , where t≤N. These segmented frames Y1, Y2, ..., Y t-1 、Y t Assign classes from a predefined classification to the corresponding video frames X1, X2, ..., X t-1 、X t pixels or other components of a .

[0035] In step 110, video frames X1, X2, ..., X t-1 、X t and the segmented frames Y1, Y2, ..., Y determined for this purpose by the machine learning system 1 t-1 、Y t In addition, for video frame X t Provide at least one rated split frame S t .

[0036] In step 120, it is determined that the camera used for recording the video data is the same as the one used for recording the video data in the video frames X1, X2, ..., X N The relative motion between scenes represented in 2.

[0037] In step 130, the determined relative motion 2 is used to extract the image from at least one segmented frame Y. t-1 Determine the expected segmentation frame in

[0038] According to block 131, the expected segmented frame is determined This may include segmenting frame Y according to the determined relative motion 2 t-1 deformation.

[0039] In step 140, the ground truth consistency 3 is determined, which describes the actual segmentation frame Y t and / or expected split frames To what extent is the video frame X t The pre-set rated split frame S t consistent.

[0040] According to block 141, the ground truth consistency 3 may be determined as the actual segmented frame Y t and / or expected split frames of frame with rated split S t A reference truth consistency set 3a of pixels or other components whose corresponding pixels or other components together satisfy a predetermined consistency criterion.

[0041] According to block 141a, the cardinality of the reference truth consistency set 3a may be Determine the N and rated split frame S t A measure of ground truth consistency of composed training examples.

[0042] In step 150, for the actual segmented frame Y t for which this is the case or for the intended segmented frame The temporal consistency 4 is determined for the pixels or other components corresponding thereto. The temporal consistency 4 indicates to what extent these pixels or other components are consistent with the expected segmented frame. Or actual segmentation frame Y t The corresponding pixel or other component is consistent.

[0043] According to block 151 , temporal consistency 4 may be determined for pixels or other components of the reference truth consistency set 3a.

[0044] According to block 151a, the actual segmentation of frame Y t The element-wise product with the binary mask that describes the actual segmented frame Y can be fed to the temporal consistency check. t Whether the pixel or other component belongs to the reference truth consistency set 3a.

[0045] According to block 152, the temporal consistency 4 may be determined as the actual segmented frame Y t or expected split frame Expected segmentation frame Or actual segmentation frame Y t A temporal consistency set 4a of those pixels or other components whose corresponding pixels or other components together satisfy a predetermined consistency criterion.

[0046] In step 160, the sought evaluation 5 of the machine learning system 1 is evaluated from the temporal consistencies 4. As long as the temporal consistencies 4 exist as a temporal consistency set 4a according to block 152, the sought evaluation 5 of the machine learning system 1 can be evaluated according to block 161 based on the cardinality of the temporal consistency set 4a.

[0047] The evaluation 5 determined by the machine learning system 1 can be assigned as a confidence value to the segmented frames Y1, Y2, ..., Y provided by the machine learning system 1. t-1 、Y t Alternatively or in combination therewith, the machine learning system 1 can be released for use in response to the determined evaluation 5 exceeding a predetermined threshold value.

[0048] In step 190, the evaluation 5 of the machine learning system can be used as feedback to optimize the parameters 1a that characterize the behavior of the machine learning system 1. The state of the completed optimization of the parameters 1a is denoted by reference numeral 1a*, and also specifies the state 1* of the completed training of the machine learning system 1.

[0049] In step 200, the trained machine learning system 1* may be fed with video frames X1, X2, ..., X recorded by at least one camera. t-1 、X t Then, in step 210, the semantic segmentation frames Y1, Y2, ..., Y t-1 、Y t In step 220 , the vehicle 50 , the driver assistance system 51 , the robot 60 , the system 70 for quality control, the system 80 for monitoring the area and / or the system 90 for medical imaging can then be controlled using the control signal 210 a.

[0050] Figure 2 Illustrate video frame X t-1 、X t The example processing process is used for evaluation 5.

[0051] exist Figure 2 In the example shown in , there are video frames X1, X2, ..., X t-1 、X t A video sequence showing only X t-1 and X t For these video frames X t-1 and X t , rated split frame S t-1 and S t Available separately. Two video frames X t-1 and X t The difference lies in the position C of the camera used for recording relative to the scene. t-1 or C t In step 120 of method 100 , a relative motion 2 is determined, which represents the relative motion between these positions C t-1 and C t The difference between.

[0052] Machine learning system 1 for video frame X t-1 Determine the segmentation frame Y t-1 , and for video frame X t Determine the segmentation frame Y t Using the relative motion 2, in step 130 of method 100 and according to block 131, the segmented frame Y is transformed t Determine the expected segmentation frame in In step 140 and according to block 141, the expected segmented frame is determined Which component of the time point t is the same as the nominal segmentation frame S t The ground truth consistency 3 is delivered in the form of a ground truth consistency set 3a of those pixels for which the consistency is given.

[0053] In step 150 and according to block 151, it is determined only for the pixels that are part of the ground truth consistency set 3a: to what extent these pixels are consistent with the actual segmented frame Y t The pixel corresponding thereto is consistent. Thus, the sought evaluation 5 of the machine learning system is then evaluated in step 160.

[0054] exist Figure 3 Segment frame Y t-1 and Y t With the corresponding rated split frame S t-1 and S t How the different consistencies of affect the evaluation 5 of the machine learning system 1 determined according to the method proposed here.

[0055] The sub-images a and b relate to the first pair of times t-1 on the one hand and t on the other hand, and therefore also to the camera pose C on the one hand. t-1 On the other hand, C t The second pair of sub-images c and d relates to the time instant t-1 on the one hand and t on the other hand, and thus also to the camera pose C on the one hand. t-1 On the other hand, C t The second pair.

[0056] The segmented frame Y shown in sub-image a t-1 The segmented frame Y shown in sub-image b t The pure temporal consistency between them is 0.909. The changes in the segmented frames therefore basically correspond to the changes in the camera poses C. t-1 and C t That's to be expected.

[0057] The segmented frame Y shown in sub-image a t-1 With the rated split frame S t-1The similarity measured by the mean Intersection over Union (mIoU) is 0.911. For the segmented frame Y shown in sub-image b t , with rated split frame S t The similarity is 0.899.

[0058] According to the method proposed here, only frame Y is segmented t-1 and Y t The component of the ground truth consistency 3 for which there is a basis for the present is incorporated into the evaluation 5 of the machine learning system by means of the temporal consistency 4. In the example of the sub-images a and b, this evaluation 5 is assigned a good value of 0.848.

[0059] Different from this example, the segmented frame Y shown in sub-images c and d t-1 and Y t Only the rated split frame S corresponding to t-1 or S t The mIoU similarity is 0.431 or 0.436. Figure 3 As shown in , the reason for this is in particular the large window area of ​​the illustrated living room, which was not correctly detected by the machine learning system 1.

[0060] A conventional evaluation of the machine learning system 1 based solely on temporal consistency 4 would still yield a good value of 0.826. However, the method proposed here takes into account that the window area is almost completely excluded from this evaluation, and temporal consistency 4 is determined on a significantly weaker basis. Therefore, the evaluation 5 determined using this method still only yields a value of 0.401.

Claims

1. A computer-implemented method (100) for evaluating a machine learning system (1) for semantic segmentation of video data, the video data comprising video frames X1, X2, ..., X N , wherein the semantic segmentation includes segmenting frames Y1, Y2, ..., Y t-1 、Y t , where t≤N, the segmented frame assigns classes from a predetermined classification to the corresponding video frames X1, X2, ..., X t-1 、X t pixels or other components of the image element, and wherein the method comprises the steps of: Provide (110) video frames X1, X2, ..., X t-1 、X t and the segmented frames Y1, Y2, ..., Y determined for this purpose by the machine learning system (1) t-1 、Y t And for video frame X t At least one rated segmentation frame S t ; Determine (120) the camera used on the one hand for recording the video data and the video frames X1, X2, ..., X N Relative motion between scenes represented in (2); From at least one segmented frame Y t-1 , using the determined relative motion (2) to determine (130) the expected segmented frame Determine (140) the ground truth consistency (3), which states: the actual segmented frame Y t and / or expected split frames To what extent is the video frame X t The pre-set rated split frame S t consistent; For the actual segmented frame Y t for which this is the case or for the intended segmented frame The pixels or other components corresponding thereto determine (150) a temporal consistency (4), which describes to what extent these pixels or other components are consistent with the expected segmented frame Or the actual segmented frame Y t The corresponding pixel or other component is consistent; and • Evaluating (160) the sought evaluation (5) of the machine learning system (1) from the temporal consistency (4).

2. The method (100) according to claim 1, wherein the ground truth consistency (3) is determined (141) as the actual segmented frame Y t and / or the expected segmented frame The rated split frame S t A reference truth consistency set (3a) of those pixels or other components whose corresponding pixels or other components together satisfy a predetermined consistency criterion.

3. The method (100) according to claim 2, wherein the cardinality of the reference truth consistency set (3a) is determined (141a) for the video frames X1, X2, ..., X N and rated split frame S t A measure of ground truth consistency of composed training examples.

4. A method (100) according to any one of claims 2 to 3, wherein the temporal consistency (4) is determined (151) for pixels or other components of the reference truth consistency set (3a).

5. The method (100) according to claim 4, wherein the actual segmented frame Y t The element-wise product with the binary mask describing the actual segmented frame Y is fed (151a) to the check with respect to temporal consistency. t Whether the pixel or other component belongs to the ground truth consistency set (3a).

6. The method (100) according to any one of claims 1 to 5, wherein the temporal consistency (4) is determined (152) as the actual segmented frame Y t or the expected segmented frame The expected segmentation frame Or the actual segmented frame Y t A temporally consistent set of pixels or other components whose corresponding pixels or other components together satisfy a predetermined consistency criterion (4a).

7. The method (100) according to claim 6, wherein the sought evaluation (5) of the machine learning system (1) is evaluated (161) according to the cardinality of the temporal consistency set (4a).

8. The method (100) according to any one of claims 1 to 7, wherein determining (130) the expected segmented frame The method comprises: making the segmented frame Y t-1 deformation.

9. The method (100) according to any one of claims 1 to 8, wherein Assigning (170) the determined evaluation (5) of the machine learning system (1) as a confidence score to the segmented frames Y1, Y2, ..., Y provided by the machine learning system (1) t-1 、Y t , and / or In response to the determined evaluation (5) exceeding a predetermined threshold, the machine learning system (1) is released (180) for use.

10. A method (100) according to any one of claims 1 to 9, wherein the evaluation (5) of the machine learning system is used as feedback (190) to optimize parameters (1a) characterizing the behavior of the machine learning system (1).

11. The method (100) according to claim 10, wherein · Use at least one camera to record video frames X1, X2, ..., X t-1 、X t Feed (200) to the trained machine learning system (1*), From the semantic segmentation frames Y1, Y2, ..., Y immediately provided by the machine learning system (1) t-1 、Y t determining (210) a control signal (210a), and Using the control signal (210a) to control (220) a vehicle (50), a driver assistance system (51), a robot (60), a system for quality control (70), a system for monitoring an area (80) and / or a system for medical imaging (90).

12. A computer program comprising machine-readable instructions which, when executed on one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method (100) according to any one of claims 1 to 11.

13. A machine-readable data carrier and / or a download product having a computer program according to claim 12.

14. One or more computers and / or computing instances having a computer program according to claim 12 and / or having a machine-readable data carrier and / or a download product according to claim 13.