Training neural networks with acquisition of superordinate training contributions

Through a neural network method that processes measurement data in two steps, combined with the design of feature extractors and task heads, the problem that neural networks in the prior art are difficult to adapt to the changes in the measurement data domain or distribution of different machines or industrial facilities is solved, and more efficient training and adaptability are achieved.

CN120069002APending Publication Date: 2025-05-30ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411754947.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-12-02
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing neural networks process measurement data from different machines or industrial facilities, it is difficult to adapt to changes in data domains or distributions, resulting in excessive training consumption and inefficiency.

Method used

A neural network method that processes measurement data in two steps is adopted, first converting the measurement data into a representation in the feature space through a feature extractor, and then the output related to the pre-given task is determined by the task header. For samples that do not belong to the training example domain and/or distribution, the neural network is marked and further trained by obtaining the rated output.

Benefits of technology

Improves the adaptability and efficiency of neural networks when processing measurement data from different machines or industrial facilities, reduces training consumption, and saves computing time and energy through distributed training mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069002A_ABST
    Figure CN120069002A_ABST
Patent Text Reader

Abstract

A neural network is trained with acquisition of superordinate training contributions. The invention relates to a method for training a neural network having a feature extractor for converting measurement data into a representation in a feature space and a task header for determining an output related to a predefined task from the representation, comprising the following steps: processing one or more samples of the measurement data into an output by means of the neural network; evaluating, on the basis of the output, whether the respective sample belongs to the domain and / or distribution of the previously used training examples, on the basis of predefined criteria; if not: acquiring a rated output for one or more samples and marking the samples with the rated output; further training the neural network in a monitored manner using one or more newly labeled samples; checking whether the performance of the further trained neural network has been improved with respect to the predefined task compared to the state prior to the further training on the basis of the test data set with the rated output marking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the training of a neural network, which is used, for example, to evaluate measurement data during quality control or monitoring of industrial facilities. Background Art

[0002] Measurement data generated during quality control of finished products or during monitoring of machines or industrial facilities is generally evaluated using neural networks. These networks have strong generalization capabilities. If the networks encounter various situations during training, they can thus also make statements about situations not seen during training.

[0003] This ability is limited to situations that essentially still belong to the same domain or distribution as the training examples used during training. If the measurement data of a whole fleet of machines or industrial facilities is to be evaluated, there are often variations between the domains or distributions of the measurement data even between similar machines or facilities, especially when the machines or industrial facilities are operated in different ways or for different purposes. A single neural network in a particular training state then does not fully match the measurement data generated in all machines or industrial facilities. In order not to multiply the training effort exactly by the number of machines or industrial facilities, the neural network can, for example, be pre-trained generally and then specifically completed (fine-tuned) for each individual machine or industrial facility. Summary of the Invention

[0004] The present invention provides a method for training a neural network. The neural network has a feature extractor for converting measurement data into a representation in a feature space and a task head for determining an output related to a pre-given task from the representation. Thus, the neural network processes the measurement data in two steps.

[0005] The method assumes a state in which the neural network has either not been trained yet or has been trained using a specific set of training examples. If the neural network has not been trained yet, the parameters characterizing its behavior may, for example, have been randomly initialized. If the neural network has been trained, then the parameters of the neural network can, for example, be trained within the scope of monitoring the training so that the neural network processes the training examples as close as possible to the rated output, using the rated output to label the training examples.

[0006] Within the scope of the method, one or more measurement data samples are processed into outputs using the neural network. From these outputs, it is evaluated whether the corresponding samples belong to the domain and / or distribution of the training examples used so far. If the neural network has just been initialized, the set of training samples used so far is still empty.

[0007] It depends on the corresponding specific application to what extent it can be seen from the output of the neural network whether the samples processed for this output belong to the domain and / or distribution of the training examples used so far. For example, if the task head is constructed to process the representation into classification scores related to one or more categories of a pre-given classification, this can be determined particularly easily. Then, for example, the level of these classification scores provides information on how safe the neural network is when assigning to one or more categories. Therefore, the "One-Hot" classification scores that are 1 for only one category and 0 for all other categories can indicate that the sample belongs to the domain and / or distribution of the training examples and can be reliably assigned to exactly one category.

[0008] However, a classification score distribution that is ambiguous across all categories, has only a weakly expressed maximum value, or no maximum value at all indicates that the sample does not belong to the domain and / or distribution of the training examples and that the neural network is unlikely to process this sample.

[0009] In the latter case, actions need to be taken to supplement the training of the neural network so that it can better process such samples in the future. Therefore, rated outputs are obtained for samples that do not belong to the domain and / or distribution of the training examples used so far, and these samples are labeled with these rated outputs. In this case, the rated outputs can come from any source. For example, domain experts can identify measurement data and / or can conduct further inspections on the objects on which the measurement data are recorded. Therefore, for example, when inspecting the quality of a workpiece, obtaining the rated output can include sawing the workpiece and conducting a metallurgical inspection of its interior.

[0010] Alternatively or in combination with the classification scores, the task head can also determine, for example, the regression value of any searched parameter. If the measurement data relate to, for example, a machine and / or an industrial facility, the searched parameter can, for example, describe a process parameter or a quality indicator (key performance indicator, KPI).

[0011] Optionally, a further distinction can be made between two types of samples, namely samples that are clearly outside the domain and / or distribution of the training examples used so far and samples for which it is a borderline issue whether they belong to this domain and / or distribution. For example, this distinction can be incorporated into the priority for obtaining the rated output for the samples. If obtaining the rated task is complex and expensive, for example, in the case of the metallurgical inspection, assigning such priorities may be particularly meaningful.

[0012] The neural network is further trained in a monitored manner using one or more newly labeled samples. This can in particular include, for example, optimizing the parameters characterizing the behavior of the neural network such that the now newly labeled samples are also each processed into the respective nominal output. A new training level of the neural network is thereby formed. In particular, during further training, for example, a working copy of the neural network can be created such that the network remains consistent and its parameters are not changed, for example, simultaneously by training and by updating the feature extractor at a higher level.

[0013] The performance of the further trained neural network is now checked according to a test data set labeled with nominal outputs to determine whether it has improved in a predefined task compared to its state before further training. The test data set contains test examples that were not used during training and thus the neural network could not "memorize" these test examples during training without actually internalizing (overfitting) the knowledge contained therein.

[0014] When the performance of the trained neural network has improved, the extractor parameters characterizing the behavior of the feature extractor and / or the changes to these parameters are provided as training contributions for the distributed training of the feature extractor.

[0015] It has been recognized that when training a team of neural networks, especially for applications that are not identical but at least similar, the progress made in the field of feature extractors can be generalized across the team. It is therefore advantageous that these advances also benefit other members of the team. In contrast, the progress made during the training of the corresponding task heads is more specific to a particular application and cannot be used or can only be used to a limited extent for other applications.

[0016] It has also been recognized that by checking the progress according to the test data set, the different speeds at which the progress can be fully realized in the corresponding application can also be taken into account. In a team of neural networks all storing measurement data from different sources, some networks can make progress very quickly, while other networks take longer to collect newly labeled samples before making progress.

[0017] One possible reason for this is the different levels of difficulty of the corresponding tasks. If, for example, multiple neural networks can be used for the optical quality control of different workpieces, anomalies on certain types of workpieces can be physically identified with better contrast than on other types of workpieces.

[0018] In addition, the availability of samples with certain characteristics can also affect the speed of training progress. If a neural network, for example, is to learn to recognize certain anomalies, a certain number of samples with exactly these anomalies are required. However, in industrial manufacturing, anomalies are rare events - if this were not the case, the entire manufacturing process would be unusable. Therefore, it may take more or less time to collect a sufficient number of samples with the corresponding anomalies. However, when the corresponding samples are collected later, these samples also benefit other machines or industrial facilities with the same or similar structure in the team by sharing the training contributions of the feature extractor. Then, for example, even without a task head, a general feature extractor can be used for anomaly recognition in such a way that the general feature extractor improves the performance of anomaly recognition based on, for example, statistical outliers in the latent space of the representation or an autoencoder.

[0019] Finally, the availability of computing time for training the neural network can also vary. For example, the hardware platforms used for training can be constructed with varying degrees of goodness. For example, depending on the time of day, the energy costs for training at different locations may also vary.

[0020] It may also happen that the performance of a specific neural network in the team first deteriorates before it improves again if it encounters particular difficulties with newly labeled samples. That is to say, simply put, it must go through a trough before it can later acquire the ability to process new samples. But this is only a matter for that neural network or the node on which it is executed. By only passing the extractor parameters for the distributed training of the feature extractor when the corresponding neural network makes progress, the possible "trough" is kept away from the other parts of the team.

[0021] The advantage of distributed training compared to centralized training that collects all samples from all applications is that it eliminates the work required to transfer samples to a central entity that performs the training. Depending on the type of samples, this may require a very large amount of network bandwidth. At the same time, the measurement data collected in machines and industrial facilities are exactly confidential data. For example, the type and quantity of the manufactured products or even the manufacturing method can be reconstructed based on these data. This is especially a huge problem for centralized data collection that extends across neural networks located in different enterprises. In applications that process personal data, such as images with recognizable individuals or license plates, centralized collection across geopolitical boundaries may also fail due to data protection laws and regulations.

[0022] The distributed training of the feature extractor can be, for example, group training, which is also called "decentralized training". In this case, the different nodes where the neural network is trained directly exchange the extractor parameters and / or the changes to these extractor parameters with each other without a central orchestration entity.

[0023] In a particularly advantageous configuration, samples of measurement data are processed into outputs on client nodes C1, ..., CN for the joint training of the feature extractors. In this case, the feature extractors are used in a state characterized by the extractor parameters received from the server node Q. The determined training contributions are then sent back to the server node Q. The server node Q then combines the training contributions received from all client nodes C1, …, CN into respective new versions of the extractor parameters, which can then be distributed to all client nodes C1, …, CN. In particular, the server node Q can determine when to transmit the new versions of the aggregated extractor parameters to the client nodes C1, …, CN. For example, the client nodes C1, …, CN can adopt the new extractor parameters immediately or after the currently running training has ended.

[0024] It is particularly advantageous to use a separate labeled test data set on each client node C1, ..., CN to check whether the performance of the neural network to be further trained has improved. In this way, the checks can be adapted to the respective applications of the client nodes C1, ..., CN. Thus, in particular, there can be a respective special specification of test examples that the neural network has to master in each application.

[0025] In another particularly advantageous configuration

[0026] · a software implementation of the neural network on at least one client node C1, ..., CN, and / or

[0027] · the hardware platform of the client nodes C1, ..., CN on which the neural network is executed, which is suitable for the measurement data specifically generated at the client nodes C1, ..., CN. In this way, the separately generated measurement data samples can be processed in a particularly efficient manner in each application, so that overall, a large amount of computing time and / or energy can be saved across the board. If, for example, the respective neural network in the application can also be operated reasonably using fixed-point or integer arithmetic, a hardware platform specifically designed for this purpose can be used, which then works significantly more energy-efficiently and faster than a general-purpose (“General Purpose”) hardware platform, such as a CPU or GPU.

[0028] In another particularly advantageous configuration, the further training of at least one client node C1, …, CN is carried out according to a time program which is determined at the location of the client nodes C1, …, CN on the basis of the time dependence of the energy costs and / or environmental impact of this training. This saves energy costs and environmental impact overall. Thus, for example, the electricity price depends highly on the time of day and in particular on the supply of electricity from the sun and wind. If active cooling of the client nodes C1, …, CN is required, the energy required for this can also be taken into account in the time program. For example, at night, when the outside temperature is lowest, the active cooling can be turned down or switched off completely, since passive cooling is sufficient. Overall, this is cheaper than obtaining the cheapest electricity at noon but having to use even more electricity for cooling for this purpose. In general, smaller hardware platforms (such as edge devices in industrial facilities) can operate and be cooled with lower energy consumption than large data centers which are necessary for processing all samples from the team in one go.

[0029] In another particularly advantageous configuration, samples to be newly labeled and / or newly labeled are collected in batches. Further training is then carried out on such batches. This can in particular ensure, for example, that further training is only attempted if the samples collected completely permit the identification of the desired connection. If the neural network receives newly labeled samples in a manner that always serves them individually, the further training will be based only on the information from each individual sample. The connections between the samples and thus the main part of the information recorded by the measurement technique remain unconsidered.

[0030] In particular, for example, the average value of the deviation of the output generated from the newly labeled samples from the corresponding rated output can be used as an indicator of an improvement in the performance of the neural network. Alternatively or in combination herewith, the average value of at least a cost function that evaluates this deviation can be used. If the corresponding indicator drops below a pre-given threshold, it can be determined that the performance of the further trained neural network has improved. Here, the cost function can in particular also measure whether the performance of the neural network is still given for the training examples used so far, or whether “catastrophic forgetting” has occurred here. However, the performance can also be measured, for example, using any measure of accuracy in the context of a pre-given task.

[0031] In another particularly advantageous configuration, during the first stage of further training, the extractor parameters are recorded and only the task head parameters that characterize the behavior of the task head are optimized. The extractor parameters are only optimized together with the task head parameters during the second stage of further training. In this way, it is possible to avoid the neural network from hastily suggesting changes to the extractor parameters that also affect the rest of the team, while changes to the task head parameters that only have a local impact are actually more appropriate. The transition from the first stage to the second stage can be controlled according to any criterion. For example, the transition can be initiated in response to the task head parameters and / or the performance of the neural network no longer changing significantly.

[0032] The measurement data can in particular include, for example, tabular data, time series, images, and / or point clouds. For example, an image can assign values of at least one measurement parameter to the pixels in a regular grid. A point cloud can assign values of at least one measurement parameter to the points in three-dimensional space, and these points do not have to be regularly arranged and do not have to be connected.

[0033] In another particularly advantageous configuration, a neural network that is configured to control and / or monitor a machine and / or an industrial facility is selected. Obtaining the rated outputs includes requesting these rated outputs from the operator of the machine and / or the industrial facility. The distributed training of the feature extractor is extended to other identical or similar machines and / or industrial facilities.

[0034] In this way, the knowledge that exists at the operator of the machine and / or the industrial facility and that may only exist there as unwritten empirical knowledge can be transferred into the neural network and thus into a form that is usable by the machine.

[0035] As mentioned above, in addition to the already mentioned advantages of distributed training, this advantage is particularly important for machines and / or industrial facilities. Exactly machines and / or industrial facilities are often used in large numbers in the same or similar form, but with different configurations and / or for different application purposes. Therefore, there are teams in which the method proposed here can both save the cost for training and achieve overall better training results due to the cooperation of many machines and / or industrial facilities.

[0036] The ultimate goal of training is to obtain better tools for evaluating measurement data in a given application using the trained neural network. Therefore, in another particularly advantageous configuration, the measurement data is fed into the neural network that has completed training. A control signal is determined based on the output generated by the trained neural network. The control signal is used to control a vehicle, a driver assistance system, a robot, a quality control system, a machine, and / or an industrial facility. In this way, the probability that the reaction performed by the separately controlled system in response to the control signal is suitable for the reaction represented by the measurement data is increased.

[0037] The method can be, in particular, fully or partially computer-implemented. Accordingly, the present invention also relates to a computer program having machine-readable instructions which, when executed on one or more computers and / or computing instances, cause the computer and / or computing instance to carry out the described method. In this sense, vehicle control devices and technical device embedded systems which are also capable of executing machine-readable instructions can also be regarded as computers. Computing instances can be, for example, virtual machines, containers or serverless execution environments, which can in particular be provided in the cloud.

[0038] The present invention also relates equally to a machine-readable data carrier and / or a download product having a computer program. A download product is a digital product which can be transmitted via a data network, i.e. can be downloaded by a user of the data network, and which can be sold, for example, in an online store for immediate download.

[0039] Furthermore, one or more computers and / or computing instances can be equipped with a computer program, a machine-readable data carrier or a download product.

[0040] Other measures for improving the present invention are shown in more detail below together with the description of the preferred embodiments of the present invention with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1a and 1b show an embodiment of a method 100 for training a neural network 1;

[0042] Figure 2 show an exemplary training of only the task head according to block 143 using the first acquired rated output 6a;

[0043] Figure 3 show an exemplary semi-supervised training of the complete neural network 1 according to block 144 using the training contribution 2a passed to the feature extractor 2. DETAILED DESCRIPTION

[0044] Figure 1a and Figure 1b show a schematic flow chart of an embodiment of a method 100 for training a neural network 1, wherein method 100 is shown up to step 160 and continues with the method steps shown in Figure 1a and continues with the method steps shown in Figure 1b The flow chart is divided into Figure 1a and Figure 1b in order to make the representation of the flow chart clearer. In the representation of Figure 1a , the method steps explicitly listed in Figure 1b are indicated by the description of Figure 1b . Similarly, at the start of the flow chart in Figure 1b is indicated Figure 1a。The neural network 1 has a feature extractor 2 and a task head 3. The feature extractor 2 is used to transfer the measurement data 4 into a representation 5 in the feature space, and the task head 3 is used to determine an output 6 related to a pre-given task according to the representation 5.

[0045] According to block 105, in particular, for example, the neural network 1 can be selected, which is configured to control and / or monitor a machine and / or an industrial facility.

[0046] In step 110, one or more samples 4a of the measurement data 4 are processed into an output 6 by means of the neural network. This process is guided internally by the representation 5 in the feature space.

[0047] According to block 111, the samples 4a of the measurement data 4 can be processed into an output 6 on the client nodes C1,..., CN for the joint training of the feature extractor 2. Then according to block 112, the feature extractor 2 can be used in a state characterized by the extractor parameters 2a received from the server node Q.

[0048] In step 120, it is evaluated according to a pre-given criterion whether the corresponding sample 4a belongs to the domain and / or distribution of the training examples used so far from the obtained output 6. If this is not the case (truth value is 0), the training of the neural network 1 is updated accordingly.

[0049] For this purpose, in step 130, a rated output 6a is obtained for one or more samples 4a, and these samples 4a are labeled with these rated outputs 6a.

[0050] According to block 131, in particular, for example, the samples 4a to be newly labeled and / or newly labeled can be collected in a batch.

[0051] According to block 132, the obtaining of the rated output 6a can in particular, for example, include requesting these rated outputs 6a from the operator of the machine and / or the industrial facility, and configuring the neural network 1 according to block 105 for its control and / or monitoring.

[0052] In step 140, the neural network 1 is further trained in a monitored manner using one or more newly labeled samples 4a. The neural network 1 is then in a further trained state 1'.

[0053] According to block 141, this further training of at least one of the client nodes C1,..., CN can be carried out according to a time program, which is determined at the locations of the client nodes C1,..., CN based on the time dependence of the energy cost and / or environmental impact of this training.

[0054] According to block 142, the further training can be performed on the batch collected according to block 131.

[0055] According to block 143, during the first stage of further training, the extractor parameters 2a can be recorded and only the task head parameters 3a that characterize the behavior of the task head 3 can be optimized. Then, according to block 144, the extractor parameters 2a can be optimized together with the task head parameters 3a during the second stage of further training.

[0056] In step 150, it is checked according to the test dataset 7 marked with the rated output whether the performance of the further trained neural network 1' in a pre-given task has improved compared to the state before further training.

[0057] According to block 151, a separate marked test dataset 7 can be used for this test on each client node C1,..., CN.

[0058] According to block 152, in response to

[0059] · the average value of the deviation between the output 6 generated from the newly marked sample 4a and the corresponding rated output 6a and / or

[0060] · at least the average value of the cost function evaluating this deviation

[0061] being lower than a pre-given threshold, it can be determined that the performance of the further trained neural network 1' has improved.

[0062] If the performance of the further trained neural network 1' does not improve further (the truth value is 0 at step 150), then according to block 131, it can return to collect a new batch of samples for further training.

[0063] Whereas if the performance of the further trained neural network 1' improves further (the truth value is 1 at step 150), then in step 160, the extractor parameters 2a that characterize the behavior of the feature extractor 2 and / or the changes in these parameters 2a are provided as the training contribution for the distributed training of the feature extractor 2. These extractor parameters 2a can also be used again during the further processing of the sample 4a in step 110.

[0064] According to block 161, in the case of joint training according to blocks 111 and 112, the determined training contribution 2a can be transmitted to the server node Q.

[0065] According to block 162, when applied to a machine and / or an industrial facility (blocks 105 and 132), the distributed training of the feature extractor 2 can be extended to other identical or similar machines and / or industrial facilities.

[0066] When any abort condition is reached, the training can end. The optimized states of the extractor parameters 2a of the feature extractor 2 or the task head parameters 3a of the task head 3 are denoted by the reference signs 2a* and 3a*, respectively. These states 2a* and 3a* determine the trained state 1* of the neural network 1.

[0067] In Figure 1a and 1b In the example shown, the measurement data 4 is provided to the trained neural network 1* in step 170. In step 180, a control signal 180a is determined based on the output 6 generated by the trained neural network 1*. In step 190, the vehicle 50, the driver assistance system 51, the robot 60, the system 70 for quality control, the machine 80, and / or the industrial facility 90 are controlled using the control signal 180a.

[0068] Figure 2 Shows how the training of the task head 3 can be started according to block 142 of method 100 in a state where the rated output 6a of the sample 4a of the measurement data 4 is not yet available.

[0069] In Figure 2 In the example shown, the server node Q coordinates joint training with three client nodes C1, C2, and C3. Joint training means the feature extractor 2, the parameters 2a of which are passed from the server node Q to all client nodes C1, C2, and C3.

[0070] For better readability, Figure 2 In is shown enlarged the process running on the client node C1. However, the process runs completely analogously on the other client nodes C2 and C3.

[0071] Since the rated outputs 6a of the samples 4a of the measurement data 4 are initially not available, these rated outputs are initially obtained according to step 130 of method 100, here: by the operator. Then the labeled samples 4a + 6a are used in step 140 of method 100 to train the task head 3, while the feature extractor 2 remains unchanged. That is, according to block 142, only the task head parameters 3a are optimized. Therefore, no training contributions to the joint training of the feature extractor 2 are also reported back to the server node Q yet.

[0072] Figure 3 Shows how the task head 3 can be trained together with the feature extractor 2 immediately afterwards according to block 143.

[0073] With Figure 2Differently, sample 4a is now fed to neural network 1 in its final training state, resulting in output 6. Based on these outputs 6, a decision is made, according to step 120 (not shown for clarity), to label sample 4a with the rated output 6a obtained according to step 130. Using the sample 4a+6a labeled in this way, then in step 140 not only the task head 3 but also the feature extractor 2 is trained. This means that both the extractor parameters 2a and the task head parameters 3a are optimized. This process takes place in the working copy of neural network 1. A further training state 1′ is formed. After checking (not drawn for clarity) according to step 150 whether the performance of neural network 1 has improved, the further training state 1′ is used as the new training state for testing new samples 4a. At the same time, the extractor parameters 2a are provided as a training contribution to the joint training of the feature extractor 2 according to step 160.

Claims

1. A method (100) for training a neural network (1) having a feature extractor (2) for converting measurement data (4) into a representation (5) in a feature space and a task head (3) for determining an output (6) associated with a predetermined task based on the representation (5), the method comprising the following steps: Processing one or more samples (4a) of the measurement data (4) into outputs (6) (110) using a neural network: Based on these outputs (6), an evaluation (120) is made based on predefined criteria: whether the corresponding sample (4a) belongs to the domain and / or distribution of the training examples used so far; and if this is not the case: Obtaining rated outputs (6a) for one or more samples (4a) and marking (130) the samples (4a) with the rated outputs (6a); Further training the neural network (1) (140) in a supervised manner using one or more newly labeled samples (4a); Using a test data set (7) labeled with a nominal output, a check (150) is performed to determine whether the performance of the further trained neural network (1') has improved on a predetermined task compared to the state before the further training; if this is the case: • Extractor parameters (2a) characterizing the behavior of the feature extractor (2) and / or changes to these parameters (2a) are provided as training contributions (160) for the distributed training of the feature extractor (2).

2. The method (100) according to claim 1, wherein A sample (4a) of the measurement data (4) is processed into an output (6) (111) for joint training of the feature extractor (2) on the client nodes C1, ..., CN; In this case the feature extractor (2) is used (112) in a state characterized by the extractor parameters (2a) received from the server node Q, and • The determined training contribution (2a) is transmitted to the server node Q (161).

3. The method (100) according to claim 2, wherein a separate labeled test data set (7) (151) is used on each client node C1, ..., CN in order to check whether the performance of the further trained neural network (1') is improved.

4. The method (100) according to any one of claims 2 to 3, wherein A software implementation of the neural network (1) on at least one client node C1, ..., CN, and / or The hardware platform of the client node C1, ..., CN, on which the neural network (1) is executed, is adapted to the measurement data (4) generated specifically at the client node C1, ..., CN.

5. A method (100) according to any one of claims 1 to 4, wherein further training (141) of at least one client node C1, ..., CN is performed according to a time program, wherein the time program is determined at the location of the client node C1, ..., CN based on the time dependency of the energy cost and / or environmental influence of the training.

6. The method (100) according to any one of claims 1 to 5, wherein Collecting the samples to be newly labeled and / or newly labeled in a batch (4a)(131), and Such batches are further trained (142).

7. The method (100) of claim 6, wherein in response to the performance of the neural network (1') not improving upon further training, additional samples (4a) are collected in the batch.

8. The method (100) according to any one of claims 1 to 7, wherein the task head (3) is configured to process the representation (5) into a classification score associated with one or more categories of a predetermined classification.

9. The method (100) according to any one of claims 1 to 8, wherein in response to the average value of the deviations of the outputs (6) generated from the newly labeled samples (4a) from the corresponding nominal outputs (6a), and / or the average of the cost functions evaluating at least this deviation, Below a predetermined threshold value, it is determined (152) that the performance of the further trained neural network (1') has improved.

10. The method (100) according to any one of claims 1 to 9, wherein In a first phase of further training, the extractor parameters (2a) are recorded and only the task head parameters (3a) (143) that characterize the behavior of the task head (3) are optimized and In a second phase of further training, the extractor parameters (2a) are optimized (144) together with the task head parameters (3a).

11. The method according to any one of claims 1 to 10, wherein the measurement data comprises tabular data, time series, images and / or point clouds.

12. The method according to any one of claims 1 to 11, wherein Selecting a neural network (1) (105) configured to control and / or monitor a machine and / or an industrial facility; Obtaining rated outputs (6a) includes requesting these rated outputs (6a) (132) from operators of machines and / or industrial facilities; and Distributed training of the feature extractor (2) is extended to other identical or similar machines and / or industrial facilities (162).

13. The method (100) according to any one of claims 1 to 12, wherein The measurement data (4) is fed to the trained neural network (1*) (170); determining a control signal (180a) (180) based on the output (6) generated by the trained neural network (1*), and · Using the control signal (180a) to control a vehicle (50), a driving assistance system (51), a robot (60), a quality control system (70), a machine (80) and / or an industrial facility (90) (190).

14. A computer program comprising machine-readable instructions which, when executed on one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method (100) according to any one of claims 1 to 13.

15. Machine-readable data carrier and / or download product having a computer program according to claim 14.

16. One or more computers and / or computing instances having a computer program according to claim 14 and / or having a machine-readable data carrier and / or a download product according to claim 15.