Pseudo-marking based on uncertainty or probability

By using the self-labeling process to generate pseudo-labeling at the edge deployment agent, the problem of performance degradation in machine learning models in edge deployment is solved, achieving more accurate model training and prediction.

CN120046679APending Publication Date: 2025-05-27INFINEON TECHNOLOGIES AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411689531.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

At the agent for edge deployment, the performance of machine learning models tends to decline due to mismatch between the scenarios considered in the initial training and the real-world deployment, resulting in reduced prediction accuracy.

Method used

The self-labeling process is used to determine pseudo-labels at the edge, avoiding creating inaccurate pseudo-labels by taking into account the uncertainty and probability of predictions provided by the ML model, and populating the training dataset used to retrain the ML model.

Benefits of technology

By automatically generating high-reliability pseudo-labels at the edges, a more comprehensive training data set can be built, reducing prediction errors, and improving the accuracy of the second training state of the ML model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046679A_ABST
    Figure CN120046679A_ABST
Patent Text Reader

Abstract

The invention relates to pseudo-marking based on uncertainty or probability. Techniques for retraining a machine learning model, such as a deep neural network, at an agent deployed on site, i.e., at an edge, are disclosed. Retraining is performed in consideration of multiple training data sets to distinguish between tags and false tags. Techniques for determining false tags based on uncertainty and / or probability are disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of the present disclosure generally relate to retraining machine learning models. Aspects of the present disclosure specifically relate to retraining at the edge. Aspects relate to a self-labeling process for determining pseudo-labels for populating a training data set used in the above retraining. Background Art

[0002] Machine learning (ML) models (such as deep neural networks) are typically used to make inferences based on measurement data obtained from sensors (such as radar sensors). Such inferences can be performed at an agent deployed at the edge. "Edge deployment" means that the agent is not a central institution implemented, for example, by a remote central server; rather, it is a field device that obtains measurement data on-site and is generally controlled to some extent by a central institution. Typically, a central institution will deploy an ML model in an initial training state to multiple agents.

[0003] According to a reference implementation, in the development phase, large-scale measurement activities are carried out to establish a training data set at the central institution, which is intended to densely sample the entire input space expected on-site, that is, to have good coverage under a wide range of scenarios and users. In the measurement activities, sensors of the same type as those later deployed on-site are used to obtain feature vectors for the training data set.

[0004] To enable training, ground truth is obtained to determine the labels forming the output samples; the output samples are paired with input samples (i.e., measurement feature vectors).

[0005] There are multiple options available for determining ground truth. For example, manual annotation by experts is an option. Alternatively or additionally, alternative sensors such as cameras can be used to collect images based on which ground truth can be reliably determined. Such alternative sensors are not available in field-deployed agents. This is called supervised learning.

[0006] Based on a training data set including input sample and output sample pairs, the ML model is trained and then deployed to agents for inference.

[0007] Despite the initial efforts for accurate data collection and labeling in the central training process, due to the mismatch between the scenarios considered in training and real-world deployment, the performance on-site often degrades. Summary of the Invention

[0008] Therefore, advanced techniques of ML models are needed to infer predictions based on measurement feature vectors, which are determined based on measurement data from sensors such as radar sensors or other depth sensors. Techniques are needed to mitigate at least some of the limitations and disadvantages identified above.

[0009] The features of the independent claims meet this need. The features of the dependent claims define embodiments.

[0010] Hereinafter, a self-labeling process is disclosed. The disclosed self-labeling process can determine pseudo-labels, which can be used as output samples in the training process for (re-)training an ML model such as a deep neural network for solving a classification task. It is particularly possible (but not mandatory) to use this self-labeling process at the edge (i.e., to execute the self-labeling process at a field-deployed agent).

[0011] The self-labeling process disclosed herein takes into account the uncertainty of the predictions provided by the ML model. Alternatively or additionally, the self-labeling process disclosed herein takes into account the probability (e.g., class probability) of the predictions provided by the ML model. By taking into account the uncertainty and / or probability, the creation of inaccurate pseudo-labels can be avoided. Pseudo-labels with higher reliability can be determined.

[0012] A computer-implemented method includes: using an ML model in a first training state to infer a prediction based on a plurality of measurement feature vectors. The method further includes: filling a first training dataset based on the prediction. Filling the first training dataset with labels of a first subset of the plurality of measurement feature vectors. The method further includes: filling a second training dataset different from the first training dataset based on the prediction. Filling the second training dataset with pseudo-labels for a second subset of the plurality of measurement feature vectors. The method further includes: determining a second training state of the ML model based on the first training dataset and the second training dataset.

[0013] For example, the method can be executed by an agent. Retraining at the edge is possible. The agent can be deployed in the field. The method can further include: obtaining an ML model in a first training state from a central institution. The method can also be executed at least partially at the central institution.

[0014] If the uncertainty of the prediction is lower than a predefined threshold, pseudo-labels can be selectively determined. Such uncertainty can be determined using dropout sampling and / or ensemble sampling and / or a probabilistic ML model and / or the evidence distribution of the corresponding prediction.

[0015] By taking into account that the uncertainty is small enough, unreliable pseudo-labels can be avoided. The data quality of the second training dataset is improved. Therefore, the second training state of the ML model will provide more accurate results.

[0016] A computer-implemented method for a central agency includes: providing a machine learning model in a first training state to a plurality of agents. The computer-implemented method further includes: obtaining, from each of the plurality of agents, information indicating a machine learning model in a second training state. The computer-implemented method further includes: determining a third training state of the machine learning model based on the plurality of second training states.

[0017] For example, the method may further include: obtaining, from each of the plurality of agents, context information for the machine learning model in the respective second training state. The above determination of the third training state of the machine learning model may also be based on the context information. Such context information may indicate the size of a first training dataset that is used for determining the second training state and filled with labels. Alternatively or additionally, such context information may indicate the size of a second training dataset that is also used for determining the second training state and filled with pseudo-labels. Such context information may also indicate the relative size of the first training dataset with respect to the second training dataset.

[0018] For example, a third training dataset may be determined based on a weighted combination of the weights of the ML models associated with each of the plurality of second training states. Then, a second training state having a higher influence (based on labels rather than pseudo-labels) of the corresponding first training dataset may be considered more prominent compared to such a second training state that is more affected (based on pseudo-labels rather than labels) by the corresponding second training state.

[0019] In addition, such second training states obtained from larger training datasets may be mainly considered.

[0020] An agent includes computing circuitry, such as a processor and a memory. The computing circuitry is configured to obtain a machine learning model in a first training state from a central agency. The computing circuitry is further configured to use the machine learning model in the first training state to infer a prediction based on a plurality of measurement factors obtained at the agent. The computing circuitry is further configured to fill a first training dataset with labels of a first subset of the plurality of measurement feature vectors based on the prediction. The computing circuitry is further configured to fill a second training dataset with pseudo-labels of a second subset of the plurality of measurement feature vectors based on the prediction. The computing circuitry is further configured to determine a second training state of the machine learning model based on the first training dataset and the second training dataset.

[0021] A central authority includes computing circuitry, such as a processor and a memory, which stores program code that can be loaded and executed by the processor. The computing circuitry is configured to provide a machine learning model in a first training state to a plurality of agents. The computing circuitry is further configured to obtain, from each of the plurality of agents, information indicating a machine learning model in a corresponding second training state. The computing circuitry is further configured to determine a third training state of the machine learning model based on the plurality of second training states.

[0022] As described above, a system includes a central authority and one or more agents.

[0023] It should be understood that, without departing from the scope of the present invention, the above features and the features not yet explained below can be used not only in the corresponding combinations indicated, but also in other combinations or in isolation. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematically illustrates agents according to various examples;

[0025] Figure 2 Schematically illustrates radar sensors according to various examples;

[0026] Figure 3 Schematically illustrates a plurality of gesture classes that can be predicted according to various examples;

[0027] Figure 4 Schematically illustrates a system including a central authority implemented by a central server and a plurality of agents according to various examples; and

[0028] Figure 5 is a flowchart of a method according to various examples. DETAILED DESCRIPTION

[0029] Some examples of the present disclosure generally provide multiple circuits or other electrical devices. All references to circuits and other electrical devices and the functionality provided by each circuit and device are not intended to be limited to only the content shown and described herein. Although specific labels may be assigned to the various circuits or other electrical devices disclosed, such labels are not intended to limit the scope of operation of the circuits and other electrical devices. Such circuits and other electrical devices may be combined and / or separated from each other in any manner based on the particular type of electrical implementation desired. It should be recognized that any circuit or other electrical device disclosed herein may include any number of microcontrollers, graphics processing unit (GPU), integrated circuits, memory devices (such as FLASH, random access memory (RAM), read only memory (ROM), electrically programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM) or other suitable variants thereof) and software that cooperate with each other to perform the operations disclosed herein. In addition, any one or more electrical devices may be configured to execute program code embodied in a non-transitory computer-readable medium, the program code being programmed to perform any number of the functions disclosed.

[0030] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be understood that the following description of the embodiments should not be considered restrictive. The scope of the present invention is not intended to be limited by the embodiments or the drawings described below, and the drawings are for illustrative purposes only.

[0031] The accompanying drawings should be regarded as schematic, and the elements shown in the drawings are not necessarily shown to scale. The various elements are rather represented such that their functionality and general purpose become clear to those skilled in the art. Any connection or coupling between the functional blocks, devices, components or other physical or functional units shown in the figures or described herein may also be implemented by an indirect connection or coupling. The coupling between components may also be established by a wireless connection. The functional blocks may be implemented in hardware, firmware, software or a combination thereof.

[0032] Hereinafter, aspects associated with ML models are disclosed. Various kinds and types of ML models may benefit from the techniques disclosed herein. Examples include classification and regression tasks performed by ML models. Examples of ML models include deep neural networks such as convolutional neural networks, support vector machines, multi-layer perceptrons, and the like.

[0033] ML models can be trained to perform different tasks. Example tasks include, for example, gesture class prediction based on a measurement feature vector obtained from radar measurements. Other example tasks include, for example, people counting based on a measurement feature vector obtained from radar measurements. Another task is, for example, people presence detection inside a vehicle. Other tasks include image processing, for example, to detect certain structures or anomalies in a two-dimensional (2-D) image, such as for wafer inspection, medical image processing, etc.

[0034] Depending on the specific task or use case for which the ML model is employed, the type of measurement data on which the measurement feature vector, which is the input sample to be input into the ML model, is based can also vary widely. Various measurement modalities and sensors can be used to obtain the measurement data. Depth sensors can be used to obtain the measurement data, such as optical time-of-flight (ToF) sensors (e.g., time-of-flight cameras or light detection and ranging (LIDAR) sensors). Another type of depth sensor is a radar sensor, such as a millimeter-wave or ultra-wideband radar sensor. A time-of-flight (ToF) camera performs ranging measurements based on the duration of a light pulse's round trip to a scene. ToF cameras typically use light-emitting diodes as the light source. LIDAR sensors use laser diodes or lasers as the light source. LIDAR sensors can use ToF ranging or the measurement of the interference of the continuous wave of the primary light and the continuous wave of the reflected secondary light. The light sources of ToF cameras and LIDAR are typically in the infrared or deep infrared region. 2-D images can be obtained using cameras or medical imaging devices (e.g., magnetic resonance imaging devices or computed tomography devices).

[0035] The various techniques disclosed herein will focus on use cases for determining a measurement feature vector based on measurement data obtained using radar sensors; however, similar techniques can also be applied to other types of sensors.

[0036] In the following, techniques for retraining an ML model are disclosed. Retraining means updating the first training state of the ML model to obtain a second training state. That is, the weights of the ML model are updated or specifically optimized.

[0037] Next, the training process for training or retraining an ML model is explained. During the training process, the input samples included in the training dataset are input into an ML model (e.g., a deep neural network), which processes the input samples through its hierarchical structure to generate predictions (forward propagation). The predictions are compared with the output samples associated with the input samples in the training dataset. The output samples can be labels or pseudo-labels. A loss function is used to perform this comparison, and the loss function is a measure that quantifies the prediction error. Then, backpropagation is employed to calculate the gradient of the loss function with respect to each weight, essentially measuring how a change in the weight will affect the error. An optimization algorithm (usually one of stochastic gradient descent or its variants) uses the gradients provided by backpropagation to adjust the weights and biases of the network. This process is repeated for multiple epochs for each entry (i.e., each pair of input sample and output sample) in the training dataset.

[0038] Various publicly available examples are based on the discovery that an accurate ML model needs to be trained on a large training dataset (i.e., a training dataset that includes a large number of input-output sample pairs). The training dataset needs to sample the entire input space of the expected input samples with sufficient density. Only in this way can the ML model accurately generalize to the input samples encountered in the field. For example, an ML model can be used to make predictions about the number of people present in a scene based on measurement feature vectors obtained using a radar sensor. To provide reliable person-counting predictions, the ML model should be trained on measurement feature vectors (here implementing the input samples) obtained for people in different positions in front of the radar sensor, different environments, different backgrounds, different people, different people's movement patterns, etc. Using a reference implementation to obtain the corresponding input sample-output sample pairs to populate the training dataset may require a large amount of time and effort. It requires a data acquisition laboratory, the availability of many users, and a significant investment of time.

[0039] To alleviate these problems, retraining can be performed at the edge. That is, the ML model can be deployed from a central agency to multiple field deployment agents, and then the agents can execute the training process to retrain the ML model based on local agent-specific training datasets. These agent-specific training datasets are populated locally. That is, the training dataset is populated based on the input sample-output sample pairs obtained at the agent.

[0040] By populating the training dataset based on the input samples and output sample pairs determined at the agent, the distribution of the input samples observed in the field in the corresponding input space can be captured. For example, certain parts of the input space that were only sparsely sampled in a central measurement activity can be densely sampled based on the measurement feature vectors and the relevant predictions of the ML model observed in the field at the agent. Therefore, this retraining at the edge can train the ML model more accurately.

[0041] However, the availability of ground truth may be limited at the agent. Typically, ground truth will be derived from user input (such as confirming a certain prediction made by an ML model or providing a correction to a prediction). In this case, a label can be determined based on the ground truth. However, the ground truth may only be sparsely available at the agent.

[0042] To mitigate this situation, according to various aspects, a self-labeling process is disclosed that is capable of determining reliable pseudo-labels in an automated manner. The pseudo-labels can form output samples and are added to the training dataset together with the corresponding input samples formed by the measured feature vectors.

[0043] A pseudo-label is an output sample for which the ground truth is not available. Thus, pseudo-labels are different from labels. A label is an output sample for which the ground truth is available. For example, even at the agent, the ground truth can be obtained from a user interaction process (such as where a verified user provides reliable results). For example, a user (such as an authenticated user) can directly or indirectly confirm a prediction inferred by an ML model. The user can change or correct a prediction inferred by the ML model. These processes can generate ground truth and thus generate labels. Since pseudo-labels do not require ground truth, pseudo-labels can be automatically generated based on the predictions of the ML model (self-labeling process). Thus, there is a tendency to have more pseudo-labels available compared to labels. Labels are typically associated with a relatively high reliability. Pseudo-labels tend to have a relatively low reliability because the ground truth is not available for determining pseudo-labels. Nevertheless, techniques for obtaining still relatively reliable pseudo-labels are disclosed.

[0044] By generating pseudo-labels at the in-field deployed agent, a more comprehensive training dataset can be constructed, sampling the input space comprehensively in the regions where actual input samples are observed. This enables tracking of distribution drift, such as where the input samples encountered by the agent are in regions of the input space that were not densely sampled in the training dataset used for initial training at the central institution. This helps reduce the error rate observed when making predictions by the ML model.

[0045] To reliably determine pseudo-labels, the uncertainty of the predictions of an ML model can be considered through a self-labeling process. The uncertainty can be achieved through at least one of aleatoric uncertainty or epistemic uncertainty. Generally, aleatoric uncertainty (also known as statistical uncertainty) refers to the concept of randomness, i.e., the variability of prediction results, which is caused by random effects inherent in the measured feature vectors. Example sources of aleatoric uncertainty are measurement or data noise, drift in sensor operation, etc. Aleatoric uncertainty is different from epistemic uncertainty (also known as systematic uncertainty). Epistemic uncertainty refers to the uncertainty caused by the limitations of the ML model (such as limited model accuracy, etc.). Thus, epistemic uncertainty can stem from the incomplete sampling of the input space with training data during the training process of the ML model. For example, consider an ML model trained to distinguish cats and dogs; however, the training data only captures certain dog breeds (such as Swiss St. Bernards and German Shepherds), while other dog breeds (such as Mexican Chihuahuas) are not captured. Then, if cats and dogs are to be distinguished based on pictures of dog breeds (such as Chihuahuas) not captured by the trained model, the epistemic uncertainty will increase. In contrast to aleatoric uncertainty, epistemic uncertainty can be reduced by retraining the ML model.

[0046] As an alternative or supplement to the self-labeling process that considers the uncertainty of predictions, the self-labeling process can consider the probability of predictions, e.g., if the ML model is a classifier, the class probability associated with the prediction; or if the ML model is a regression model (where variance corresponds to uncertainty), the Gaussian probability distribution associated with the prediction.

[0047] The predicted probability is different from the predicted uncertainty. Probability measures the likelihood or confidence in the inferred result, while uncertainty characterizes the range of possible outcomes due to incomplete information. For example, the class probabilities for multiple classes distinguished by a deep neural network (implemented as an example of an ML model) can be obtained from the softmax activation function at the output layer of the deep neural network. Softmax is just one example of an activation function. Other activation functions can be used at the output layer of the deep neural network, and the class probabilities for multiple classes can also be obtained from other types of activation functions. To give a specific example of the difference between probability and uncertainty: A deep neural network can be trained to classify an image as depicting a "dog" or a "cat". When the deep neural network processes an image that depicts only a dog, the corresponding class probability for the "dog" class will be high, while the class probability for the "cat" class will be low. The uncertainty may also be relatively low, unless the image is noisy or includes a strange perspective. On the other hand, if the deep neural network processes an image that shows both a dog and a cat, the class probability for the "dog" class will be approximately 50%, while the class probability for the "cat" class will be approximately 50%. This uncertainty is not affected by the picture showing both a cat and a dog; that is, it may also be relatively low, unless the image is noisy or includes a strange perspective.

[0048] According to the example, such aspects of the self-labeling process for determining pseudo-labels for filling a training dataset at an in-field deployment agent can be combined with techniques of federated learning. Federated learning helps to retrain an ML model at multiple decentralized agents (also called clients or in-field deployment devices) based on their local training datasets; while ensuring that the training datasets remain distributed and private. Federated learning does not transfer the training datasets from the agents to a central institution for training, but allows the ML model to be retrained locally at each agent, where only the ML model updates (weights) are sent back to the central institution. The central institution can then combine the updated weights obtained from multiple agents.

[0049] By inputting a measure of uncertainty and / or a measure of the predicted probability into the self-labeling process, unreliable pseudo-labels can be detected and not added to the training dataset, i.e., they will not be considered in subsequent retraining processes. This helps to reduce the potential negative impact of mislabeled samples and ensures better utilization of output samples for which the ground truth is not available.

[0050] According to various disclosed examples, the training process for a training dataset to be populated based on labels and / or pseudo-labels respectively can be flexibly implemented. In particular, on the one hand, a separate training dataset can be maintained for output samples associated with labels, and on the other hand, a separate training dataset can be maintained for output samples associated with pseudo-labels. Thus, the training process can be flexibly triggered in all of the following cases: only labels are available; only pseudo-labels are available; a combination of labels and pseudo-labels is available.

[0051] The disclosed technology can be combined with the concept of federated learning. Here, multiple agents can report their respective updated agent-specific training states of an ML model to a central agency, and then the central agency can merge these agent-specific training states to obtain an updated merged training state of the ML model.

[0052] Figure 1 Agent 65 is schematically illustrated. Agent 65 is a device that can be deployed on-site. Thus, Agent 65 can be referred to as an edge-deployed agent 65. For example, Agent 65 can be a people counting statistics system, a gesture recognition system, a medical imaging device, a surveillance system, etc.

[0053] Agent 65 includes a sensor 70 and a processing device 60. Hereinafter, sensor 70 is exemplified in a non-limiting manner as a radar sensor and will be so referred to, but sensor 70 can be any sensor (such as a depth sensor, a camera, etc.).

[0054] The processing device 60 can obtain measurement data 64 from the radar sensor 70. A processor 62 (such as a general-purpose processor (central processing unit CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC)) can receive the measurement data 64 via an interface 61 and process the measurement data 64. For example, the measurement data 64 can include data frames that include samples of an analog-to-digital converter. Further preprocessing can also be implemented at the radar sensor 70; for example, the radar sensor 70 can output a 2-D spectrogram, such as a range-Doppler spectrogram, or an azimuth-elevation spectrogram, or a range-time spectrogram, or a Doppler-time spectrogram, or an azimuth-time spectrogram, or an elevation-time spectrogram. Based on this preprocessing, a measurement feature vector (hereinafter simply referred to as a feature vector) encoding the measurement data can be determined.

[0055] For example, typical feature vectors determined based on the measurement data of the radar sensor 70 include one or more of the following dimensions: the range of a gesture object; the speed of a gesture object (sometimes also referred to as the Doppler shift); the angular orientation of a gesture object; the azimuth of a gesture object; and the elevation of a gesture object.

[0056] Figure 2FIG. illustrates aspects related to a radar sensor 70. The radar sensor 70 includes a processor 72 (labeled as a digital signal processor DSP) coupled to a memory 73. Based on program code stored in the memory 73, the processor 72 can perform various functions related to transmitting radar pulses 86 using a transmit antenna 77 and a digital-to-analog converter (DAC) 75 or a voltage-controlled oscillator. After the radar pulse 86 is reflected by the scene 80, the processor 72 can detect the corresponding reflected radar pulses 87 (e.g., sorted in an L-shape at a half-wavelength distance; see Figure 2 illustration) using an ADC 76 and a plurality of receive antennas 78-1, 78-2, 78-3. The processor 72 can process the raw data samples obtained from the ADC 76 to a greater or lesser extent. For example, a data frame can be determined and output. In addition, a spectrogram can also be determined.

[0057] The radar measurement can be implemented based on the fundamental frequency-modulated continuous wave (FMCW) principle. A frequency chirp can be used to implement the radar pulse 86. The frequency of the chirp can be adjusted between a frequency range of 57 GHz to 64 GHz. The transmitted signal is backscattered and has a time delay corresponding to the distance of the reflecting object captured by all three receive antennas. Then, the received signal is mixed with the transmitted signal and then low-pass filtered to obtain an intermediate signal. The frequency of this signal is significantly lower than the frequency of the transmitted signal, and thus the sampling rate of the ADC 76 can be reduced accordingly. The ADC can operate at a sampling frequency of 2 MHz and an accuracy of 12 bits.

[0058] As shown in the figure, the scene 80 includes a plurality of objects 81-83. For example, the objects 81, 82 can correspond to the background, while the object 83 can belong to the user's hand. Based on the radar measurement, the gestures performed by the hand can be classified. Figure 3 Some gesture classes that can be distinguished are illustrated in

[0059] Figure 3 Schematically illustrates the corresponding labels of such gestures 501-510 and gesture classes 520, but other gestures are also possible. According to the techniques described herein, the gesture class prediction of such gestures 501-510 can be reliably inferred.

[0060] Such gesture class inference can be implemented on an edge-deployed processor (such as the processor 62 of the proxy 65 (see Figure 1 ). However, as will be explained in connection with Figure 4 , the central server can also be part of the overall data processing.

[0061] Figure 4Schematically illustrates system 130. System 130 includes a central server 139 that implements the functionality of a central authority. For example, central server 139 can be located in a server farm. Central server 139 can be an Internet server connected to the Internet.

[0062] Central server 139 includes a processor 311 coupled to a memory 312. Processor 311 can load program code from memory 312 and execute the program code to perform the techniques disclosed herein. Processor 311 is also coupled to a communication interface 313. Processor 311 can transmit messages or receive messages via the communication interface. For example, the server can be coupled to the Internet via interface 313.

[0063] System 130 also includes a plurality of agents 131 - 134, which can also be referred to as clients. Each of agents 131 - 134 can be implemented according to agent 65 discussed in conjunction with Figure 1 Agents 131 - 134 can be automotive control units (sometimes also referred to as vehicle head units). Agents 131 - 134 can be vending machines (e.g., for dispensing train tickets, etc.). Agents 131 - 134 can be gaming machines.

[0064] As Figure 4 shown, agents 131 - 134 are communicating with central server 139 via the Internet, for example. For example, cellular radio access technology can be employed.

[0065] According to an example, central server 139 can provide a downlink message to agents 131 - 134 that indicates an ML model 121 in a given training state. For example, the downlink message can indicate the architecture (e.g., number of layers, skip connections, type of layers, etc.). Weights can be indicated.

[0066] Then, ML model 121 can be used at each of the plurality of agents 131 - 134 for inferring predictions, such as gesture class predictions (see Figure 3 ). The inference is based on a measurement feature vector encoding locally obtained measurement data. For example, a radar sensor can be used to obtain radar measurements.

[0067] The ML model 121 can also be retrained at each of the agents 131 - 134 based on the respective locally populated training datasets. Thus, the initial training state of the ML model 121 can be updated at the agents 131 - 134 through edge - deployed retraining. Then, the agents 131 - 134 can provide an uplink message to the central server 139, which indicates the weights of the ML model 121 when such retraining is completed (i.e., in the updated training state). This enables the central server 139 to collect such updates from multiple agents 131 - 134 and combine or specifically average the respective weights of the updated training state of the ML model 121 to determine another training state of the ML model 121. Then, this further updated training state of the ML model 121 can be provided to the agents 131 - 134 in further iterations. A corresponding downlink message can include an update of the weights; it may not be necessary to signal all the parameters of the ML model. For example, the architecture can be fixed, the number of layers can be fixed, and the type of layers can remain unchanged. From the above, it can be understood that techniques of federated learning can be adopted.

[0068] Figure 4 The figure illustrates that messages indicating the ML model 121 are exchanged between the agents 131 - 134 and the central server 139. Additionally, the message can also indicate other information transmitted between the agents 131 - 134 and the central server 139, such as auxiliary information and / or context information. For example, the central server 139 can provide auxiliary information to the agents 131 - 134, which configures the training process performed at the agents 131 - 134 and / or configures the self - labeling process performed at the agents 131 - 134. For example, the agents 131 - 134 can provide context information to the central server 139, which indicates the attributes of the agent - specific training datasets populated at the agents 131 - 134.

[0069] Furthermore, in a scenario where the training process is not performed at each of the agents 131 - 134 (but at the central server 139), the agents 131 - 134 can provide their training datasets to the central server 139. Generally, preferably, the training datasets are retained at the agents 131 - 134, for example, to minimize data transfer and maintain privacy.

[0070] Figure 5 is a flowchart of a method according to various examples. Figure 5 Generally relates to performing inference at the edge using an ML model. Figure 5 Also relates to edge retraining of an ML model. Figure 5 Includes aspects of a self - labeling process. Figure 5 Also relates to aspects of federated learning.

[0071] Figure 5 The method includes an iterative process. Multiple iterations 3085 of the entire process can be performed. Each iteration in the iterations 3085 is associated with a corresponding training state of the distribution of the ML model to multiple agents. After initializing the central authority and the multiple agents, the ML model in the initial (e.g., first) training state is distributed to the agents. The agents retrain the ML model based on the labels and pseudo-labels obtained from the self-labeling process. This results in an ML model in an updated agent-specific training state (e.g., second training state), which is then reported back to the central authority (e.g., using an encrypted uplink message). Then, the central authority can combine the multiple second training states to determine an updated combined training state (e.g., third training state). This serves as the basis for subsequent iterations 3085, i.e., the distribution to the agents. Next, the iterative process will be explained in conjunction with Figure 5 a specific implementation.

[0072] In Figure 5 , the branch 3098 is executed at each of the multiple agents (although only a single representative agent among the multiple agents will be explained hereinafter). For example, when loading program code from a memory, the boxes associated with the branch 3098 can be executed by the processor of the corresponding agent (such as the processor 62 (see Figure 1 ).

[0073] In addition, the branch 3099 is executed by the central authority, such as the central authority implemented by a server (such as the server 139). A processor (such as the processor 311) can load program code from a memory (such as the memory 312) and execute the program code to execute the boxes of the branch 3099.

[0074] At box 3105, the central authority provides the ML model in the current training state to the multiple agents. This training state is associated with the current iteration 3085.

[0075] In the first iteration 3085, the initial training state can be determined, for example, based on centralized measurements and training activities. In addition, randomized weights can be used to obtain the initial training state. The specific choice of the initialization process is irrelevant to subsequent operations.

[0076] Box 3015 can include selecting an agent from among multiple candidate agents to which the ML model in the current training state is provided. That is, a subset of all candidate agents can be selected. This selection can be based on the availability of the battery and / or communication efficiency and / or power consumption and / or charge state and / or health state of the battery. For example, agents with a particularly low battery charge state or agents without a direct power source can be avoided. Doing so can avoid further depletion of the battery or shortening of the lifespan of such agents.

[0077] At block 3005, an ML model is obtained at the agent accordingly. Blocks 3001 and 3005 can include: transmitting a downlink message or a plurality of downlink messages from a central authority to each agent via the Internet, for example. Such messages can be encrypted.

[0078] At block 3010, a measurement feature vector is obtained at the agent. This defines the current inference iteration 3080. The measurement feature vector is obtained based on one or more measurements performed at the agent. The sensor can be a depth sensor (such as a radar sensor).

[0079] At block 3015, an inference can be performed based on the measurement feature vector of block 3010. Thus, a prediction can be inferred based on the measurement feature vector and using the ML model. The measurement feature vector or its representation is input into the ML model. A forward pass of the measurement feature vector through an artificial neural network is performed.

[0080] The prediction can be used to control the functions of the agent. For example, a graphical user interface can be controlled. Information associated with the prediction can be output. Technical functionality can be controlled. A control signal can be emitted.

[0081] The prediction can also be used as part of a self-labeling process. This will be explained next.

[0082] At block 3020, it is determined whether a ground truth is available for the prediction inferred at block 3015. For example, block 3020 can include determining whether user input indicating the ground truth associated with the measurement feature vector obtained at block 3010 is available.

[0083] For example, the user input can confirm the prediction. For example, the prediction is associated with a certain gesture class, and the user can confirm that the gesture class has been correctly predicted. The user input can also correct the prediction. For example, the prediction can involve the segmentation of a region of interest in an image. Then, the user can change some of its partitions or parts. Then, the corrected segmentation can be regarded as the ground truth.

[0084] The ground truth may be available not only through user input. For example, sometimes the ground truth can be determined based on alternative measurements (i.e., using another measurement modality other than the measurement modality used to determine the measurement feature vector at block 3010). Such alternative measurements may only be available in certain cases. For example, a radar sensor can be used for depth measurement; depth measurement based on a stereo vision camera is also possible, but only in a daytime scene. At night, alternative depth measurements may not be available.

[0085] As a general rule, the ground truth can be determined based on at least one of user input or reference measurement.

[0086] In response to the ground truth being available, box 3025 is executed. Here, a label is determined based on the ground truth. The label can be equivalent to the ground truth or can be derived from the ground truth. The measured feature vector and the associated label are added to the first training dataset as an input-output sample pair. This first training dataset collects input sample and output sample pairs, where the output sample corresponds to the label.

[0087] In response to the ground truth not being available at box 3020, a self-labeling process is executed at box 3039.

[0088] At box 3039, there are various options for implementing the self-labeling process. Figure 5 An example is illustrated in.

[0089] In Figure 5 the self-labeling process includes box 3040. Box 3040 is optional. At box 3040, it is determined whether to determine a pseudo-label based on the prediction inferred for the measured feature vector at box 3015. For box 3040, various decision criteria can be envisioned.

[0090] For example, if the prediction inferred at box 3015 is associated with significant uncertainty, it can be judged that the pseudo-label cannot be reliably determined. For example, if the uncertainty (epistemic uncertainty and / or aleatoric uncertainty) of the corresponding prediction determines that the current iteration 3080 of box 3015 is below a predefined threshold, the pseudo-label can be selectively determined. For example, it can be determined whether the uncertainty of the prediction determined for sample i is less than the threshold τ, i.e.,

[0091] As an alternative or supplement to considering uncertainty at box 3040, the probability of the corresponding prediction inferred at the current iteration 3080 of box 3015 can be considered. For example, it can be considered whether the class probability of the classifier ML model has a change exceeding a specific threshold, i.e., a given class is more likely than other classes. In another example, it can be determined whether one of the class probabilities associated with the current prediction exceeds a certain predefined threshold. For example, it can be determined whether max c′ is greater than the threshold, i.e., max c′ Here, is the predicted probability that the input sample i (measured feature vector) belongs to class c based on the prediction (i.e., the output of the ML model). max c′ is the maximum predicted probability of sample i across all classes. It is determined whether this maximum predicted probability exceeds the threshold.

[0092] The ML model can perform classification. This classification is associated with multiple class probabilities. The ML model is trained to distinguish multiple classes. To this end, the class probability for each of the multiple classes is determined. Then, the specific class with the highest class probability is selected as the predicted class. For example, for an ML model implemented through a deep neural network, such class probabilities can be obtained at the output layer of the deep neural network where the softmax activation function is applied. The softmax activation function takes a vector as input and converts its individual values to probabilities based on their magnitudes. High numerical values result in high class probabilities. The sum of the class probabilities is normalized to 1.

[0093] For example, such thresholds for uncertainty (τ u ) and / or probability (τ p ) can be predefined. The thresholds can be determined by a central authority. Information indicating one or more such thresholds can be provided to the agent by the central authority as auxiliary information.

[0094] For example, such thresholds for uncertainty (τ u ) and / or probability (τ p ) can be agent-specific; i.e., different agents can have different thresholds.

[0095] For example, the uncertainty (τ u ) and / or probability (τ p ) such thresholds can be determined based on the credibility associated with each agent. For example, some agents are known to likely produce high-quality data; this can be reflected in the thresholds being configured to accept more pseudo-labels.

[0096] Such thresholds for uncertainty (τ u ) and / or probability (τ p ) can depend on iteration 3085 (which can be referred to as a dynamic threshold). The (multiple) thresholds can vary according to iteration 3085. An example linear dependence of the threshold from value τ start to value τ end above a specific decay_duration is shown in Equation 1. Current_step represents iteration 3085.

[0097]

[0098] After determining that the pseudo-label can be reliably determined, box 3045 is executed; otherwise, box 3045 is skipped (i.e., the measured feature vectors can be discarded in the context of retraining). At box 3045, the pseudo-label is determined based on the prediction of box 3015, and the second training dataset is populated by adding the current measured feature vector of the current iteration 3080 of box 3010 and the pseudo-label determined at box 3045 as an input sample and output sample pair to the second training dataset.

[0099] The pseudo-label of sample i can be determined as follows:

[0100]

[0101] argmax c′ is the index of the class with the highest predicted probability for sample i. That is, the pseudo-label is determined based on the class probabilities by selecting the class with the maximum class probability.

[0102] The second training dataset filled at the multiple inference iterations 3080 of block 3045 is different from the first training dataset filled at the multiple iterations 3080 of block 3025. Thus, two separate agent-specific training sets are reserved for the labels and pseudo-labels respectively at each agent. These two training datasets can have different sizes. Typically, the labels are available more sparsely compared to the pseudo-labels. Thus, there can be a tendency that the first training dataset filled at the multiple iterations 3080 of block 3025 has a smaller size compared to the second training dataset filled at the multiple iterations 3080 of block 3045.

[0103] By maintaining separate training datasets filled based on the labels and pseudo-labels respectively, multiple effects can be achieved. As a first effect, the cases where only labels are available, or only pseudo-labels are available, or both labels and pseudo-labels are available can be handled flexibly. In particular, the retraining of the ML model does not need to be delayed until both labels and pseudo-labels are available. This facilitates accurate and flexible retraining of the ML model in various situations encountered in the field. As a second effect, by maintaining different training datasets filled based on the labels and pseudo-labels respectively, the impact of the pseudo-labels on the training process can be adjusted flexibly. For example, depending on the confidence level associated with the respective agent, the pseudo-labels can be considered more prominently or less prominently. For example, the pseudo-labels can be considered more prominently for later iterations 3085; this is based on the assumption that for later iterations 3085, the overall accuracy of the ML model increases, such that as the overall accuracy of the ML model increases, the reliability of the pseudo-labels also increases. Such an impact will become apparent in combination with subsequent blocks.

[0104] At block 3050, it is determined whether retraining is possible. For example, block 3050 can include: determining the size of at least one of the first training datasets filled at the multiple iterations 3080 of block 3025, or determining the size of the second training datasets filled at the multiple iterations 3080 of block 3045. A retraining trigger provided by a central agency can be monitored.

[0105] For example, for agents with a particularly low state of charge of their batteries, the retraining at box 3050 can also be skipped. For agents not connected to the grid or without other direct power sources, the training process can be skipped. Such techniques are based on the following discovery: due to the large number of computational operations required, performing the training process is particularly energy-consuming. Therefore, in cases where energy resources are limited, the retraining of the ML model can be postponed or skipped entirely.

[0106] If retraining is not possible, further inference iterations 3080 are performed, i.e., a new current measurement feature vector is obtained at box 3010.

[0107] Otherwise, the training process is performed at box 3054. The training process produces an updated training state of the ML model.

[0108] The retraining is based on a first training dataset and a second training dataset.

[0109] Multiple options for performing the retraining are possible.

[0110] For example, separate backpropagation optimizations can be performed based on the first training dataset and the second training dataset. This is illustrated in box 3055 and box 3060. At box 3055, a first updated weight of the ML model in the current training state obtained in the current iteration 3085 of box 3005 is determined based on the first training dataset. Then the first training dataset can be refreshed.

[0111] The loss function for box 3055 can be:

[0112]

[0113] Here, N is the number of labeled samples in the first training dataset, C is the number of classes in the classification problem, y i,c is the label of sample i and class c. It indicates whether sample i belongs to class c (1 if true, 0 otherwise). is the predicted probability that sample i belongs to class c. It is the output of the softmax function of the model and represents the confidence of the model in the predicted class probability. It should be understood that the loss function given by Equation 3 is only an example, and other examples are possible. Various loss functions are known in the art and can be utilized according to various public scenarios.

[0114] At box 3060, a second updated weight of the ML model is determined, which is in the current training state obtained in the current iteration 3085 of box 3005 or updated based on the first updated weight in box 3055. This optimization in box 3060 is based on the second training dataset. Then the second training dataset can be refreshed.

[0115] The loss function of box 3060 can be expressed as follows:

[0116]

[0117] is the generated pseudo-label for sample i and class c. It represents the model prediction for sample i belonging to class c. is the predicted probability that sample i belongs to class c, which is based on the output of the ML model.

[0118] At optional box 3065, if there are separate sets of updated weights (from box 3055 and box 3060), these weights can be combined. For example, a weighted combination can be performed. For example, a weighted average can be determined between the weights determined at box 3055 and the weights determined at box 3060.

[0119] More generally, the relative weights of the influence of the first training dataset (filled in multiple inference iterations 3080 based on labels at box 3025) on the second training state can be considered relative to the influence of the second training dataset (filled in multiple inference iterations 3080 based on pseudo-labels at box 3045) on the second training state to determine the second training state that is determined as the output of box 3054. The respective weights that adjust the influence of the labels and pseudo-labels can be fixedly predefined, can be provided by a central agency (such as auxiliary information), and / or can be dynamic (i.e., depending on iteration 3085). The relative weighting can gradually increase the influence of the pseudo-labels on subsequent iterations 3085 (and the increased accuracy of the ML model).

[0120] For example, let ∈ be the weighting factor that adjusts the influence of the second training dataset filled based on pseudo-labels, then ∈ can be given by the following formula:

[0121]

[0122] ∈(start) is the initial value of ∈ (at the first iteration 3085). ∈(end) is the final value of ∈ (at the last iteration 3085). growth_duration is the number of iterations 3085 during which ∈ will grow. current_step is the current iteration 3085. Throughout the growth period, Equation 5 gradually increases from ∈(start) to ∈(end).

[0123] Then, at box 3070, the ML model in the updated second training state obtained from box 3065 can be provided to the central agency.

[0124] For example, an incremental update can be provided that indicates the difference between the weights of the ML model in an updated training state obtained from the retraining process performed at block 3054 and the weights of the ML model in a training state obtained at block 3005. More generally, information indicating the weights obtained at block 3055, and / or information indicating the weights obtained at block 3060, and / or information indicating the final updated training state (e.g., the combined weights obtained from block 3065) can be provided to a central authority.

[0125] Optionally, context information can be provided to the central authority at block 3070. Such context information can include an indication of the size of the first training dataset, and / or an indication of the size of the second training dataset, and / or an indication of the ratio of these sizes, and / or a relative weighting factor of the relative weighting and influence of the first training dataset with respect to the second training dataset.

[0126] At block 3110, the central authority obtains an ML model in an updated training state. More specifically, at block 3110, the central authority obtains ML models in multiple updated training states from multiple agents. Thus, in block 3115, these multiple updated training states and agent-specific training states can be combined (e.g., averaged), which is referred to as federated learning, to determine a new and improved current training state of the ML model in a further iteration 3085 of block 3105.

[0127] For example, the combination at block 3115 can take into account the context information provided by each agent. For example, depending on the size of the underlying training dataset and / or depending on whether the corresponding updated training state is based primarily on labels or pseudo-labels, the influence of the corresponding updated training state of the transmitted weights can be relatively weighted. Similarly, such weighting can depend on iteration 3085.

[0128] In summary, Figure 5 Aspects regarding retraining of an ML model considering pseudo-labels determined during a self-labeling process at block 3039 are illustrated. The self-labeling process performed at block 3039 takes into account the probability and / or uncertainty of the predictions of block 3015. Conventional self-labeling processes are less effective because poor network calibration and poor data quality can erroneously generate pseudo-labeled samples, resulting in a noisy training dataset and, in turn, poor generalization of the ML model. To address this issue, selecting predictions with low uncertainty and / or high probability of pseudo-labels can greatly reduce the impact of poor calibration and misassigned pseudo-labels. Uncertainty classification attempts to maintain the simplicity, generality, and ease of implementation of classical pseudo-labeling while addressing the calibration problem to significantly improve pseudo-labeling performance.

[0129] There are different options available for determining uncertainty. These options are summarized in Table 1:

[0130] 1 Sensor data 2 Randomly discard sampling 3 Overall sampling 4 Evidence distribution 5 Bayesian neural network

[0131] Table 1: Options for determining uncertainty.

[0132] Based on the sensor information provided by the sensor, uncertainty can be determined, more specifically aleatoric uncertainty. This is shown in Example 1 of Table 1.

[0133] Regarding Example 2 of Table 1, uncertainty can be determined based on random dropout sampling (e.g., Monte Carlo dropout sampling). Such random dropout sampling includes multiple forward passes through the ML model, where for each of the multiple forward passes, a random subset of the weights or interconnections (neurons in a deep neural network) is disabled / set to zero, i.e., temporarily removed from the ML model. By performing multiple forward passes, multiple predictions are obtained. The variation in the predictions is a measure of uncertainty.

[0134] Furthermore, uncertainty (more specifically epistemic uncertainty) can be determined based on ensemble sampling of the distribution of the corresponding predictions. See Example 3 of Table 1. Here, through random variation of the weights, an ensemble of multiple instances of the ML model is determined based on the current training state of the ML model. A single input sample (measurement feature vector) is passed forward through these multiple instances, which yields multiple predictions. The variation in the predictions is a measure of uncertainty.

[0135] For such scenarios of random dropout sampling or ensemble sampling according to Example 2 and Example 3 of Table 1, uncertainty can be determined as follows:

[0136]

[0137]

[0138] Again, it is the predicted probability that sample i of the output of the ML model belongs to class c. K is the number of ML models in the ensemble or the number of Monte Carlo dropout masks. is the probability predicted by the k-th classifier, or the number of Monte Carlo samples (dropout masks) generated during the inference of sample i belonging to class c. represents the estimated uncertainty of sample i.

[0139] Furthermore, as shown in Example 4 of Table 1, uncertainty can be determined based on the evidence distribution of the corresponding predictions. For the corresponding measurement feature vector, the evidence distribution can be predicted by the ML neural network. Determining uncertainty based on the evidence distribution may only require a single forward pass of the ML neural network; therefore, this uncertainty is sometimes referred to as "one-pass uncertainty".

[0140] Evidential Deep Learning (EDL) is a concept disclosed in "Evidential deep learning to quantify classification uncertainty" by Sensoy, Murat, Lance Kaplan, and Melih Kandemir, Advances in neural information processing systems 31 (2018), which can be combined with pseudo-labeling to utilize uncertainty estimation for more reliable and efficient semi-supervised learning. The concept of EDL originates from Dempster-Shafer evidence theory. It assigns credibility masses to subsets within a frame of discernment, representing possible states or class labels. And these credibility assignments can be estimated using the Dirichlet distribution, thereby quantifying the credibility masses and uncertainties within a defined frame. Each frame consists of K mutually exclusive singletons (e.g., class labels), and each singleton is assigned a credibility mass (b k ), while determining the overall uncertainty mass (u). The credibility mass (b k ) of a singleton is calculated using the evidence (e k ), and the uncertainty (u) is inversely proportional to the total evidence. The formulas for credibility and uncertainty are as follows:

[0141] Credibility mass of a singleton:

[0142] Uncertainty:

[0143] Here, S represents the sum of the evidence values of all singletons in the frame, given by

[0144]

[0145] It is worth noting that when no evidence is available, the credibility of each singleton is zero, and the uncertainty reaches its maximum value (one).

[0146] In EDL, a loss function is introduced, which specifically combines two terms.

[0147]

[0148] where λ t = min(1.0, t / 10) ∈ [0, 1] is the annealing coefficient, t is the exponent of the current training epoch, D(p i |<1,..., 1>) is the uniform Dirichlet distribution, and are the parameters of the Dirichlet distribution after removing misleading evidence from the prediction parameter α of sample i i The first term is inspired by the sum of squares loss and is expressed as

[0149] This loss helps to minimize the prediction error and variance of the Dirichlet distribution generated by the neural network. This term also gives priority to model fitting over variance calculation. The second term of the loss function is a penalty factor represented by the Kullback-Leibler (KL) divergence term. This term is responsible for penalizing instances that do not contribute to the overall data fitting.

[0150]

[0151] Probabilistic ML (such as Bayesian neural network-based (see Table 1: Example 5) or overall ML) can provide uncertainty estimates for its predictions.

[0152]

[0153] Various modifications can be made to the method. For example, box 3070 is optional. Not all scenarios require the agent to report to the central agency. The technique of federated learning is optional.

[0154] The method of Figure 5 can be variously modified. For example, box 3070 is optional. Not all scenarios require the agent to report to the central agency. The technique of federated learning is optional.

[0155] In addition, the training process can also be performed at the central agency instead of at the agent. In this scenario, the training dataset can be transmitted from the agent to the central agency. Then, box 3054 can be performed at the central agency.

[0156] In summary, at least the following examples have been disclosed.

[0157] Example

[0158] Example 1. A computer-implemented method used in agents (65, 131, 132, 133, 134), including:

[0159] Obtaining a machine learning model (121) in a first training state from a central agency (139),

[0160] Using the machine learning model (121) in the first training state to infer (3015) a prediction based on a plurality of measured feature vectors obtained at the agents (65, 131, 132, 133, 134),

[0161] Based on the prediction, filling (3025) a first training dataset with labels of a first subset of the plurality of measured feature vectors,

[0162] Populate (3045) a second training dataset with pseudo-labels using a second subset of the plurality of measurement feature vectors, based on the prediction, and

[0163] Determine (3054) a second training status of the machine learning model (121) based on the first training dataset and the second training dataset.

[0164] Example 2. The computer-implemented method of Example 1, further comprising:

[0165] If the uncertainty of the prediction is below a threshold, selectively (3040) determine the pseudo-labels.

[0166] Example 3. The computer-implemented method of Example 2,

[0167] wherein the machine learning model (121) is obtained from a central authority (139) during an iterative update process and subsequently provides a plurality of training statuses of the machine learning model (121),

[0168] wherein the threshold depends on the iteration (3085) of the iterative update process.

[0169] Example 4. The computer-implemented method of Example 2 or 3, further comprising:

[0170] Determine the uncertainty based at least on the corresponding predicted evidence distribution, which is predicted by the machine learning model for the corresponding measurement feature vector.

[0171] Example 5. The computer-implemented method of any one of Examples 2 to 4, further comprising:

[0172] Determine the uncertainty based on at least one of dropout sampling or overall sampling of the corresponding predicted distribution.

[0173] Example 6. The computer-implemented method of any one of the foregoing examples, further comprising:

[0174] Selectively (3040) determine the pseudo-labels according to the predicted probability.

[0175] Example 7. The computer-implemented method of Example 6,

[0176] wherein if the predicted probability exceeds a threshold, the determination of the pseudo-labels is performed,

[0177] wherein the machine learning model (121) is obtained from a central authority during an iterative update process and subsequently provides a plurality of training statuses of the machine learning model (121),

[0178] wherein the threshold depends on the iteration (3085) of the iterative update process.

[0179] Example 8. The computer-implemented method of any of the foregoing examples further comprises:

[0180] Determining first update weights of a machine learning model based on a first training dataset,

[0181] Determining second update weights of the machine learning model based on a second training dataset, and

[0182] Determining a second training state by performing a combination of the first update weights and the second update weights.

[0183] Example 9. The computer-implemented method of any of the foregoing examples,

[0184] wherein the second training state is determined taking into account a relative weight of the influence of the first training dataset on the second training state compared to the influence of the second training dataset on the second training state.

[0185] Example 10. The computer-implemented method of Example 9,

[0186] wherein the relative weight depends on auxiliary information provided by a central authority.

[0187] Example 11. The computer-implemented method of Example 9 or 10,

[0188] wherein the machine learning model (121) is obtained from a central authority during an iterative update process and subsequently provides multiple training states of the machine learning model (121),

[0189] wherein the relative weight depends on an iteration (3085) of the iterative update process.

[0190] Example 12. The computer-implemented method of Example 11,

[0191] wherein for subsequent iterations (3085), the relative weight gradually increases the influence of the second training dataset on the second training state.

[0192] Example 13. The computer-implemented method of any of the foregoing examples further comprises:

[0193] Providing to the central authority at least one of the following: a first weight, a second weight, or information indicating the second training state.

[0194] Example 14. The computer-implemented method of any of the foregoing examples further comprises:

[0195] Providing to the central authority information indicating at least one of the following: a first size of the first training dataset, a second size of the second training dataset, or a weighting factor that regulates the relative influence of the first training dataset and the second training dataset on the second training state.

[0196] Example 15. The computer-implemented method of any of the foregoing examples further includes:

[0197] Determining (3020) whether ground truth associated with a given measurement feature vector is available, and:

[0198] In response to the ground truth being available: determining (3025) a corresponding label for the given measurement feature vector based on the ground truth, and adding the given measurement feature vector and the corresponding label to a first training data set,

[0199] In response to the ground truth being unavailable, and further in response to determining that a pseudo-label can be determined for the given measurement feature vector: adding the given measurement feature vector and the corresponding pseudo-label to a second training data set, and

[0200] In response to the ground truth being unavailable, and further in response to determining that a pseudo-label cannot be determined for the given measurement feature vector: discarding the given measurement feature vector.

[0201] Example 16. A computer-implemented method used in a central agency, including:

[0202] Providing (3105) a machine learning model in a first training state to a plurality of agents (65, 131, 132, 133, 134), and each agent among the plurality of agents (65, 131, 132, 133, 134) executes the method of any of the examples,

[0203] Obtaining (3110) information indicating the machine learning model in a second training state from each agent among the plurality of agents (65, 131, 132, 133, 134), and

[0204] Determining (3115) a third training state of the machine learning model based on the plurality of second training states.

[0205] Example 17. The computer-implemented method of Example 16 further includes:

[0206] Obtaining context information of the machine learning model in the corresponding second training state from each agent among the plurality of agents,

[0207] wherein the determination of the third training state of the machine learning model is further based on the context information.

[0208] Example 18. The computer-implemented method of Example 17

[0209] wherein the context information indicates at least one of the size of the first training data set used to determine the second training state and filled with labels, the size of the second training data set used to determine the second training state and filled with pseudo-labels, and the relative size of the first training data set relative to the second training data set.

[0210] Example 19. An agent (65, 131, 132, 133, 134) includes computing circuitry configured to:

[0211] obtain a machine learning model (121) in a first training state from a central authority (139),

[0212] infer (3015) a prediction using the machine learning model (121) in the first training state based on a plurality of measured feature vectors obtained at the agent (65, 131, 132, 133, 134),

[0213] fill (3025) a first training data set with labels of a first subset of the plurality of measured feature vectors based on the prediction,

[0214] fill (3045) a second training data set with pseudo-labels of a second subset of the plurality of measured feature vectors based on the prediction, and

[0215] determine (3054) a second training state of the machine learning model (121) based on the first training data set and the second training data set.

[0216] Example 20. The agent of Example 19, wherein the computing circuitry is configured to perform the method of any one of Examples 1 to 15.

[0217] Example 21. A central authority including computing circuitry configured to:

[0218] provide (3105) a machine learning model in a first training state to a plurality of agents (65, 131, 132, 133, 134), preferably a plurality of agents configured according to the agent of Example 19,

[0219] obtain (3110) from each of the plurality of agents (65, 131, 132, 133, 134) information indicating a machine learning model in a corresponding second training state, and

[0220] determine (3115) a third training state of the machine learning model based on the plurality of second training states.

[0221] Example 22. The central authority of Example 21, wherein the computing circuitry is configured to perform the method of any one of Examples 16 to 18.

[0222] Example 23. A system includes a plurality of agents configured according to the agent of Example 19 and a central authority of Example 21.

[0223] Although the invention has been shown and described with respect to certain preferred embodiments, other persons skilled in the art will be able to conceive of equivalents and modifications upon reading and understanding the specification. The invention includes all such equivalents and modifications and is limited only by the scope of the appended claims.

[0224] For illustration, various techniques have been disclosed above, in which a self-labeling process for determining pseudo-labels is performed at a field-deployed agent. However, such a self-labeling process can also be performed at a central authority, taking into account the uncertainty of the corresponding prediction for determining the pseudo-labels and / or taking into account the class probability of the corresponding prediction for determining the pseudo-labels.

[0225] In addition, various examples have been disclosed according to which measurement feature vectors encoding radar measurements can be processed. However, the techniques disclosed herein can also be applied to other types and kinds of sensors, such as other types of depth sensors.

Claims

1. A computer-implemented method for use in an agent (65, 131, 132, 133, 134), comprising: obtaining a machine learning model (121) in a first training state from a central authority (139), inferring (3015) a prediction based on a plurality of measured feature vectors acquired at the agent (65, 131, 132, 133, 134) using the machine learning model (121) in the first training state, populating (3025) a first training data set using labels for a first subset of the plurality of measured feature vectors based on the predictions, populating (3045) a second training data set using pseudo labels for a second subset of the plurality of measured feature vectors based on the predictions, and A second training state of the machine learning model (121) is determined (3054) based on the first training data set and the second training data set.

2. The computer-implemented method of claim 1 , further comprising: If the uncertainty of the prediction is below a threshold, the pseudo-label is selectively (3040) determined.

3. The computer-implemented method of claim 2, wherein the machine learning model (121) is obtained from the central authority (139) in an iterative updating process, and then a plurality of training states of the machine learning model (121) are provided, Wherein the threshold value depends on the iteration (3085) of the iterative updating process.

4. The computer-implemented method of claim 2 or 3, further comprising: The uncertainty is determined based at least on the corresponding predicted evidence distribution, which is predicted by the machine learning model for the corresponding measured feature vector.

5. The computer-implemented method according to any one of claims 2 to 4, further comprising: The uncertainty is determined based on at least one of a random discard sampling or an overall sampling of the corresponding predicted distribution.

6. The computer-implemented method of any preceding claim, further comprising: The pseudo-label is selectively (3040) determined based on the predicted probability.

7. The computer-implemented method of claim 6, wherein if the probability of the prediction exceeds a threshold, the determination of the pseudo-label is performed, wherein the machine learning model (121) is obtained from the central authority in an iterative updating process, and then a plurality of training states of the machine learning model (121) are provided, Wherein the threshold value depends on the iteration (3085) of the iterative updating process.

8. The computer-implemented method of any preceding claim, further comprising: determining a first update weight for the machine learning model based on the first training data set, determining a second updated weight for the machine learning model based on the second training data set, and The second training state is determined by performing a combination of the first update weight and the second update weight.

9. A computer-implemented method according to any one of the preceding claims, The second training state is determined taking into account a relative weight of an influence of the first training data set on the second training state compared to an influence of the second training data set on the second training state.

10. The computer-implemented method of claim 9, The relative weights depend on auxiliary information provided by the central authority.

11. A computer-implemented method according to claim 9 or 10, wherein the machine learning model (121) is obtained from the central authority in an iterative updating process, and then a plurality of training states of the machine learning model (121) are provided, Wherein the relative weight depends on the iteration (3085) of the iterative updating process.

12. The computer-implemented method of claim 11, Wherein for subsequent iterations (3085), the relative weight gradually increases the influence of the second training data set on the second training state.

13. The computer-implemented method of any preceding claim, further comprising: At least one of the following is provided to the central authority: the first weight, the second weight, or information indicative of the second training status.

14. The computer-implemented method of any preceding claim, further comprising: Information indicating at least one of: a first size of the first training data set, a second size of the second training data set, or a weighting factor that adjusts the relative influence of the first training data set and the second training data set on the second training state is provided to the central authority.

15. The computer-implemented method of any preceding claim, further comprising: It is determined (3020) whether ground truth values ​​associated with a given measured feature vector are available, and: In response to the ground truth being available: determining (3025) the corresponding label for the given measured feature vector based on the ground truth, and adding the given measured feature vector and the corresponding label to the first training data set, In response to the ground truth being unavailable, and further in response to determining that a pseudo label can be determined for the given measured feature vector: adding the given measured feature vector and the corresponding pseudo label to the second training data set, and In response to the ground truth being unavailable, and further in response to determining that a pseudo label cannot be determined for the given measured feature vector: discarding the given measured feature vector.