Method and system for automatically annotating sensor data
By grouping sensor data frames by ambient conditions and selectively retraining neural networks, the method addresses the inefficiencies of manual annotation, enhancing scalability and quality in sensor data annotation processes.
Patent Information
- Application Number
- JP2024517021
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-17
- Filing Date
- 2022-09-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-15
AI Technical Summary
Existing methods for annotating sensor data, particularly for autonomous driving, require significant human labor and time-consuming quality checks, limiting the scalability and efficiency of data annotation processes.
A method that groups sensor data frames based on ambient conditions, uses a neural network for initial annotation, selects random samples for quality measurement, and re-trains the network when quality thresholds are not met, thereby reducing manual effort and improving annotation quality.
This approach significantly reduces manual labor and computational resources by focusing retraining on specific ambient conditions, ensuring high-quality annotations without extensive manual checks, thus accelerating large-scale annotation projects.
Smart Images

Figure 0007708969000001 
Figure 0007708969000002 
Figure 0007708969000003
Abstract
Description
Technical Field
[0001] The present invention relates to a method and a computer system for automatically annotating sensor data frames, in particular data frames of imaging sensors.
Background Art
[0002] Autonomous driving does not promise existing levels of comfort and safety in daily traffic. Despite significant investment by various companies, existing approaches can only be used in limited situations and / or assume only a subset of truly autonomous behavior. The reason is the lack of a sufficient amount and variety of available driving scenarios. Therefore, further progress is limited by the need for a huge amount of sufficiently different training data and validation data (i.e., independent ground truth data). To prepare training data, it is generally necessary to record a number of different driving scenarios by vehicles equipped with a series of sensors, in particular imaging sensors such as one or more cameras, LiDAR sensors and / or radar sensors. Before using these recorded scenarios as training data, it is necessary to annotate them.
[0003] This annotation is often carried out by annotation service providers, who receive the recorded sensor data and divide it into work packets for a large number of human workforces, also referred to as labelers. The exact annotation required (e.g., object classification to be distinguished) depends on each project and is described in a detailed labeling specification. Customers supply raw data to the annotation service provider and expect high-quality annotation in a short period according to the customer's information. The number of labelers required for the completion of the annotation project increases with the increase in the amount of data supplied and also with the decrease in the time frame for a fixed amount of data. For these reasons, relatively large-scale annotation projects that would supply sufficient ground truth data to verify, for example, autonomous vehicles cannot be achieved only by using human labor and require the automation of the annotation process.
[0004] The automated approach uses neural networks to label the recorded sensor data. The first set of received data is labeled manually and then used to train a dedicated neural network. Once the dedicated neural network is sufficiently trained, it can annotate a large amount of recorded imaging sensor data. This significantly reduces the labor compared to a purely manual approach. However, in order to maintain high annotation quality, time-consuming quality checks by humans are still required. Since the quality assurance process still has to be applied to all annotations, there is a linear relationship between the project volume and the labor required to meet the project requirements.
[0005] Therefore, there is a need for an improved method for automatically annotating sensor data, especially imaging sensor data, and it would be particularly desirable to guarantee high annotation quality while reducing the number of manual quality checks.
Summary of the Invention
Problems to be Solved by the Invention
[0006] An object of the present invention is to provide a method and a computer system for automatically annotating sensor data frames, particularly video frames or LiDAR point clouds.
Means for Solving the Problems
[0007] In a first aspect of the present invention, there is provided a computer-implemented method for automatically annotating sensor data frames, the method comprising: receiving a plurality of sensor data frames; grouping the frames into a plurality of packets based on at least one conditional attribute, the conditional attribute representing ambient conditions that existed during the recording of the sensor data frames; using a neural network to annotate frames from a first packet, the annotation including assigning at least one data point to each frame, the first packet including frames such that at least one conditional attribute is within a selected range of values; selecting a first random sample of one or more frames from the first packet and determining a quality measure for the data points; and when the computer identifies that the quality measure for at least one frame in the first random sample is below a predetermined threshold, the method comprising: receiving a corrected annotation for the frames in the first random sample; retraining the neural network based on the frames in the first random sample; selecting a second random sample of one or more frames from the frames of the first packet that were not included in the first random sample; Annotating frames of a second random sample using the retrained neural network, and receiving a quality measure for the data points and verifying that the quality measure for the frames in the second random sample exceeds a predetermined threshold, and annotating the remaining frames of the first packet using the retrained neural network, and exporting the annotated frames of the first packet is further provided.
[0008] The host computer can be implemented as a single standard computer including a processor such as a general-purpose microprocessor, a display device, and an input device. Alternatively, the host computer system can include one or more servers including a plurality of processing elements, and the servers are connected via a network to a client including a display device and an input device. Thus, the annotation software can be implemented partially or fully on a remote server, such as on a computer cloud, and thus only a graphical user interface needs to be implemented locally. The export of the annotated frames can include, for example, storing the frames on an external data medium and / or converting or integrating them into a predetermined data format.
[0009] By grouping sensor data frames based on condition attributes representing the ambient conditions existing at the time of recording, the possible correlation between the condition attributes and the annotation accuracy can be considered. The ambient conditions existing during the recording of the frames may affect the annotation accuracy. In an annotation containing multiple data points, this influence may vary depending on the data points. When the sensor data includes camera images taken at night, it may be more difficult to determine the position and / or classification of objects. However, attributes of a vehicle, such as the state of an indicator light, are more perceptible than in bright daylight. The present invention makes it possible to identify ambient conditions that inhibit annotation accuracy and improve the neural network under these conditions by selective retraining. Since the retraining is targeted at the problematic ambient conditions, the overall training effort is reduced. This further reduces the computational performance required for training and, consequently, the energy consumption.
[0010] The term "neural network" may relate to a single neural network, a combination of different neural networks depending on a given architecture, or any type of technology based on machine learning that learns from sample data with a teacher, semi-supervised, or unsupervised. Different neural networks can be used for different data points, i.e., the position and / or classification of an object can be determined using a first neural network, while the attributes of the object can be determined using at least one additional neural network.
[0011] Manual work is only used to create training data, test data, and / or validation data for the purpose of systematically improving a neural network for annotating frames or other automated components based on machine learning, so the effort for large-scale annotation projects can be significantly reduced. Typically, the quality level will be sufficient for providing automated results without further manual checking after several iterations of retraining the neural network, i.e., for annotation by the neural network. The method of the present invention further reduces the manual effort and time required by focusing the retraining on conditions where the annotation quality is still insufficient.
[0012] As a quality measure, for example, the area overlap between an automatically created bounding box and a manually created bounding box within the quality control framework can be used. It is also possible to require the maximum number and / or maximum rate of misassigned object classifications and / or false positives and / or false negatives. In that case, for example, if the area overlap of the bounding box is too small, the quality measure will fall below a predetermined threshold. It is also possible to assume, as a quality measure, that in a random sample from a predetermined number of frames, at most a predetermined number of false positives or misidentified objects and / or false negatives or objects not misidentified may occur. In that case, for example, if the number of unidentified objects in the random sample exceeds the maximum allowable number, the quality measure will fall below a predetermined threshold.
[0013] The step of selecting a second random sample of frames from the first packet and the step of annotating the remaining frames of the first packet using the retrained network may be swapped. For example, before the second random sample is selected, all the remaining frames of the first packet can be annotated using the retrained network. If only the annotation of the frames of the second random sample using the retrained network is performed and the annotation of further frames is postponed until sufficient annotation quality is confirmed, this reduces the computational effort in cases where the neural network has to be retrained more than once, thereby accelerating the retraining process and the annotation process.
[0014] In one embodiment, the received sensor data includes frames of at least one imaging sensor, such as, for example, one or more cameras, LiDAR sensors, and / or radar sensors. The received sensor data can also include additional sensor data recorded simultaneously with the imaging sensor data, such as, for example, GPS location, vehicle acceleration, or data from a rain sensor. In the case of an image frame, i.e., a frame having image data or a data frame from an imaging sensor, the conditional attributes are preferably geographical location, time, weather conditions, visibility conditions, road type, distance to an object, and / or traffic density. The distance to an object can be the distance to the nearest object, the distance to the farthest object, or the average distance to a plurality of objects identified within the frame, and by considering the distance of the object as an ambient condition at the time of recording, the influence on the performance of object detection and / or object classification of a neural network can be quantified. In the case of an image frame, at least one data point preferably includes the position of an object, the classification of the object, the position of the edge of the bounding box, the degree of overlap of the object with other objects, the correlation of the object within the image frame with an object within a preceding or subsequent image frame (as a result of object tracking), and / or the operation of an indicator light, such as, for example, a turn signal or a brake light. The number of data points can depend on the content of the image frame, such as, for example, a large number of automobiles and pedestrians in an urban scenario having a corresponding number of object positions, object classifications, and possible attributes for the corresponding object classifications.
[0015] In one embodiment, the received sensor data includes audio frames recorded by at least one microphone. In the case of an audio frame, i.e., a frame with audio data, the conditional attributes are preferably the geographical location, the gender and / or age of the recorded speaker, the size of the room and / or the scale of background noise. In the case of an audio frame, at least one data point includes one or more text words identified from the audio frame. The words are distinguishable from a plurality of subsequent audio frames, and thus data points can be derived from the plurality of audio frames. The difficulty in identifying speech may depend, for example, on the frequency range emitted by the speaker, the presence of reverberation or echo from the room and / or the level of background noise present.
[0016] Preferably, the step of receiving a plurality of sensor data frames includes a step of preprocessing the frames, and at least one of the conditional attributes for a frame is determined by a dedicated neural network based on the frame and / or at least one of the conditional attributes for a frame is determined based on additional sensor data recorded simultaneously with the frame. The additional sensor data can be combined and / or used for inquiries regarding various different services that present, for example, the type of lighting conditions based on weather conditions, time and geographical location.
[0017] In one embodiment, the first random sample includes two or more frames selected from the first packet. Preferably, as soon as the computer identifies that the quality metric for the first random sample is below a predetermined threshold, further calculations on the frames from the first packet cease until a modified annotation for the frames in the first random sample is received. Additional frames from the first group can be manually annotated and added to the frames of the first random sample, thereby enabling the use of a relatively large dataset for retraining the model. By deferring further processing until the neural network is retrained, a significant amount of time and energy is saved.
[0018] Advantageously, the selection of the frame for the first random sample depends on the data points for which the quality measure is to be determined, and in particular, in the case of object detection, a single frame is randomly selected, and / or in the case of object tracking, a batch of consecutive frames is randomly selected. By using an intelligent strategy for extracting random samples, the improvement achievable by retraining is maximized. For example, an object detector such as for identifying traffic signs benefits from training data with a large variance, and thus, a random selection of a single frame is a useful first random sample. On the other hand, the tracking component benefits from consecutive data. Because only in that case can the tracking of the same object between consecutive frames be carried out. In such cases, to be useful as a random sample, a series of consecutive (e.g., always 10) frames will be randomly selected for a wide variety of objects. As an example, intelligent sample extraction will extract frames 10 - 20, frames 100 - 110, and frames 235 - 245 for the first random sample when determining the quality measure for the tracking component. To obtain a large variance in the random samples, the software component performing the sample extraction can define a minimum temporal interval between random samples to ensure that the various different frames are taken under different ambient conditions. Additionally or alternatively, one or more attributes can be considered during sample extraction. For example, when random samples are selected to quantify the ability of an object detector at night, various different environments such as a metropolis, the countryside, or a highway can be defined. In that case, the random selection will be carried out among all random samples that meet the defined criteria.
[0019] In one embodiment, the steps of selecting a current random sample from one or more frames from a first packet, determining a quality measure for the data point, receiving a modified annotation for the frames in the current random sample, and retraining a neural network based on the frames in the current random sample are repeated until the quality measure for the frames in the current random sample exceeds a predetermined threshold or until the first packet contains no remaining frames. Advantageously, the neural network is retrained until it is also possible to correctly handle ambient conditions that are disadvantageous to the annotation process.
[0020] Preferably, annotating the sensor data and recording the sensor data are performed alternately or simultaneously. When it is identified that the quality measure for at least one frame in the first random sample is below a predetermined threshold, the computer requests to record additional sensor data such that at least one conditional attribute is within the selected value range of the first packet. By equipping the test vehicle with an automated recording device that runs a selection program to trigger recording as soon as a predetermined recording condition is met, or by the test driver requesting to drive under predetermined conditions, such as at night, the value range of the conditional attribute can be selected. Thereby, new data is recorded mainly with respect to ambient conditions where the neural network requires further training. By carefully selecting the training data, the improvement per training effort is maximized. Accordingly, the computational performance required for training and the energy consumption are also reduced.
[0021] In a second aspect of the invention, a computer-implemented method for automatically annotating sensor data that includes frames, such as video frames or audio frames, is envisioned. The method is implemented by at least one processor of a host computer, and the method a) receiving a plurality of sensor data frames, b) grouping frames into packets based on at least one conditional attribute, the conditional attribute representing ambient conditions that were present during the recording of the sensor data frames, and c) annotating frames from a first packet using a neural network, the annotation including assigning at least one data point to each frame, the first packet including frames such that at least one conditional attribute is within a selected range of values, and d) selecting a first random sample of one or more frames from the first packet and determining a quality measure for the data points, and e) identifying that a quality measure for at least one frame in the first random sample is below a predetermined threshold, and f) receiving a corrected annotation for the frames in the first random sample and retraining the neural network using the frames in the first random sample, and g) annotating at least one of the remaining frames of the first packet using the retrained neural network, and h) selecting a second random sample of one or more frames from at least one of the annotated remaining frames of the first packet and determining a quality measure for the data points, and i) identifying that a quality measure for the frames in the second random sample is above a predetermined threshold, and j) annotating the remaining frames from the first packet using the retrained neural network, and k) exporting the annotated frames including.
[0022] One aspect of the invention also relates to a non - volatile computer - readable medium including instructions that, when executed by a microprocessor of a computer system, cause the computer system to perform a method according to the invention as described above or as claimed in the appended claims.
[0023] In a further aspect of the present invention, a computer system including a host computer is contemplated, where the host computer includes a processor, main memory, a display, a device for human input, and non-volatile memory, particularly a hard disk or a solid state drive. The non-volatile memory includes instructions that, when executed by the processor, cause the computer system to implement the method according to the present invention.
[0024] The processor may be a general-purpose microprocessor commonly used as the central unit of a personal computer, or the processor may include one or more processing elements configured to perform special calculations, such as a graphics processor. In an alternative embodiment of the present invention, instead of or in addition to the processor, a programmable logic device such as an FPGA and / or an FPGA including an IP core microprocessor configured to provide, for example, a fixed range of functions may be used.
[0025] Brief Description of the Drawings A better understanding of the present invention can be obtained by considering the following detailed description of the preferred embodiments in combination with the following drawings.
Brief Description of the Drawings
[0026]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Mode for Carrying Out the Invention
[0027] In the drawings, like elements are assigned the same reference numerals. The present invention can take various different forms and alternative forms, but in the drawings, specific embodiments are shown as examples and are described in detail herein. However, it is self-evident that the drawings and the detailed description of specific embodiments do not limit the present invention to the disclosed particular forms. On the contrary, the present invention encompasses all the following variations, equivalents, and alternative forms within the inventive concept and the scope of validity of the present invention as defined by the appended claims.
[0028] FIG. 1 shows an exemplary embodiment of a computer system.
[0029] The illustrated embodiment includes a host computer PC having a display ANZ and a user interface device such as a keyboard TAS and a mouse MAU, and further, an external server can be connected via a network as indicated by the cloud symbol.
[0030] The host computer PC includes at least one processor CPU having one or more cores, a main memory RAM, and a plurality of devices connected to a local bus (such as PCI-Express) that exchanges data with the CPU via a bus controller BS. These devices include, for example, a graphics processor GPU for controlling a display, a control unit USB for connecting peripheral devices, a non-volatile memory HDD such as a hard disk or a solid state drive, and a network interface NC. Further, the host computer can include a dedicated accelerator KI for a neural network. The accelerator may be configured as a programmable logic device such as an FPGA, or may be configured as a graphics processor suitable for general computing, or may be configured as an application-specific integrated circuit. Preferably, the non-volatile memory includes instructions that cause a computer system to implement the method according to the present invention when executed by one or more cores of the processor CPU.
[0031] In an alternative embodiment, the host computer can include one or more servers that include one or more processing elements, as shown as a cloud in the drawings, and the servers are connected via a network to a client that includes a display device and an input device. Thus, the annotation environment can be realized partially or completely on a remote server, for example, within a cloud computer device. A personal computer can be used as a client that includes a display device and an input device via a network. Alternatively, the graphical user interface of the annotation environment can be displayed on a portable computer system, such as on a smartphone or a tablet having a touch screen user interface in particular.
[0032] FIG. 2 shows an exemplary video frame together with a schematic diagram of possible data points in the upper left insertion diagram.
[0033] The drawings show a photograph or a frame of a metropolitan landscape. Such a frame may be part of a video recording. Generally, the recordings provided by customers consist of video data or audio data, and these video data or audio data are in a continuous context such as a 5-minute drive recorded via, for example, a camera and a LiDAR sensor, or a 10-minute voice recording. The video recording may consist of, for example, a series of consecutive frames, and these frames themselves capture a series of objects. The neural network processes the recording to create an annotation that can include a plurality of data points, and each data point represents one specific aspect.
[0034] A data point is a parameter that represents a specific characteristic of the recording and is applicable at all levels of detail. The level of detail may be the entire recording, a series of consecutive or random frames, a single frame, or an object on a frame. Specific examples are annotations for a car consisting of a bounding box that represents the position of the car with a certain degree of accuracy, vertical lines marking the edges of the car, a classification for representing the type of the car, attributes regarding trimming or occlusion, winkers, brake lights, color, etc. A data point may be a classification, a frame, a segment, a polygon, a polyline, an attribute such as a winker, a brake light, a color, a sub-classification, tracking information, degree of occlusion, degree of trimming, a complex classification representing the importance of an object / frame / clip, sound, text, sensation, or any other automatically identifiable information.
[0035] In the upper left inset in the drawing, various different data points regarding automobiles are shown. The automobiles can be of various different types, for example, a delivery vehicle, an SUV, or a sports car. The position or rather the dimensions of the automobile are generally indicated by a bounding box, that is, a rectangular frame or cuboid surrounding the automobile. The vertical lines indicate the boundaries of the automobile. A further possible data point regarding the automobile is the activation of an indicator light, for example, the direction indicator shown in the inset.
[0036] There are a number of automobiles within the frame, each surrounded by a boundary frame. The automobiles may be fully visible, for example, like a car driving straight in front of a camera, or they may be obscured. The traffic density in a metropolitan landscape may inhibit the annotation quality, for example, by making it difficult to accurately determine the boundary lines of the boundary frame due to occlusion.
[0037] Figure 3 shows a schematic diagram of an exemplary packet of video frames.
[0038] A normal method for creating sensor data for training or validating autonomous vehicles is to drive a test driver around everywhere, during which all sensor data of interest, such as camera data, LiDAR data, and / or GPS data, are recorded. Since these data are not sorted, the first recording (Recording 1) may be taken at midday on a highway, while on the other hand, the next recording (Recording 2) may also be taken during the day, but possibly during rainfall. The next recording 3 (Recording 3) may be taken at night. In subsequent recordings, the surrounding conditions may change in an unexpected manner.
[0039] Figure 4 shows a schematic diagram of an exemplary packet of video frames grouped according to additional time information and weather information. Since the visibility of objects strongly depends on the time and weather conditions, the annotation quality for object detectors is correlated with these surrounding conditions.
[0040] It is advantageous to group the recorded frames into packets or clusters according to time and weather conditions. In the illustrated example, recordings 1, 5, and 6 were made during the day on sunny days, while recordings 2, 4, and 7 were made during the day on rainy days due to rainfall. Recording 3 was made at night under rainy conditions.
[0041] Additional criteria may be used to group the recorded frames into packets or clusters. As an example in the context of autonomous driving, the data provided by the customer can be bundled not only based on day / night and rainy / sunny, but also based on road type, e.g., urban roads versus highways.
[0042] To provide a group of frames with uniform annotation quality, frames recorded between similar ambient conditions are processed together. In one embodiment, various different neural networks can be used to annotate the frames based on at least one ambient condition that was present at the time of recording of each frame.
[0043] Figure 5 shows an exemplary diagram showing the correlation between ambient conditions and annotation quality.
[0044] The frames of the first group of cluster 1, which includes recordings 2, 4, and 7, were recorded during the day with a rain pattern or on a rainy day. The accuracy of cluster 1 is in the vicinity of 90% based on manual quality checks. Thus, although this annotation still has to be checked, the neural network is capable of creating sufficiently accurate data after repeating retraining several times.
[0045] Frames of the second group of cluster 2 containing records 1 and 5 were recorded during sunny days. The accuracy of cluster 2 is 99% based on manual quality checks. Since this is sufficiently accurate, the quality checks for groups of frames recorded under the same ambient conditions can be completely omitted.
[0046] Frames of the third group of cluster 3 containing records 3 and 8 were recorded during rainy nights. The accuracy of cluster 3 is 50% based on manual quality checks and is thus clearly unacceptable. Frames recorded under the same ambient conditions require thorough manual checks and improved training of the neural network.
[0047] Since the frames are grouped according to ambient conditions, human labor is invested in the groups of frames where it is most needed. Frames recorded under suitable conditions are fully automatable. The computational performance or energy required for new training of the neural network is also invested where it has a significant impact on the quality of the annotation.
[0048] Figure 6 is a schematic diagram of an automated system implementing the method according to the invention. The automated system implements the various different steps of the method in dedicated components and is well adapted for implementation within a cloud computing environment.
[0049] In the first step, "data capture", records not sorted by the customer are received. To enable uniform processing, the records can be normalized, for example, split into multiple frames.
[0050] In the second step, "enrichment", the frames from the recording are analyzed and automatically enriched by metadata related to the measurement of the automated quality. This step is presented as a prerequisite for automation, but in alternative embodiments, this enrichment can also be performed after automation, for example, based on information collected during annotation such as traffic density or the distance from an object to a sensor, depending on the desired metadata. In the context of autonomous driving, the metadata or conditional attributes related to the annotation quality may be geography, weather conditions, road type, lighting conditions, and / or time of day. For the efficiency of automation, it is beneficial to process groups of frames as a whole in subsequent steps. In the case of projects involving nested recording and processing of multiple frames, it may be advantageous to add frames recorded under the same ambient conditions until a predetermined cluster size is achieved before proceeding with further processing steps. Thus, enrichment and clustering include techniques for adding static or dynamic metadata to the recording and techniques for inserting individual recordings into larger clusters of definable size based on metadata enrichment.
[0051] In the third step, "scheduler", different groups of frames are sorted for annotation by the automation engine, which drives one or more automation components to annotate the frames using one or more data points. The scheduler selects a group of frames for processing based on the availability of a new version of the automation component. The automation component may generate a single data point, such as a vertical line, or may generate multiple corresponding data points, such as a bounding box and object classification. The automation component may be a neural network or any other type of technology based on machine learning that learns with, semi-supervised, or without a teacher from data samples.
[0052] In the fourth step, the "Automation Engine", a group of frames is processed by at least one automation component that assigns data points to the frames. The automation system generates any kind of data points respectively via the automation components, that is, the automation components are the central part of the workflow of the annotation system. Preferably, the data points convey metadata that details the version of the automation component used to generate the result. The automation engine includes techniques for accurately storing relevant metadata via the automation components.
[0053] In the fifth step, the "Sample Check", a random sample of frames is selected for quality control. In quality control, for example, frames with corresponding annotations such as bounding boxes can be presented to a human annotator, and the human annotator can be asked whether this bounding box is correct. Alternatively, a user interface for adjusting and / or adding bounding boxes can be presented to the human annotator in case an object is missed by the neural network. The automation system determines a quality measure from the type and number of corrections made by the human annotator.
[0054] In the 6th step, "Did the sample check pass?", the system determines whether the annotated quality or quality measure exceeds a predetermined threshold. When the automated system confirms that the predetermined threshold is exceeded (Yes), the group of frames containing the selected random sample is exported and supplied to the customer. If at least one group of frames recorded under a particular set of ambient conditions passes the sample check, the automated system can determine that the quality of the annotation for all groups of frames having the same ambient conditions may be exported without further quality checks, and thus, steps 5 and 6 should be skipped. In one embodiment, the automated system can count the number of groups having ambient conditions with sufficient annotation quality and skip the check of the random sample as soon as a predetermined number of groups pass the sample check. When the automated system confirms that the group of frames did not pass the sample check (No), the implementation continues in the 8th step, where the automated system identifies whether the frames recorded under the ambient conditions of the selected random sample are required for the dataset. Whether it is required for the dataset may depend on the number of frames recorded under the same conditions that have already been used for training the model. If a sufficient number of frames have already been used for training, the group of frames can simply be inserted into the 3rd step, "Scheduler", for re - processing as soon as the retrained neural network becomes available.
[0055] In the 7th step, "Customer Sample Check", the customer can check a random sample of the exported frames to confirm that the annotation complies with the customer's settings and the required annotation quality. If the customer rejects a group of frames, in the "Correction" step, either the random sample or the entire group of frames is processed manually. Preferably, the automation system enforces sample checks on all subsequent groups with the same ambient conditions until a new group of frames passes the sample check of the 6th step and / or the customer sample check of the 7th step.
[0056] In the 9th step, "Correction", manual annotation is performed on a random sample of frames that failed the test, or on a random sample or entire group of frames rejected by the customer. The manually annotated frames are exported and supplied to the customer for the 7th step, i.e., for customer sample check. The manually annotated frames are also used for retraining the neural network by supplying the corrected data to the training dataset, the validation dataset, or the test dataset. These datasets are symbolically illustrated by cylinders.
[0057] In the 10th step, "Flywheel", retraining of at least one neural network or automation component that generated the data points rejected in the sample check is performed. The retraining of the neural network improves the automation quality. Preferably, the automation component is improved to a level where as few manual inspections as possible are required for as many metadata clusters (i.e., frames recorded in a specific set of ambient conditions). To enable a rapid improvement in efficiency, the iteration time for retraining must be as short as possible.
[0058] The flywheel includes techniques for efficiently storing a training dataset for each respective automation component (each data point) in order to monitor changes in the training dataset and automatically trigger retraining as soon as a predetermined or automatically determined threshold for changes in the training dataset is detected. Further, the flywheel includes techniques for automatically deploying the retrained model to the automation components and for notifying planners of version changes in the automation components.
[0059] When new data is recorded simultaneously with or nested within the annotation of a frame, an additional step of targeted data capture can be performed. The automation components are improved by a large number of training iterations in the dataset that is continuously refined, which better reflects real-world dispersion over time. The confidence level for each metadata cluster enables a systematic approach for capturing exactly those data samples where the automation results are most affected. Referring to FIG. 5, the frames of cluster 3 are recorded at night, and the automatic annotation results in an annotation quality that is actually unacceptable. As soon as this is discovered in the check step, targeted data capture can be requested, and during this targeted data capture, night-time samples are specifically recorded under these surrounding conditions to improve the training dataset of the automation components.
[0060] In a preferred embodiment, the level and amount of additional training data of a particular type (cluster) are determined according to the confidence level. All data recorded under the same conditions can be used for retraining. As soon as the correction of mis-annotated frames is carried out, these frames are directly fed into the training data set of a special automation component. However, usually, it is not necessary to manually correct all data for a particular cluster and data point. Instead, only samples up to the next retraining threshold level are extracted and corrected. The remaining part of the data is automatically scheduled for re-execution using a higher version of the automation component. Targeted data capture includes techniques for selecting samples of interest based on a metadata cluster up to a predetermined amount for manual correction. Further, targeted data capture preferably includes techniques for marking low-quality samples that are not required for retraining to perform automation in a higher version of each automation component.
[0061] By utilizing the correlation between the ambient conditions when the frame is recorded and the resulting annotation quality, the method according to the invention enables manual work to be invested, especially in the rapid improvement of neural networks, where this neural network is used to create automatic annotations for supply to customers, thus significantly accelerating relatively large-scale annotation projects required, for example, for verification.
[0062] A person skilled in the art will recognize that the order of at least some of the steps of the method according to the invention may be changed without departing from the scope of protection of the claimed invention. Although the invention has been described with respect to a limited number of embodiments, a person skilled in the art will recognize numerous modifications and variations of these embodiments. The appended claims are intended to cover all such modifications and variations that are within the true spirit and scope of the invention.
Claims
**Claim 1** A method implemented by a computer for automatically annotating sensor data frames, the method comprising: Receiving a plurality of frames; Grouping the frames into a plurality of packets based on at least one conditional attribute, the conditional attribute representing ambient conditions that existed during the recording of the frames; Using a neural network to annotate frames from a first packet of the plurality of packets, the annotation including assigning at least one data point to each frame, the first packet including frames such that the at least one conditional attribute is within a selected range of values; Selecting a first sample of one or more frames from the first packet and determining a quality measure for the data point, the method further comprising the computer receiving a manually corrected annotation for the frames in the first sample when the computer identifies that the quality measure for at least one frame in the first sample is below a predetermined threshold, and retraining the neural network based on the frames in the first sample; Selecting a second sample of one or more frames from the frames of the first packet that were not included in the first sample; Annotating the frames of the second sample using the retrained neural network to determine a quality measure for the data point; Identifying that the quality measure for the frames in the second sample exceeds a predetermined threshold; Annotating the remaining frames of the first packet using the retrained neural network; Exporting the annotated frames of the first packet; comprising In the case of image frames, the at least one data point includes the position of an object, the classification of the object, the position of the edges of a bounding box, the correlation between the object in the image frame and an object in a preceding or subsequent image frame, and / or the operation of an indicator light, and / or In the case of an audio frame, the at least one data point includes one or more text words identified from the audio frame. Method. **Claim 2** In the case of a frame having image data, the conditional attributes are geographical location, time, weather conditions, visibility conditions, road type, distance to an object, and / or traffic density, and / or In the case of an audio frame, the conditional attributes are geographical location, gender and / or age of the speaker, room size, and / or scale of background noise. The method according to claim 1. **Claim 3** The step of receiving a plurality of frames includes a step of preprocessing the frames. At least one of the conditional attributes for a frame is determined by a dedicated neural network based on the frame, and / or At least one of the conditional attributes for a frame is determined based on additional sensor data recorded simultaneously with the frame. The method according to claim 1. **Claim 4** The first sample includes two or more frames selected from the first packet. As soon as the computer determines that the quality measure for the first sample is below the predetermined threshold, further calculations on the frames from the first packet are not performed until a corrected annotation for the frames in the first sample is received. The method according to claim 1. **Claim 5** The selection of frames for the first sample depends on the data points for which the quality measure is to be determined. In particular, in the case of object detection, a single frame is randomly selected, and / or in the case of object tracking, a batch of consecutive frames is randomly selected. The method according to claim 1. **Claim 6** Selecting a current sample from one or more frames from the first packet; Determining a quality measure for the data point; Receiving a corrected annotation for the frames in the current sample; Retraining the neural network based on the frames in the current sample. is repeated until the quality measure for the frame in the current sample exceeds the predetermined threshold or until the first packet contains no remaining frames, The method according to claim 1.
7. Annotating sensor data and recording sensor data are performed alternately or simultaneously, When it is determined that the quality measure for at least one frame in the first sample is below a predetermined threshold, the computer requests to record additional sensor data such that the at least one conditional attribute is within the selected value range of the first packet. The method according to claim 1.
8. A method for automatically annotating sensor data including frames such as video frames or audio frames, The method is performed by at least one processor of a host computer, The method is, a) receiving a plurality of frames; b) grouping the frames into a plurality of packets based on at least one conditional attribute, the conditional attribute representing ambient conditions that were present during the recording of the frames; c) using a neural network to annotate frames from a first packet of the plurality of packets, the annotation including assigning at least one data point to each frame, the first packet including frames such that the at least one conditional attribute is within a selected value range; d) selecting a first sample of one or more frames from the first packet and determining a quality measure for the data point; e) identifying that the quality measure for at least one frame in the first sample is below a predetermined threshold; f) the host computer receiving a manually corrected annotation for the frames in the first sample and retraining the neural network using the frames in the first sample; g) using the retrained neural network to annotate at least one of the remaining frames of the first packet; h) selecting second samples of one or more frames from at least one annotated remaining frame of the first packet and determining a quality measure for the data points; i) identifying that the quality measure for the frame in the second sample exceeds a predetermined threshold; j) annotating the remaining frames from the first packet using the retrained neural network; k) exporting the annotated frames; comprising in the case of an image frame, the at least one data point includes the position of an object, the classification of the object, the position of the edge of the bounding box, the correlation between the object in the image frame and an object in a preceding or succeeding image frame, and / or the operation of an indicator light, and / or in the case of an audio frame, the at least one data point includes one or more text words identified from the audio frame, a method. **Claim 9** A non-volatile computer-readable medium comprising instructions that, when executed by a microprocessor of a computer system, cause the computer system to perform the method according to any one of claims 1 to 8. **Claim 10** A computer system including a host computer, wherein the host computer includes a microprocessor, direct access memory, a display, a device for human input, and non-volatile memory, especially a hard disk or a solid state drive, wherein the non-volatile memory includes instructions that, when executed by the microprocessor, cause the computer system to perform the method according to any one of claims 1 to 8.