Unmanned model training sample labeling method and device, equipment and storage medium
Patent Information
- Application Number
- CN202211361780.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2042-11-02
AI Technical Summary
然而,由于样本数据的量非常大,需要耗费大量的人力资源去进行标注,使得模型训练的效率低下
[0041]The present invention provides a method for labeling training samples of autonomous driving models, comprising: acquiring a mind map, wherein the mind map has a preset sample labeling process; acquiring a sample image and determining the object to be labeled in the sample image; popping up multiple label options for labelers to label according to the labeling process; responding to the labeler's labeling operation; selecting a target label from the multiple label options to label the object. By utilizing the pre-constructed mind map, when the labeler is labeling, multiple label options pop up for the labeler to label according to the labeling process in the mind map, so that the labeler can select the appropriate target label from the multiple label options to label the object to be labeled based on the content displayed in the sample image, thereby improving labeling efficiency and labeling quality and saving labor costs.
Smart Images

Figure CN115601724B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to autonomous driving technology, and more particularly to a method, apparatus, device, and storage medium for labeling training samples for autonomous driving models. Background Technology
[0002] With the development and transformation of the automotive industry, vehicle technology is also constantly evolving, and a new transportation ecosystem based on autonomous driving technology is facing reconstruction.
[0003] Autonomous driving technology uses onboard sensors to perceive its surrounding environment, collect environmental information, and then uses a pre-trained autonomous driving model in the control device (i.e., the vehicle's intelligent brain) to perform precise calculations and analyses of the environmental information. Finally, it sends commands to the ECU (Electronic Control Unit) to control different devices in the autonomous vehicle, thereby achieving fully automatic operation and the goal of autonomous driving.
[0004] Currently, supervised learning is commonly used to train autonomous driving models. However, this method requires collecting and labeling a large amount of sample data. The sheer volume of data necessitates significant human resources for labeling, leading to inefficient model training. Furthermore, the quality of manually labeled data is difficult to guarantee. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and storage medium for labeling training samples of autonomous driving models, so as to improve labeling efficiency and quality and save labor costs.
[0006] In a first aspect, the present invention provides a method for labeling training samples for an autonomous driving model, comprising:
[0007] Obtain a mind map, which contains a pre-defined sample annotation process;
[0008] Acquire sample images and identify the objects to be labeled in the sample images;
[0009] Following the described annotation process, multiple label options will pop up for the annotator to annotate.
[0010] In response to the annotator's annotation operation, a target label is selected from a plurality of the label options to annotate the object.
[0011] Optionally, determining the objects to be labeled in the sample image includes:
[0012] The sample images are input into a pre-trained target detection model for processing, and the objects to be labeled are detected from the sample images.
[0013] Optionally, the sample images are input into a pre-trained object detection model for processing to detect the objects to be labeled from the sample images, including:
[0014] Extract candidate regions including the object from the sample image;
[0015] Extract image features from the candidate regions;
[0016] The category of the object is determined based on the image features.
[0017] Optionally, following the described annotation process, a variety of label options will pop up for the annotator to annotate, including:
[0018] In response to the annotator's current operation, the process node corresponding to the current operation is determined from the mind map, and multiple label options to be annotated under the process node are displayed on the annotation interface.
[0019] Optionally, in response to the annotator's current operation, the process node corresponding to the current operation is determined from the mind map, and multiple label options to be annotated under the process node are displayed on the annotation interface, including:
[0020] In response to the annotator's image loading operation, the sample image is loaded into the annotation interface;
[0021] The process node corresponding to the image loading operation is determined to be a type labeling node;
[0022] The annotation interface displays multiple type labels to be annotated under the type annotation node, and the type labels are used to represent the type of the object.
[0023] Optionally, after selecting a target label from a plurality of label options to label the object in response to the labeler's labeling operation, the method further includes:
[0024] Based on the target label, determine the process node corresponding to the annotation operation from the mind map;
[0025] Determine the next-level process node corresponding to the process node of the annotation operation, and display the multiple label options to be annotated under the next-level process node on the annotation interface.
[0026] Optionally, the labeling operation is a vehicle type labeling operation, and the multiple label options to be labeled under the next level process node include vehicle facing forward, vehicle facing backward, front wheels, rear wheels, headlights on, headlights off, door open, door closed, trunk open, and trunk closed.
[0027] Optionally, after selecting a target label from a plurality of label options to label the object in response to the labeler's labeling operation, the method further includes:
[0028] From the mind map, identify two target process nodes that have a subordinate relationship;
[0029] Obtain the target labels already marked under the two target process nodes;
[0030] Based on the subordinate relationship between the two target process nodes, verify whether the two target labels are incorrectly labeled.
[0031] Secondly, the present invention also provides an unmanned driving model training sample annotation device, comprising:
[0032] The mind map acquisition module is used to acquire mind maps, which have a pre-set sample annotation process.
[0033] The object identification module is used to acquire sample images and identify the objects to be labeled in the sample images.
[0034] The label option display module is used to pop up multiple label options for labelers to label according to the labeling process;
[0035] The annotation module is used to select a target label from a plurality of label options to annotate the object in response to the annotation operation of the annotator.
[0036] Thirdly, the present invention also provides an electronic device, comprising:
[0037] One or more processors;
[0038] Storage device for storing one or more programs;
[0039] When the one or more programs are executed by the one or more processors, the one or more processors implement the autonomous driving model training sample annotation method provided in the first aspect of the present invention.
[0040] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the autonomous driving model training sample annotation method provided in the first aspect of the present invention.
[0041] The present invention provides a method for labeling training samples of autonomous driving models, comprising: acquiring a mind map, wherein the mind map has a preset sample labeling process; acquiring a sample image and determining the object to be labeled in the sample image; popping up multiple label options for labelers to label according to the labeling process; responding to the labeler's labeling operation; selecting a target label from the multiple label options to label the object. By utilizing the pre-constructed mind map, when the labeler is labeling, multiple label options pop up for the labeler to label according to the labeling process in the mind map, so that the labeler can select the appropriate target label from the multiple label options to label the object to be labeled based on the content displayed in the sample image, thereby improving labeling efficiency and labeling quality and saving labor costs.
[0042] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 A flowchart of a method for labeling training samples for an autonomous driving model provided in an embodiment of the present invention;
[0045] Figure 2 This is a partial schematic diagram of a mind map provided in an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram illustrating a label option display according to an embodiment of the present invention;
[0047] Figure 4 This is another schematic diagram of a mind map provided in an embodiment of the present invention;
[0048] Figure 5 This is a schematic diagram illustrating another label option provided in an embodiment of the present invention;
[0049] Figure 6 This is a schematic diagram of the structure of an unmanned driving model training sample annotation device provided in an embodiment of the present invention;
[0050] Figure 7 This is a schematic diagram of the structure of an electronic device provided as an embodiment of the present invention.
[0051] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0052] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0053] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0054] Figure 1 This is a flowchart illustrating a method for labeling training samples for an autonomous driving model, provided in an embodiment of the present invention. This embodiment is applicable to labeling training samples for autonomous driving models. The method can be executed by an autonomous driving model training sample labeling device provided in this embodiment. This device can be implemented in software and / or hardware, and is typically configured in an electronic device, such as... Figure 1 As shown, the method for labeling training samples for autonomous driving models specifically includes the following steps:
[0055] S101. Obtain the mind map, which contains a pre-defined sample annotation process.
[0056] In this embodiment of the invention, a mind map including the sample annotation process is pre-constructed. The mind map records the object to be annotated and the label options representing the attributes of the object to be annotated. For example, if the object to be annotated is a car, the label options representing the attributes of the object to be annotated may include vehicle model, vehicle orientation, whether the headlights are on, whether the doors are open, whether the trunk is open, etc. If the object to be annotated is a person, the label options representing the attributes of the object to be annotated may include whether the person is riding a bicycle, posture, whether the person is occluded, etc.
[0057] In this embodiment of the invention, when it is necessary to annotate sample images, a pre-constructed mind map can be read, and then the annotator can be assisted in annotating according to the sample annotation process in the mind map.
[0058] S102. Obtain the sample image and determine the objects to be labeled in the sample image.
[0059] In this embodiment of the invention, the sample image can be a real-world image collected by the unmanned vehicle during its daily tasks. When it is necessary to annotate the sample image, the collected sample image is acquired, and the object to be annotated in the sample image is determined. For example, in this embodiment of the invention, the object to be annotated may include cars, trucks, pedestrians, two-wheeled vehicles, three-wheeled vehicles, trolleys, construction areas, stationary obstacles, etc., and this embodiment of the invention does not impose any limitations.
[0060] In some embodiments of the present invention, the annotator can manually determine the objects to be annotated in the sample image, for example, the annotator selects the objects to be annotated in the sample image.
[0061] In other embodiments of the present invention, a pre-trained object detection model can be used to process the sample images and detect the objects to be labeled from the sample images, thereby saving labor costs and improving labeling efficiency. The object detection model can be a one-stage object detection model or a two-stage object detection model. One-stage object detection models can include YOLO and SSD models, while two-stage object detection models can include R-CNN, SPPNet, Fast R-CNN, Faster R-CNN, FPN, etc., and the embodiments of the present invention are not limited thereto. The embodiments of the present invention use a pre-trained object detection model to perform object detection on the sample images, automatically detecting the objects to be labeled in the sample images, eliminating the need for manual detection by labelers and improving labeling efficiency.
[0062] For example, Fast R-CNN is used as an example in this embodiment of the invention to illustrate the object detection process of the object detection model. The processing procedure of the Fast R-CNN model is as follows:
[0063] 1. Extract candidate regions containing objects from sample images.
[0064] For example, a Fast R-CNN model typically includes a backbone network and region proposal networks.
[0065] After acquiring the sample images, they are input into the backbone network for processing to obtain feature maps of the in-vehicle images. Specifically, the backbone network can be a deep convolutional network, including multiple convolutional layers, multiple activation function layers, and multiple pooling operation layers. In one embodiment of the present invention, the backbone network includes 13 convolutional layers, 13 activation function layers, and 4 pooling operation layers. The kernel size and stride of each convolutional layer, and the pooling window and stride of the pooling operation layers can be set as needed. The pooling operation can be max pooling or average pooling, which is not limited in this embodiment of the present invention. Max pooling uses the largest number in a local region selected within the pooling window to represent that region, retaining the largest feature. Average pooling uses the average value in a local region selected within the pooling window to represent that region. Pooling is used to reduce the number of training parameters, reduce the dimensionality of the features output by the convolutional layers, reduce overfitting, retain only the most useful feature information, and reduce noise propagation. The activation functions of each activation function layer can be ReLU, Sigmoid, or Tanh functions, and this embodiment of the invention is not limited thereto. It should be noted that in other embodiments of the invention, the backbone network can also be a residual network (ResNet) or a VGG network, etc., and this embodiment of the invention is not limited thereto.
[0066] After extracting the feature map from the sample image, the feature map is input into a pre-defined region candidate network for processing to determine candidate regions of interest (RoIs). Specifically, for each element in the feature map, multiple anchor boxes of different scales are generated, centered on that element. The matrix representation of the anchor box is [x, y, w, h], where x and y are the coordinates of the anchor box's center point, and w and h are the width and height of the anchor box. Object detection is performed on each anchor box obtained in the above steps to determine whether the anchor box contains a detected object. When the anchor box contains a detected object (only the existence of a detected object is known, but the specific category is unknown), the anchor box is considered a target anchor box containing the detected object. Simultaneously, a maximum suppression algorithm is used to exclude target anchor boxes representing the same object, ultimately ensuring that one object corresponds to one target anchor box. Linear regression is performed on the width, height, and center point offset of the target anchor box to achieve translation and scaling processing, obtaining candidate regions where the detected object is completely located within the candidate regions.
[0067] 2. Extract image features from candidate regions.
[0068] For example, the RoI Pooling layer receives the features of the candidate regions output by the region candidate network, maps the features of the candidate regions to the corresponding positions in the feature map to obtain the mapped regions, then divides the mapped regions into several sub-regions, and then performs max pooling and full convolution operations on the sub-regions to obtain high-dimensional image features.
[0069] 3. Determine the category of the object based on image features.
[0070] For example, high-dimensional image features are input into a pre-set classifier to obtain the type of the object, such as a car, truck, pedestrian, two-wheeled vehicle, three-wheeled vehicle, cart, construction area, stationary obstacle, etc.
[0071] S103. Following the annotation process, multiple label options will pop up for the annotator to annotate.
[0072] In this embodiment of the invention, the mind map records the sample annotation process, as well as objects to be annotated without process nodes and label options representing the attributes of the objects to be annotated. This embodiment of the invention displays multiple label options for annotators according to the annotation process recorded in the mind map, allowing annotators to select the correct label based on the content displayed in the sample image, thereby achieving annotation of the sample image and improving annotation efficiency and accuracy.
[0073] In some embodiments of the present invention, each operation of the annotator corresponds to a process node in the mind map. When the annotator performs the current operation, the electronic device responds to the annotator's current operation, determines the process node corresponding to the current operation from the mind map, and displays multiple label options to be annotated under the process node on the annotation interface.
[0074] S104. In response to the annotator's annotation operation, select the target label from multiple label options to annotate the object.
[0075] After the annotation interface displays multiple label options for the annotator to choose from, the annotator can select a label that matches the object to be annotated as the target label based on the content displayed in the sample image, and then label the object to be annotated.
[0076] The autonomous driving model training sample annotation method provided in this embodiment of the invention includes: acquiring a mind map, in which a sample annotation process is preset; acquiring a sample image and determining the object to be annotated in the sample image; popping up multiple label options for annotators to annotate according to the annotation process; responding to the annotator's annotation operation; selecting a target label from the multiple label options to annotate the object; by utilizing the pre-constructed mind map, when the annotator annotates, multiple label options pop up for the annotator to annotate according to the annotation process in the mind map, so that the annotator can select a suitable target label from the multiple label options to annotate the object to be annotated based on the content displayed in the sample image, thereby improving annotation efficiency and annotation quality and saving labor costs.
[0077] In some embodiments of the present invention, after selecting a target label from multiple label options to label the object in response to the labeler's labeling operation, the method further includes:
[0078] Based on the target labels, determine the process nodes corresponding to the annotation operations from the mind map.
[0079] Determine the next-level process node corresponding to the annotation operation, and display the multiple label options to be annotated under the next-level process node on the annotation interface.
[0080] In this embodiment of the invention, each annotation process node has a certain annotation order. After annotating the object to be annotated according to the current process node, the next-level process node corresponding to the annotation operation is determined, and multiple label options to be annotated under the next-level process node are displayed on the annotation interface.
[0081] Figure 2 This is a partial schematic diagram of a mind map provided in an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating a label option display provided in an embodiment of the present invention, exemplarily, as shown below. Figure 2 , 3As shown, each rectangle in the mind map represents a node. The annotator's current operation can be an image loading operation. The electronic device responds to the annotator's image loading operation by loading the sample image onto the annotation interface. After loading the sample image onto the annotation interface, the process node corresponding to the image loading operation is determined to be the object type annotation node, which indicates the type of object in the sample image that needs to be annotated. At this time, the annotation interface displays multiple type labels to be annotated under the object type annotation node. The object type labels are used to represent the type of object, such as car, truck, pedestrian, two-wheeled, three-wheeled, cart, construction area, stationary obstacle, etc. The type of object can be detected by the object detection network in the aforementioned embodiment. After the annotator selects one of the labels, it indicates that objects of that type need to be annotated. For example, after the annotator selects the "car" label, the object in the sample image will be labeled "car". If there are multiple objects of the same type in the sample image, the annotator can select a specific object in the sample image for annotation.
[0082] like Figure 2 , 3 As shown, the annotator's current operation can be an annotation operation, such as selecting the "Car" label. After selecting the "Car" label, the next-level process node corresponding to this annotation operation is determined. The next-level process node is the Car Type Annotation Node, which means that the selected car type needs to be annotated. At this time, the annotation interface displays multiple car type labels to be annotated under the Car Type Annotation Node. Car type labels are used to indicate the type of car, such as Car (Ordinary), Van, Police Car, Ambulance, etc. The annotator selects the label that matches the object to be annotated in the sample image based on the content displayed in the sample image. For example, if the object to be annotated in the sample image is an ordinary car, the annotator can select the "Car (Ordinary)" label and label the car in the sample image as "Car (Ordinary)".
[0083] Figure 4 This is another schematic diagram of the mind map provided in an embodiment of the present invention. Figure 5 This is an illustration of another label option provided in an embodiment of the present invention, such as... Figure 4 , 5As shown, the annotator's current operation can be an annotation operation, such as selecting the "Car (Normal)" label. After selecting the "Car (Normal)" label, the next-level process node corresponding to this annotation operation is determined. The next-level process node is the car attribute annotation node, which means that the attributes of the selected car (e.g., doors, wheels, etc.) need to be annotated. At this time, the annotation interface displays multiple car attributes to be annotated under the car attribute annotation node, such as normal door opening, wheels, and special door opening. Among them, normal door opening can refer to whether the two side doors are open or not, special door opening can refer to whether the trunk is open or not, and wheels can refer to the front wheels or the rear wheels. The annotator selects the label that matches the object to be annotated in the sample image based on the content displayed in the sample image. For example, if the normal car to be annotated in the sample image shows wheels, the annotator can select the "Wheels" label and label the car in the sample image with "Wheels".
[0084] Furthermore, such as Figure 4 , 5 As shown, the annotator's current operation can be an annotation operation, such as selecting the "Wheel" label. After selecting the "Wheel" label, the next-level process node corresponding to this annotation operation is determined. The next-level process node is the wheel attribute annotation node, which means that the attributes of the selected wheel need to be annotated (e.g., front wheel or rear wheel). At this time, the annotation interface displays multiple wheel attributes to be annotated under the wheel attribute annotation node, such as front wheel and rear wheel. The annotator selects the label that matches the object to be annotated in the sample image based on the content displayed in the sample image. For example, if the sample image shows the rear wheel of a regular car, the annotator can select the "Rear Wheel" label and label the car in the sample image with "Rear Wheel".
[0085] In this embodiment of the invention, the multiple label options to be labeled at the next-level process node include vehicle facing forward, vehicle facing backward, front wheels, rear wheels, headlights on, headlights off, door open, door closed, trunk open, and trunk closed. The labeling process can be referred to the foregoing. Figures 2-5 The process shown in this embodiment of the invention will not be described again here.
[0086] In some embodiments of the present invention, after the target objects are labeled, the labels can be verified to improve the accuracy of the labeling. Specifically, two target process nodes with a subordinate relationship are identified from the mind map, and the target labels already labeled under the two target process nodes are obtained. Based on the subordinate relationship between the two target process nodes, it is verified whether the two target labels are labeled incorrectly. For example, taking an object type labeling node and a car type labeling node as an example, the car type labeling node is a subordinate node of the object type labeling node. If the label labeled by the object type labeling node is "car", then the label labeled by the car type labeling node must be a subordinate label of the car, which is a further description of the car, such as "car (ordinary)", "van", "police car", "ambulance". If the label labeled by the car type labeling node is not a further description of the car, it indicates that the labeling is incorrect. At this time, an alarm can be issued to the labeler. The alarm form can include sound, light, and a combination of both. The embodiments of the present invention are not limited here.
[0087] Annotation often requires selecting the object to be annotated. In some embodiments of this invention, when annotating an object, a rectangular box can be used to select the object in the sample image, as well as a local area within the object. For example, when annotating the object type, a large rectangular box is generated to select the car to be annotated. When annotating other attributes of the car, a smaller rectangular box is generated within the large rectangular box to select the local area to be annotated, such as the rear wheel. The large rectangular box that selects the entire object is called the main frame, and the rectangular box that selects the local area to be annotated is called the sub-frame. Sub-frames are always within the main frame. There may be overlapping main frames in the image, so generally, a main frame is selected, and then a sub-frame is generated within the selected frame. Each frame is a rectangle with four coordinate points. The x-coordinate of the sub-frame is controlled between the minimum and maximum x-coordinates of the main frame, and the y-coordinate of the sub-frame is controlled between the minimum and maximum y-coordinates of the main frame. Assume that the minimum point in the x-direction of the main frame is Xmin, the maximum point is Xmax, the minimum point in the y-direction is Ymin, and the maximum point in the y-direction is Ymax. The coordinates of each x-point in the sub-frame are greater than or equal to Xmin and less than or equal to Xmax, and the coordinates of each y-point in the sub-frame are greater than or equal to Ymin and less than or equal to Ymax.
[0088] This invention also provides a device for labeling training samples for autonomous driving models. Figure 6 This is a schematic diagram of the structure of an unmanned driving model training sample annotation device provided in an embodiment of the present invention, as shown below. Figure 6 As shown, the autonomous driving model training sample annotation device includes:
[0089] Mind map acquisition module 201 is used to acquire mind maps, wherein the mind maps have a preset sample annotation process;
[0090] The annotation object determination module 202 is used to acquire a sample image and determine the objects to be annotated in the sample image;
[0091] The label option display module 203 is used to pop up multiple label options for labelers to label according to the labeling process;
[0092] The annotation module 204 is used to select a target label from a plurality of label options to annotate the object in response to the annotation operation of the annotator.
[0093] In some embodiments of the present invention, the annotation object determination module 202 includes:
[0094] The object identification submodule is used to input the sample image into a pre-trained target detection model for processing, and to detect the object to be labeled from the sample image.
[0095] In some embodiments of the present invention, the annotation object determination submodule includes:
[0096] A candidate region extraction unit is used to extract candidate regions including the object from the sample image;
[0097] The feature extraction unit is used to extract image features of the candidate region;
[0098] An object category determination unit is used to determine the category of the object based on the image features.
[0099] In some embodiments of the present invention, the label option display module 203 includes:
[0100] The label option display submodule is used to respond to the annotator's current operation, determine the process node corresponding to the current operation from the mind map, and display multiple label options to be annotated under the process node on the annotation interface.
[0101] In some embodiments of the present invention, the label option display submodule includes:
[0102] An image loading unit is used to load the sample image into the annotation interface in response to the image loading operation of the annotator;
[0103] A node determination unit is used to determine that the process node corresponding to the image loading operation is a type labeling node;
[0104] The label display unit is used to display multiple type labels to be labeled under the type label node on the labeling interface. The type labels are used to indicate the type of the object.
[0105] In some embodiments of the present invention, the unmanned driving model training sample annotation device further includes:
[0106] The node determination module is used to determine the process node corresponding to the annotation operation from the mind map based on the target label after selecting a target label from multiple label options to annotate the object in response to the annotator's annotation operation.
[0107] The label display module is used to determine the next-level process node of the process node corresponding to the labeling operation, and to display multiple label options to be labeled under the next-level process node on the labeling interface.
[0108] In some embodiments of the present invention, the labeling operation is a vehicle type labeling operation, and the multiple label options to be labeled under the next level process node include vehicle facing forward, vehicle facing backward, front wheels, rear wheels, headlights on, headlights off, door open, door closed, trunk open, and trunk closed.
[0109] In some embodiments of the present invention, the unmanned driving model training sample annotation device further includes:
[0110] The target process node determination module is used to determine two target process nodes with a subordinate relationship from the mind map after selecting a target label from multiple label options to label the object in response to the labeler's labeling operation;
[0111] The target label acquisition module is used to acquire the target labels that have been marked under the two target process nodes;
[0112] The verification module is used to verify whether the two target labels are incorrectly labeled based on the subordinate relationship between the two target process nodes.
[0113] The above-mentioned unmanned driving model training sample annotation device can execute the unmanned driving model training sample annotation method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of executing the unmanned driving model training sample annotation method.
[0114] This application provides an electronic device. Figure 7This is a schematic diagram of an electronic device provided for an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0115] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0116] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0117] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the autonomous driving model training sample annotation method.
[0118] In some embodiments, the autonomous driving model training sample annotation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the autonomous driving model training sample annotation method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the autonomous driving model training sample annotation method by any other suitable means (e.g., by means of firmware).
[0119] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0123] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0124] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0125] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the autonomous driving model training sample annotation method provided in any embodiment of this application.
[0126] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0127] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0128] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for labeling training samples for an autonomous driving model, characterized in that, include: Obtain a mind map, which contains a pre-defined sample annotation process; Acquire sample images and identify the objects to be labeled in the sample images; The annotation process will bring up multiple label options for the annotator to annotate; each action of the annotator corresponds to a process node in the mind map. In response to the annotator's annotation operation, a target label is selected from multiple label options to annotate the object; The process of determining the objects to be labeled in the sample image includes: The sample images are input into a pre-trained target detection model for processing, and the objects to be labeled are detected from the sample images. After selecting a target label from a plurality of label options to label the object in response to the labeler's labeling operation, the method further includes: From the mind map, identify two target process nodes that have a subordinate relationship; Obtain the target labels already marked under the two target process nodes; Based on the subordinate relationship between the two target process nodes, verify whether the two target labels are incorrectly labeled.
2. The method for labeling training samples for an autonomous driving model according to claim 1, characterized in that, The sample images are input into a pre-trained object detection model for processing to detect the objects to be labeled from the sample images, including: Extract candidate regions including the object from the sample image; Extract image features from the candidate regions; The category of the object is determined based on the image features.
3. The method for labeling training samples for autonomous driving models according to any one of claims 1-2, characterized in that, Following the described annotation process, several label options will pop up for the annotator to annotate, including: In response to the annotator's current operation, the process node corresponding to the current operation is determined from the mind map, and multiple label options to be annotated under the process node are displayed on the annotation interface.
4. The method for labeling training samples for an autonomous driving model according to claim 3, characterized in that, In response to the annotator's current action, the process node corresponding to the current action is determined from the mind map, and multiple label options to be annotated under the process node are displayed on the annotation interface, including: In response to the annotator's image loading operation, the sample image is loaded into the annotation interface; The process node corresponding to the image loading operation is determined to be a type labeling node; The annotation interface displays multiple type labels to be annotated under the type annotation node, and the type labels are used to represent the type of the object.
5. The method for labeling training samples for an autonomous driving model according to claim 3, characterized in that, After selecting a target label from a plurality of label options to label the object in response to the labeler's labeling operation, the method further includes: Based on the target label, determine the process node corresponding to the annotation operation from the mind map; Determine the next-level process node corresponding to the process node of the annotation operation, and display the multiple label options to be annotated under the next-level process node on the annotation interface.
6. The method for labeling training samples for an autonomous driving model according to claim 5, characterized in that, The labeling operation is a vehicle type labeling operation. The multiple label options to be labeled under the next level process node include vehicle facing forward, vehicle facing backward, front wheels, rear wheels, headlights on, headlights off, door open, door closed, trunk open, and trunk closed.
7. A training sample annotation device for an unmanned driving model, characterized in that, include: The mind map acquisition module is used to acquire mind maps, which have a pre-set sample annotation process. The object identification module is used to acquire sample images and identify the objects to be labeled in the sample images. The label option display module is used to pop up multiple label options for the labeler to label according to the labeling process; wherein, each operation of the labeler corresponds to a process node in the mind map; The annotation module is used to select a target label from a plurality of label options to annotate the object in response to the annotation operation of the annotator; The annotation object determination module includes: The object identification submodule is used to input the sample image into a pre-trained target detection model for processing, and to detect the object to be labeled from the sample image; The unmanned driving model training sample annotation device also includes: The target process node determination module is used to determine two target process nodes with a subordinate relationship from the mind map after selecting a target label from multiple label options to label the object in response to the labeler's labeling operation; The target label acquisition module is used to acquire the target labels that have been marked under the two target process nodes; The verification module is used to verify whether the two target labels are incorrectly labeled based on the subordinate relationship between the two target process nodes.
8. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the autonomous driving model training sample annotation method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the autonomous driving model training sample annotation method as described in any one of claims 1-6.
Citation Information
Patent Citations
Sample hierarchical annotation and model training method and device and electronic equipment
CN109815978A
Combined label labeling method and system
CN115098712A