Image labeling method and device, electronic equipment and medium

By using pre-trained image annotation model to automatically annotate real-time and offline image data in infrastructure scenarios, the problems of low efficiency and low accuracy of manual annotation in the prior art are solved, and more efficient and accurate image annotation is achieved.

CN120070939APending Publication Date: 2025-05-30BEIJING CHINA POWER INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411964846.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, illegal operations and safety hazard identification of infrastructure scene images rely on manual annotation, resulting in low labeling accuracy and low efficiency.

Method used

An image annotation method is proposed, by obtaining real-time and offline image data and inputting it into a pre-trained image annotation model, the model is used to identify and annotate the target objects, and realize automated annotation.

Benefits of technology

Through automated labeling, the efficiency and accuracy of image labeling are improved, the omission of manual labeling is avoided, and the ability to identify safety hazards in infrastructure scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070939A_ABST
    Figure CN120070939A_ABST
Patent Text Reader

Abstract

The invention provides an image annotation method and device, electronic equipment and a medium, and the method comprises the steps: obtaining initial image data which comprises real-time image data and offline image data; and inputting the initial image data into an image annotation model obtained by pre-training, and identifying and annotating a target object in the initial image data by using the image annotation model to obtain annotated image data. Through the trained image annotation model, automatic identification and annotation of the target object are realized, manual annotation by a user is not needed, the annotation efficiency is improved, meanwhile, the phenomenon of omission caused by manual annotation can be avoided, and the annotation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and particularly to an image annotation method, apparatus, electronic device and medium. Background Art

[0002] In infrastructure scenarios, there may be situations where operators perform illegal operations or equipment has potential safety hazards. If not discovered and resolved in a timely manner, serious safety problems may occur. Therefore, it is necessary to identify illegal operations or potential safety hazards in images of infrastructure scenarios.

[0003] Currently, most unstructured image data annotation methods for image recognition object detection mainly rely on manual use of tools to mark the rectangular areas of interest in the image and generate label data, with low annotation accuracy and efficiency. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to propose an image annotation method, apparatus, electronic device and medium to solve or partially solve the above problems.

[0005] Based on the above purpose, the first aspect of the present disclosure provides an image annotation method, the method comprising:

[0006] Obtaining initial image data, where the initial image data includes real-time image data and offline image data;

[0007] Inputting the initial image data into a pre-trained image annotation model, and using the image annotation model to identify and annotate the target object in the initial image data to obtain annotated image data.

[0008] Based on the same inventive concept, the second aspect of the present disclosure proposes an image annotation apparatus, comprising:

[0009] A data acquisition module configured to obtain initial image data, where the initial image data includes real-time image data and offline image data;

[0010] A model processing module that inputs the initial image data into a pre-trained image annotation model, and uses the image annotation model to identify and annotate the target object in the initial image data to obtain annotated image data.

[0011] Based on the same inventive concept, the third aspect of the present disclosure proposes an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, where the processor implements the above-mentioned image annotation method when executing the computer program.

[0012] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the image annotation method as described above.

[0013] Based on the same inventive concept, a fifth aspect of the present disclosure provides a vehicle including the image annotation device described in the second aspect, the electronic device described in the third aspect, or the storage medium described in the fourth aspect.

[0014] As can be seen from the above, the present disclosure provides an image annotation method, device, electronic device, and medium. Initial image data is obtained, where the initial image data includes real-time image data and offline image data. The initial image data is input into a pre-trained image annotation model, where the image annotation model is used to identify and annotate target objects in the image data. After being processed by the image annotation model, annotated image data is output. Through the trained image annotation model, automatic identification and annotation of target objects are achieved, eliminating the need for manual annotation by users, improving the annotation efficiency, and at the same time avoiding the omission phenomenon in manual annotation, thus improving the annotation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 is a flowchart of the image annotation method according to an embodiment of the present disclosure;

[0017] Figure 2 is a corresponding flowchart of the image annotation device according to an embodiment of the present disclosure;

[0018] Figure 3 is a schematic diagram of the underlying interface of the image annotation device according to an embodiment of the present disclosure;

[0019] Figure 4 is a structural block diagram of the image annotation device according to an embodiment of the present disclosure;

[0020] Figure 5 is a schematic structural diagram of the electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the objectives, technical solutions, and advantages of the present disclosure more clearly understood, the following further elaborates on the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0022] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0023] The following are the explanations of the terms related to the present disclosure:

[0024] YOLOv3: YOLOv3 is the third version of the YOLO series of object detection algorithms, released by Joseph Redmon and Alexey Bochkovskiy in 2018. It is an improvement based on YOLOv2, introducing a series of changes to improve detection performance.

[0025] URL: URL (Uniform Resource Locator). It is the unified resource location identifier for the www. Simply put, a URL is the web address, commonly known as the "website address". A URL is a concise representation of the location and access method of resources obtained on the Internet and is the address of standard resources on the Internet.

[0026] HTTP: The Hypertext Transfer Protocol (HTTP) is a simple request-response protocol that usually runs on top of TCP. It specifies what messages the client may send to the server and what responses it will receive.

[0027] FTP: The File Transfer Protocol (FTP) is a widely used standard protocol for file transfer in a network. As a basic tool in network communication, FTP allows users to interact with the server through client software to achieve file upload, download and other file operations.

[0028] In the infrastructure scenario, there may be situations where operators perform illegal operations or equipment has potential safety hazards. If these situations are not discovered and resolved in a timely manner, serious safety problems may occur. Therefore, it is necessary to identify illegal operations or potential safety hazards in the images of the infrastructure scenario.

[0029] Currently, most of the annotation methods for unstructured image data in object detection of image recognition mainly rely on manual use of tools to mark the rectangular regions of interest in the image and generate label data. The annotation accuracy is not high and the efficiency is low.

[0030] The mainstream technologies and methods in image recognition are all supervised learning algorithms. Compared with unsupervised learning, the important difference is that supervised learning is a machine learning task of learning target features from a labeled training set and inferring the feature function. Therefore, the prerequisite is to annotate the data.

[0031] Practice has proved that in the development process of ordinary image recognition applications based on deep learning, the quantity and quality of unstructured image materials have an unexpectedly large impact on the final model effect. Just having the "quantity" is not enough, and the quality of data annotation also needs to be ensured. The quantity of data and the quality of annotation determine the upper limit of the task, and the algorithm is just approaching this upper limit continuously.

[0032] Currently, most of the annotation methods for unstructured image data in object detection of image recognition mainly rely on manual use of tools to mark the rectangular regions of interest in the image and generate label data. There are many annotation tools for image data. The main tool for marking rectangular regions is the open-source labelImg. By dragging and dropping, each region of interest in the image is drawn, and finally an annotation file for a single image is generated.

[0033] However, data annotation is a time-consuming and laborious task. There are relevant papers reporting the time for annotating target boxes on ImageNet (an open-source image dataset for computer vision tasks). The time for drawing a box is 25.5s, the time for verifying its quality is 9.0s, and the time for checking whether there are other similar objects to be annotated is 7.8s.

[0034] Therefore, there are many specialized unstructured data annotation companies in the market now, including image data annotation companies, which specifically provide image annotation services. In addition, there are also Internet companies with mature artificial intelligence platforms that have developed intelligent annotation modules based on their platforms. On the one hand, after uploading the data to their platforms, automated data annotation tasks need to be completed, mainly for target annotation in some common scenarios, and data security issues are involved; on the other hand, the background support models of the intelligent annotation functions of artificial intelligence platforms are basically models with high computational complexity and large network volumes, which limit their flexible application on ordinary PC terminals.

[0035] Based on the above description, this embodiment proposes an image annotation method, as Figure 1 shown, the method includes:

[0036] Step 101, obtain initial image data, where the initial image data includes real-time image data and offline image data.

[0037] Step 102, input the initial image data into a pre-trained image annotation model, and use the image annotation model to identify and annotate the target object in the initial image data to obtain annotated image data.

[0038] Specifically in implementation, obtain initial image data, and the initial image data is data collected from the infrastructure site. Specifically, the initial image data includes real-time image data and offline image data.

[0039] Among them, the real-time image data is data collected in real time. For the image data collected in real time, it is transmitted over the network through a URL, and the online image can be directly accessed through the URL link. In an application, a network request library or framework can be used to send an HTTP request to obtain the data of the online image. These requests can be GET requests for obtaining image data or POST requests for uploading image data.

[0040] The offline image data is image data obtained offline. The image can be stored in the local file system or a remote server and then downloaded to the required place through the network. The transmission method can include using protocols such as HTTP and FTP for file downloading.

[0041] Pre-train the initial image annotation model, and after the training is completed, obtain the image annotation model, where the initial image annotation model is a model using a neural network structure, and the specific model structure is not limited.

[0042] Input the initial image data into the pre-trained image annotation model, and after being processed by the image annotation model, output the annotated image data, where the annotated image data includes the annotated target object, so that subsequent users can determine the specific category of safety violations according to the annotated image data, and then adjust the corresponding target object to avoid the occurrence of safety accidents.

[0043] The image annotation model is used to identify and annotate target objects in image data. The target object is an object corresponding to the object type included during the training of the image annotation model. The object type is the category of safety violations at the infrastructure construction site. The object categories specifically include at least one of the following: not wearing a safety helmet, improper wearing of the safety helmet's chin strap, improper dressing (wearing short sleeves), improper dressing (wearing shorts), red vests for supervisors, green vests for safety officers, crossing fixed rigid fences, crossing movable rigid fences, crossing warning lines, untidy placement of formwork, oxygen cylinder placed upside down, acetylene cylinder placed upside down, oxygen cylinder and acetylene cylinder placed adjacent to each other, cable reel placed upside down, improper laying of power cables, missing hook safety device, damaged hook safety device, improper covering of holes, no wooden pads under the crane outriggers, rain protection facilities not protecting against rain, etc.

[0044] Through the above solution, initial image data is obtained, where the initial image data includes real-time image data and offline image data. The initial image data is input into a pre-trained image annotation model. The image annotation model is used to identify and annotate target objects in image data. After being processed by the image annotation model, annotated image data is output. Through the trained image annotation model, automatic identification and annotation of target objects are achieved, eliminating the need for users to manually annotate, improving the annotation efficiency, and at the same time avoiding the omission phenomenon in manual annotation, thus improving the accuracy of annotation.

[0045] In some embodiments, the training process of the image annotation model specifically includes:

[0046] Step 10A: Obtain an annotation data set and an initial image annotation model, where the annotation data set includes original images and the actual annotated images corresponding to the original images;

[0047] Step 10B: Input the original images into the initial image annotation model for training to obtain initial annotation results;

[0048] Step 10C: Compare the initial annotation results with the actual annotated images to determine the target loss function;

[0049] Step 10D: In response to the target loss function converging to a preset convergence threshold, determine that the training of the initial image annotation model is completed to obtain an image annotation model.

[0050] In specific implementation, an annotated dataset and an initial image annotation model are obtained. The annotated dataset includes original images and actual annotated images corresponding to the original images. The actual annotated images are annotated images obtained by manually annotating or using semi-automatic annotation tools to annotate 20 types of infrastructure site safety violation category images according to predefined rules. The predefined rules adopt a bounding box annotation rule: by selecting the target object and annotating the position and category information of each object in the image.

[0051] The original image is input into the initial image annotation model for training to obtain an initial annotation result, where the initial annotation result is an annotated image obtained by the model annotating the original image.

[0052] The initial annotation result is compared with the actual annotated image, that is, the training annotated image output by the model is compared with the actual annotated image to determine the target loss function. The types of the target loss function include at least one of the following: mean square error loss function, cross-entropy loss function, logarithmic loss function, exponential loss function, square loss function, or absolute value loss function, etc.

[0053] In response to the target loss function converging to a preset convergence threshold, it is determined that the training of the initial image annotation model is completed, and an image annotation model is obtained.

[0054] Through the above solution, by using the annotated dataset to train the initial image annotation model and determining the target loss function according to the initial annotation result and the actual annotated image, the determination of the target loss function is more accurate. Furthermore, when the target loss function converges, a trained image annotation model is obtained, and subsequently, the initial image input into the image annotation model can be annotated, and the obtained annotated image is more accurate.

[0055] In some embodiments, step 10B specifically includes:

[0056] Step 10B1, input the original image into the initial image annotation model;

[0057] Step 10B2, obtain a preset plurality of object categories, and perform recognition processing on the original image according to the object categories to obtain a plurality of annotation boxes;

[0058] Step 10B3, determine a target annotation box from the plurality of annotation boxes, and use the target object category and target annotation area corresponding to the target annotation box as the initial annotation result.

[0059] In specific implementation, the original image is input into the initial image annotation model to obtain multiple preset object categories, where the object categories specifically include at least one of the following: the category of not wearing a safety helmet, the category of the safety helmet's chin strap being irregular, the category of improper dressing (wearing short sleeves), the category of improper dressing (wearing shorts), the category of supervisors' red vests, the category of safety officers' green vests, the category of crossing a fixed rigid fence, the category of crossing a movable rigid fence, the category of crossing a warning line, the templates being placed untidily, the oxygen cylinder being placed upside down, the acetylene cylinder being placed upside down, the oxygen cylinder and the acetylene cylinder being placed adjacent to each other, the cable reel being placed upside down, the power cord being laid irregularly, the hook safety device being missing, the hook safety device being damaged, the hole covering being irregular, the crane outrigger without a wooden pad, the rain protection facility not providing rain protection, etc.

[0060] The original image is identified and processed according to the object categories to obtain multiple annotation frames, where each annotation frame corresponds to one object category.

[0061] The target annotation frame is determined from the obtained multiple annotation frames, where the number of the target annotation frames may be one or more. The target object category and the target annotation area corresponding to the target annotation frame are determined, and the target object category and the target annotation area are used as the initial annotation result. Among them, the target annotation area is the area containing the target object corresponding to the object category.

[0062] In some embodiments, before step 10B3 determines the target annotation frame from the multiple annotation frames, it further includes:

[0063] Step 10A, obtaining a preset confidence threshold;

[0064] Step 10B, comparing the confidence corresponding to each annotation frame with the confidence threshold respectively, and removing the annotation frames with the confidence less than the confidence threshold.

[0065] In specific implementation, a preset confidence threshold is obtained, and the confidence threshold is the minimum value of the preset confidence. The confidence corresponding to each annotation frame is compared with the confidence threshold respectively. If there are annotation frames with the confidence less than the confidence threshold, the annotation frames with the confidence less than the confidence threshold are removed.

[0066] Through the above solution, by removing the annotation frames with the confidence lower than the confidence threshold, the influence of the annotation frames with too low confidence on the subsequent training results is avoided, and the accuracy of the image annotation model is further improved.

[0067] In some embodiments, step 10B4 determines the target annotation frame from the multiple annotation frames, specifically including:

[0068] The target annotation frame is determined from the multiple annotation frames through at least one round of iterative operation, and each round of iterative operation is performed as follows:

[0069] Step 10B41, select the annotation box with the highest confidence from all the annotation boxes as the initial target annotation box;

[0070] Step 10B42, denote the annotation boxes other than the initial target annotation box among all the annotation boxes as other annotation boxes;

[0071] Step 10B43, calculate the first intersection over union (IoU) between the initial target annotation box and each of the other annotation boxes according to the annotation area of the initial target annotation box and the annotation areas of the other annotation boxes;

[0072] Step 10B44, delete the other annotation boxes corresponding to the first IoU greater than the preset IoU threshold from all the annotation boxes;

[0073] Step 10B45, determine whether there are any other annotation boxes among all the annotation boxes after deletion;

[0074] Step 10B46, in response to the existence of other annotation boxes, select the annotation box with the highest confidence from all the existing other annotation boxes as the initial target annotation box for the next round of iteration, and enter the next round of iterative operation;

[0075] Step 10B47, in response to the non - existence of other annotation boxes, exit at least one round of iterative operation, and use all the initial target annotation boxes as the target annotation boxes.

[0076] In specific implementation, the target annotation box is determined from multiple annotation boxes through at least one round of iterative operation, and each round of iterative operation is performed as follows:

[0077] Determine the confidence corresponding to each annotation box, and the confidence is expressed by the formula:

[0078]

[0079] where confidence is the confidence, representing the likelihood of the object category contained in the annotation box and the accuracy of the position, lou is the intersection over union, and the overlap rate between the real bounding box of the annotation and the bounding box obtained by the model is calculated.

[0080] Select the annotation box with the highest confidence from all the annotation boxes as the initial target annotation box, and denote the annotation boxes other than the initial target annotation box among all the annotation boxes as other annotation boxes. Determine the annotation area of the initial target annotation box and the annotation areas of the other annotation boxes, and calculate the first intersection over union between the initial target annotation box and each of the other annotation boxes according to the annotation area of the initial target annotation box and the annotation areas of the other annotation boxes.

[0081] Obtain a preset intersection over union (IoU) threshold, and compare each first IoU with the preset IoU threshold. Determine other bounding boxes corresponding to the first IoU being greater than the preset IoU threshold, and delete such other bounding boxes.

[0082] Determine whether there are still other bounding boxes other than the initial target bounding box among all the bounding boxes after deletion. If there are other bounding boxes, at this time, select the bounding box with the highest confidence from all the existing other bounding boxes as the initial target bounding box for the next round, and perform the next round of iterative operation.

[0083] Repeat the iterative operation until there are no other bounding boxes other than the initial target bounding box among all the bounding boxes after deletion. At this time, the iteration ends, exit at least one round of iterative operation, and use all the initial target bounding boxes as the target bounding boxes.

[0084] Exemplarily, the number of bounding boxes is 8. Determine the bounding box A with the highest confidence among the 8 bounding boxes, and use the bounding box A as the initial target bounding box. Calculate the first IoU between the bounding box A and each of the other 7 bounding boxes. If there are 3 first IoUs that are all greater than the preset IoU threshold, then delete the bounding boxes corresponding to the 3 first IoUs.

[0085] At this time, there are still 4 bounding boxes left among all the bounding boxes after deletion. Determine the bounding box B with the highest confidence among the 4 bounding boxes, and use the bounding box B as the initial target bounding box for the next round of iteration. Calculate the first IoU between the bounding box B and each of the other 3 bounding boxes. If there are 3 first IoUs that are all greater than the preset IoU threshold, then delete the bounding boxes corresponding to the 3 first IoUs.

[0086] At this time, there are no other bounding boxes among all the bounding boxes after deletion. At this time, use the bounding box A and the bounding box B as the target bounding boxes.

[0087] Through the above solution, by using non-maximum suppression of the intersection over union to eliminate redundant bounding boxes and suppressing and filtering redundant bounding boxes, the accuracy of the image annotation model is further improved.

[0088] In some embodiments, step 10C specifically includes:

[0089] Step 10C1, recognize the actual annotation image to obtain the image label corresponding to the actual annotation image;

[0090] Step 10C2, compare the object category in the initial annotation result with the image label to determine the target loss function.

[0091] In specific implementation, the actual labeled image is recognized to obtain the image label corresponding to the actual labeled image, and the image label is the label of the object category corresponding to the target object in the actual labeled image.

[0092] Compare the object category in the initial annotation result with the image label, that is, compare the object category output by the model with the actual object category, and determine the target loss function, and the determined target loss function is more accurate.

[0093] In some embodiments, step 101 specifically includes:

[0094] Step 1011, obtain the initial annotation dataset;

[0095] Step 1012, perform image scaling processing on the data in the initial annotation dataset to obtain the first annotation dataset;

[0096] Step 1013, perform region division processing on the data in the first annotation dataset to obtain the annotation dataset.

[0097] In specific implementation, obtain the initial annotation dataset, perform image shrinking processing on the data in the initial annotation dataset, and scale the data in the initial annotation dataset to a preset size to obtain the first annotation dataset.

[0098] Exemplarily, the preset size is 448*448, and the initial annotation dataset includes image data A and image data B. The size of image data A is 144*144, then the size of image A is enlarged to the preset size 448*448. The size of image data B is 600*600, then the size of image B is reduced to the preset size 448*448.

[0099] Perform region division processing on the data in the first annotation dataset according to the preset number of regions to obtain the annotation dataset.

[0100] Exemplarily, the preset number of regions is 7*7, then the data in the first annotation dataset is logically divided according to the 7*7 regions to obtain the annotation dataset.

[0101] Through the above solution, by preprocessing the data in the initial annotation dataset, the data in the initial annotation dataset is made more unified, which is convenient for the subsequent processing of the image annotation model.

[0102] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by the cooperation of multiple devices. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0103] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0104] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an image annotation device, which includes an image acquisition module, a model training module, and an image annotation interface module. Refer to Figure 2 , Figure 2 the flowchart corresponding to the image annotation device of the embodiment, including:

[0105] The image acquisition module is configured to obtain an annotation data set, and the annotation data set includes original images and actual annotation images corresponding to the original images. The actual annotation images are artificial annotation image sets formed by manually annotating or using semi-automatic annotation tools to annotate 20 types of infrastructure site safety violation category images according to predefined rules for further model training. The predefined rules adopt the bounding box annotation rule to select the target object by framing and annotate the position and category information of each object in the image.

[0106] The 20 types of infrastructure site safety violation categories include not wearing a safety helmet, the safety helmet chin strap being non-standard, improper dressing (wearing short sleeves), improper dressing (wearing shorts), supervisors' red vests, safety officers' green vests, crossing fixed hard fences, crossing movable hard fences, crossing warning lines, uneven template placement, oxygen cylinder being placed upside down, acetylene cylinder being placed upside down, oxygen cylinder and acetylene cylinder being placed adjacent to each other, cable reel being placed upside down, non-standard power cord laying, missing hook safety device, damaged hook safety device, non-standard hole covering, no wooden pad under the crane outrigger, rain protection facilities not providing rain protection, etc.

[0107] The model training module is configured to input the original image (unmanually labeled image) into the initial image annotation model for training to obtain an initial annotation result. Identify the actual annotation image (i.e., the manually labeled image) to obtain the image label corresponding to the actual annotation image; compare the object categories in the initial annotation result with the image label to determine the target loss function. In response to the target loss function converging to a preset convergence threshold, determine that the training of the initial image annotation model is completed to obtain an image annotation model.

[0108] The model training module is specifically configured to perform image processing on the obtained infrastructure manual annotation image set: image shrinking and region division. To ensure image consistency, it is uniformly scaled to a size of 448*448 and divided into S*S regions for logical division. In this embodiment, the region is 7*7 regions.

[0109] To identify each object in the image, such as safety helmets, upper garments, materials, etc., based on the center position of the annotated object, predictions are made according to the divided regions it falls into. The positions and confidences of multiple bounding boxes are predicted for each region. The confidence is calculated as follows:

[0110]

[0111] Among them, confidence is the confidence, indicating the likelihood that a bounding box contains a certain infrastructure object and the accuracy of its position. Lou is the intersection over union, which calculates the overlap rate between the annotated true bounding box and the predicted bounding box.

[0112] The model training module is specifically configured to remove the bounding boxes with relatively low credibility based on the confidence threshold and use the NMS algorithm to remove the bounding boxes with relatively high overlap rates. Predict 20 categories of infrastructure site safety violation categories for each region, calculate the probability value of each category to obtain the final prediction result for annotation.

[0113] The image annotation interface module is specifically configured to provide an image recognition interface, access infrastructure offline images and infrastructure real-time images, and automatically annotate.xml format image dataset standard annotation files using the trained image annotation model.

[0114] For the real-time acquired image information, it is transmitted over the network through the URL. Online images can be directly accessed through the URL link. In the application program, a network request library or framework can be used to send HTTP requests to obtain the data of online images. These requests can be GET requests for obtaining image data or POST requests for uploading image data.

[0115] For the offline acquired image information, the schematic diagram of the underlying interface of the image annotation device is asFigure 3 As shown, the image can be stored in the local file system or on a remote server and then downloaded via the network to the required location. The transmission method may include file downloads using protocols such as HTTP and FTP.

[0116] To ensure the accuracy of the annotation tool, the self - established infrastructure scene dataset covers twenty categories, involving a large dataset. To improve feasibility, the remote server stores the offline manually annotated image set, and the camera and detection device can detect the infrastructure site image data, which is batch - imported as the annotation source data.

[0117] The above - mentioned data is transmitted to the hardware dependencies of the annotation tool through the Bluetooth module and the storage module, and is available for further storage modules and manual verification modules. The hardware dependencies of the intelligent image annotation tool include, but are not limited to, the hardware devices of mobile phones, tablets, and computers. The storage module can store the standard annotation files of the image dataset. The manual verification module corrects the annotated images and then transmits them to the storage module.

[0118] Based on the same inventive concept, corresponding to the method of any of the above - mentioned embodiments, the present disclosure also provides an image annotation device.

[0119] Refer to Figure 4 , Figure 4 For the image annotation device of the embodiment, it includes:

[0120] A data acquisition module 401, configured to acquire initial image data, where the initial image data includes real - time image data and offline image data;

[0121] A model processing module 402, which inputs the initial image data into a pre - trained image annotation model, and uses the image annotation model to identify and annotate the target objects in the initial image data to obtain annotated image data.

[0122] In some embodiments, the device includes a model training module, and the model training module specifically includes:

[0123] A data acquisition unit, configured to acquire an annotation dataset and an initial image annotation model, where the annotation dataset includes original images and the actual annotated images corresponding to the original images;

[0124] A training unit, configured to input the original images into the initial image annotation model for training to obtain an initial annotation result;

[0125] A comparison unit, configured to compare the initial annotation result with the actual annotated images to determine the target loss function;

[0126] A training completion unit, configured to determine that the training of the initial image annotation model is completed and obtain an image annotation model in response to the target loss function converging to a preset convergence threshold.

[0127] In some embodiments, the training unit is specifically configured to:

[0128] Input the original image into the initial image annotation model;

[0129] Obtain a preset plurality of object categories, and perform recognition processing on the original image according to the object categories to obtain a plurality of annotation boxes;

[0130] Determine a target annotation box from the plurality of annotation boxes, and use the target object category and the target annotation area corresponding to the target annotation box as the initial annotation result.

[0131] In some embodiments, the training unit is further specifically configured to:

[0132] Obtain a preset confidence threshold;

[0133] Compare the confidence of each annotation box with the confidence threshold respectively, and eliminate the annotation boxes with a confidence less than the confidence threshold.

[0134] In some embodiments, the training unit is further specifically configured to:

[0135] Determine a target annotation box from the plurality of annotation boxes through at least one round of iterative operations. Each round of iterative operations is performed as follows:

[0136] Select the annotation box with the highest confidence from all the annotation boxes as the initial target annotation box;

[0137] Denote the annotation boxes other than the initial target annotation box among all the annotation boxes as other annotation boxes;

[0138] According to the annotation area of the initial target annotation box and the annotation areas of the other annotation boxes, calculate the first intersection over union (IoU) between the initial target annotation box and each other annotation box respectively;

[0139] Delete the other annotation boxes corresponding to the first IoU greater than a preset IoU threshold from all the annotation boxes;

[0140] Determine whether there are other annotation boxes among all the annotation boxes after deletion;

[0141] In response to the existence of other annotation boxes, select the annotation box with the highest confidence from all the existing other annotation boxes as the initial target annotation box for the next round, and enter the next round of iterative operations;

[0142] In response to the absence of other annotation frames, exit at least one round of iterative operations, and use all the initial target annotation frames as the target annotation frames.

[0143] In some embodiments, the comparison unit is specifically configured to:

[0144] Identify the actual annotation image to obtain the image label corresponding to the actual annotation image;

[0145] Compare the object category in the initial annotation result with the image label to determine the target loss function.

[0146] In some embodiments, the data acquisition module 401 is specifically configured to:

[0147] Obtain an initial annotation data set;

[0148] Perform image scaling processing on the data in the initial annotation data set to obtain a first annotation data set;

[0149] Perform region division processing on the data in the first annotation data set to obtain an annotation data set.

[0150] For the sake of convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0151] The device in the above embodiment is used to implement the corresponding image annotation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0152] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the image annotation method in any of the above embodiments.

[0153] Figure 5 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0154] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0155] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0156] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0157] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0158] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0159] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and do not necessarily include all the components shown in the figure.

[0160] The electronic device of the above embodiment is used to implement the corresponding image annotation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0161] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the image annotation method as described in any of the foregoing embodiments.

[0162] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0163] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the image annotation method as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0164] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a vehicle including the image annotation device, the electronic device, and the computer-readable storage medium in the above embodiments, and the vehicle device implements the image annotation method as described in any of the foregoing embodiments.

[0165] The vehicle of the above embodiment is used to implement the image annotation method as described in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0166] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.

[0167] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure technical solution based on the prompt message.

[0168] As an optional but non-limiting implementation manner, when responding to an active request from a user, the manner of sending a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0169] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0170] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the idea of the present disclosure, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.

[0171] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation manner of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0172] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0173] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An image annotation method, characterized in that: include: Acquiring initial image data, wherein the initial image data includes real-time image data and offline image data; The initial image data is input into a pre-trained image annotation model, and the image annotation model is used to identify and annotate the target object in the initial image data to obtain annotated image data.

2. The method according to claim 1, characterized in that The training process of the image annotation model specifically includes: Acquire an annotated data set and an initial image annotation model, wherein the annotated data set includes an original image and an actual annotated image corresponding to the original image; Inputting the original image into the initial image annotation model for training to obtain an initial annotation result; Comparing the initial annotation result with the actual annotated image to determine a target loss function; In response to the target loss function converging to a preset convergence threshold, it is determined that the training of the initial image annotation model is completed, and an image annotation model is obtained.

3. The method according to claim 2, characterized in that The step of inputting the original image into the initial image annotation model for training to obtain an initial annotation result includes: Inputting the original image into the initial image annotation model; Acquire a plurality of preset object categories, and perform recognition processing on the original image according to the object categories to obtain a plurality of annotation boxes; A target marking frame is determined from the multiple marking frames, and a target object category and a target marking area corresponding to the target marking frame are used as an initial marking result.

4. The method according to claim 3, characterized in that Before determining the target annotation box from the multiple annotation boxes, the method further includes: Get the preset confidence threshold; The confidence corresponding to each marked box is compared with the confidence threshold, and the marked boxes whose confidence is less than the confidence threshold are eliminated.

5. The method according to claim 3, characterized in that: The determining a target annotation frame from the multiple annotation frames comprises: The target annotation box is determined from multiple annotation boxes through at least one round of iterative operation, and each round of iterative operation is performed as follows: Select the annotation box with the highest confidence from all annotation boxes as the initial target annotation box; Recording the annotation frames in all annotation frames except the initial target annotation frame as other annotation frames; According to the marked area of ​​the initial target marked box and the marked areas of the other marked boxes, respectively calculating a first intersection-over-union ratio between the initial target marked box and each other marked box; Delete the other marked boxes whose first intersection-and-union ratio is greater than the preset intersection-and-union ratio threshold from all marked boxes; Determine whether there are other annotation boxes among all the annotation boxes after deletion; In response to the existence of other annotation boxes, the annotation box with the highest confidence is selected from all other existing annotation boxes as the initial target annotation box of the next round, and the next round of iterative operation is entered; In response to the absence of other annotation boxes, at least one round of iterative operation is exited, and all initial target annotation boxes are used as target annotation boxes.

6. The method according to claim 3, characterized in that: The comparing the initial annotation result with the actual annotated image to determine the target loss function includes: Recognize the actual annotated image to obtain an image label corresponding to the actual annotated image; The target object category in the initial annotation result is compared with the image label to determine a target loss function.

7. The method according to claim 2, characterized in that The step of obtaining the labeled data set includes: Get the initial labeled dataset; Performing image scaling processing on the data in the initial annotated data set to obtain a first annotated data set; Performing region division processing on the data in the first annotated data set to obtain an annotated data set.

8. An image annotation device, characterized in that: include: A data acquisition module is configured to acquire initial image data, wherein the initial image data includes real-time image data and offline image data; The model processing module inputs the initial image data into a pre-trained image annotation model, and uses the image annotation model to identify and annotate the target object in the initial image data to obtain annotated image data.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program. 10 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to claim 1 .