Object labeling method and apparatus

By acquiring image annotation points and forming a coordinate set through an object annotation device, the problem of low efficiency in manual annotation is solved, and efficient annotation of training images for autonomous driving is achieved.

CN118015577BActive Publication Date: 2026-08-04CHERY AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHERY AUTOMOBILE CO LTD
Filing Date
2023-12-29
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Manual image annotation is inefficient and cannot meet the annotation requirements of a large number of training images in autonomous driving.

Method used

A method and apparatus for object annotation are provided, which acquires an image to be annotated and obtains annotation points based on user annotation instructions, forms a coordinate set, and finally generates annotation data.

Benefits of technology

It improves the efficiency of image annotation, making it faster and more efficient than manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118015577B_ABST
    Figure CN118015577B_ABST
Patent Text Reader

Abstract

The application discloses an object labeling method and device, and belongs to the field of information processing. The method comprises the following steps: acquiring a to-be-labeled image, wherein the to-be-labeled image comprises a to-be-labeled object; labeling the to-be-labeled image based on a labeling instruction triggered by a user for the to-be-labeled image to obtain at least one labeling point, wherein the at least one labeling point is located on the to-be-labeled object; acquiring a coordinate set of the to-be-labeled object based on the at least one labeling point, wherein the coordinate set comprises the coordinates of a plurality of pixel points, the plurality of pixel points are located in a target region of the to-be-labeled image, and the coincidence degree of the to-be-labeled object in the to-be-labeled image and the target region is greater than a preset coincidence degree; and labeling the to-be-labeled object based on the coordinate set to obtain labeling data. The application can improve the labeling efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing, and in particular to a method and apparatus for object annotation. Background Technology

[0002] Recognizing road markings such as lane lines using a recognition model is an essential part of autonomous driving. Therefore, it requires training the recognition model with a large number of labeled training images to improve its accuracy in recognizing road markings. Currently, training images are typically labeled manually to obtain labeled training images. However, manual labeling is inefficient. Summary of the Invention

[0003] This application provides an object annotation method and apparatus, which can improve annotation efficiency. The technical solution of this application is as follows.

[0004] Firstly, a method for object annotation is provided, the method comprising:

[0005] Obtain an image to be labeled, the image including objects to be labeled;

[0006] The image to be labeled is labeled based on the labeling instruction triggered by the user to obtain at least one labeling point, and the at least one labeling point is located on the object to be labeled;

[0007] The coordinate set of the object to be labeled is obtained based on the at least one annotation point. The coordinate set includes the coordinates of multiple pixels. The multiple pixels are located within the target area of ​​the image to be labeled. The overlap between the object to be labeled in the image to be labeled and the target area is greater than a preset overlap.

[0008] The object to be labeled is labeled based on the coordinate set to obtain the labeling data.

[0009] Optionally, after obtaining the coordinate set of the object to be labeled based on the at least one labeled point, the method further includes:

[0010] Based on the coordinate set, the color of the object to be labeled is set to a first color to obtain a preprocessed image, the preprocessed image including the object to be labeled with the first color;

[0011] Based on the user's operation command triggered on the preprocessed image, determine whether the user is satisfied with the preprocessed image;

[0012] The step of annotating the object to be annotated based on the coordinate set to obtain annotation data includes:

[0013] If the user is satisfied with the preprocessed image based on the operation instructions, the object to be labeled is labeled based on the coordinate set to obtain labeling data.

[0014] Optionally, the method further includes: if it is determined based on the operation instruction that the user is not satisfied with the preprocessed image, obtaining the coordinate set of the object to be labeled based on the annotation points obtained by the user in re-triggered annotation instruction for the image to be labeled, and annotating the object to be labeled based on the re-acquired coordinate set.

[0015] Optionally, the step of annotating the object to be annotated based on the coordinate set to obtain annotation data includes:

[0016] Based on the coordinate set, determine multiple target pixels that meet the conditions among the multiple pixels;

[0017] The object to be labeled is labeled based on the multiple target pixels to obtain labeling data.

[0018] Optionally, the object to be labeled is a linear object, the minimum distance between each of the plurality of target pixels and the target detection curve is less than a preset distance, the target detection curve is determined based on a subset of the plurality of pixels, and the number of the plurality of target pixels is greater than a preset number; or,

[0019] The object to be labeled is a non-linear object, and the plurality of target pixels includes at least one of the following: the pixel with the largest horizontal coordinate, the pixel with the smallest horizontal coordinate, the pixel with the largest vertical coordinate, and the pixel with the smallest vertical coordinate.

[0020] Secondly, an object annotation device is provided, the device comprising:

[0021] The first acquisition module is used to acquire an image to be labeled, the image to be labeled including objects to be labeled;

[0022] The first annotation module is used to annotate the image to be annotated based on the annotation command triggered by the user for the image to be annotated, so as to obtain at least one annotation point, wherein the at least one annotation point is located on the object to be annotated;

[0023] The second acquisition module is used to acquire the coordinate set of the object to be labeled based on the at least one annotation point. The coordinate set includes the coordinates of multiple pixels, which are located within the target area of ​​the image to be labeled. The overlap between the object to be labeled in the image to be labeled and the target area is greater than a preset overlap.

[0024] The second annotation module is used to annotate the object to be annotated based on the coordinate set to obtain annotation data.

[0025] Optionally, the device further includes: a setting module, configured to, after the second acquisition module acquires the coordinate set of the object to be labeled based on the at least one annotation point, set the color of the object to be labeled to a first color based on the coordinate set to obtain a preprocessed image, the preprocessed image including the object to be labeled with the first color;

[0026] The determination module is used to determine whether the user is satisfied with the preprocessed image based on the operation command triggered by the user for the preprocessed image;

[0027] The second annotation module is used to annotate the object to be annotated based on the coordinate set to obtain annotation data when the determining module determines that the user is satisfied with the preprocessed image based on the operation instruction.

[0028] Optionally, the second acquisition module is further configured to: when the determining module determines, based on the operation instruction, that the user is not satisfied with the preprocessed image, acquire the coordinate set of the object to be labeled based on the annotation points obtained by the user in re-triggered annotation instruction for the image to be labeled;

[0029] The second annotation module is also used to: annotate the object to be annotated based on the coordinate set re-acquired by the second acquisition module.

[0030] Optionally, the second annotation module is used for:

[0031] Based on the coordinate set, determine multiple target pixels that meet the conditions among the multiple pixels;

[0032] The object to be labeled is labeled based on the multiple target pixels to obtain labeling data.

[0033] Optionally, the object to be labeled is a linear object, the minimum distance between each of the plurality of target pixels and the target detection curve is less than a preset distance, the target detection curve is determined based on a subset of the plurality of pixels, and the number of the plurality of target pixels is greater than a preset number; or,

[0034] The object to be labeled is a non-linear object, and the plurality of target pixels includes at least one of the following: the pixel with the largest horizontal coordinate, the pixel with the smallest horizontal coordinate, the pixel with the largest vertical coordinate, and the pixel with the smallest vertical coordinate.

[0035] Thirdly, an object annotation device is provided, including a memory and a processor;

[0036] The memory is used to store computer programs;

[0037] The processor is configured to execute a computer program stored in the memory to cause the object annotation device to perform the object annotation method provided by the first aspect or any optional implementation thereof.

[0038] Fourthly, a computer device is provided, which may be a terminal device such as a smartphone, tablet computer, laptop computer or desktop computer, the computer device including the object annotation device provided as in the second aspect or any optional implementation of the second aspect, or the computer device including the object annotation device provided as in the third aspect.

[0039] Fifthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed, the computer program implements the object annotation method provided as described in the first aspect or any alternative method of the first aspect.

[0040] In a sixth aspect, a computer program product is provided, the computer program product comprising a program or code, which, when executed, implements the object annotation method provided as described in the first aspect or any alternative method of the first aspect.

[0041] The beneficial effects of the technical solution provided in this application are:

[0042] This application provides an object annotation method and apparatus, wherein the object annotation method is executed by the object annotation apparatus. After acquiring an image to be annotated, the object annotation apparatus annotates the image to obtain at least one annotation point based on an annotation command triggered by a user for the image to be annotated. The object annotation apparatus then obtains a set of coordinates of the object to be annotated based on the at least one annotation point, and annotates the object to be annotated based on the set of coordinates. Therefore, in this embodiment of the application, the object annotation apparatus annotates the objects to be annotated in the image to be annotated, which can improve annotation efficiency compared to manual image annotation. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart of an object annotation method provided in an embodiment of this application;

[0045] Figure 2 This is a schematic diagram of a display interface provided in an embodiment of this application;

[0046] Figure 3 This is a schematic diagram of another display interface provided in an embodiment of this application;

[0047] Figure 4 This is a schematic diagram of another display interface provided in an embodiment of this application;

[0048] Figure 5 This is a schematic diagram of yet another display interface provided in an embodiment of this application;

[0049] Figure 6 This is a flowchart of another object annotation method provided in an embodiment of this application;

[0050] Figure 7 This is a schematic diagram of an object labeling device provided in an embodiment of this application;

[0051] Figure 8 This is a schematic diagram of another object labeling device provided in an embodiment of this application.

[0052] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0054] In recent years, with the increasing application of artificial intelligence technology in the driving field, autonomous driving technology has attracted much attention. For safety reasons, autonomous vehicles need to perceive their surroundings; therefore, using recognition models to identify road markings such as lane lines is an essential part of autonomous driving. This requires training the recognition model with a large number of labeled training images to improve its accuracy in recognizing road markings. Currently, training images are typically labeled manually to obtain labeled training images. However, manual labeling is inefficient.

[0055] This application provides an object annotation method and apparatus, wherein the object annotation method is executed by the object annotation apparatus. After acquiring an image to be annotated, the object annotation apparatus annotates the image based on annotation instructions triggered by the user to obtain at least one annotation point. Based on the at least one annotation point, the object annotation apparatus obtains a set of coordinates of the object to be annotated, and annotates the object to be annotated based on the coordinate set to obtain annotation data. Therefore, this application embodiment uses the object annotation apparatus to annotate the objects in the image to be annotated, which can improve annotation efficiency compared to manual image annotation.

[0056] Please refer to Figure 1 The diagram illustrates a flowchart of an object annotation method provided in an embodiment of this application. This object annotation method is executed by an object annotation device. The object annotation device can be a computer device or a functional component within a computer device, such as a smartphone, tablet, laptop, or desktop computer. Figure 1 As shown, the object annotation method includes the following steps S101 to S104.

[0057] S101. Obtain the image to be labeled, which includes the object to be labeled.

[0058] The image to be labeled can be any image in the image set. The images in the image set are images captured of the road environment and may include road markings such as lane lines, as well as objects such as vehicles, pedestrians, and traffic lights. The objects to be labeled can be road markings such as lane lines, or objects such as vehicles, pedestrians, and traffic lights.

[0059] In optional embodiments, the object annotation device acquires a set of images to be annotated, and obtains images to be annotated from the set of images to be annotated. In one embodiment, the object annotation device acquires the set of images to be annotated from a server. In another embodiment, the object annotation device stores the set of images to be annotated, and acquires the set of images to be annotated stored by the object annotation device. The set of images to be annotated may be obtained by the object annotation device from the server in advance, or it may be generated by the object annotation device.

[0060] In an optional embodiment, while the vehicle is traveling on the road, the image acquisition device on the vehicle acquires road environment images and sends the acquired road environment images to the server. The server receives road environment images sent by multiple vehicles and determines the set of road environment images sent by multiple vehicles as the image set to be labeled. For example, the image acquisition device is a vehicle-mounted camera, a dashcam, or other similar device.

[0061] S102. Based on the annotation instruction triggered by the user for the image to be annotated, annotate the image to be annotated to obtain at least one annotation point, the at least one annotation point being located on the object to be annotated.

[0062] After acquiring an image to be labeled, the object labeling device displays the image on a display interface. The user can trigger a labeling command on the image displayed by the object labeling device. The object labeling device receives the user's labeling command and labels the image based on the command to obtain at least one labeling point. The at least one labeling point is located on the object to be labeled within the image. For example, the object labeling device includes a human-computer interaction component, through which the user triggers a labeling command on the image displayed by the object labeling device. This human-computer interaction component includes, but is not limited to, a keyboard, mouse, and touchscreen.

[0063] In one embodiment, a user selects at least one pixel in the image to be labeled using the human-computer interaction component and triggers a labeling instruction based on the at least one pixel. The object labeling device receives the labeling instruction and labels the at least one pixel in the image to obtain at least one labeled point. The at least one pixel is located on the object to be labeled in the image. For example, Figure 2 This is a schematic diagram of a display interface provided in an embodiment of this application, in which the object marking device is used... Figure 2 The displayed interface shows the image to be annotated. For example... Figure 2 As shown, the image to be labeled includes the object to be labeled, and the display interface also includes a confirm button and a cancel button. Users can click as shown... Figure 2 The image to be annotated shows pixels A1, A2, A3, and A4. Users can select pixels A1, A2, A3, and A4, and then click the "Confirm" button to trigger an annotation command. The object annotation device then annotates the selected pixels A1, A2, A3, and A4 based on this command to obtain annotated points A1, A2, A3, and A4. Optionally, after selecting at least one pixel on the image, the user can also cancel the selection by clicking the "Cancel" button. For example, clicking pixels A1, A2, A3, and A4 sequentially and then clicking the "Cancel" button will cancel the selection of pixels A1, A2, A3, and A4. For example, a user clicks the cancel button once to cancel the selection of pixel A4, a user clicks the cancel button twice to cancel the selection of pixels A4 and A3, a user clicks the cancel button three times to cancel the selection of pixels A4, A3 and A2, and a user clicks the cancel button four times to cancel the selection of pixels A4, A3, A2 and A1. This application embodiment does not limit this.

[0064] In another embodiment, the image to be labeled includes multiple points to be labeled. The user selects at least one point located on the object to be labeled from among the multiple points to be labeled, and triggers a labeling command based on the at least one point to be labeled. The object labeling device receives the labeling command and determines the at least one point to be labeled as a labeled point based on the labeling command. For example, Figure 3 This is a schematic diagram of another display interface provided in an embodiment of this application, in which the object marking device is used as... Figure 3 The displayed interface shows the image to be annotated. For example... Figure 3 As shown, the image to be labeled includes the object to be labeled and multiple points to be labeled (small white boxes on the image indicate the points to be labeled). The display interface also includes a confirm button and a cancel button. Users can click as shown... Figure 3 The image to be labeled shows points B1, B2, B3, and B4. Users can select these points and then click the "Confirm" button to trigger a labeling command. The object labeling device then labels the selected points based on this command, resulting in labeled points B1, B2, B3, and B4. Optionally, after selecting at least one point on the image, the user can also cancel the selection by clicking the "Cancel" button. For example, clicking points B1, B2, B3, and B4 sequentially and then clicking the "Cancel" button will cancel the selection of all points. These multiple points in the image can be generated by the object labeling device based on the image itself. For instance, the object labeling device can determine a point at a preset distance on the image to obtain these multiple points.

[0065] S103. Obtain the coordinate set of the object to be labeled based on the at least one annotation point. The coordinate set includes the coordinates of multiple pixels. The multiple pixels are located within the target area of ​​the image to be labeled. The overlap between the object to be labeled in the image to be labeled and the target area is greater than a preset overlap.

[0066] In one embodiment, the object annotation device identifies the object to be annotated in the image based on at least one annotation point in the image to be annotated (identifying the area occupied by the object to be annotated in the image to be annotated). After identifying the object to be annotated, the object annotation device obtains the coordinates of multiple pixels of the object to be annotated in the image to be annotated. Specifically, if the multiple pixels are located within a target area of ​​the image to be annotated, the object annotation device obtains the coordinates of the multiple pixels within the target area, and determines the set of coordinates of the multiple pixels as the coordinate set of the object to be annotated.

[0067] In another embodiment, the object annotation device inputs the image to be annotated, including the at least one annotation point, into an image recognition model. The image recognition model identifies the object to be annotated in the image based on the at least one annotation point, obtains the coordinates of each pixel of the object, and outputs the coordinates of multiple pixels of the object. The object annotation device obtains the coordinates of the multiple pixels output by the image recognition model. Specifically, if the multiple pixels are located within a target region of the image, the object annotation device obtains the coordinates of the multiple pixels within the target region, and determines the set of coordinates of the multiple pixels as the coordinate set of the object to be annotated. The image recognition model is used to identify objects in an image and obtain the coordinates of pixels within those objects. The image recognition model can be a deep learning model or a machine learning model, such as a model using deep neural networks, convolutional neural networks, recurrent neural networks, Tensorflow (a deep learning framework), etc.

[0068] Before inputting the image to be labeled into the image recognition model, the object labeling device acquires the image processing model. For example, the object labeling device acquires a pre-trained image recognition model, or the object labeling device trains the image recognition model. This application embodiment uses the object labeling device training the image recognition model as an example. In one embodiment, the object labeling device trains the image recognition model based on a training image set. The training image set includes multiple training images. For each training image, the object labeling device inputs it into an initial recognition model, causing the initial recognition model to perform image recognition on the training image and obtain the coordinates of the object's pixels in the training image based on the recognition result. The object labeling device adjusts the parameters of the initial recognition model according to the pixel coordinates of the object output by the initial recognition model until a termination condition is met. The object labeling device determines the recognition model obtained when the termination condition is met as the trained image recognition model.

[0069] S104. Based on the coordinate set of the object to be labeled, label the object to obtain the labeling data.

[0070] In an optional embodiment, after the object annotation device obtains the coordinate set of the object to be annotated, it sets the color of the object to be annotated to a first color based on the coordinate set to obtain a preprocessed image. The preprocessed image includes the object to be annotated with the first color. For example, the object annotation device sets the color of the pixels indicated by each coordinate in the coordinate set to the first color to obtain the preprocessed image. After obtaining the preprocessed image, the object annotation device displays the preprocessed image on a display interface. The user can trigger an operation command based on the preprocessed image displayed by the object annotation device. The operation command indicates whether the user is satisfied with the preprocessed image. The object annotation device receives the operation command and determines whether the user is satisfied with the preprocessed image based on the operation command. If the object annotation device determines that the user is satisfied with the preprocessed image based on the operation command, the object annotation device annotates the object to be annotated based on the coordinate set to obtain annotation data. If the object annotation device determines, based on the operation command, that the user is dissatisfied with the preprocessed image, the object annotation device obtains the coordinate set of the object to be annotated based on the annotation points obtained from the annotation command retried by the user for the image to be annotated, and annotates the object based on the re-acquired coordinate set. Here, the first color is a pre-set color, which can be a single color or a composite color. For example, the first color is red, green, or blue. Alternatively, the first color is a composite color composed of the three primary colors.

[0071] For example, object annotation devices in such Figure 4 The displayed interface shows the preprocessed image. For example... Figure 4 As shown, the display interface includes a pre-processed image, as well as "Correct" and "Incorrect" buttons. The pre-processed image includes an object to be labeled, and the color of the object to be labeled in the pre-processed image is set to a first color. Figure 4 (Using black as an example). The user determines whether the coordinate set is correct based on the preprocessed image. If the user confirms that the coordinate set is correct, they click the "Correct" button on the display interface to trigger a confirmation command. The object annotation device then confirms the coordinate set is correct based on this confirmation command, and subsequently jumps to the following... Figure 5 The display interface shown is as follows. Figure 5As shown, the display interface includes a preprocessed image and "Complete" and "Incomplete" buttons. The user determines whether the coordinate set is complete based on the preprocessed image. If the user determines that the coordinate set is complete, clicking the "Complete" button triggers a confirmation command. The object annotation device then confirms the coordinate set is complete based on this confirmation command. After receiving the operation command (specifically, the confirmation command), the object annotation device determines that the user is satisfied with the preprocessed image and then annotates the object to be annotated based on the coordinate set to obtain annotation data. Optionally, if the user determines that the coordinate set is incorrect, the user can click the "Complete" button. Figure 4 An incorrect button in the displayed interface triggers a cancellation command, or, if the user determines that the coordinate set is incomplete, the user clicks an incorrect button. Figure 5 The incomplete button in the displayed interface triggers a cancellation command. After receiving this command (i.e., the cancellation command), the object annotation device determines that the user is not satisfied with the preprocessed image based on the command, and then jumps to the following... Figure 2 The displayed interface shows that the user re-triggers the annotation command for the image to be annotated. The object annotation device re-annotates the image based on the user's re-triggered annotation command to obtain at least one annotation point. Based on the re-annotated at least one annotation point, the object annotation device obtains the coordinate set of the object to be annotated through an image recognition model. Based on the re-obtained coordinate set, the object annotation device sets the color of the object to be annotated to a first color to obtain a pre-processed image, and then displays the re-obtained pre-processed image.

[0072] In an optional embodiment, if the object annotation device determines that the user is satisfied with the preprocessed image based on the user's operation command triggered for the preprocessed image, the object annotation device annotates the object to be annotated based on the coordinate set of the object to be annotated to obtain annotation data. In a specific embodiment, the coordinate set includes the coordinates of multiple pixels. The object annotation device determines multiple target pixels among the multiple pixels that meet certain conditions based on the coordinate set, and annotates the object to be annotated based on the multiple target pixels to obtain annotation data.

[0073] In one embodiment, the object to be labeled is a linear object. The minimum distance between each of the plurality of target pixels and the target detection curve is less than a preset distance. The target detection curve is determined based on a subset of the plurality of pixels, and the number of the plurality of target pixels is greater than a preset number. The object labeling device detects the plurality of pixels to determine a plurality of target pixels that meet certain conditions and the target detection curve. For example, the object labeling device uses a random sample consensus algorithm to detect the plurality of pixels. In a specific embodiment, the object labeling device randomly selects a plurality of pixels from the plurality of pixels. The object labeling device determines a detection curve based on the plurality of pixels, and all the plurality of pixels lie on the detection curve (for example, the object labeling device sequentially connects the plurality of pixels to obtain the detection curve). The object labeling device verifies the reliability of the detection curve based on the pixels other than the plurality of pixels. If the verification determines that the reliability of the detection curve is high, the object labeling device determines the detection curve as the target detection curve, and then determines the plurality of target pixels based on the target detection curve. If the verification determines that the reliability of the detection curve is low, the object labeling device randomly selects several new pixels from the plurality of pixels. Based on these newly selected pixels, the object labeling device determines a new detection curve and verifies the reliability of the newly determined detection curve using the pixels other than the newly selected pixels. This process continues until the object labeling device determines the target detection curve.

[0074] For ease of description, the pixels other than the pixels used to determine the detection curve among the plurality of pixels are referred to as the pixels to be verified. The object labeling device verifies the reliability of the detection curve based on the pixels other than the specified pixels among the plurality of pixels, including: the object labeling device determining the minimum distance between each pixel to be verified among the plurality of pixels and the detection curve; the object labeling device determining the number of pixels among the plurality of pixels whose minimum distance to the detection curve is less than a preset distance based on the minimum distance between the pixels to be verified among the plurality of pixels and the detection curve; if the number of pixels to be verified whose minimum distance to the detection curve is less than the preset distance is greater than or equal to a preset number, the object labeling device determines that the reliability of the detection curve is high, and the object labeling device identifies the detection curve as the target detection curve; if the number of pixels to be verified whose minimum distance to the detection curve is less than the preset distance is less than the preset number, the object labeling device determines that the reliability of the detection curve is low, and the object labeling device does not identify the detection curve as the target detection curve. If the number of test pixels among the plurality of pixels whose minimum distance to the detection curve is less than a preset distance is greater than or equal to the preset number, the object labeling device determines the test pixels among the plurality of pixels whose minimum distance to the detection curve is less than the preset distance as target pixels, and the object labeling device can determine multiple target pixels.

[0075] Optionally, when the object to be labeled is a linear object, the object labeling device labels the object based on the multiple target pixels to obtain labeling data, including: the object labeling device sequentially connects the multiple target pixels to obtain a labeling line, which is the labeling data. In one embodiment, the object labeling device sequentially connects the multiple target pixels in ascending or descending order of their horizontal coordinates to obtain a labeling line. During the connection process, for target pixels with the same horizontal coordinate, the object labeling device sequentially connects the target pixels with the same horizontal coordinate in ascending or descending order of their vertical coordinates. In another embodiment, the object labeling device sequentially connects the multiple target pixels in ascending or descending order of their vertical coordinates to obtain a labeling line. During the connection process, for target pixels with the same vertical coordinate, the object labeling device sequentially connects the target pixels with the same vertical coordinate in ascending or descending order of their horizontal coordinates.

[0076] It should be noted that the above process of determining the target detection curve and target pixels is described by the object annotation device sequentially determining the detection curve. That is, the object annotation device determines a detection curve based on a subset of pixels from a set of coordinates, verifies the reliability of the detection curve, and if the reliability of the detection curve is high, the object annotation device determines this detection curve as the target detection curve, and then determines the multiple target pixels based on the target detection curve. If the reliability of the detection curve is low, the object annotation device re-determines the detection curve and performs reliability verification until the target detection curve and target pixels are determined. In some embodiments, the object annotation device can determine multiple detection curves at once, simultaneously verify the reliability of these multiple detection curves, and determine the target detection curve from the multiple detection curves based on the reliability verification results. For example, the object annotation device determines at least one detection curve with high reliability from the multiple detection curves and randomly selects one detection curve from the at least one detection curve as the target detection curve, or the object annotation device determines the detection curve with the highest reliability from the at least one detection curve as the target curve. Among them, the detection curve with the highest reliability refers to the detection curve with the largest number of pixels to be tested whose minimum distance to the detection curve is less than a preset distance among all pixels to be tested. In other words, the detection curve with the highest reliability refers to the detection curve with the largest number of target pixels determined based on the at least one detection curve.

[0077] In an optional embodiment, the object to be labeled is a non-linear object, and the plurality of target pixels satisfying the conditions include at least one of the following: the pixel with the largest x-coordinate, the pixel with the smallest x-coordinate, the pixel with the largest y-coordinate, and the pixel with the smallest y-coordinate. The object labeling device sorts the plurality of pixels according to their x-coordinates, and determines the pixel with the largest and smallest x-coordinates from the plurality of pixels according to the sorting order. Similarly, the object labeling device sorts the plurality of pixels according to their y-coordinates, and determines the pixel with the largest and smallest y-coordinates from the plurality of pixels according to the sorting order. For example, the object labeling device sorts the plurality of pixels using a bubble sort algorithm, a selection sort algorithm, or an insertion sort algorithm based on their x-coordinates. The object labeling device also sorts the plurality of pixels using a bubble sort algorithm, a selection sort algorithm, or an insertion sort algorithm based on their y-coordinates.

[0078] In an optional embodiment, when the object to be labeled is a non-linear object, the object labeling device sequentially connects the multiple target pixels to obtain a labeling box, which is the labeling data. For example, the multiple target pixels include the pixel with the largest horizontal coordinate, the pixel with the smallest horizontal coordinate, the pixel with the largest vertical coordinate, and the pixel with the smallest vertical coordinate. The object labeling device sequentially connects the pixel with the largest horizontal coordinate, the pixel with the smallest horizontal coordinate, the pixel with the largest vertical coordinate, and the pixel with the smallest vertical coordinate to obtain the labeling box.

[0079] In specific embodiments, such as Figure 6 As shown, after the object annotation device annotates the image to be annotated to obtain at least one annotation point, the object annotation device obtains the coordinate set of the object to be annotated based on the at least one annotation point. Based on the coordinate set, the object annotation device sets the color of the object to be annotated to a first color to obtain a preprocessed image. For example, if the overlap between the range of pixels of the first color in the preprocessed image and the range of the object to be annotated reaches a first degree of overlap, the user considers the obtained coordinate set correct. If the overlap between the range of pixels of the first color in the preprocessed image and the range of the object to be annotated reaches a second degree of overlap, the user considers the obtained coordinate set complete. The second degree of overlap is greater than the first degree of overlap. The user, through methods such as... Figure 4 The preprocessed image displayed in the interface determines whether the obtained coordinate set is correct. If the obtained coordinate set is confirmed to be correct, the user clicks as shown. Figure 4 The correct button on the displayed interface triggers a confirmation command. Based on this confirmation command, the object annotation device determines that the coordinate set is correct, and then the object annotation device jumps to... Figure 5 The displayed interface shows that the user can... Figure 5 The preprocessed image displayed in the interface determines whether the acquired coordinate set is complete. If the acquired coordinate set is confirmed to be complete, the user clicks as shown. Figure 5The complete button on the displayed interface triggers a confirmation command. Based on this confirmation command, the object annotation device confirms the coordinate set is complete and then determines whether the object to be annotated is a linear object. If the object is linear, the object annotation device annotates the object based on the coordinate set to obtain annotation lines. If the object is non-linear, the object annotation device annotates the object based on the coordinate set to obtain annotation boxes. After annotating the object, the object annotation device determines whether all objects in the image to be annotated have been annotated. If all objects in the image to be annotated have been annotated, the object annotation device annotates the unannotated objects in the image to obtain at least one annotation point. The object annotation device then annotates the unannotated objects based on this at least one annotation point to obtain annotation data.

[0080] If the obtained coordinate set is incomplete, the user clicks as shown below. Figure 4 If an incorrect button on the displayed interface triggers a cancellation command, the device will mark the pixels in the image where coordinates were not obtained, thus re-obtaining at least one mark. Based on this re-marked mark, the object marking device uses an image recognition model to obtain the coordinates of multiple pixels of the object to be marked, adds these coordinates to the coordinate set, and sets the color of the object to be marked to a first color based on the newly determined coordinate set, resulting in a new preprocessed image. The device then uses this new preprocessed image to determine if the obtained coordinate set is correct. If the obtained coordinate set is correct, the device determines if the obtained coordinate set is complete. If the obtained coordinate set is incorrect, the object marking device deletes the at least one mark, re-marks the image to be marked to obtain at least one mark, and uses an image recognition model to obtain the coordinate set of the object to be marked. Based on this coordinate set, the object marking device sets the color of the object to be marked to a first color, resulting in a preprocessed image. The user then uses this preprocessed image to determine if the obtained coordinate set is correct.

[0081] In summary, the object annotation method provided in this application involves the object annotation device acquiring an image to be annotated, then annotating the image based on a user-triggered annotation command to obtain at least one annotation point. The object annotation device then obtains the coordinate set of the object to be annotated based on these at least one annotation point, and annotates the object based on the coordinate set to obtain annotation data. Therefore, this application embodiment uses the object annotation device to annotate objects in the image to be annotated, which improves annotation efficiency compared to manual image annotation.

[0082] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0083] Please refer to Figure 7 The diagram illustrates an object annotation device 700 provided in an embodiment of this application. The object annotation device 700 is used to perform actions such as... Figure 1 The method provided in the illustrated embodiment. The object annotation device 700 includes a first acquisition module 701, a first annotation module 702, a second acquisition module 703, and a second annotation module 704.

[0084] The first acquisition module 701 is used to acquire an image to be labeled, which includes objects to be labeled.

[0085] The first annotation module 702 is used to annotate the image to be annotated based on the annotation command triggered by the user for the image to be annotated to obtain at least one annotation point, the at least one annotation point being located on the object to be annotated;

[0086] The second acquisition module 703 is used to acquire a set of coordinates of the object to be labeled based on the at least one annotation point. The set of coordinates includes the coordinates of multiple pixels, which are located within the target area of ​​the image to be labeled. The overlap between the object to be labeled in the image to be labeled and the target area is greater than a preset overlap.

[0087] The second annotation module 704 is used to annotate the object to be annotated based on the coordinate set to obtain annotation data.

[0088] Optional, please continue to refer to Figure 7 The object annotation device 700 also includes:

[0089] The setting module 705 is used to set the color of the object to be labeled to a first color based on the coordinate set of the at least one annotation point after the second acquisition module 703 acquires the coordinate set of the object to be labeled to obtain a preprocessed image, the preprocessed image including the object to be labeled with the first color.

[0090] The determination module 706 is used to determine whether the user is satisfied with the preprocessed image based on the operation command triggered by the user for the preprocessed image;

[0091] The second annotation module 704 is used to annotate the object to be annotated based on the coordinate set to obtain annotation data when the determining module 706 determines that the user is satisfied with the preprocessed image based on the operation instruction.

[0092] Optionally, the second acquisition module 703 is further configured to: when the determination module 706 determines, based on the operation instruction, that the user is not satisfied with the preprocessed image, acquire the coordinate set of the object to be labeled based on the annotation points obtained by the annotation instruction retried by the user for the image to be labeled;

[0093] The second annotation module 704 is also used to: annotate the object to be annotated based on the coordinate set re-acquired by the second acquisition module 703.

[0094] Optionally, the second annotation module 704 is used for:

[0095] Based on this coordinate set, determine multiple target pixels that meet the conditions among these multiple pixels;

[0096] The object to be labeled is labeled based on these multiple target pixels to obtain the labeling data.

[0097] Optionally, the object to be labeled is a linear object, the minimum distance between each of the plurality of target pixels and the target detection curve is less than a preset distance, the target detection curve is determined based on a subset of the plurality of pixels, and the number of the plurality of target pixels is greater than a preset number; or,

[0098] The object to be labeled is a non-linear object, and the multiple target pixels include at least one of the following: the pixel with the largest horizontal coordinate, the pixel with the smallest horizontal coordinate, the pixel with the largest vertical coordinate, and the pixel with the smallest vertical coordinate.

[0099] In summary, the technical solution provided in this application involves an object annotation device acquiring an image to be annotated, then annotating the image based on a user-triggered annotation command to obtain at least one annotation point. The object annotation device then obtains the coordinate set of the object to be annotated based on these at least one annotation point, and annotates the object based on the coordinate set to obtain annotation data. Therefore, this application embodiment uses the object annotation device to annotate objects in the image to be annotated, which improves annotation efficiency compared to manual image annotation.

[0100] This application provides an object annotation apparatus, including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, causing the object annotation apparatus to perform the object annotation method provided in the above embodiments.

[0101] As an example, please refer to Figure 8 This illustration shows a schematic diagram of another object annotation device 800 provided in an embodiment of this application. The object annotation device 800 is a computer device or a functional component deployed in a computer device. The computer device can be a smartphone, tablet, laptop, or desktop computer, etc. The object annotation device 800 is used to perform... Figure 1 The method provided in the illustrated embodiment.

[0102] Typically, the object labeling device 800 includes a processor 801 and a memory 802.

[0103] Processor 801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 801 may be implemented using at least one hardware form selected from digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). Processor 801 may include, but is not limited to, a central processing unit (CPU). In some embodiments, processor 801 may integrate a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. Processor 801 may also include an artificial intelligence (AI) processor to handle computational operations related to machine learning.

[0104] The memory 802 includes one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one instruction, which is executed by the processor 801 to implement the object annotation method provided in the embodiments of this application.

[0105] In some embodiments, the object marking device 800 may also optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal line, or circuit board. The peripheral device may include at least one of: a radio frequency circuit 804, a display screen 805, a camera 806, an audio circuit 807, a positioning component 808, and a power supply 809.

[0106] Peripheral device interface 803 can be used to connect at least one input / output (I / O) related peripheral device to processor 801 and memory 802. In some embodiments, processor 801, memory 802 and peripheral device interface 803 are integrated on the same chip or circuit board; in some embodiments, any one or two of processor 801, memory 802 and peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0107] The radio frequency (RF) circuit 804 is used to receive and transmit radio frequency (RF) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, etc. The RF circuit 804 can communicate with other devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), and wireless local area networks; this application embodiment does not limit this to any particular protocol.

[0108] Display screen 805 is used to display a user interface (UI). The UI may include graphics, text, icons, video, and any combination thereof. When display screen 805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 801 for processing. In this case, display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 805 may be a flexible display screen. Furthermore, display screen 805 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 805 may be a liquid crystal display (LCD), an organic light-emitting diode (OLED) display screen, etc.

[0109] The camera assembly 806 is used to acquire images or videos. Optionally, the camera assembly 806 includes a front-facing camera and a rear-facing camera. Optionally, the computer device is a terminal device such as a smartphone or tablet. Typically, the front-facing camera is located on the front panel of the computer device, and the rear-facing camera is located on the back of the computer device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, virtual reality (VR) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.

[0110] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 801 for processing, or to the radio frequency circuit 804 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, positioned in different parts of the computer device. The microphone can also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker can be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement.

[0111] The positioning component 808 is used to locate the geographic location of a computer device to enable navigation or location-based services (LBS). The positioning component 808 can be a positioning component based on the Global Positioning System (GPS), BeiDou system, or Galileo system.

[0112] Power supply 809 is used to supply power to various components in a computer device. Power supply 809 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 809 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil.

[0113] In some embodiments, the object labeling device 800 further includes one or more sensors 810. The one or more sensors 810 include, but are not limited to: a fingerprint sensor 811, an optical sensor 812, a proximity sensor 813, a pressure sensor 814, an acceleration sensor 815, and a gyroscope sensor 816.

[0114] The fingerprint sensor 811 is used to collect a user's fingerprint. The processor 801 identifies the user based on the fingerprint collected by the fingerprint sensor 811, or vice versa. When the user's identity is verified as trusted, the processor 801 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 811 can be located on the front, back, or side of the computer device. When the computer device has physical buttons or a manufacturer's logo, the fingerprint sensor 811 can be integrated with the physical buttons or the manufacturer's logo.

[0115] An optical sensor 812 is used to collect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 807 based on the ambient light intensity collected by the optical sensor 812. Specifically, when the ambient light intensity is high, the display brightness of the display screen 807 is increased; when the ambient light intensity is low, the display brightness of the display screen 807 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 807 based on the ambient light intensity collected by the optical sensor 812.

[0116] The proximity sensor 813, also known as a distance sensor, is typically located on the front panel of the display screen 805 of a computer device. The proximity sensor 813 is used to detect the distance between the user and the display screen 805. In one embodiment, when the proximity sensor 813 detects that the distance between the user and the display screen 805 is gradually decreasing, the processor 801 controls the display screen 805 to switch from a screen-on state to a screen-off state; when the proximity sensor 813 detects that the distance between the user and the display screen 805 is gradually increasing, the processor 801 controls the display screen 805 to switch from a screen-off state to a screen-on state.

[0117] The pressure sensor 814 can be disposed on the side bezel of the computer device and / or the lower layer of the touch display screen 805. When the pressure sensor 814 is disposed on the side bezel of the computer device, it can detect the user's grip signal on the computer device, and the processor 801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 814. When the pressure sensor 814 is disposed on the lower layer of the touch display screen 805, the processor 801 can control the operable controls on the UI interface based on the user's pressure operation on the touch display screen 805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0118] Accelerometer 815 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by a computer device. For example, accelerometer 815 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 801 can control touchscreen 805 to display the user interface in landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 815. Accelerometer 815 can also be used for games or for acquiring user motion data.

[0119] The gyroscope sensor 816 can detect the orientation and rotation angle of the computer device. The gyroscope sensor 816, in conjunction with the accelerometer sensor 815, can collect 3D motion data from the user on the computer device. Based on the data collected by the gyroscope sensor 812, the processor 801 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0120] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the object labeling device 800. The object labeling device 800 may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0121] This application provides a computer-readable storage medium storing a computer program that, when executed (e.g., by a computer device, an object annotation device, one or more processors, etc.), implements all or part of the steps of the object annotation method provided in the above embodiments.

[0122] This application provides a computer program product, which includes a program or code. When the program or code is executed (e.g., by a computer device, an object annotation device, one or more processors, etc.), it implements all or part of the steps of the object annotation method provided in the above embodiments.

[0123] It should be understood that the term "at least one" in this application refers to one or more, and "multiple" refers to two or more. The term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Furthermore, for clarity, the terms "first," "second," and "third" are used in this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," and "third" do not limit the quantity or order of execution.

[0124] The method embodiments and device embodiments provided in this application can be referenced interchangeably, and this application does not limit them. The order of operations in the method embodiments provided in this application can be appropriately adjusted, and operations can be added or removed as needed. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and therefore will not be elaborated further.

[0125] In the corresponding embodiments provided in this application, it should be understood that the disclosed devices, etc., can be implemented by other configurations. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0126] The modules described as separate components may or may not be physically separate, and the components described as modules may or may not be physical modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0127] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any equivalent modifications or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An object annotation method, characterized in that, include: Obtain the image to be labeled, including the object to be labeled; The image to be labeled is labeled based on the labeling command triggered by the user to obtain at least one labeling point on the object to be labeled; The coordinate set of the object to be labeled is obtained based on the at least one annotation point. The coordinate set includes the coordinates of multiple pixels. The multiple pixels are located within the target area of ​​the image to be labeled. The overlap between the object to be labeled in the image to be labeled and the target area is greater than a preset overlap. Based on the coordinate set, the color of the object to be labeled is set to a first color to obtain a preprocessed image, the preprocessed image including the object to be labeled in the first color; Based on the user's operation command triggered on the preprocessed image, determine whether the user is satisfied with the preprocessed image; If the user is satisfied, multiple target pixels that meet certain conditions are determined from the set of coordinates. The object to be labeled is then labeled based on these multiple target pixels. The object to be labeled is a linear object, and the minimum distance between the target pixels and the target detection curve is less than a preset distance. The target detection curve is determined based on a subset of the multiple pixels, and the number of the multiple target pixels is greater than a preset number. Alternatively, if the object to be labeled is a non-linear object, the multiple target pixels include at least one of the following: the pixel with the largest horizontal coordinate, the pixel with the smallest horizontal coordinate, the pixel with the largest vertical coordinate, and the pixel with the smallest vertical coordinate. If the user is dissatisfied, the coordinate set of the object to be labeled is obtained based on the annotation points obtained by the user in re-triggering the annotation command for the image to be labeled, and the object to be labeled is labeled based on the re-obtained coordinate set.

2. An object marking device, characterized in that, include: The first acquisition module is used to acquire the image to be labeled, including the object to be labeled; The first annotation module is used to annotate the image to be annotated based on the annotation command triggered by the user for the image to be annotated, so as to obtain at least one annotation point located on the object to be annotated; The second acquisition module is used to acquire the coordinate set of the object to be labeled based on the at least one annotation point. The coordinate set includes the coordinates of multiple pixels, which are located within the target area of ​​the image to be labeled. The overlap between the object to be labeled in the image to be labeled and the target area is greater than a preset overlap. A setting module is used to set the color of the object to be labeled to a first color based on the coordinate set to obtain a preprocessed image, wherein the preprocessed image includes the object to be labeled in the first color; The determining module is used to determine whether the user is satisfied with the preprocessed image based on the operation command triggered by the user for the preprocessed image; The second annotation module is used to, when the user is satisfied, determine multiple target pixels that meet certain conditions from the multiple pixels based on the coordinate set, and annotate the object to be annotated based on the multiple target pixels. The object to be annotated is a linear object, the minimum distance between the target pixels and the target detection curve is less than a preset distance, the target detection curve is determined based on a portion of the multiple pixels, and the number of the multiple target pixels is greater than a preset number; or, the object to be annotated is a non-linear object, and the multiple target pixels include at least one of the following: the pixel with the largest horizontal coordinate, the pixel with the smallest horizontal coordinate, the pixel with the largest vertical coordinate, and the pixel with the smallest vertical coordinate. The second acquisition module is further configured to, in the event that the user is dissatisfied, acquire the coordinate set of the object to be labeled based on the annotation points obtained by the user in re-triggered annotation command for the image to be labeled; The second annotation module is also used to annotate the object to be annotated based on the re-acquired set of coordinates.