An automatic annotation method, device, equipment and readable storage medium

Through the automatic labeling method, the object recognition model and WebGL technology are used to automatically identify and label the target objects in the image, and the labeling box is adjusted by voice, the problems of low efficiency and low accuracy of existing data labeling are solved, and efficient and accurate automatic labeling is achieved.

CN112685998BActive Publication Date: 2025-06-10GLODON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110003636.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-04
Publication Date
2025-06-10
Estimated Expiration
2041-01-04

AI Technical Summary

Technical Problem

Existing data annotation methods rely on manual annotation, are inefficient and have low accuracy, and commonly used tools are complex to use and have high learning costs.

Method used

An automatic labeling method is provided, using the object recognition model to identify the target object from the image to be processed, and the labeling box is drawn through WebGL technology. Users can adjust the labeling box through voice commands to achieve automated labeling.

Benefits of technology

It improves the efficiency and accuracy of data labeling, reduces the error rate and operation complexity of manual labeling, and has a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112685998B_ABST
    Figure CN112685998B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic annotation method, apparatus, device and readable storage medium. The method includes: responding to the triggering of identifying a target object in the image to be processed, and drawing an annotation box in the image to be processed for bounding the target object; responding to the triggering of starting the voice adjustment of the annotation box operation, obtaining a user voice instruction, and according to the annotation instruction for annotating the processed image, obtaining an object recognition model corresponding to the annotation instruction; using the object recognition model to adjust the target annotation box in the image to be processed according to the user voice instruction, so as to complete the annotation of the image to be processed; the present invention can improve the efficiency and accuracy of data annotation at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data annotation, and particularly relates to an automatic annotation method, device, equipment and readable storage medium. Background Art

[0002] In the field of data annotation, the current annotation method mainly relies on manual annotation, that is, data annotators rely on visual recognition to annotate target objects in the dataset to be annotated. Usually, the amount of data to be annotated can reach tens of thousands, hundreds of thousands or more. Therefore, the error probability of manual annotation is relatively high and the efficiency is relatively low, which cannot meet the needs of daily production. In addition, data annotators can also use open-source tools such as labelme or modelArts for data annotation. However, labelme needs to configure a python environment to be used, and modelArts is a tool integrated on the cloud platform and requires corresponding permissions to be used. Therefore, when using the above tools for data annotation, it is relatively cumbersome and has a high learning cost, which is not convenient and the user experience is poor. Summary of the Invention

[0003] The purpose of the present invention is to provide an automatic annotation method, device, equipment and readable storage medium, which can improve the efficiency and accuracy of data annotation at the same time.

[0004] According to one aspect of the present invention, an automatic annotation method is provided, and the method includes:

[0005] In response to a triggered annotation instruction for annotating a to-be-processed image, obtain an object recognition model corresponding to the annotation instruction;

[0006] Use the object recognition model to identify a target object from the to-be-processed image, and draw an annotation box for bounding the target object in the to-be-processed image;

[0007] In response to a triggered operation of starting to adjust the annotation box by voice, obtain a user voice instruction, and adjust the target annotation box in the to-be-processed image according to the user voice instruction to complete the annotation of the to-be-processed image.

[0008] Optionally, the obtaining of the object recognition model corresponding to the annotation instruction includes:

[0009] Obtain an object recognition model corresponding to the annotation instruction from a preset model database, where the model database includes multiple object recognition models pre-trained for recognizing different objects; or,

[0010] Receive an object recognition model pre-trained for recognizing the target object uploaded from a preset interface.

[0011] Optionally, drawing a bounding box for bounding the target object in the image to be processed includes:

[0012] Obtaining the position information of the target object in the image to be processed;

[0013] Determining the two-dimensional coordinates of the starting point and the border size information of the bounding box according to the position information;

[0014] Converting the two-dimensional coordinates of the starting point into three-dimensional coordinates of the starting point;

[0015] Rendering the bounding box in the image to be processed by using WebGL according to the three-dimensional coordinates of the starting point and the border size information.

[0016] Optionally, adjusting the target bounding box in the image to be processed according to the user voice instruction includes:

[0017] Identifying the identification information and adjustment information from the user voice instruction by using a pre-trained voice recognition model;

[0018] Adjusting the attributes of the target bounding box corresponding to the identification information according to the adjustment information; wherein, the attributes include: position and size.

[0019] Optionally, adjusting the attributes of the target bounding box corresponding to the identification information according to the adjustment information specifically includes:

[0020] When the adjustment information is used to adjust the position of the bounding box, obtaining the initial horizontal width and initial vertical height of the target bounding box and the horizontal width and vertical height of the image to be processed;

[0021] Calculating a first horizontal movement step size according to the initial horizontal width of the target bounding box and the horizontal width of the image to be processed;

[0022] Calculating a first vertical movement step size according to the initial vertical height of the target bounding box and the vertical height of the image to be processed;

[0023] Adjusting the position of the target bounding box according to the first horizontal movement step size and the first vertical movement step size according to the adjustment information.

[0024] Optionally, adjusting the attributes of the target bounding box corresponding to the identification information according to the adjustment information specifically includes:

[0025] When the adjustment information is used to adjust the size of the annotation box, obtain the areas of all annotation boxes in the image to be processed, calculate the mean and variance of the annotation box areas according to the probability of each area occurrence, and determine the probability density function for the annotation box area based on the mean and variance;

[0026] Obtain the first area of the target annotation box, and calculate the first probability value corresponding to the first area according to the probability density function;

[0027] Add the first probability value to the preset adjustment step length to obtain a second probability value, and calculate the second area corresponding to the second probability value according to the probability density function;

[0028] Calculate a first adjustment ratio based on the first area and the second area;

[0029] Adjust the size of the target annotation box according to the adjustment information based on the first adjustment ratio.

[0030] Optionally, the method further includes:

[0031] After drawing an annotation box for framing the target object in the image to be processed, generate identification information and an adjustment record queue for the annotation box;

[0032] Add the identification information to the adjustment record queue;

[0033] Record the adjustment information for each time of the annotation box and the attributes after adjustment according to each adjustment information in chronological order into the adjustment record queue.

[0034] Optionally, adjusting the attributes of the target annotation box corresponding to the identification information according to the adjustment information includes:

[0035] Obtain the adjustment record queue corresponding to the identification information;

[0036] When the adjustment information is a rollback, adjust the annotation box according to the attribute of the previous time sequence of the latest attribute recorded in the adjustment record queue.

[0037] Optionally, adjusting the position of the target annotation box according to the adjustment information based on the first horizontal movement step length and the first vertical movement step length specifically includes:

[0038] Obtain the adjustment record queue corresponding to the identification information of the target annotation box;

[0039] When the adjustment information is used to adjust the horizontal position of the annotation box, it is determined whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the horizontal position of the annotation box. If so, when the historical adjustment information of the current time series is inconsistent with that of the subsequent time series and the historical adjustment information of the previous time series is consistent with the adjustment information, the first horizontal movement step length is adjusted to the second horizontal movement step length according to the preset compensation parameter, and the horizontal position of the target annotation box is adjusted according to the second horizontal movement step length and the adjustment information. If not, the horizontal position of the target annotation box is adjusted according to the first horizontal movement step length and the adjustment information;

[0040] When the adjustment information is used to adjust the vertical position of the annotation box, it is determined whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the vertical position of the annotation box. If so, when the historical adjustment information of the current time series is inconsistent with that of the subsequent time series and the historical adjustment information of the previous time series is consistent with the adjustment information, the first vertical movement step length is adjusted to the second vertical movement step length according to the preset compensation parameter, and the vertical position of the target annotation box is adjusted according to the second vertical movement step length and the adjustment information. If not, the vertical position of the target annotation box is adjusted according to the first vertical movement step length and the adjustment information.

[0041] Optionally, the adjusting the size of the target annotation box according to the first adjustment ratio and the adjustment information specifically includes:

[0042] Obtain an adjustment record queue corresponding to the identification information of the target annotation box;

[0043] When the adjustment information is used to adjust the size of the annotation box, it is determined whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the size of the annotation box;

[0044] If so, when the historical adjustment information of the current time series is inconsistent with that of the subsequent time series and the historical adjustment information of the previous time series is consistent with the adjustment information, the preset adjustment step length is adjusted according to the preset compensation parameter, the second adjustment ratio is recalculated according to the adjusted adjustment step length, and the size of the target annotation box is adjusted according to the second adjustment ratio and the adjustment information;

[0045] If not, the vertical position of the target annotation box is adjusted according to the first adjustment ratio and the adjustment information.

[0046] To achieve the above object, the present invention further provides an automatic annotation device, and the device specifically includes the following components:

[0047] An acquisition module, configured to acquire an object recognition model corresponding to the annotation instruction in response to a triggered annotation instruction for annotating a to-be-processed image.

[0048] A drawing module, configured to identify a target object from the to-be-processed image by using the object recognition model, and draw an annotation box for bounding the target object in the to-be-processed image.

[0049] An adjustment module, configured to acquire a user voice instruction in response to a triggered operation of starting to adjust the annotation box by voice, and adjust a target annotation box in the to-be-processed image according to the user voice instruction, so as to complete the annotation of the to-be-processed image.

[0050] To achieve the above object, the present invention further provides a computer device, which specifically includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the automatic annotation method introduced above are implemented.

[0051] To achieve the above object, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the automatic annotation method introduced above are implemented.

[0052] The automatic annotation method, device, equipment and readable storage medium provided by the present invention can intelligently identify a target object from a to-be-processed drawing according to an object recognition model uploaded by a user, and complete the preliminary drawing of an annotation box according to the identified target object by using WebGL technology; in addition, the user can also control the annotation action by voice to adjust the annotation box without manual operation; compared with the prior art, it can improve the efficiency and accuracy of data annotation at the same time. Description of the Drawings

[0053] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0054] Figure 1 An optional flowchart of the automatic annotation method provided for Embodiment 1;

[0055] Figure 2 An optional structural composition diagram of the automatic annotation device provided for Embodiment 2;

[0056] Figure 3 An optional hardware architecture diagram of the computer device provided for Embodiment 3. Specific Embodiments

[0057] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0058] Embodiment 1

[0059] The embodiment of the present invention provides an automatic annotation method, as Figure 1 shown, the method specifically includes the following steps:

[0060] Step S101: In response to a triggered annotation instruction for annotating a to-be-processed image, obtain an object recognition model corresponding to the annotation instruction.

[0061] In this embodiment, the annotation instruction is used to frame a target object in the to-be-processed image through an annotation box to implement the one-key annotation function of the to-be-processed image; for example, when the target object is a human face, all human face areas are framed in the to-be-processed image through an annotation box.

[0062] Specifically, the obtaining of the object recognition model corresponding to the annotation instruction includes:

[0063] Obtain an object recognition model corresponding to the annotation instruction from a preset model database, where the model database includes multiple object recognition models pre-trained for recognizing different objects; or,

[0064] Receive an object recognition model pre-trained for recognizing the target object uploaded from a preset interface.

[0065] In this embodiment, after the user triggers the annotation instruction, according to the target object to be annotated, the user can select a corresponding object recognition model from the preset model database, or upload a custom object recognition model through the preset interface to recognize the target object through the object recognition model. In this embodiment, multiple object recognition models for recognizing different objects will be pre-trained in advance. When it is detected that the user triggers the annotation instruction, the object recognition model including the object identification information can be selected from the model database according to the object identification information included in the annotation instruction.

[0066] It should also be noted that the object recognition model is a model trained in advance for recognizing target objects. The process of training the object recognition model can adopt existing machine learning algorithms, so it will not be elaborated here. For example, through machine learning algorithms, a face recognition model for recognizing faces or a steel bar recognition model for recognizing steel bars can be trained.

[0067] Step S102: Use the object recognition model to recognize the target object from the to-be-processed image, and draw a bounding box in the to-be-processed image for bounding the target object.

[0068] Specifically, the drawing of the bounding box in the to-be-processed image for bounding the target object includes:

[0069] Step A1: Obtain the position information of the target object in the to-be-processed image;

[0070] In this embodiment, when using the object recognition model to recognize the target object from the to-be-processed image, the object recognition model will output the position coordinate information of the target object in the to-be-processed image.

[0071] Step A2: Determine the two-dimensional coordinates of the starting point and the border size information of the bounding box according to the position information;

[0072] Among them, the two-dimensional coordinates of the starting point are the two-dimensional coordinates of the starting point for drawing the bounding box in the to-be-processed image; since the bounding box is usually rectangular, the starting point is any vertex or the center point of the bounding box, and the border size information includes the length and width of the bounding box.

[0073] Step A3: Convert the two-dimensional coordinates of the starting point into three-dimensional coordinates of the starting point;

[0074] Since in this embodiment, WebGL (Web Graphics Library) is used to draw the bounding box, and WebGL is a tool for drawing three-dimensional images, it is necessary to convert the two-dimensional coordinates of the starting point into three-dimensional coordinates of the starting point; preferably, the two-dimensional coordinates of the starting point are converted into three-dimensional coordinates of the starting point by adding a Z-axis coordinate with a value of 0.

[0075] Step A4: According to the three-dimensional coordinates of the starting point and the border size information, use WebGL to render the bounding box in the to-be-processed image.

[0076] It should be noted that in existing data annotation tools, the annotation boxes are all drawn using canvas. However, when the amount of data to be annotated is large, it will lead to a high CPU occupancy rate, resulting in the lag of the visualization interface. In this embodiment, WebGL is used to render the annotation boxes in the visualization interface. Since WebGL renders the annotation boxes based on the open-source framework pixijs and relies on the principle of GPU hardware acceleration for rendering, it greatly reduces the CPU occupancy rate and ensures the speed of data annotation.

[0077] Through the above steps S101 to S102, different object recognition models can be used to complete the annotation work of different objects, providing users with a one-key intelligent annotation service. Users can select the corresponding object recognition model required for one-key annotation according to the target object to be annotated, so as to identify the target object and use WebGL to draw the annotation box to complete one-key annotation. For example, there is currently a steel bar dataset consisting of 5000 images. Users can trigger the annotation instruction for one-key annotation to achieve one-key annotation of 5000 steel bar images. Imagine if manual annotation is carried out, both the annotation accuracy and the time consumed will be greatly reduced.

[0078] Step S103: In response to the triggered operation of starting to adjust the annotation box by voice, obtain the user voice instruction, and adjust the target annotation box in the image to be processed according to the user voice instruction to complete the annotation of the image to be processed.

[0079] In this embodiment, users can trigger an annotation instruction to complete the operation of one-key annotating the target object in the image to be processed, and then draw an annotation box in the image to be processed to frame all the target objects in the image to be processed; however, due to the recognition accuracy of the object recognition model and certain deviations when using WebGL to draw the annotation box, users can also trigger the operation of adjusting the annotation box by voice according to their needs, so as to adjust the drawn annotation box through the user voice instruction to improve the accuracy of data annotation; in addition, since a voice interaction service for adjusting the annotation box according to the user voice instruction is provided in this embodiment, there is no need for users to operate manually, which is more convenient and fast.

[0080] Specifically, the adjustment of the annotation box in the image to be processed according to the user voice instruction includes:

[0081] Step B1: Use a pre-trained voice recognition model to identify the identification information and adjustment information from the user voice instruction;

[0082] Among them, the identification information is the information used to uniquely identify the annotation box. For example, a preset ID can be used as the identification information, or the position coordinates of the annotation box can be used as the identification information;

[0083] The adjustment information is the information used to adjust the attributes of the annotation box. Among them, the attributes include: position and size; the adjustment information specifically includes: adding, deleting, backing, forwarding, moving up, moving down, moving left, moving right, zooming in, zooming out, widening, heightening, narrowing, and shortening.

[0084] Step B2: Adjust the attributes of the target annotation box corresponding to the identification information according to the adjustment information; among them, the attributes include: position and size.

[0085] It should be noted that the speech recognition model can be a model trained based on the open-source TensorFlow.js front-end framework's speech_commands; users only need to record voice sample data in the early stage and set the training data adjustment iteration times, overlap rate, and probability threshold in different scenarios. When the accuracy and loss rate of the trained model reach the preset threshold, a speech recognition model applied to the data annotation tool can be obtained, and users can complete the annotation work without manual operation and only by sending voice commands.

[0086] Further, the specific steps of Step B2 include:

[0087] Step B21: When the adjustment information is used to adjust the position of the annotation box, obtain the initial horizontal width L w and the initial vertical height L h of the target annotation box, as well as the horizontal width R w and the vertical height R h ;

[0088] In this embodiment, when the adjustment information is moving up, moving down, moving left, or moving right, obtain the initial size information of the target annotation box and the size information of the image to be processed.

[0089] Step B22: Calculate the first horizontal movement step step w according to the initial horizontal width L W of the target annotation box and the horizontal width R X of the image to be processed;

[0090] Preferably,

[0091] Step B23: Calculate according to the initial vertical height L h of the target annotation box and the vertical height R W, calculate the first vertical movement step step Y ;

[0092] Preferably,

[0093] Step B24: According to the first horizontal movement step step X and the first vertical movement step step Y , adjust the position of the target annotation box according to the adjustment information.

[0094] For example, when the adjustment information is to move upward, move the target annotation box upward according to the first vertical movement step step Y , and when the adjustment information is to move leftward, move the target annotation box leftward according to the first horizontal movement step step X .

[0095] Furthermore, step B2 specifically further includes:

[0096] Step B21': When the adjustment information is used to adjust the size of the annotation box, obtain the areas of all annotation boxes in the image to be processed, calculate the mean and variance of the annotation box areas according to the probability of each area occurring, and determine the probability density function for the annotation box areas based on the mean and variance;

[0097] In this embodiment, when the adjustment information is to enlarge, reduce, widen, heighten, narrow, or lower, execute step B21'.

[0098] Preferably, the areas of all annotation boxes in the image to be processed conform to a normal distribution, X ∼ N(μ, σ 2 ); by counting the total number Q L of all annotation boxes, and the number i = 1, 2, 3... n of annotation boxes with each area, using to represent the probability of the X i th area occurring, this well describes the probability distribution of the annotation boxes. According to the characteristics of the random variable, the mean μ and variance σ 2 of the annotation box area distribution can be obtained. Therefore, the probability density function for the annotation box areas can be obtained as:

[0099] Step B22': Obtain the first area of the target annotation box and calculate the first probability value corresponding to the first area according to the probability density function;

[0100] Preferably, calculate the first area X = L according to the current horizontal width L w and the current vertical height L h of the target annotation boxw *L h 。

[0101] Step B23': Add the first probability value to the preset adjustment step length to obtain a second probability value, and calculate a second area corresponding to the second probability value according to the probability density function;

[0102] Regarding the supplementary description of the adjustment step length step P , since the area of the annotation box follows a normal distribution, this adjustment step length refers to the probability of the adjustment range of the probability density function P(X) for each adjustment, that is, 0 < step P < 1.

[0103] Step B24': Calculate a first adjustment ratio according to the first area and the second area;

[0104] Preferably, use the ratio of the second area to the first area as the first adjustment ratio.

[0105] Step B25': Adjust the size of the target annotation box according to the first adjustment ratio and the adjustment information.

[0106] Preferably, in step B25', with the starting point of the target annotation box as a reference, adjust the border sizes in the horizontal and vertical directions respectively according to the first adjustment ratio to enlarge or reduce the target annotation box.

[0107] It should also be noted that when the adjustment information is to widen, heighten, narrow, or shorten, adjust the corresponding border size according to the first adjustment ratio.

[0108] Furthermore, the method further includes:

[0109] Step C1: After drawing an annotation box for bounding the target object in the image to be processed, generate identification information and an adjustment record queue for the annotation box;

[0110] Step C2: Add the identification information to the adjustment record queue;

[0111] Step C3: Record the adjustment information for the annotation box each time and the attributes after adjustment according to the adjustment information each time in chronological order in the adjustment record queue.

[0112] In this embodiment, the adjustment record queue mainly records each annotation action during the period from the start of annotation to the completion of annotation by the user and the attribute information of the annotation box after each annotation action.

[0113] Furthermore, step B2 specifically includes:

[0114] Step B21: Obtain an adjustment record queue corresponding to the identification information;

[0115] Step B22: When the adjustment information is a rollback, adjust the bounding box according to the attribute of the previous time sequence of the latest attribute recorded in the adjustment record queue.

[0116] For example, for bounding box A, the user sequentially triggers the following adjustment information: Shift the annotation box A to the left, <ii>Move the annotation box A upward, <iii>Shrink the annotation box A; when the adjustment information triggered by the user again is to go back, then according to the status <ii>Attribute adjustment annotation box

[0117] Further, the step B24 specifically includes:

[0118] Obtain an adjustment record queue corresponding to the identification information of the target annotation box;

[0119] When the adjustment information is used to adjust the horizontal position of the annotation box, determine whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the horizontal position of the annotation box. If so, when the historical adjustment information of the current time sequence is inconsistent with the historical adjustment information of the next time sequence and the historical adjustment information of the previous time sequence is consistent with the adjustment information, adjust the first horizontal movement step length to the second horizontal movement step length according to the preset compensation parameter, and adjust the horizontal position of the target annotation box according to the second horizontal movement step length according to the adjustment information. If not, adjust the horizontal position of the target annotation box according to the first horizontal movement step length according to the adjustment information;

[0120] When the adjustment information is used to adjust the vertical position of the annotation box, determine whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the vertical position of the annotation box. If so, when the historical adjustment information of the current time sequence is inconsistent with the historical adjustment information of the next time sequence and the historical adjustment information of the previous time sequence is consistent with the adjustment information, adjust the first vertical movement step length to the second vertical movement step length according to the preset compensation parameter, and adjust the vertical position of the target annotation box according to the second vertical movement step length according to the adjustment information. If not, adjust the vertical position of the target annotation box according to the first vertical movement step length according to the adjustment information.

[0121] In this embodiment, the horizontal movement step length and the vertical movement step length are adjusted in the following manner:

[0122]

[0123]

[0124] where β and γ are preset compensation parameters; are the horizontal widths of three consecutive time sequences.

[0125] In this embodiment, if three consecutive adjustment messages are [Move Left] -> [Move Left] -> [Move Left], it indicates that the annotation box has not reached the required position, and at this time, the adjustment step size remains unchanged; if three consecutive adjustment messages are [Move Left] -> [Move Right] -> [Move Left], it indicates that the second [Move Right] has moved too far, and at this time, the horizontal movement step size needs to be reduced by a compensation parameter, so that the position after the third [Move Left] is between the position after the first [Move Left] and the position after the second [Move Right].

[0126] Furthermore, step B25’ specifically includes:

[0127] Obtain an adjustment record queue corresponding to the identification information of the target annotation box;

[0128] When the adjustment message is used to adjust the size of the annotation box, determine whether the two latest historical adjustment messages in the adjustment record queue are both used to adjust the size of the annotation box;

[0129] If so, when the historical adjustment message of the current time sequence is inconsistent with the historical adjustment message of the next time sequence and the historical adjustment message of the previous time sequence is consistent with the adjustment message, adjust the preset adjustment step size according to a preset compensation parameter, recalculate the second adjustment ratio according to the adjusted adjustment step size, and adjust the size of the target annotation box according to the second adjustment ratio according to the adjustment message;

[0130] If not, adjust the vertical position of the target annotation box according to the first adjustment ratio according to the adjustment message.

[0131] In this embodiment, the adjustment step size is adjusted in the following manner:

[0132]

[0133] where α is a preset compensation parameter; preferably α = 0.1; A i-2 、A i-1 、A i are the areas of three consecutive time sequences.

[0134] In this embodiment, if three consecutive adjustment messages are [Enlarge] -> [Enlarge] -> [Enlarge], it indicates that the annotation box has not reached the required area size, and at this time, the adjustment step size remains unchanged; if three consecutive adjustment messages are [Enlarge] -> [Shrink] -> [Enlarge], it indicates that the annotation box has exceeded the required area size according to the second [Shrink], and the adjustment step size needs to add a compensation parameter α = 0.1, that is, α * step P1 , so that the area after the third [Enlarge] is between the area after the first [Enlarge] and the area after the second [Shrink].

[0135] It should be noted that, in the data annotation tools in the prior art, when the return operation is triggered, the previously formed annotation box will be directly deleted. However, in this embodiment, by setting an adjustment record queue for each annotation box to record various adjustment operations for the annotation box, when the return operation needs to be executed, the position and style of the annotation box can be adjusted to the previous state according to the records in the adjustment record queue; this is more in line with the software development habits and ideas of users.

[0136] Furthermore, the method further includes:

[0137] Before step S101, obtain a training sample set for training the object recognition model, and determine the to-be-processed image from the training sample set; wherein, the training sample set includes multiple images to be annotated;

[0138] Each image in the training sample set is annotated in the manner of the above-mentioned steps S101 to S103;

[0139] When the annotation of all images in the training sample set is completed, perform model training on the object recognition model according to the training sample set to optimize the object recognition model.

[0140] In this embodiment, the above-introduced automatic annotation method can also be used for the annotation work of sample data in the early stage of model training, thereby improving the accuracy of the model trained using the sample data.

[0141] In addition, the automatic annotation method introduced in this embodiment can be inherited into various hardware devices in the form of a js package, and can be used across platforms and multiple terminals, being more general and flexible. The above method does not need to pay attention to environmental problems. The services of identifying the target object through the object recognition model, drawing the annotation box of the target object through WebGL, and adjusting the annotation box through the user's voice command can be like components, being transparent to the outside while supporting configurability and detachable.

[0142] Embodiment 2

[0143] The embodiment of the present invention provides an automatic annotation device, as Figure 2 shown, the device specifically includes the following components:

[0144] An acquisition module 201, configured to obtain an object recognition model corresponding to the annotation instruction in response to the triggered annotation instruction for annotating the to-be-processed image;

[0145] A drawing module 202, configured to use the object recognition model to identify a target object from the to-be-processed image, and draw an annotation box for framing the target object in the to-be-processed image;

[0146] An adjustment module 203, configured to obtain a user voice instruction in response to a triggered start voice adjustment annotation box operation, and adjust a target annotation box in the image to be processed according to the user voice instruction, so as to complete the annotation of the image to be processed.

[0147] Specifically, an acquisition module 201 is configured to:

[0148] Obtain an object recognition model corresponding to the annotation instruction from a preset model database, where the model database includes a plurality of object recognition models pre-trained for recognizing different objects; or, receive an object recognition model pre-trained for recognizing the target object uploaded from a preset interface.

[0149] In addition, a drawing module 202 is specifically configured to:

[0150] Obtain position information of the target object in the image to be processed; determine two-dimensional coordinates of a starting point and border dimension information of the annotation box according to the position information; convert the two-dimensional coordinates of the starting point into three-dimensional coordinates of the starting point; and render the annotation box in the image to be processed by using WebGL according to the three-dimensional coordinates of the starting point and the border dimension information.

[0151] In addition, the adjustment module 203 specifically includes:

[0152] An identification unit, configured to identify identification information and adjustment information from the user voice instruction by using a pre-trained voice recognition model;

[0153] An adjustment unit, configured to adjust attributes of a target annotation box corresponding to the identification information according to the adjustment information; where the attributes include: position and size.

[0154] Further, the adjustment unit is specifically configured to:

[0155] When the adjustment information is used to adjust the position of the annotation box, obtain an initial horizontal width and an initial vertical height of the target annotation box and a horizontal width and a vertical height of the image to be processed; calculate a first horizontal movement step according to the initial horizontal width of the target annotation box and the horizontal width of the image to be processed; calculate a first vertical movement step according to the initial vertical height of the target annotation box and the vertical height of the image to be processed; and adjust the position of the target annotation box according to the first horizontal movement step and the first vertical movement step according to the adjustment information.

[0156] In addition, the adjustment unit is further specifically configured to:

[0157] When the adjustment information is used to adjust the size of the annotation box, obtain the areas of all annotation boxes in the image to be processed, calculate the mean and variance of the annotation box areas according to the probabilities of each area occurrence, and determine the probability density function for the annotation box areas based on the mean and variance; obtain the first area of the target annotation box, and calculate the first probability value corresponding to the first area according to the probability density function; add the first probability value to a preset adjustment step length to obtain a second probability value, and calculate the second area corresponding to the second probability value according to the probability density function; calculate a first adjustment ratio based on the first area and the second area; and adjust the size of the target annotation box according to the adjustment information according to the first adjustment ratio.

[0158] Further, the apparatus further includes:

[0159] A recording module, configured to, after drawing an annotation box for bounding the target object in the image to be processed, generate identification information and an adjustment record queue for the annotation box; add the identification information to the adjustment record queue; and record the adjustment information for the annotation box each time and the attributes after adjustment according to each adjustment information in chronological order into the adjustment record queue.

[0160] In addition, the adjustment unit is further specifically configured to:

[0161] Obtain the adjustment record queue corresponding to the identification information; when the adjustment information is a rollback, adjust the annotation box according to the attribute of the previous chronological order of the latest attribute recorded in the adjustment record queue.

[0162] Furthermore, when the adjustment unit implements the function of adjusting the position of the target annotation box according to the first horizontal movement step length and the first vertical movement step length according to the adjustment information, it specifically includes:

[0163] Obtain the adjustment record queue corresponding to the identification information of the target annotation box;

[0164] When the adjustment information is used to adjust the horizontal position of the annotation box, determine whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the horizontal position of the annotation box. If so, when the historical adjustment information of the current chronological order is inconsistent with the historical adjustment information of the next chronological order and the historical adjustment information of the previous chronological order is consistent with the adjustment information, adjust the first horizontal movement step length to a second horizontal movement step length according to a preset compensation parameter, and adjust the horizontal position of the target annotation box according to the second horizontal movement step length according to the adjustment information. If not, adjust the horizontal position of the target annotation box according to the first horizontal movement step length according to the adjustment information;

[0165] When the adjustment information is used to adjust the vertical position of the annotation box, it is determined whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the vertical position of the annotation box. If so, when the historical adjustment information of the current time series is inconsistent with that of the next time series and the historical adjustment information of the previous time series is consistent with the adjustment information, the first vertical movement step is adjusted to the second vertical movement step according to the preset compensation parameter, and the vertical position of the target annotation box is adjusted according to the second vertical movement step according to the adjustment information. If not, the vertical position of the target annotation box is adjusted according to the first vertical movement step according to the adjustment information.

[0166] In addition, when the adjustment unit implements the function of adjusting the size of the target annotation box according to the first adjustment ratio according to the adjustment information, it specifically includes:

[0167] Obtain an adjustment record queue corresponding to the identification information of the target annotation box;

[0168] When the adjustment information is used to adjust the size of the annotation box, it is determined whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the size of the annotation box;

[0169] If so, when the historical adjustment information of the current time series is inconsistent with that of the next time series and the historical adjustment information of the previous time series is consistent with the adjustment information, the preset adjustment step is adjusted according to the preset compensation parameter, the second adjustment ratio is recalculated according to the adjusted adjustment step, and the size of the target annotation box is adjusted according to the second adjustment ratio according to the adjustment information;

[0170] If not, the vertical position of the target annotation box is adjusted according to the first adjustment ratio according to the adjustment information.

[0171] Embodiment III

[0172] This embodiment also provides a computer device, such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a rack server, a blade server, a tower server or a cabinet server (including an independent server or a server cluster composed of multiple servers) that can execute programs. As Figure 3 shown, the computer device 30 of this embodiment at least includes, but is not limited to, a memory 301 and a processor 302 that can communicate with each other through a system bus. It should be noted that Figure 3 only the computer device 30 with components 301-302 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0173] In this embodiment, the memory 301 (i.e., the readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 301 may be an internal storage unit of the computer device 30, such as the hard disk or memory of the computer device 30. In other embodiments, the memory 301 may also be an external storage device of the computer device 30, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 30. Of course, the memory 301 may also include both the internal storage unit and the external storage device of the computer device 30. In this embodiment, the memory 301 is generally used to store the operating system and various application software installed in the computer device 30. In addition, the memory 301 may also be used to temporarily store various types of data that have been output or will be output.

[0174] In some embodiments, the processor 302 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 302 is generally used to control the overall operation of the computer device 30.

[0175] Specifically, in this embodiment, the processor 302 is used to execute the program of the automatic annotation method stored in the memory 301. When the program of the automatic annotation method is executed, the following steps are implemented:

[0176] In response to a triggered annotation instruction for annotating the image to be processed, obtain an object recognition model corresponding to the annotation instruction;

[0177] Use the object recognition model to identify the target object from the image to be processed, and draw an annotation box for bounding the target object in the image to be processed;

[0178] In response to a triggered start of the voice adjustment annotation box operation, obtain a user voice instruction, and adjust the target annotation box in the image to be processed according to the user voice instruction to complete the annotation of the image to be processed.

[0179] For the specific implementation process of the above method steps, reference may be made to the first embodiment, and this embodiment will not be repeated here.

[0180] Embodiment Four

[0181] This embodiment also provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, a server, an App application store, etc., on which a computer program is stored. When the computer program is executed by a processor, the following method steps are implemented:

[0182] In response to a triggered annotation instruction for annotating a to-be-processed image, obtain an object recognition model corresponding to the annotation instruction;

[0183] Use the object recognition model to identify a target object from the to-be-processed image, and draw an annotation box for bounding the target object in the to-be-processed image;

[0184] In response to a triggered start of the voice adjustment annotation box operation, obtain a user voice instruction, and adjust the target annotation box in the to-be-processed image according to the user voice instruction to complete the annotation of the to-be-processed image.

[0185] For the specific implementation process of the above method steps, reference can be made to the first embodiment, and this embodiment will not be repeated here.

[0186] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0187] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0189] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.< / ii> < / iii> < / ii>

Claims

1. An automatic annotation method, characterized in that, the method includes: In response to a triggered annotation instruction for annotating the image to be processed, obtaining an object recognition model corresponding to the annotation instruction; Using the object recognition model to identify a target object from the image to be processed, and drawing an annotation box in the image to be processed for bounding the target object; In response to a triggered operation of starting to adjust the annotation box by voice, obtaining a user voice instruction, and adjusting the target annotation box in the image to be processed according to the user voice instruction to complete the annotation of the image to be processed; wherein, the method further includes: Generating identification information and an adjustment record queue for the annotation box, adding the identification information to the adjustment record queue, and recording each adjustment information for the annotation box and the attributes after adjustment according to each adjustment information in chronological order into the adjustment record queue; The adjusting the target annotation box in the image to be processed according to the user voice instruction includes: Using a pre-trained speech recognition model to identify identification information and adjustment information from the user voice instruction, obtaining an adjustment record queue corresponding to the identification information, and when the adjustment information is a rollback, adjusting the annotation box according to the attribute of the previous time sequence of the latest attribute recorded in the adjustment record queue.

2. The automatic annotation method according to claim 1, characterized in that, the obtaining the object recognition model corresponding to the annotation instruction includes: Obtaining an object recognition model corresponding to the annotation instruction from a preset model database, wherein the model database includes multiple object recognition models pre-trained for recognizing different objects; or, Receiving an object recognition model pre-trained for recognizing the target object uploaded from a preset interface.

3. The automatic annotation method according to claim 1, characterized in that, the drawing an annotation box in the image to be processed for bounding the target object includes: Obtaining the position information of the target object in the image to be processed; Determining the two-dimensional coordinates of the starting point and the border size information of the annotation box according to the position information; Converting the two-dimensional coordinates of the starting point into three-dimensional coordinates of the starting point; Rendering the annotation box in the image to be processed using WebGL according to the three-dimensional coordinates of the starting point and the border size information.

4. The automatic annotation method according to claim 1, characterized in that, the adjusting the target annotation box in the image to be processed according to the user voice instruction further includes: Using a pre-trained speech recognition model to identify identification information and adjustment information from the user voice instruction; Adjusting the attributes of the target annotation box corresponding to the identification information according to the adjustment information; wherein, the attributes include: position and size.

5. The automatic annotation method according to claim 4, characterized in that, the adjusting the attributes of the target annotation box corresponding to the identification information according to the adjustment information specifically includes: When the adjustment information is used to adjust the position of the annotation box, obtain the initial horizontal width and initial vertical height of the target annotation box, as well as the horizontal width and vertical height of the image to be processed; Calculate a first horizontal movement step according to the initial horizontal width of the target annotation box and the horizontal width of the image to be processed; Calculate a first vertical movement step according to the initial vertical height of the target annotation box and the vertical height of the image to be processed; Adjust the position of the target annotation box according to the adjustment information according to the first horizontal movement step and the first vertical movement step.

6. The automatic annotation method according to claim 4, wherein, The adjustment of the attributes of the target annotation box corresponding to the identification information according to the adjustment information specifically includes: When the adjustment information is used to adjust the size of the annotation box, obtain the areas of all annotation boxes in the image to be processed, calculate the mean and variance of the annotation box areas according to the probabilities of each area occurrence, and determine the probability density function for the annotation box area according to the mean and variance; Obtain the first area of the target annotation box and calculate a first probability value corresponding to the first area according to the probability density function; Add the first probability value to a preset adjustment step to obtain a second probability value, and calculate a second area corresponding to the second probability value according to the probability density function; Calculate a first adjustment ratio according to the first area and the second area; Adjust the size of the target annotation box according to the adjustment information according to the first adjustment ratio.

7. The automatic annotation method according to claim 5, wherein, The adjustment of the position of the target annotation box according to the adjustment information according to the first horizontal movement step and the first vertical movement step specifically includes: Obtain an adjustment record queue corresponding to the identification information of the target annotation box; When the adjustment information is used to adjust the horizontal position of the annotation box, determine whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the horizontal position of the annotation box. If so, when the historical adjustment information of the current time sequence is inconsistent with the historical adjustment information of the next time sequence and the historical adjustment information of the previous time sequence is consistent with the adjustment information, adjust the first horizontal movement step to a second horizontal movement step according to a preset compensation parameter, and adjust the horizontal position of the target annotation box according to the second horizontal movement step according to the adjustment information. If not, adjust the horizontal position of the target annotation box according to the first horizontal movement step according to the adjustment information; When the adjustment information is used to adjust the vertical position of the annotation box, it is determined whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the vertical position of the annotation box. If so, when the historical adjustment information of the current time series is inconsistent with the historical adjustment information of the next time series and the historical adjustment information of the previous time series is consistent with the adjustment information, the first vertical movement step is adjusted to the second vertical movement step according to the preset compensation parameter, and the vertical position of the target annotation box is adjusted according to the second vertical movement step according to the adjustment information. If not, the vertical position of the target annotation box is adjusted according to the first vertical movement step according to the adjustment information.

8. The automatic annotation method according to claim 6, wherein, The adjusting the size of the target annotation box according to the first adjustment ratio according to the adjustment information specifically includes: Obtaining an adjustment record queue corresponding to the identification information of the target annotation box; When the adjustment information is used to adjust the size of the annotation box, it is determined whether the latest two historical adjustment information in the adjustment record queue are both used to adjust the size of the annotation box; If so, when the historical adjustment information of the current time series is inconsistent with the historical adjustment information of the next time series and the historical adjustment information of the previous time series is consistent with the adjustment information, the preset adjustment step is adjusted according to the preset compensation parameter, a second adjustment ratio is recalculated according to the adjusted adjustment step, and the size of the target annotation box is adjusted according to the second adjustment ratio according to the adjustment information; If not, the vertical position of the target annotation box is adjusted according to the first adjustment ratio according to the adjustment information.

9. An automatic annotation device, wherein, The device includes: An acquisition module, configured to acquire an object recognition model corresponding to the annotation instruction in response to a triggered annotation instruction for annotating a to-be-processed image; A drawing module, configured to identify a target object from the to-be-processed image by using the object recognition model, and draw an annotation box for bounding the target object in the to-be-processed image; An adjustment module, configured to acquire a user voice instruction in response to a triggered operation of starting to adjust the annotation box by voice, and adjust the target annotation box in the to-be-processed image according to the user voice instruction to complete the annotation of the to-be-processed image; Wherein, the device further includes: A recording module, configured to generate identification information and an adjustment record queue for the annotation box, add the identification information to the adjustment record queue, and record each adjustment information for the annotation box and the attributes after adjustment according to each adjustment information in chronological order to the adjustment record queue; The adjustment module is specifically configured to: Use a pre-trained voice recognition model to identify identification information and adjustment information from the user voice instruction, obtain an adjustment record queue corresponding to the identification information, and when the adjustment information is a rollback, adjust the annotation box according to the attribute of the previous time series of the latest attribute recorded in the adjustment record queue.

10. A computer device, the computer device Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Image annotation method and terminal device

    CN108573279A

  • Information labeling method based on neural network model

    CN111444746A