Sampling method, computing device and computer readable storage medium

By sorting and giving priority to sampling based on the accuracy of the image marked data, the problem of missed image labeled errors in large-scale labeled data is solved, and the efficiency and accuracy of data acceptance are improved.

CN120071052APending Publication Date: 2025-05-30HUBEI QIGUANG TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411946614.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In agent creation or vertical model creation, when using labeled data to fine-tune the machine learning model, data acceptance becomes difficult as the amount of labeled data increases, and existing sampling methods cannot effectively increase the probability of labeled error images being randomly inspected.

Method used

By acquiring labeled data for multiple images, the accuracy of each image is determined and sampling from multiple images based on the accuracy, images with lower accuracy have higher sampling priority.

Benefits of technology

This improves the probability that the labeling error image is sampled, ensures priority processing of the labeling error image, and improves the efficiency and accuracy of data acceptance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071052A_ABST
    Figure CN120071052A_ABST
Patent Text Reader

Abstract

The invention relates to a sampling method, computing equipment and a computer readable storage medium. The sampling method comprises the following steps: acquiring labeled data of a plurality of images; determining the accuracy degree of the labeled data of each image; and sampling from the plurality of images based on the degree of accuracy of the labeled data, the image having a lower degree of accuracy having a higher sampling priority.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a sampling method, a computing device, and a computer-readable storage medium. Background Art

[0002] With the advent of the era of artificial intelligence, the demand for training data for AI training is increasing. Although unlabeled data can be used for training machine learning models in some scenarios, in the creation of agents or vertical domain models, it is a common technical means to fine-tune machine learning models using labeled data. Therefore, higher requirements are also placed on the quantity and quality of labeled data. After data annotation is completed by humans or machines, such as dedicated or non-dedicated algorithms, the labeled data needs to be inspected. However, the increasing volume of labeled data poses challenges to the inspection work. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a sampling method, including: Obtaining the labeled data of multiple images; Determining the accuracy of the labeled data of each image; and Sampling from the multiple images based on the accuracy of the labeled data, where images with lower accuracy have higher sampling priorities.

[0004] In at least one embodiment, determining the accuracy of the labeled data of each image includes: Performing text detection on the multiple images to obtain text detection results corresponding to each image; Determining the accuracy of the labeled data of the image according to the labeled data and the text detection results of each image.

[0005] In at least one embodiment, the labeled data includes the positions and presented texts of one or more annotation boxes of each image, the text detection results include the positions and detected texts of one or more text detection boxes, the accuracy includes a first matching degree between the one or more annotation boxes and the one or more text detection boxes, and a second matching degree between the presented text of the annotation box and the detected text of the text detection box having an association relationship according to the first matching degree.

[0006] In at least one embodiment, a text detection box without the association relationship is a missed annotation box of the image, an annotation box without the association relationship is a mislabeled annotation box of the image, an annotation box having the association relationship and whose second matching degree meets the requirements is a correct annotation box of the image, and an annotation box having the association relationship but whose second matching degree does not meet the requirements is a mislabeled annotation box of the image; The method further includes: Determining the accuracy of the labeled data of the image based on the correct bounding boxes, missed bounding boxes, and mislabeled bounding boxes of the image.

[0007] In at least one embodiment, the first matching degree is calculated according to the ratio between the intersection of the bounding box and the text detection box and the text detection box.

[0008] In at least one embodiment, in response to there being multiple text detection boxes that determine the association relationship with the bounding box, the detected texts of the multiple text detection boxes are merged according to the positions of the multiple text detection boxes.

[0009] In at least one embodiment, determining the second matching degree between the presented text of the bounding box having an association relationship and the detected text of the text detection box according to the first matching degree includes: Obtaining the minimum number of edit operations required for mutual conversion between the presented text and the detected text, where the edit operations include at least one of inserting characters, deleting characters, and replacing characters; Determining the second matching degree according to the number of inserted characters, the number of deleted characters, the number of replaced characters, and the character length of the target text, where the target text is the presented text or the detected text.

[0010] In at least one embodiment, an image with a smaller accuracy has a higher sampling priority, including: Sorting the multiple images according to the accuracy, and an image with a lower accuracy is ranked higher among the multiple images; At least one embodiment of the present disclosure provides a computing device, including: A processor; A memory including one or more computer program instructions; Wherein, the one or more computer program instructions are stored in the memory and, when executed by the processor, implement the sampling method provided in any embodiment of the present disclosure.

[0011] At least one embodiment of the present disclosure provides a computer-readable storage medium for storing non-transitory computer-readable instructions, which can implement the sampling method provided in any embodiment of the present disclosure when executed by a computer. Description of the Drawings

[0012] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0013] Figure 1 FIG. 1 shows a schematic block diagram of a computing device 100 provided by at least one embodiment of the present disclosure; Figure 2 A- Figure 2 FIGS. 6D show various possible labeled data output by a human or a machine for the image 20 provided by at least one embodiment of the present disclosure; Figure 3 FIG. 9 shows a flowchart of a sampling method provided by at least one embodiment of the present disclosure; Figure 4 FIG. 12 shows a schematic diagram of a correct labeling result output by a human or a machine for the image 20 provided by at least one embodiment of the present disclosure; Figure 5 FIG. 15 shows a schematic diagram of a text detection result shown by an OCR detection model for the image 20 provided by at least one embodiment of the present disclosure; Figure 6 FIG. 18 shows a schematic block diagram of a hardware implementation of a sampling entity 600 provided by at least one embodiment of the present disclosure; Figure 7A FIG. 21 shows a schematic block diagram of a computing device 700 provided by at least one embodiment of the present disclosure; Figure 7B FIG. 24 shows a schematic block diagram of another computing device 800 provided by at least one embodiment of the present disclosure; Figure 8 FIG. 27 shows a schematic diagram of a storage medium 900 provided by at least one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0015] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar words used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "a", "an" or "the" do not denote a quantity limitation, but mean that there is at least one. Words such as "comprising" or "including" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Words such as "upper", "lower", "left" and "right" are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0016] Typically, before using an image dataset to perform supervised training or fine-tuning on a machine learning model, the image dataset can be purposefully annotated according to the goals of the supervised training or fine-tuning. For example, if the desired machine learning model can effectively identify controls in an image (such as a control presenting text, which is referred to as a "text control" in this disclosure), the machine learning model can output, based on the input image, information such as the position of the control in the image, the text content of the control, and the type of the control.

[0017] This enables the machine learning model to understand various controls of the graphical user interface and has a good effect in, for example, using the machine learning model to imitate or reproduce the operation process of a user on the graphical user interface. For example, a machine learning model that has been subjected to supervised training or fine-tuning can clearly understand the content of one or more graphical user interfaces, and the machine learning model or another machine learning model can perform operations such as sequential operations (such as imitating the operation of clicking on various controls) to achieve various functions.

[0018] In one example, a machine learning model can be obtained by training an untrained model using a labeled image dataset. Generally, this requires a relatively large image dataset. Or it can be obtained by fine-tuning a pre-trained model using a labeled image dataset. Generally, this can be achieved using a relatively small image dataset. Each image in the labeled image dataset can include a corresponding labeled data, which can include the types, positions of all controls in this image, and the text descriptions included in the controls. The labeled data can be stored in forms such as files, texts, tables or other suitable forms.

[0019] The sources of image datasets can be diverse. In some examples, they can be obtained by taking screenshots of the graphical user interface (GUI) being displayed on the display interface. In other examples, they can be obtained by using a camera to capture the GUI being displayed on the display interface. Or they can be obtained by transferring the GUI from a network or physical memory, or by combining one or more of the above methods.

[0020] For example, when the GUI is displayed on the display interface, the GUI or a part of the GUI can be captured through the screenshot function of the device or the screenshot function of the software. After obtaining the captured and saved image, the image can be annotated, for example, through manual annotation or machine annotation (such as a machine learning model with an annotation function), so as to generate the annotated data of the image.

[0021] Whether the annotated data of the image is obtained through manual annotation, machine annotation or combined with other feasible annotation methods, there may be incorrect annotated data. For example, when all the text controls in the image need to be annotated, one possible incorrect situation is missing the annotation of a text control. For example, there is no relevant annotation in the area where there is a text control. Another possible incorrect situation is mislabeling a text control. For example, annotating an area where there is no text control, or although annotating the position of the area where there is a text control, the corresponding text description is incorrect.

[0022] It can be understood that incorrect annotated data will cause the machine learning model to learn incorrect information, which may cause the output result of the trained or fine-tuned machine learning model to deviate from the preset direction.

[0023] The annotated data of the image can be inspected for acceptance to verify the correctness of the annotation. However, in the case of a large amount of annotated data (such as having a very large number of images and corresponding annotated data), it is difficult to inspect all the annotated data for acceptance. A feasible method is proportional sequential sampling or proportional random sampling, or proportional stride sampling to inspect the annotated data of the image for acceptance, and only inspect the extracted image samples to evaluate the overall annotation quality of the annotated data. However, regardless of whether the above sampling method is random, sequential or stride sampling, the probability of detecting mislabeled images cannot be increased during the sampling process, that is, the possibility or priority of all images being sampled is the same, and there is still a problem of missed detection of mislabeled images during the sampling process.

[0024] The present disclosure exemplarily provides a sampling method, which determines the sampling priorities of all images according to the accuracy of the labeled data of each image, and images with lower labeling accuracy have higher sampling priorities or probabilities. This makes it possible for images that may have labeling errors to have higher priorities or probabilities of being sampled, reduces the probability of images that may have labeling errors being missed during inspection, and the set of sampled image samples can better reflect the overall labeling quality.

[0025] See Figure 1 as shown Figure 1 FIG. shows a schematic block diagram of a computing device 100 provided by at least one embodiment of the present disclosure. The computing device 100 may be configured to run the sampling method provided by any example of the present disclosure.

[0026] The computing device 100 may be embodied as any type of computing device. For example, without limitation, the computing device 100 may be embodied as any of the following or otherwise included in any of the following: server computers, embedded computing systems, system-on-chips (SOCs), multi-processor systems, processor-based systems, consumer computing devices, smart phones, cellular phones, desktop computers, tablet computers, notebook computers, laptop computers, network devices, routers, switches, networked computers, wearable computers, handheld devices, messaging devices, camera devices, and / or any other computing device. The computing device 100 includes a processor 102, a memory 104, an input / output (I / O) subsystem 106, a data store 108, a communication circuit 110, a graphics processor 112, a display 114, and one or more peripheral devices 116.

[0027] In some embodiments, one or more of the components of the computing device 100 may be incorporated into another component or otherwise form part of another component. For example, in some embodiments, the memory 104 or a portion thereof may be incorporated into the processor 102. In some embodiments, one or more of the components may be physically separated from another component. For example, in one embodiment, an SOC having a processor 102 and a memory 104 may be connected to a data store 108 external to the SOC via a universal serial bus (USB) connector.

[0028] Processor 102 may be embodied as any type of processor capable of performing the functions described herein. For example, processor 102 may be embodied as a single-core or multi-core processor, a single-socket or multi-socket processor, a digital signal processor, a graphics processor, a neural network computing engine, an image processor, a microcontroller, or other processor or processing / control circuitry. Similarly, memory 104 may be embodied as any type of volatile or non-volatile memory or data storage capable of performing the functions described herein. In operation, memory 104 may store various data and software used during the operation of computing device 100, such as an operating system, applications, programs, libraries, and drivers. Memory 104 may be communicatively coupled to processor 102 via I / O subsystem 106, which may be embodied as circuitry and / or components for facilitating input / output operations between processor 102, memory 104, and other components of computing device 100. For example, I / O subsystem 106 may be embodied as or otherwise include: a memory controller hub, an input / output control hub, a firmware device, a communication link (i.e., a point-to-point link, a bus link, a line, a cable, an optical waveguide, a printed circuit board trace, etc.), and / or other components and subsystems for facilitating input / output operations. I / O subsystem 106 may connect various internal and external components of computing device 100 to each other using any suitable connector, interconnect, bus, protocol, etc. (such as, a system-on-a-chip (SOC) architecture, USB2, USB3, USB4, etc.). In some embodiments, I / O subsystem 106 may form part of a system-on-a-chip (SoC) and may be incorporated, along with processor 102, memory 104, and other components of computing device 100, on a single integrated circuit chip.

[0029] Data storage 108 may be embodied as any type of one or more devices configured for short-term or long-term data storage. For example, data storage 108 may include any one or more memory devices and circuitry, memory cards, hard disk drives, solid state drives, or other data storage devices.

[0030] The communication circuitry 110 may be embodied as any type of interface capable of docking the computing device 100 with other computing devices, such as via one or more wired or wireless connections. In some embodiments, the communication circuitry 110 may be capable of docking with any suitable cable type, such as a cable or an optical fiber cable. The communication circuitry 110 may be configured to use any one or more communication technologies and associated protocols (e.g., cellular network, Ethernet, WiMAX, near field communication (NFC), WiFi, Bluetooth, etc.). The communication circuitry 110 may be located on a chip separate from the processor 102, or the communication circuitry 110 may be included with the processor 102 in a multi-chip package, or even on the same chip as the processor 102. The communication circuitry 110 may be embodied as one or more plug-in boards, daughter cards, network interface cards, controller chips, chip sets, dedicated components (such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)), or other devices that may be used by the computing device 102 to connect to another computing device. In some embodiments, the communication circuitry 110 may be embodied as part of a system-on-chip (SoC) that includes one or more processors, or the communication circuitry 110 may be included on a multi-chip package that also contains one or more processors. In some embodiments, the communication circuitry 110 may include a local processor (not shown) and / or local memory (not shown) that are both local to the communication circuitry 110. In such embodiments, the local processor of the communication circuitry 110 may be capable of performing one or more of the functions of the processor 102 described herein. Additionally, in such embodiments, the local memory of the communication circuitry 110 may be integrated into one or more components of the computing device 102 at the board level, slot level, chip level, and / or other levels.

[0031] The graphics processor 112 is configured to perform graphics computations, such as rendering graphics to be displayed on the display 114. Additionally, in some embodiments, the graphics processor 112 may perform general computing tasks, and / or may perform tasks suitable for the graphics processor 112, such as large parallel operations. The graphics processor 112 may be embodied as any type of processor capable of performing the functions described herein. For example, the graphics processor 112 may be embodied as a single-core or multi-core processor, a single-slot or multi-slot processor, a digital signal processor, a microcontroller, or other processor or processing / control circuitry.

[0032] The display 114 can be embodied as any type of display on which information can be presented to a user of the computing device 100, such as a touchscreen display, a liquid crystal display (LCD), a light-emitting diode (LED) display, an active matrix organic light-emitting diode (AMOLED) display, a cathode ray tube (CRT) display, a plasma display, an image projector (e.g., 2D or 3D), a laser projector, a heads-up display, and / or other display technologies. In some embodiments, the computing device 100 can have more than one display 114 connected to the computing device. The computing device 100 can be capable of disconnecting some or all of the displays 114, such as to reduce the power consumed by the displays 114. Similarly, in some embodiments, the computing device 100 can be capable of changing various parameters of some or all of the displays 114 to reduce power usage, such as changing the refresh rate, resolution, etc.

[0033] In some embodiments, the computing device 100 can include other or additional components (such as those commonly applied in computing devices). For example, the computing device 100 can also have peripherals 116, such as a keyboard, a mouse, speakers, a microphone, an external storage device, etc. In some embodiments, the computing device 100 can be connected to a docking device, which can dock with various devices (including the peripherals 116).

[0034] In other examples, the computing device 100 can also be configured as an image annotation device, that is, the computing device 100 can execute both the relevant annotation algorithms and the sampling methods of the present disclosure, and the use of the computing device 100 is not a limitation to the present disclosure.

[0035] See Figure 2 As shown in FIG. A-2D, various possible annotation results of the output of the artificial or machine on the image 20 are shown. The image 20 can be obtained by any of the foregoing methods in this article. It can be understood that multiple images 20 can form the image dataset for training or fine-tuning the machine learning model described above in this article.

[0036] In one example, the image 20 may include a first control 200, a second control 210, and a third control 220. The first control 200, the second control 210, and the third control 220 may be text controls that perform a certain function on the graphical user interface, such as allow, reject, or other appropriate functions, etc. Usually, the above text controls will display the relevant text description of performing the function (also referred to as "rendered text" in this article). For example, if the image 20 is an interface of a music program, the first control 200, the second control 210, and the third control 220 may be controls that perform different music functions, such as singer selection function, region selection function, today's recommendation function, etc. The relevant text description "Singer" is displayed on the first control 200, the relevant text description "Region" is displayed on the second control 200, and the relevant text description "Today's Recommendation" is displayed on the third control 200.

[0037] As shown in Figure 2 Figure A, which shows a correct labeled data output by a human or a machine for the image 20. The first control 200 is output by a human or a machine with the position of the first annotation box 202 and the corresponding first rendered text 204, the second control 210 is output by a human or a machine with the position of the second annotation box 212 and the corresponding second rendered text 214, and the third control 220 is output by a human or a machine with the position of the third annotation box 222 and the corresponding third rendered text 224.

[0038] As shown in Figure 2 Figure B, which shows a labeled data with missing labels output by a human or a machine for the image 20. The first control 200 is output by a human or a machine with the position of the first annotation box 202 and the corresponding first rendered text 204, the second control 210 is output by a human or a machine with the position of the second annotation box 212 and the corresponding second rendered text 214, and the third control 220 is not output by a human or a machine with the position of any annotation box or any rendered text. That is to say, the third control 220 is missed by a human or a machine in labeling.

[0039] As shown in Figure 2As shown in C, it shows a kind of mislabeled annotated data output by humans or machines for image 20. The first control 200 outputs the position of the first annotation box 202 and the corresponding first presentation text 204 by humans or machines. The second control 210 outputs the position of the second annotation box 212 and the corresponding second presentation text 214 by humans or machines. The third control 220 outputs the position of the third annotation box 222 and the corresponding third presentation text 224 by humans or machines. However, the third presentation text 224 output by humans or machines is "Today's recommendation", which is different from the actual presentation text "Today's recommendation" of the third control 220. That is to say, although humans or machines correctly labeled the position of the third control 220, the text content corresponding to the third annotation box 222 was incorrectly output by humans or machines, and the third annotation box 222 is still a mislabeled annotation box.

[0040] See Figure 2 As shown in D, it shows a kind of mislabeled annotated data output by humans or machines for image 20. The first control 200 outputs the position of the first annotation box 202 and the corresponding first presentation text 204 by humans or machines. The second control 210 outputs the position of the second annotation box 212 and the corresponding second presentation text 214 by humans or machines. The third control 220 outputs the position of the third annotation box 222 and the corresponding third presentation text 224 by humans or machines. And in a blank area of the image 20 or an area with content (the content is not shown in the figure) but not text content, humans or machines output the position of a fourth annotation box 232, and the presentation text corresponding to the fourth annotation box 232 is empty. That is to say, the fourth annotation box 232 is a mislabeled annotation box output by humans or machines.

[0041] It can be understood that for controls with text content, the results of the annotated data output by humans or machines may include: the annotation box position is correct and the text content is correct (such as Figure 2 A), the annotation box and / or text content is missing (such as Figure 2 B), the annotation box position is correct and the text content is wrong (such as Figure 2 C), the annotation box position is wrong (such as Figure 2 D).

[0042] In one example, as Figure 2 shown, each image has at least one annotation result output by humans or machines, that is, an annotated data. Thus, the present disclosure provides a method for sampling from multiple images according to the accuracy of the annotated data.

[0043] See Figure 3 shown, Figure 3 shows a flowchart of a sampling method provided by at least one embodiment of the present disclosure, including: S310, obtain the labeled data of multiple images; For example, if it is desired that a machine learning model can identify each text control included in a graphical user interface or a partial or entire image captured through a graphical user interface, a labeled image data set is required to train or fine-tune the machine learning model. The image data set may include multiple images and corresponding labeled data. The labeled data of each image can be output manually or by a machine, or downloaded from a network such as the cloud. Each labeled data may include the positions of one or more annotation boxes and the presented text corresponding to each annotation box, or may further include the type corresponding to each annotation box. For example, 0 indicates a non-text type, and 1 indicates a text type.

[0044] In one example, the labeled data of each image can be an xml-format file, which has a mapping relationship with the corresponding image. Each xml file records the positions of one or more annotation boxes (such as the upper-left and lower-right coordinates of the annotation box), the type corresponding to the annotation box, and the presented text corresponding to the annotation box.

[0045] It can be understood that the type corresponding to the annotation box indicates the type of the control. For example, if the annotation box type is 0, it indicates that the corresponding control is a non-text control, and if the annotation box type is 1, it indicates that the corresponding control is a text control.

[0046] In some examples, after obtaining the xml file of the image, it can be parsed into a json file and then the labeled data can be extracted from the json file. This is because xml files are easy to store and represent complex data structures, while json files are easy for computing devices to parse and read.

[0047] Take Figure 2Exemplary description is given to the annotated data of the image 20 in A. The annotated data of the image 20 includes a first annotation box 202, a second annotation box 212, and a third annotation box 222. The first annotation box 202 includes a position, a presented text, and / or a control type. The position of the first annotation box 202 can be represented by its upper left corner coordinates (a1, b1) and lower right corner coordinates (a2, b2). The presented text of the first annotation box 202 is "singer", and the type of the first annotation box 202 is 1 (not shown in the figure, represented as a text control). The second annotation box 212 includes a position, a presented text, and / or a control type. The position of the second annotation box 212 can be represented by its upper left corner coordinates (c1, d1) and lower right corner coordinates (c2, d2). The presented text of the second annotation box 212 is "region", and the control type of the second annotation box 212 is 1 (not shown in the figure, represented as a text control). The third annotation box 222 includes a position, a presented text, and / or a control type. The position of the third annotation box 222 can be represented by its upper left corner coordinates (e1, f1) and lower right corner coordinates (e2, f2). The presented text of the third annotation box 222 is "Today's Recommendation", and the control type of the third annotation box 222 is 1 (not shown in the figure, represented as a text control).

[0048] It should be noted that the annotated data uses, for example Figure 4 the upper left corner coordinates plus the lower right corner coordinates as shown to define the position of each annotation box, Figure 4 which shows a schematic diagram of the correct annotation results output by artificial or machine for the image 20 provided by at least one embodiment of the present disclosure, Figure 4 and in which the above-mentioned upper left corner coordinates and lower right corner coordinates can be obtained from the annotation boxes output when the image 20 is annotated by artificial or machine as shown in Figure 2 A. Their x coordinates and y coordinates are usually normalized values within a range of 0 - 1.

[0049] S320, determining the accuracy of the annotated data of each image; Referring to Figure 2 as shown in A - 2D, the image 20 may include various possible annotated data, and the accuracy of each kind of annotated data may be different. The accuracy can be defined according to the absence and / or error of the annotated data. For example, Figure 2 in the case shown in A, there is no absence and / or error in the annotated data, so it has the highest accuracy, while Figure 2 in the cases shown in B - 2D, there are absences and / or errors in the annotated data, so it has a relatively low accuracy.

[0050] Further, it can be based on the correct annotation boxes and mislabeled annotation boxes and the missing or mislabeled bounding boxes to define the accuracy acc, for example, using formula (1): ; to calculate the accuracy of the labeled data for each image.

[0051] For example Figure 2 In the case shown in A, there is no missing or incorrect labeled data, that is, its correct bounding box is 3, the mislabeled bounding box and the missing or mislabeled bounding box are both 0, and its accuracy can be defined as acc = 3 / (3 + 0 + 0) = 1. For example Figure 2 In the case shown in B, the position and the presented text of the third bounding box 222 corresponding to the third text control 220 are missing, that is, there is a missing or mislabeled bounding box , the correct bounding box is 2, the mislabeled bounding box is 0, and its accuracy can be defined as acc = 2 / (2 + 0 + 1) = 2 / 3. For example Figure 2 In the case shown in C, there is incorrect presented text for the third bounding box 222. Even though the position of the third bounding box 222 is correct, the third bounding box 222 is still considered a mislabeled bounding box, that is, there is a mislabeled bounding box , the correct bounding box is 2, the missing or mislabeled bounding box is 0, and its accuracy can be defined as acc = 2 / (2 + 1 + 0) = 2 / 3.

[0052] S330, sample from the multiple images based on the accuracy of the labeled data.

[0053] Among them, the images with lower accuracy have higher sampling priorities.

[0054] It can be understood that the accuracy of the labeled data defines the reliability of the labeled data output by manual or machine or other suitable labeling parties for each image. Referring to the foregoing description in this article, the lower the accuracy, the more missing and / or incorrect the labeled data of the image. That is to say, preferentially sampling the images with lower accuracy can increase the probability of selecting the mislabeled images, that is, ensure that the mislabeled images can be preferentially selected.

[0055] In one example, preferentially sampling the images with lower accuracy is achieved by setting the sampling probability for each image. For example, setting a higher sampling probability for the images with lower accuracy, so as to ensure a greater sampling possibility for the images with lower accuracy in the sampling activity.

[0056] In another example, preferentially sampling images with a lower accuracy level is achieved by sorting multiple images based on their accuracy levels. For example, multiple images can be sorted according to the accuracy levels of the labeled data, and the images with lower accuracy levels are ranked higher among the multiple images, and then sequential sampling is performed from the sorted multiple images.

[0057] For example, the sampling ratio of multiple images can be determined according to the requirements of the actual business. Suppose there are a total of 1000 images, and a suitable sampling ratio is 10%, that is, 100 images are selected for verifying the labeled data. Each image is sorted from low to high according to its accuracy level, and the image with the lowest accuracy level is at the front of the sequence. Sampling is performed from the front of the sequence according to the sample size of 100 until 100 image samples are selected, and the labeled data of the 100 image samples is verified.

[0058] It should be noted that when the number of images is large, there may be several or more images with the same accuracy level. When several or more images with the same accuracy level are not in the sampling sequence (for example, none of them are in the top 100), they can be ignored. When several or more images with the same accuracy level are all in the sampling sequence (for example, all are in the top 100), then all of them are sampled.

[0059] When a part of several or more images with the same accuracy level is in the sampling sequence (for example, in the top 100) and another part is not in the sampling sequence (for example, not in the top 100), one way is to randomly select image samples from the several or more images with the same accuracy level until the sampling meets the quantity of 100, and another way is to appropriately increase the sampling ratio or threshold so that the other images not in the sampling sequence can all be sampled.

[0060] It can be seen that the sampling method provided by the present disclosure can preferentially sample the mislabeled images among all the images. Compared with directly using random sampling or sequential sampling, it increases the discovery probability of mislabeled images, making the sampling result more intuitively reflect the reliability of the object performing the labeling task, such as humans or machines. And under the same sampling ratio, it can discover more problems existing in the labeled data.

[0061] Referring to the foregoing description of this article, the accuracy of the labeled data can be defined based on the absence and / or errors of the labeled data. The absence and / or errors of the labeled data can be determined according to the matching between the labeled data of the image and the text detection result obtained by performing a text detection on the image. The text detection result represents the positions of all text regions in the image and the corresponding text content. That is, text detection can be performed on multiple images to obtain the text detection result corresponding to each image; the accuracy of the labeled data of each image can be determined according to the labeled data and the text detection result of each image.

[0062] For example, an OCR (Optical Character Recognition) detection model can be used to obtain the above text detection result. The OCR detection model can be, for example, EasyOCR, YouTu, ChineseOCR, PaddleOCR, etc., which does not limit the present disclosure.

[0063] For example, using PaddleOCR to perform OCR processing on an image can obtain a mask of one or more text regions in the image and the corresponding detected text. The above mask represents the position of each text region in the image, and the detected text represents the text content displayed on the image in this text region.

[0064] It can be understood that for the convenience of processing by the computing device, the above mask can be converted into a text detection box. Usually, the mask of the text region includes a set of abscissas X mask and a set of ordinates Y mask , through formula (2): ; to obtain the upper left coordinate and the lower right coordinate of the corresponding text detection box. The above upper left coordinate and the lower right coordinate indicate the position of the text detection box.

[0065] Since the labeled data already includes the positions and presented texts of one or more annotation boxes, the matching between the labeled data and the text detection result can be performed through the matching of positions and the matching of presented text and detected text.

[0066] For example, the first matching degree calculated based on the positions between one or more annotation boxes and one or more text detection boxes can be determined first.

[0067] For example, the Intersection over Union (IOU) can be used to measure the overlap degree between two bounding boxes, which is the intersection of the two bounding boxes divided by their union. However, the present disclosure provides another method for calculating the overlap degree of bounding boxes.

[0068] See Figure 5 as shown Figure 5 shows a schematic diagram of the text detection result of the OCR detection model provided by at least one embodiment of the present disclosure for the image 20, including a first text detection box 206, a second text detection box 216, and a third text detection box 226. The first case of the text control is shown as the first text detection box 206 and the first annotation box 202, and the second text detection box 216 and the second annotation box 212. For the first text control 200 and the second text control 210, the area with text only occupies a small part of the first text control 200 and the second text control 210, and other areas can display, for example, images or other suitable content. For example, for a control with an image and text, the entire control area can trigger a click event, but only a small part of the control area has text. The second case of the text control is shown as the third text detection box 226 and the third annotation box 222. For the third text control 220, the area with text occupies most of the third text control 220.

[0069] That is to say, although the above three text detection boxes are all located within the corresponding annotation boxes (and may partially exceed the annotation boxes in some examples), the IOU in the first case calculated using the traditional IOU is quite different from the IOU in the second case. This is because the intersection of the text detection box and the annotation box in the first case is much smaller than their union. This makes the result of the IOU calculation a relatively small value (for example, less than 0.5); while in the second case, the intersection of the text detection box and the annotation box is close to their union, which makes the result of the IOU calculation a value close to 1 (for example, 0.97). This makes it difficult to determine an accurate IOU threshold to judge the first matching degree between the annotation box and the text detection box.

[0070] Therefore, the present disclosure exemplarily calculates the above first matching degree by dividing the intersection between the annotation box and the text detection box by the text detection box.

[0071] Taking Figure 5 as an example to calculate the first matching degree between the first annotation box 202 and the first text detection box 206, the upper left coordinate of the first annotation box 202 is (a1, b1), and the lower right coordinate is (a2, b2). The upper left coordinate of the first text detection box 206 is , and the lower right coordinate is , the first matching degree is calculated through formula (3) and formula (4). : ; ; It can be seen that the denominator of formula (4) represents the area of the first text detection box 206, and the numerator of formula (4) represents the overlapping area between the first annotation box 202 and the first text detection box 206.

[0072] By the method provided in the present disclosure for calculating the first matching degree , regardless of whether the proportion of the text area occupying the control is high or low, the intersection of the annotation box and the text detection box is compared with the text detection box (i.e., the text area). Therefore, when there is an association relationship between the text detection box and the annotation box (i.e., generally, the text detection box is considered to belong to the annotation box), the calculated first matching degree is always a relatively large value, such as greater than 0.9, and there will not be a situation where the first matching degree calculated by the traditional IOU method is relatively small, such as less than 0.5, that is, in this case, the traditional IOU method will consider that there is no association relationship between the text detection box and the annotation box (while in fact there is). The method provided in the exemplary manner of the present disclosure for calculating the first matching degree can effectively avoid the generation of the above errors and can reduce the difficulty of setting the first matching degree threshold.

[0073] After calculating the first matching degree between each text detection box and each annotation box , it is necessary to determine the annotation box to which each text detection box corresponds. First, the maximum first matching degree can be selected from the multiple first matching degrees between the text detection box and each annotation box . Then, the relationship between the maximum first matching degree and the first matching degree threshold is compared. If the maximum first matching degree is greater than the first matching degree threshold, there is an association relationship between the text detection box and the annotation box corresponding to the maximum first matching degree , that is, the text detection box belongs to the annotation box corresponding to the maximum first matching degree . If the maximum first matching degree is not greater than the first matching degree threshold, the text detection box has no corresponding annotation box, that is, there is no association relationship. Obviously, there is a text control in the area of the text detection box but there is no corresponding annotation box in the annotation data. At this time, the number of missed annotation boxes is incremented by 1.

[0074] In one example, information recording can be synchronized for the text detection box with missing labels. For example, the position of the text detection box and the detected text are recorded to facilitate subsequent verification.

[0075] In one example, a first matching degree threshold slightly less than 1, such as 0.95 or other appropriate values, can be set to determine the association relationship between each text detection box and each annotation box.

[0076] It should be noted that the reason the first matching degree threshold does not take 1 is that the text detection box recognized by OCR may be too large, resulting in, for example, the boundary of the text detection box exceeding the boundary of the annotation box, which makes the calculated first matching degree less than 1. If the first matching degree threshold takes 1, misjudgment of the association relationship may occur in this case.

[0077] In one example, it is possible to first determine whether there is an association relationship between the current text detection box and all annotation boxes according to the process described above in this article, and then determine the association relationship between the next text detection box and all annotation boxes. In another example, it is possible to determine in parallel whether there is an association relationship between multiple text detection boxes and annotation boxes.

[0078] When all text detection boxes have been traversed, all annotation boxes need to be traversed again.

[0079] In one example, according to the first matching degree calculated above in this article greater than the first matching degree threshold, obtain the text detection box corresponding to the current annotation box. If the current annotation box does not include any corresponding text detection box, that is, the current annotation box has no association relationship with any text detection box, obviously there is no text content in the current annotation box. For example, it may be that an incorrect annotation was made in an area without a text control (as shown in Figure 2 232 of D), and the current annotation box is a mislabeled annotation box. At this time, the number of mislabeled annotation boxes is incremented by 1. If the current annotation box includes at least one corresponding text detection box, then the current annotation box has an association relationship with at least one text detection box, that is, the position of the current annotation box is at least correctly labeled.

[0080] In one example, information recording can be synchronized for the currently mislabeled annotation box. For example, the position of the current annotation box and the presented text (for example, it can be recorded as a null value when there is no presented text) are recorded to facilitate subsequent verification.

[0081] After determining the association relationship between the text detection box and the annotation box according to the first matching degree it is also necessary to determine the second matching degree between the presented text of the annotation box with an association relationship and the detected text of the text detection box according to the first matching degree.

[0082] For example, obtain the presented text in the labeled data of the current annotation box, obtain the detected text of the text detection box associated with the current annotation box, and calculate the second matching degree between the presented text and the detected text based on the detected text obtained by OCR recognition.

[0083] In one example, the word accuracy rate is used to measure the second matching degree between the presented text and the detected text. When the word accuracy rate is greater than the word accuracy rate threshold (i.e., the second matching degree meets the requirements), it can be considered that the presented text and the detected text match successfully, that is, the presented text in the labeled data is correctly labeled. As previously described in this article, the position of the current annotation box is at least correctly labeled (i.e., there is an associated relationship), so the current annotation box is a correctly labeled box, and the number of correctly labeled boxes is incremented by 1.

[0084] If the word accuracy rate is not greater than the word accuracy rate threshold (i.e., the second matching degree does not meet the requirements), it can be considered that the presented text and the detected text do not match successfully, that is, the presented text of the annotation box is mislabeled. Even if the position of the current annotation box is at least correctly labeled, the current annotation box is still a mislabeled box, and the number of mislabeled boxes is incremented by 1.

[0085] In one example, information about the mislabeled current annotation box can be recorded synchronously, such as recording the position of the current annotation box and the presented text (for example, it can be recorded as a null value when there is no presented text), which is convenient for subsequent verification.

[0086] It should be noted that assuming that on the basis shown in Figure 2 B, the third control 220 has the position of an annotation box output by a human or a machine, such as Figure 2 A or Figure 2 the position of the third annotation box 222 shown in C, but lacks relevant text descriptions, then the second matching degree between the presented text of the third annotation box 222 and the text detection box with an associated relationship does not meet the requirements, that is, the third annotation box 222 is a mislabeled box .

[0087] It can be understood that when traversing the attribution relationship between all text detection boxes and annotation boxes, whether each annotation box includes at least one text detection box, and the word accuracy rate between the associated annotation box and the text detection box , the correctly labeled boxes , mislabeled boxes and missing-labeled boxes The number of labeled images is then calculated using the formula (1) mentioned above to calculate the accuracy acc of the labeled data for each image.

[0088] That is to say, the process described in this example can output the accuracy acc of the labeled data of each image. In a sense, the accuracy acc represents the matching degree between the labeled data and the text detection result obtained by the OCR detection model performing OCR on the image.

[0089] The above example of the present disclosure describes the accuracy of identifying the labeled data of an image using the output of an OCR detection model, by using the missing label box to identify the missing labeled data in the image where text controls exist , there is no text control but the labeled data is mislabeled or there is a text control but the labeled data presents the wrong labeling box , and a correctly labeled box where a text control exists and the location of the labeled data and the rendered text are correctly labeled , which describes the accuracy of the labeled data in a more refined way, making the sampling of incorrectly labeled images more accurate.

[0090] In one example, the text detection box that determines the association relationship with the annotation box is one, and in this case, it is only necessary to determine the character accuracy rate between the presented text of the annotation box and the detected text of the text detection box. That's it.

[0091] In another example, there are multiple text detection boxes that determine the association relationship with the annotation box. In this case, it is necessary to merge the detected texts of the multiple text detection boxes according to the positions of the multiple text detection boxes, and then determine the word accuracy between the presented text of the annotation box and the merged detected text. .

[0092] For example, the text splicing order can be obtained by arranging the positions of multiple text detection boxes twice in the x-axis direction and the y-axis direction. For example, the text detection boxes that determine the association relationship with the annotation box include , and , through the text detection box , and The upper left corner coordinates are used to determine the merging order. Assuming the text detection box The coordinates of the upper left corner are (0.1, 0.2), and the text detection box The coordinates of the upper left corner are (0.6, 0.1), and the text detection box The coordinates of the upper left corner are (0.4, 0.1).

[0093] First, according to the text detection box in the x-axis direction , and The upper left coordinates are sorted for the first time, and the sorting method is ascending order, that is, the order of the sorted set is { }.

[0094] Then, in the y-axis direction, according to the text detection boxes , and the upper left coordinates are sorted for the second time, and the sorting method is ascending order, that is, the order of the sorted set is { }.

[0095] Finally, according to the order of the set after the second sorting, , and are traversed, and , and the detected texts of the text detection boxes are concatenated in turn to obtain the combined detected text.

[0096] Exemplarily, the present disclosure performs two sorts on multiple text detection boxes based on the x-axis direction and the y-axis direction, and combines the detected texts in the multiple text detection boxes after sorting, effectively ensuring the reliability of subsequent text matching.

[0097] As previously described herein, the character alignment rate can be used to measure the second matching degree between the presented text and the detected text. Herein, the possible ways to calculate the character alignment rate will be described exemplarily.

[0098] Assume that the presented text to be matched is , the detected text is (for example, the text in a text detection box or the text combined from multiple text detection boxes), the character length (quantity) of the presented text is , and the character length (quantity) of the detected text is . Exemplarily, the present disclosure uses the minimum cost matching Levenshtein principle and applies the dynamic programming algorithm to obtain the presented text and the detected text The minimum number of edit operations required for mutual conversion between them is obtained to get the number of characters deleted D, the number of characters inserted I, and the number of characters replaced S. Then, according to the number of characters deleted D, the number of characters inserted I, the number of characters replaced S, and the character length of the target text, the second matching degree is calculated. It can be understood that the edit operation is a quantitative measurement of the difference degree between two texts, and the measurement method is based on at least how many times of processing (i.e., the minimum edit operation) to transform one text into another text. Additionally, the presented text is transformed into the detection text , and the detection text is transformed into the presented text . The sum of the number of edit operations required is the same, and the difference lies in that the types of edit operations required (for example, one is a character deletion operation, while the other is a character insertion operation) may be different.

[0099] It can be understood that when the second matching degree is calculated based on the minimum number of edit operations required for the presented text to be transformed into the detection text , the target text is the detection text ; when the second matching degree is calculated based on the minimum number of edit operations required for the detection text to be transformed into the presented text , the target text is the presented text .

[0100] In an example of the present disclosure, taking the text detection result output by the OCR detection model as the standard, the minimum number of edit operations required for the presented text to be transformed into the detection text and the character length of the detection text are used in formula (5): ; to calculate the second matching degree.

[0101] Next, this article will describe the Levenshtein principle and related calculation processes.

[0102] The Levenshtein distance is a type of edit distance, which generally refers to the minimum number of edit operations required to transform one text into another text using only the following specific allowed rules between two texts. The allowed edit operations include: 1. Replace a character, replacing one character with another character; 2. Insert a character, inserting a character; 3. Delete a character, deleting a character.

[0103] Referring to the foregoing description in this article, the presented text and the detection text The character lengths are respectively and , then the Levenshtein distance between the presented text and the detected text meets the formula (6): ; Among them, is an indicator function, which takes the value of 1 when and takes the value of 0 otherwise. represents the first characters of and the first characters of and are both starting from 1). If , that is, if any one of the character lengths of the presented text and the detected text is 0 (i.e., an empty string), then the Levenshtein distance between the presented text and the detected text is the maximum of the character lengths of the two texts.

[0104] If , then take the minimum value from the following three edit operations: represents deleting a character from to reach , that is, reducing the string in by 1, calculating the Levenshtein distance from the string to string, and adding 1 to the result. This is actually equivalent to deleting the character at the th position in ; represents inserting a character in to reach , that is, reducing the string in by 1, calculating the Levenshtein distance from the string to string, and adding 1 to the result. This is actually equivalent to deleting the character at the th position in ; represents replacing the character in to reach (whether to replace depends on whether holds, only when only when replacement occurs), which is actually equivalent to replacing the character at the position with the character at the position in . position.

[0105] Then, determine the minimum number of edit operations required to convert the string from to the string to based on the minimum values of the above three edit operations. When runs to and , runs to , the number of deleted characters D, the number of inserted characters I, and the number of replaced characters S required for the minimum edit operation described in the present disclosure can be obtained. Then, calculate the word accuracy rate between the presented text and the detected text according to formula (5) described in this article.

[0106] Figure 6 FIG. shows a schematic block diagram of the hardware implementation of the sampling entity 600 provided by at least one embodiment of the present disclosure.

[0107] The sampling entity 600 can be implemented using a processing system 601 that includes one or more processors 602. Examples of the processor 602 include a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable hardware configured to perform various functions described throughout the present disclosure. In various examples, the sampling entity 600 can be configured to perform any one or more of the functions described herein. That is, the processor 602 as used in the sampling entity 600 can be used to implement the processes and procedures described below and shown in Figure 3 or any one or more of the processes and procedures of the embodiments herein.

[0108] In this example, the processing system 601 can be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system 601, the bus 616 can include any number of interconnecting buses and bridges. The bus 616 communicatively couples various circuits of one or more processors (which are generally represented by the processor 602), a memory 613, and a computer-readable medium (which is generally represented by the computer-readable medium 607). The bus 616 can also connect various other circuits such as a timing source, peripheral devices, voltage regulators, and power management circuits, which are well known in the art and thus will not be described further. The bus interface 614 provides an interface between the bus 616 and the network interface 615. The network interface 615 provides a communication interface or unit for communicating with various other devices via a transmission medium. Depending on the nature of the device, a user interface 612 (e.g., a keyboard, a display, a speaker, a microphone, a joystick, a touch screen) can also be provided.

[0109] In some aspects of the present disclosure, the processor 602 can include an acquisition circuit 603 configured for various functions, such as, for example, acquiring labeled data of a plurality of images. For example, the acquisition circuit 603 can include logic circuitry coupled to a memory component (e.g., the memory 613 and / or the computer-readable medium 607), where the acquisition circuit 603 can be configured to define and / or obtain labeled data of a plurality of images (e.g., the labeled data of the plurality of images can be determined via the user interface 612). As Figure 6As shown, the processor 602 may further include a determination circuit 604 configured to determine the accuracy of the labeled data for each image. For example, the determination circuit 604 may include logic circuitry coupled to a memory component (e.g., memory 613 and / or computer-readable medium 607), wherein the determination circuit 604 may be configured to determine the accuracy of the labeled data based on any one of a plurality of parameters (e.g., such parameters may be defined via the user interface 612), such as determining the accuracy according to the position and presented text of one or more annotation boxes in the labeled data, and the position and detected text of one or more text detection boxes obtained by performing OCR on the image. The processor 602 may further include a sorting circuit 605 configured to sort a plurality of images according to the accuracy, and the execution of the sorting circuit 605 causes the images with lower accuracy to be located in the earlier positions in the sequence. For example, the sorting circuit 605 may include logic circuitry coupled to a memory component (e.g., memory 613 and / or computer-readable medium 607). The processor 602 may further include a sampling circuit 606 configured to sample from a plurality of images based on the accuracy of the labeled data. For example, the file generation circuit 606 may include logic circuitry coupled to a memory component (e.g., memory 613 and / or computer-readable medium 607), wherein the sampling circuit 606 may be configured to obtain a sampled image sample based on a plurality of parameters (e.g., such parameters may be defined via the user interface 612 and the parameters output by the sorting circuit 605).

[0110] Although not shown in Figure 6 , in one example, the processor may further include a transmission circuit configured for various functions, such as sending the sampled image sample (e.g., already including the labeled data) to a server or other suitable computing device, or receiving a plurality of images (e.g., already including the labeled data) sent by the server. It should be understood that the transmission circuit may include logic circuitry coupled to the network interface 615, and such logic circuitry may be configured to determine whether and when to send the image sample (e.g., already including the labeled data) to one or more servers via the network interface 615, and to determine whether and when to receive a plurality of images (e.g., already including the labeled data) sent by one or more servers via the network interface 615.

[0111] It should be understood that the processor 602 is responsible for managing the bus and general processing, which includes executing software stored on the computer-readable medium 607. When executed by the processor 602, the software causes the processing system 601 to perform various functions described hereinafter for any specific device. The computer-readable medium 607 and the memory 613 may also be used to store data manipulated by the processor 602 when executing the software.

[0112] One or more processors 602 in the processing system may execute software. Software should be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, processes, functions, etc., regardless of whether it is called software, firmware, middleware, microcode, hardware description language, or other terms. The software may be located on a computer-readable medium 607. The computer-readable medium 707 may be a non-transitory computer-readable medium. By way of example, non-transitory computer-readable media include magnetic storage devices (e.g., hard disks, floppy disks, magnetic tape), optical disks (e.g., compact disc (CD) or digital versatile disc (DVD)), smart cards, flash memory devices (e.g., cards, sticks, or key drives), random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, removable disks, and any other suitable medium for storing software and / or instructions that can be accessed and read by a computer. The computer-readable medium 607 may be located within the processing system 601, outside the processing system 601, or distributed across multiple entities including the processing system 601. The computer-readable medium 607 may be embodied in a computer program product.

[0113] In one or more examples, the computer-readable storage medium 607 may include acquisition software 608 configured for various functions, such as acquiring annotated data for a plurality of images. As Figure 6 shown, the computer-readable storage medium 607 may also include determination software 609 configured to determine the accuracy of the annotated data for each image. For example, the determination software 609 may be configured to determine the accuracy of the annotated data based on any one of a plurality of parameters (e.g., such parameters may be defined via a user interface 612), such as based on the position and rendered text of one or more annotation boxes in the annotated data, and the position and detected text of one or more text detection boxes obtained by performing OCR on the image to determine the accuracy. As Figure 6 shown, the computer-readable storage medium 607 may also include sorting software 610 configured to sort a plurality of images according to the accuracy, and the execution of the sorting software 610 causes images with lower accuracy to be located in a more forward position in the sequence. As Figure 6 shown, the computer-readable storage medium 607 may also include sampling software 611 configured to sample from a plurality of images based on the accuracy of the annotated data. For example, the sampling software 611 may obtain an image sample sampled from a plurality of images. Although Figure 6Not shown in the figure, the computer-readable storage medium 607 may further include transmission software configured for various functions, such as sending the sampled image samples (e.g., already including labeled data) to a server or other suitable computing device, or receiving multiple images (e.g., already including labeled data) sent by the server.

[0114] Of course, in the above example, the circuits included in the processor 602 are provided only as examples, and other units for performing the described functions may be included in various aspects of the present disclosure, including but not limited to instructions stored in the computer-readable storage medium 607, or any other suitable device or unit.

[0115] Next, at least one embodiment of the present disclosure further provides a computing device, which includes a processor and a memory. The memory includes one or more computer program modules. The one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing any one of the sampling methods described above in this article.

[0116] Figure 7A A schematic block diagram of a computing device 700 provided by at least one embodiment of the present disclosure is shown. As Figure 7A shown, the computing device 700 includes a processor 710 and a memory 720. The memory 720 is used to store non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 710 is used to run the non-transitory computer-readable instructions, and when the non-transitory computer-readable instructions are run by the processor 710, one or more steps in the communication method described above can be executed. The memory 720 and the processor 710 may be interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0117] For example, the processor 710 may be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) may be of the X86 or ARM architecture, etc. The processor 710 may be a general-purpose processor or a dedicated processor, and may control other components in the computing device 700 to perform desired functions.

[0118] For example, the memory 720 may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor 710 may run one or more computer program modules to implement various functions of the computing device 700. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.

[0119] It should be noted that in the embodiments of the present disclosure, for the specific functions and technical effects of the computing device 700, reference may be made to the description of the sampling method in the foregoing text, and details are not described herein again.

[0120] Figure 7B FIG. shows a schematic block diagram of another computing device 800 provided by at least one embodiment of the present disclosure. The computing device 800 is, for example, suitable for implementing the sampling method provided by any embodiment of the present disclosure. The computing device 800 may be a terminal device such as a mobile phone, a tablet computer, a navigation device, a vehicle (such as an in-vehicle system), a smart wearable device, etc. It should be noted that Figure 7B The illustrated computing device 800 is merely an example and will not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0121] As Figure 7B shown, the computing device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 810, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 820 or the program loaded from the storage device 880 into the random access memory (RAM) 830. In the RAM 830, various programs and data required for the operation of the computing device 800 are also stored. The processing device 810, the ROM 820, and the RAM 830 are connected to each other through a bus 840. The input / output (I / O) interface 850 is also connected to the bus 840.

[0122] Typically, the following devices can be connected to the I / O interface 850: input devices 860 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 870 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 880 including, for example, magnetic tapes, hard disks, etc.; and communication devices 890. The communication device 890 can allow the computing device 800 to communicate with other computing devices wirelessly or wiredly to exchange data. Although Figure 7B the computing device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the computing device 800 can alternatively implement or have more or fewer devices.

[0123] For example, according to an embodiment of the present disclosure, the above sampling method can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the above communication method. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 890, or installed from the storage device 880, or installed from the ROM 820. When the computer program is executed by the processing device 810, the functions defined in the sampling method provided by the embodiments of the present disclosure can be realized.

[0124] At least one embodiment of the present disclosure also provides a computer-readable storage medium for storing non-temporary computer-readable instructions, which can realize any sampling method described herein when executed by a computer.

[0125] Figure 8 The schematic diagram of a storage medium 900 provided by at least one embodiment of the present disclosure is shown. As Figure 8 shown, the storage medium 900 is used to store non-temporary computer-readable instructions 910. For example, when the non-temporary computer-readable instructions 910 are executed by a computer, one or more steps of any sampling method described herein can be executed.

[0126] For example, the storage medium 900 can be applied to the above computing device 700 or computing device 800. For example, the storage medium 900 can be Figure 7A the memory 720 in the computing device 700 shown, or, the storage medium 900 can be Figure 7B the ROM 820 or RAM 830 in the computing device 800 shown. For example, the relevant description about the storage medium 800 can refer to Figure 7A the corresponding description of the memory 720 in the computing device 700 shown, which will not be elaborated here.

[0127] The following points need to be noted: (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0128] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0129] As mentioned above, the above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A sampling method comprising: Get annotated data for multiple images; Determining the accuracy of the labeled data for each image; as well as Sampling is performed from the plurality of images based on the accuracy of the labeled data, wherein images with a lower accuracy have a higher sampling priority.

2. The method according to claim 1, wherein: Determining the accuracy of the labeled data for each image, including: Performing text detection on the multiple images to obtain a text detection result corresponding to each image; The accuracy of the labeled data of each image is determined according to the labeled data and the text detection result of the image.

3. The method according to claim 2, wherein: The annotated data includes the position and presented text of one or more annotation boxes of each image, the text detection result includes the position and detected text of one or more text detection boxes, the accuracy includes a first matching degree between the one or more annotation boxes and the one or more text detection boxes, and a second matching degree between the presented text of the associated annotation box and the detected text of the text detection box determined based on the first matching degree.

4. The method according to claim 3, wherein: The text detection box that does not have the association relationship is a missing annotation box of the image, the annotation box that does not have the association relationship is a wrongly labeled annotation box of the image, the annotation box that has the association relationship and the second matching degree meets the requirement is the correct annotation box of the image, and the annotation box that has the association relationship but the second matching degree does not meet the requirement is the wrongly labeled annotation box of the image; The method further comprises: The accuracy of the labeled data of the image is determined based on the correct labeled boxes, the missing labeled boxes, and the wrongly labeled boxes of the image.

5. The method according to claim 3, wherein: The first matching degree is calculated based on a ratio between an intersection of the annotation box and the text detection box and the text detection box.

6. The method according to claim 3, further comprising: In response to the fact that there are multiple text detection boxes that have the association relationship determined with the annotation box, the detected texts of the multiple text detection boxes are merged according to the positions of the multiple text detection boxes.

7. The method according to claim 3, wherein: Determining a second matching degree between the presented text of the annotation box and the detected text of the text detection box having an associated relationship according to the first matching degree includes: Obtaining the minimum editing operations required for mutual conversion between the presented text and the detected text, wherein the editing operations include at least one of inserting characters, deleting characters, and replacing characters; The second matching degree is determined according to the number of times the characters are inserted, the number of times the characters are deleted, the number of times the characters are replaced, and the character length of the target text, wherein the target text is a presented text or a detected text.

8. The method according to claim 1, wherein: Images with a lower stated accuracy have a higher sampling priority, including: sorting the plurality of images according to the accuracy, with images having a lower accuracy being placed at a higher position among the plurality of images; Sequential sampling is performed from the sorted plurality of images.

9. A computing device comprising: processor; a memory including one or more computer program instructions; The one or more computer program instructions are stored in the memory and, when executed by the processor, implement the sampling method according to any one of claims 1 to 8.

10. A computer-readable storage medium non-temporarily storing computer-readable instructions, wherein: When the computer-readable instructions are executed by a processor, the sampling method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method and device for assisting in detecting image annotation quality

    CN110782439A

  • Annotation detection method and device and computer readable storage medium

    CN111353555A

  • Image labeling method and device and storage medium

    CN113095444A

  • Annotation picture auditing method and device

    CN113344015A