An image data acquisition method and device, electronic equipment and medium
Patent Information
- Application Number
- CN202211005775.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-08-22
AI Technical Summary
[0003]当下通常是由人工对已标注的样本数据进行审核,但因为样本的数量多,往往需要耗费大量的时间成本和人力成本,并且人工在长时间下对大量数据进行审查,会出现疲劳误判的现象,审核的准确度降低
[0018] In summary, according to the image data acquisition method proposed in this disclosure, multiple images to be reviewed are obtained. These images are obtained by manually annotating an initial image. The multiple images to be reviewed are preprocessed to obtain preprocessed images. Preprocessing is used to remove images with failed annotations and/or inconsistent dimensions with the initial image from the multiple images to be reviewed. Then, at least one of the overall image accuracy and local accuracy is calculated on the preprocessed images. Based on at least one of the overall image accuracy and local accuracy, data for training the image annotation model is determined. Compared with existing manual review methods, this solution provides an efficient and low-cost automatic review method for annotated images. It uses at least one review indicator to automatically review manually annotated images, and uses the reviewed data as training data for the image annotation model, improving the accuracy of the training data. Automatic review significantly reduces the amount of manual review, saving manpower while improving the efficiency and accuracy of the review process.
Smart Images

Figure CN117671249B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to an image data acquisition method, apparatus, electronic device, and medium. Background Technology
[0002] Image semantic segmentation algorithms have a wide range of applications, including autonomous driving, image quality enhancement, security, and remote sensing monitoring. Currently, most mainstream image semantic segmentation algorithms employ artificial intelligence (AI) algorithms, which use neural network models to predict and segment semantic information in images. AI-based image semantic segmentation algorithms require a large number of labeled samples for training, and the accuracy of these samples significantly impacts the model's performance; therefore, the accuracy of the sample data needs to be verified.
[0003] Currently, the labeled sample data is usually reviewed manually. However, due to the large number of samples, it often requires a lot of time and manpower. Furthermore, manual review of a large amount of data over a long period of time can lead to fatigue and misjudgment, reducing the accuracy of the review. Summary of the Invention
[0004] This disclosure provides an image data acquisition method, apparatus, electronic device, chip, and medium to solve problems in related technologies, save manpower, and improve the accuracy of review.
[0005] A first aspect of this disclosure provides an image data acquisition method, comprising: acquiring a plurality of images to be reviewed, wherein the images to be reviewed are images obtained by manually annotating an initial image; preprocessing the plurality of images to be reviewed to obtain a preprocessed image, wherein the preprocessing is used to remove images with failed annotations and / or inconsistent with the size of the initial image from the plurality of images to be reviewed; calculating at least one of full-image accuracy and local accuracy on the preprocessed image; and determining data for training an image annotation model based on at least one of full-image accuracy and local accuracy.
[0006] In some embodiments of this disclosure, preprocessing multiple images to be reviewed includes: when a first pixel exists in the image to be reviewed, determining the image to be reviewed as a first image, wherein the label of the first pixel does not belong to a preset label range; and / or when the size of the image to be reviewed is inconsistent with the initial image, determining the image to be reviewed as a second image; and using the image to be reviewed after removing the first image and / or the second image as a preprocessed image.
[0007] In some embodiments of this disclosure, the method further includes: re-annotating the first image and / or the second image; and preprocessing the re-annotated first image and / or the second image.
[0008] In some embodiments of this disclosure, calculating at least one of global accuracy and local accuracy for a preprocessed image includes: obtaining a prediction image, wherein the prediction image is an annotated image obtained by predicting an initial image using an image annotation model, and the prediction image has the same number of pixels as the preprocessed image; and calculating the global accuracy and / or local accuracy of the preprocessed image based on the preprocessed image and the prediction image.
[0009] In some embodiments of this disclosure, calculating the overall image accuracy of the preprocessed image based on the preprocessed image and the predicted image includes: comparing corresponding pixels in the preprocessed image and the predicted image to obtain the number of pixels with the same label in the preprocessed image and the predicted image; and using the ratio between the number of pixels with the same label and the number of pixels in the predicted image or the preprocessed image as the overall image accuracy.
[0010] In some embodiments of this disclosure, calculating the local precision of the preprocessed image based on the preprocessed image and the predicted image includes: obtaining the number of pixels in the preprocessed image and the predicted image that are labeled with a preset label; calculating the intersection-union ratio (IU) of the preprocessed image and the predicted image; and using the IU as the local precision.
[0011] In some embodiments of this disclosure, determining the data for training the image annotation model based on at least one of global precision and local precision includes: determining preprocessed images with global precision greater than or equal to a global precision threshold, and / or preprocessed images with local precision greater than or equal to a local precision threshold, as the data for training the image annotation model.
[0012] In some embodiments of this disclosure, the above method further includes re-annotating preprocessed images with global precision less than a global precision threshold and / or preprocessed images with local precision less than a local precision threshold.
[0013] A second aspect of this disclosure provides an image data acquisition apparatus, comprising: an acquisition module for acquiring multiple images to be reviewed, wherein the images to be reviewed are images obtained by manually annotating an initial image; a preprocessing module for preprocessing the multiple images to be reviewed to obtain a preprocessed image, wherein the preprocessing is used to remove images with failed annotations and / or inconsistent sizes with the initial image from the multiple images to be reviewed; a calculation module for calculating at least one of full-image accuracy and local accuracy on the preprocessed image; and a filtering module for determining data for training an image annotation model based on at least one of full-image accuracy and local accuracy.
[0014] A third aspect of this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described in the first aspect of this disclosure.
[0015] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the first aspect of this disclosure.
[0016] A fourth aspect of this disclosure provides a computer program product including a computer program that is executed by a processor using the methods described in the first aspect of this disclosure.
[0017] A fifth aspect of this disclosure provides a chip including one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory, which, when executed by the processors, cause the electronic device to perform the methods described in the first aspect of this disclosure.
[0018] In summary, according to the image data acquisition method proposed in this disclosure, multiple images to be reviewed are obtained. These images are obtained by manually annotating an initial image. The multiple images to be reviewed are preprocessed to obtain preprocessed images. Preprocessing is used to remove images with failed annotations and / or inconsistent dimensions with the initial image from the multiple images to be reviewed. Then, at least one of the overall image accuracy and local accuracy is calculated on the preprocessed images. Based on at least one of the overall image accuracy and local accuracy, data for training the image annotation model is determined. Compared with existing manual review methods, this solution provides an efficient and low-cost automatic review method for annotated images. It uses at least one review indicator to automatically review manually annotated images, and uses the reviewed data as training data for the image annotation model, improving the accuracy of the training data. Automatic review significantly reduces the amount of manual review, saving manpower while improving the efficiency and accuracy of the review process.
[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0021] Figure 1 A flowchart of an image data acquisition method provided in this embodiment of the disclosure;
[0022] Figure 2 A detailed flowchart of an image data acquisition method provided in this embodiment of the disclosure;
[0023] Figure 3 A schematic diagram of an image data acquisition method provided in an embodiment of this disclosure;
[0024] Figure 4 This is a schematic diagram of a calculation scheme in an image data acquisition method provided in an embodiment of the present disclosure;
[0025] Figure 5 This is a schematic diagram of the structure of an image data acquisition device provided in an embodiment of the present disclosure;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0027] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0028] The development of artificial intelligence (AI) is rapid. Its intelligence lies in simulating human intelligence, mimicking the processes of human recognition, analysis, and judgment. Computer vision, as the "eyes" of machines, is the foundation and core of AI development. Semantic segmentation is a fundamental task in computer vision, enabling the segmentation and recognition of entities in a scene, helping machines to deeply understand their surroundings. Image semantic segmentation algorithms have wide applications in various fields such as autonomous driving, image quality enhancement, security, and remote sensing monitoring. Currently, mainstream image semantic segmentation algorithms are primarily based on deep learning, using neural network models to predict and segment semantic information in images. This requires a large number of labeled samples for training, and the accuracy of these labeled samples greatly affects the model's performance. Therefore, verifying the input data is a crucial aspect of semantic segmentation algorithms.
[0029] Currently, relevant technologies for image annotation mainly include: methods for image annotation and review by annotators, and methods for image annotation review using recognition models. Methods based on annotators are manpower-intensive, and due to the large volume of image annotation data to be reviewed, they often require significant manpower and time. Furthermore, annotators may experience fatigue under continuous work, leading to judgment errors and reduced accuracy. Current model-based recognition methods typically use a single standard to review image annotations, failing to comprehensively address the review needs of multiple categories and scenarios.
[0030] To address the problems existing in related technologies, this disclosure proposes an image data acquisition scheme. This scheme implements batch review of image annotation data based on the design of an algorithm model. After model processing, the number of images requiring manual review will be greatly reduced, saving time and manpower costs and improving review efficiency. Furthermore, this scheme uses at least one of full-image accuracy and local accuracy as review indicators, which broadens the scope of reviewed images and can also be applied to the review of complex images of multiple categories.
[0031] It should be noted that the image data acquisition method of this invention can be applied to an image data acquisition device of this invention, which can be configured on an electronic device. The electronic device can be a personal computer, a personal digital assistant, a mobile terminal, etc., and the mobile terminal can be a mobile phone, tablet computer, remote server, or other hardware device with various operating systems. This application does not limit the specific form of the computer device.
[0032] Figure 1 This is a flowchart illustrating an image data acquisition method provided in an embodiment of this disclosure. Figure 1 As shown, the image processing method includes steps 101-104.
[0033] Step 101: Obtain multiple images to be reviewed. The images to be reviewed are those obtained by manually annotating the initial images.
[0034] The primary purpose of this disclosure is to review manually annotated image data in order to obtain more accurate annotation data for image annotation model training. Therefore, in the embodiments of this disclosure, the image to be reviewed is an image obtained by manually annotating an initial image. The initial image is an image suitable for the corresponding model training scenario and can be obtained through any means, which is not limited in this disclosure.
[0035] Step 102: Preprocess the multiple images to be reviewed to obtain preprocessed images. The preprocessing is used to remove images with failed annotations and / or inconsistent with the initial image size from the multiple images to be reviewed.
[0036] In this disclosure, multiple images awaiting review are first preprocessed. Preprocessing involves filtering multiple images from the same batch to remove those with failed annotations and / or inconsistent dimensions with the initial image. Failed annotations refer to obvious errors in manual annotation. For example, when manually annotating each pixel in an image, the preset image annotation labels are 1-10, each label corresponding to a specific annotation type in the image. For instance, a person in the image is labeled as 1, a car as 2, etc. When undefined labels such as 11, 10a, etc., which are not within the preset label range, or when label format errors occur, it is considered an error in the manual annotation process, and such images are considered failed annotations.
[0037] Furthermore, because the segmentation algorithm needs to ensure that the original image and the labeled image are the same size, and because subsequent accuracy calculations require this, the preprocessing stage also needs to remove images whose size is inconsistent with the initial image. For example, when the size of the manually labeled image is inconsistent with the size of the initial image, such images will also be removed during the preprocessing stage.
[0038] This solution processes multiple images simultaneously, saving time and improving efficiency compared to manual review. This step is only a preliminary screening; this solution can also consider other factors or choose other preprocessing methods. No specific processing method is limited here.
[0039] Step 103: Calculate at least one of the overall image accuracy and local accuracy for the preprocessed image.
[0040] In this disclosure, at least one of the overall image accuracy and local accuracy is calculated for the image to be reviewed, and used as a subsequent review indicator. The overall image accuracy and local accuracy are obtained by comparing the preprocessed image to be reviewed with an image predicted by an image annotation model. In this disclosure, since the purpose is to review manually annotated images, the preprocessed image to be reviewed is used as the comparison object, and the image predicted by the image annotation model is used as the comparison benchmark.
[0041] In other words, the image used for comparison can be obtained through a model with more accurate judgment capabilities. The semantic segmentation result of the input original image is predicted, and the predicted result is used as a more accurate comparison image. In this embodiment, the specific type of image annotation model is not limited.
[0042] It is understandable that the calculation methods for full-image precision and local precision can be the same or different. The full-image precision index takes all the pixels of the image as the calculation object, while the local precision index takes a portion of the pixels of the image as the calculation object.
[0043] Step 104: Determine the data for training the image annotation model based on at least one of full-image accuracy and local accuracy.
[0044] In this disclosure, the overall image accuracy is calculated for the preprocessed image to detect the accuracy of the annotation information in the image to be reviewed, and the local accuracy is calculated for the image to be reviewed to detect the accuracy of the object location region in the image to be reviewed. Therefore, by detecting and filtering the image to be reviewed based on at least one of the overall image accuracy and local accuracy, more accurate image annotation data can be obtained, which can then be used for image annotation model training, ensuring the quality of model training.
[0045] In summary, the image data acquisition method proposed in this disclosure includes acquiring multiple images to be reviewed, which are images obtained by manually annotating an initial image; preprocessing the multiple images to be reviewed to obtain preprocessed images; removing images with failed annotations and / or inconsistent sizes with the initial image from the multiple images to be reviewed; calculating at least one of the overall image accuracy and local accuracy on the preprocessed image; and determining the data for training the image annotation model based on at least one of the overall image accuracy and local accuracy. Applying this scheme will significantly reduce the number of images requiring manual review, avoid the phenomenon of fatigue-induced misjudgment in manual review, improve detection accuracy while saving manpower, and, by using at least one of the overall image accuracy and local accuracy as the review metric, the scope of reviewed images is broader and can also be applied to the review of complex images of multiple categories.
[0046] based on Figure 1 The embodiment shown, Figure 2 A detailed flowchart of an image data acquisition method proposed in this disclosure is further shown. For ease of understanding, Figure 3 A schematic diagram of the overall scheme is shown.
[0047] Step 201: Obtain multiple images to be reviewed. The images to be reviewed are images obtained by manually annotating the initial images.
[0048] Specifically, the objective of this invention is to review manually labeled image data to obtain more accurate labeled data for image labeling model training. In the embodiments of this disclosure, the image to be reviewed is an image obtained by manually labeling an initial image. The initial image is an image suitable for the corresponding model training scenario. It is generally believed that manual labeling may be incorrect, and model training requires more accurate image labeling data as training samples. Therefore, it is necessary to review the manually labeled image data.
[0049] For example, in an image semantic segmentation model applied to autonomous driving, road marking images during vehicle movement are needed as training samples. Unlabeled road images are the original images. Objects that may appear on the road are assigned numbers and labels. Objects are manually identified, their edges are bounded, and corresponding labels are selected for annotation. Pixels within the bounded area are then labeled with the corresponding preset labels. The annotated image serves as the image to be reviewed in this solution.
[0050] Step 202: Preprocess the multiple images to be reviewed to obtain preprocessed images. The preprocessing is used to remove images with failed annotations and / or inconsistent with the initial image size from the multiple images to be reviewed.
[0051] In the embodiments of this disclosure, step 202 includes: when a first pixel exists in the image to be reviewed, the image to be reviewed is determined as a first image, wherein the label of the first pixel does not belong to the preset label range; and / or when the size of the image to be reviewed is inconsistent with the initial image, the image to be reviewed is determined as a second image; and the image to be reviewed after removing the first image and / or the second image is used as a preprocessed image.
[0052] The preprocessing in this scheme involves a simple screening of the image to be processed, detecting specific pixels. When the pixel label does not belong to the preset label range, the image is classified as the first image. When the size of the image to be reviewed is inconsistent with the initial image, the image to be reviewed is classified as the second image. Finally, the image to be reviewed after removing the first image and / or the second image is used as the preprocessed image.
[0053] Among them, if the pixel labels in the image do not belong to the preset label range, it may be due to obvious errors in the labeling information, such as incorrect or illegal labels. For example, when manually labeling each pixel in an image, the preset image labeling labels are 1-10, and each label corresponds to the specific labeling type in the image. For example, a person in the image is labeled as 1, and a car in the image is labeled as 2, etc. When other undefined labels such as 11, 10a, etc., which do not belong to the preset labeling range appear in the labeled image, or when there are label format errors, it is considered that an error has occurred in the manual labeling process, and such images are labeled as failed images.
[0054] The removal of non-compliant second images is to ensure that the image to be reviewed is consistent with the original image. This is because the segmentation algorithm needs to guarantee that the original image and the labeled image are the same size, and subsequent accuracy calculations also require this. Therefore, the preprocessing stage also removes images whose dimensions are inconsistent with the initial image. For example, when the size of the manually labeled image is inconsistent with the size of the initial image, such images will also be removed during the preprocessing stage.
[0055] For example, if a pixel's label doesn't fall within the preset label range—for instance, if the pixel region labeled as "person" in an image to be reviewed should be labeled 1 but is labeled 2—this pixel region is the first pixel. The image to be reviewed containing this first pixel will be identified as the first image. If preprocessing includes this detection, the first image will be discarded during preprocessing. Similarly, if preprocessing includes size detection, images with dimensions inconsistent with the original image will be identified as the second image and discarded.
[0056] In embodiments of this disclosure, step 202 further includes: re-annotating the first image and / or the second image; and preprocessing the re-annotated first image and / or the second image, such as... Figure 3 As shown.
[0057] Specifically, the first and / or second images that were rejected during preprocessing can be re-processed as images awaiting review after being re-annotated.
[0058] In this disclosure, at least one of the overall image accuracy and local accuracy is calculated for the preprocessed image, and detection and screening are performed based on at least one of the overall image accuracy and local accuracy to determine the data for training the image annotation model.
[0059] The following describes in detail the calculation of at least one of the overall image precision and local precision for the preprocessed image, requiring at least one of the following steps 203 and 204:
[0060] Step 203: Calculate the full-image accuracy of the preprocessed image based on the preprocessed image and the predicted image.
[0061] Specifically, step 203 includes: comparing the corresponding pixels of the preprocessed image and the predicted image to obtain the number of pixels with the same label in the preprocessed image and the predicted image; and using the ratio between the number of pixels with the same label and the number of pixels in the predicted image or the preprocessed image as the overall image accuracy.
[0062] In this step, corresponding pixels of the preprocessed image and the predicted image are compared. The predicted image is obtained by using an AI model to predict the semantic segmentation results of the original image. The predicted result is shown below. Figure 4 As shown. The AI model selected in this step is a neural network model with relatively accurate judgment capabilities, trained on a large amount of image data. This solution currently uses the Segformer network, which has stronger learning capabilities and robustness compared to UNet structures, making it more suitable as an annotation and review tool. The formula for calculating the overall image accuracy is:
[0063]
[0064] Where OA represents the overall map precision, and N...T N represents the number of pixels in the preprocessed image that match the labeled values in the predicted image. F This represents the number of pixels whose labels do not match those in the preprocessed image and the predicted image, with N as the denominator. T +N F This refers to the number of pixels in the predicted or preprocessed image. The overall image accuracy calculated in this way can be used to measure the degree of similarity between the preprocessed and predicted images. The closer the preprocessed and predicted images are, the higher the accuracy of their labeled data.
[0065] Step 204: Calculate the local accuracy of the preprocessed image based on the preprocessed image and the predicted image.
[0066] Specifically, step 204 includes: obtaining the number of pixels labeled with preset labels in the preprocessed image and the predicted image respectively; calculating the cross-union ratio (CUP) of the preprocessed image and the predicted image; and using the CUP as the local precision.
[0067] The accuracy of object localization in the preprocessed image can be measured by calculating the intersection-union ratio (IUU) of the corresponding bounded regions of the same object in the preprocessed image and the predicted image.
[0068] For example, in this step, the intersection-over-union ratio (IoU), commonly used in segmentation, is typically chosen to evaluate class accuracy. The formula for IoU is:
[0069]
[0070] Where A and B are two regions to be compared. A represents a bounding box region of an object in the preprocessed image, i.e., a region with the same label, and B represents the corresponding bounding box region of the same object in the predicted image, i.e., a region with the corresponding preset label. The Intersection over Union (IoU) is the ratio between the intersection of A and B and the union of A and B. In this formula, A∩B represents the number of pixels with the preset label in the intersection of A and B, and A∪B represents the number of pixels with the preset label in the union of A and B.
[0071] Step 205: Determine the data for training the image annotation model based on at least one of full-image accuracy and local accuracy.
[0072] Specifically, step 205 includes: determining preprocessed images with global precision greater than or equal to a global precision threshold, and / or preprocessed images with local precision greater than or equal to a local precision threshold, as data for training the image annotation model.
[0073] Specifically, based on the above steps, this scheme filters image annotation data by comparing preprocessed and predicted images, and calculating at least one of the overall image accuracy and local accuracy. Figure 3 As shown, the accuracy calculation refers to the calculation of overall image accuracy and local accuracy. The input image annotation model is trained based on the filtered results.
[0074] When applying this scheme, specific thresholds can be preset, and at least one of the global accuracy and local accuracy can be compared with the corresponding set threshold. Preprocessed images with global accuracy greater than or equal to the global accuracy threshold, and / or preprocessed images with local accuracy greater than or equal to the local accuracy threshold, i.e., images that pass the screening, are determined as data for training the image annotation model, while images that do not meet the conditions are images that do not pass the screening.
[0075] The overall image accuracy threshold is typically set between 0.7 and 0.9. A threshold that achieves the best detection results can be set based on the specific application. A threshold higher than this indicates a better match between the preprocessed and predicted images, higher accuracy of the image annotation data, and therefore the detection result is considered correct. The local image accuracy threshold is typically set to 0.5. A threshold higher than this also indicates a better match between the preprocessed and predicted images, higher accuracy of the image annotation data, and therefore the detection result is considered correct.
[0076] For example, at least one of the global precision and local precision of a preprocessed image is calculated. For instance, in one implementation process, both global and local precision are calculated simultaneously. A pre-set global precision threshold is 0.8, and the calculated value is 0.9 (meaning it's greater than or equal to the global precision threshold). A pre-set local precision threshold is 0.5, and the calculated value is 0.4 (meaning it's less than or equal to the global precision threshold). Because this implementation process simultaneously calculates both global and local precision for screening, the preprocessed image fails the screening and cannot be used as training data for the image annotation model. Using the filtered data as training data for the image annotation model can improve the accuracy of the training data. Figure 3 As shown, the input image annotation model is obtained by filtering the data. Figure 3 The splitter in the middle.
[0077] Specifically, step 205 further includes: re-annotating preprocessed images with global precision less than a global precision threshold, and / or preprocessed images with local precision less than a local precision threshold, such as... Figure 3 As shown.
[0078] In this disclosure, preprocessed images with global precision less than the global precision threshold and / or local precision less than the local precision threshold, i.e., images that fail the screening, cannot be determined as data for training the image annotation model, but can be re-annotated and re-entered into the preprocessing step as images to be processed.
[0079] Using the algorithm model disclosed herein, in a task involving the annotation and review of 4000 images, 3203 images passed the review, and only 797 images required final manual screening. The experiments demonstrate that this solution reduces the manual workload for image data annotation and review by over 80%, while significantly improving detection accuracy and quality.
[0080] In summary, current image annotation data review requires manual labor, consuming significant time and resources. Furthermore, continuous work can lead to fatigue and errors in judgment, reducing the accuracy of the review. This disclosure proposes an algorithm-based model for batch review of image annotation data. After processing by the algorithm, the number of images requiring manual review is significantly reduced, saving time and manpower costs and improving review efficiency. Addressing the issue that current model-based recognition methods typically use a single standard to review image annotations, failing to comprehensively address the review of multiple categories and scenarios, this solution uses at least one of full-image accuracy and local accuracy as the review metric. This broadens the scope of image review and is applicable to the review of complex images across multiple categories. Experimental results demonstrate that this method significantly improves detection accuracy and quality.
[0081] In summary, the solution disclosed herein can achieve the following beneficial effects:
[0082] 1. This disclosure implements batch review of image annotation data based on an algorithm model. After processing by the algorithm model, the number of images requiring manual review will be greatly reduced, saving time and manpower costs and improving review efficiency.
[0083] 2. This disclosure uses at least one of full-image accuracy and local accuracy as the review criteria, which broadens the scope of image review and can also be applied to the review of complex images of multiple categories.
[0084] 3. Experimental verification shows that the method disclosed herein can improve the accuracy and quality of detection.
[0085] Figure 5 This is a schematic diagram of the structure of an image data acquisition device 500 provided in an embodiment of this disclosure. Figure 5 As shown, the image data acquisition device includes:
[0086] The acquisition unit 510 is used to acquire multiple images to be reviewed, which are images obtained by manually annotating the initial images.
[0087] The preprocessing unit 520 is used to preprocess multiple images to be reviewed in order to obtain preprocessed images. The preprocessing is used to remove images with failed annotations and / or inconsistent with the initial image size from the multiple images to be reviewed.
[0088] The calculation unit 530 is used to calculate at least one of the overall image accuracy and local accuracy for the preprocessed image;
[0089] The filtering unit 540 is used to determine the data for training the image annotation model based on at least one of full-image accuracy and local accuracy.
[0090] In some embodiments, the preprocessing unit 520 is specifically used to: when a first pixel exists in the image to be reviewed, determine the image to be reviewed as a first image, wherein the label of the first pixel does not belong to the preset label range; and / or when the size of the image to be reviewed is inconsistent with the initial image, determine the image to be reviewed as a second image; and take the image to be reviewed after removing the first image and / or the second image as the preprocessed image.
[0091] In some embodiments, the preprocessing unit 520 is specifically used to: re-annotate the first image and / or the second image, and then preprocess the re-annotated first image and / or the second image.
[0092] In some embodiments, the calculation unit 530 is specifically used to: acquire a prediction image, wherein the prediction image is an labeled image obtained by predicting the initial image using an image annotation model, and the prediction image has the same number of pixels as the preprocessed image; and calculate the overall accuracy and / or local accuracy of the preprocessed image based on the preprocessed image and the prediction image.
[0093] In some embodiments, the calculation unit 530 is further configured to: compare the corresponding pixels of the preprocessed image and the predicted image to obtain the number of pixels in the preprocessed image and the predicted image that have the same label; and use the ratio between the number of pixels with the same label and the number of pixels in the predicted image or the preprocessed image as the overall image accuracy.
[0094] In some embodiments, the calculation unit 530 is further configured to: obtain the number of pixels labeled with preset labels in the preprocessed image and the predicted image respectively; calculate the intersection-union ratio (IU) of the preprocessed image and the predicted image; and use the IU as the local precision.
[0095] In some embodiments, the filtering unit 540 is specifically used to: determine preprocessed images with global precision greater than or equal to a global precision threshold, and / or preprocessed images with local precision greater than or equal to a local precision threshold, as data for training the image annotation model.
[0096] In some embodiments, the filtering unit 540 is further configured to: re-annotate preprocessed images with global precision less than a global precision threshold, and / or preprocessed images with local precision less than a local precision threshold.
[0097] Corresponding to the methods provided in the above embodiments, this disclosure also provides an image data acquisition device. Since the device provided in this disclosure corresponds to the methods provided in the above embodiments, the implementation of the methods is also applicable to the device provided in this embodiment, and will not be described in detail in this embodiment.
[0098] The methods and apparatus provided in the embodiments of this application have been described above. To implement the functions of the methods provided in the embodiments of this application, the electronic device may include a hardware structure and software modules, and may implement the above functions in the form of a hardware structure, software modules, or a hardware structure plus software modules. One of the above functions may be executed in the form of a hardware structure, software modules, or a hardware structure plus software modules.
[0099] Figure 6 This is a block diagram illustrating an electronic device 600 for implementing the above-described image data acquisition method, according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, computer, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0100] Reference Figure 6 The electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0101] Processing component 602 typically controls the overall operation of electronic device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.
[0102] Memory 604 is configured to store various types of data to support the operation of electronic device 600. Examples of this data include instructions for any application or method operating on electronic device 600, contact data, phonebook data, messages, pictures, videos, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0103] Power supply component 606 provides power to various components of electronic device 600. Power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 600.
[0104] Multimedia component 608 includes a screen that provides an output interface between electronic device 600 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When electronic device 600 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0105] Audio component 610 is configured to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) configured to receive external audio signals when electronic device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.
[0106] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0107] Sensor assembly 614 includes one or more sensors for providing state assessments of various aspects of electronic device 600. For example, sensor assembly 614 may detect the on / off state of electronic device 600, the relative positioning of components such as the display and keypad of electronic device 600, changes in position of electronic device 600 or a component of electronic device 600, the presence or absence of user contact with electronic device 600, orientation or acceleration / deceleration of electronic device 600, and temperature changes of electronic device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0108] Communication component 616 is configured to facilitate wired or wireless communication between electronic device 600 and other devices. Electronic device 600 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio), or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0109] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0110] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of an electronic device 600 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0111] Embodiments of this disclosure also propose a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the image data acquisition method described in the above embodiments of this disclosure.
[0112] Embodiments of this disclosure also provide a computer program product, including a computer program that is executed by a processor using the image data acquisition method described in the above embodiments of this disclosure.
[0113] Embodiments of this disclosure also propose a chip including one or more interface circuits and one or more processors; the interface circuits are used to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory, which, when executed by the processor, cause the electronic device to perform the image data acquisition method described in the above embodiments of this disclosure.
[0114] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0115] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0116] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0117] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (control method), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic device, and portable optical disc read-only memory (CDROM). In addition, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning paper or other media, followed by editing, interpreting or otherwise processing as necessary, and then stored in computer memory.
[0118] It should be understood that various parts of the embodiments of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0119] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0120] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc.
[0121] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for acquiring image data, characterized in that, The method includes: Multiple images to be reviewed are obtained, which are images obtained by manually annotating the initial images. The plurality of images to be reviewed are preprocessed to obtain preprocessed images. The preprocessing is used to remove images that have failed to be labeled and / or whose size is inconsistent with the initial image from the plurality of images to be reviewed. Calculate at least one of full-image precision and local precision for the preprocessed image; The data used for training the image annotation model is determined based on at least one of the overall image accuracy and the local accuracy. Wherein, the calculation of at least one of the overall image accuracy and local accuracy for the preprocessed image includes: Obtain a predicted image, which is an labeled image obtained by predicting the initial image using the image annotation model, and the predicted image has the same number of pixels as the preprocessed image; Based on the preprocessed image and the predicted image, calculate the overall accuracy and / or local accuracy of the preprocessed image; The step of calculating the full-image accuracy of the preprocessed image based on the preprocessed image and the predicted image includes: Compare the corresponding pixels of the preprocessed image and the predicted image to obtain the number of pixels in the preprocessed image and the predicted image that have the same label. The ratio of the number of pixels with the same label to the number of pixels in the predicted image or the preprocessed image is used as the overall image accuracy.
2. The method according to claim 1, characterized in that, The preprocessing of the plurality of images to be reviewed includes: When a first pixel exists in the image to be reviewed, the image to be reviewed is identified as the first image, wherein the label of the first pixel does not belong to the preset label range; and / or When the size of the image to be reviewed is inconsistent with the initial image, the image to be reviewed is determined as the second image; The image to be reviewed after removing the first image and / or the second image is used as the preprocessed image.
3. The method according to claim 2, characterized in that, The method further includes: The first image and / or the second image are re-annotated; the re-annotated first image and / or the second image are then subjected to the aforementioned preprocessing.
4. The method according to claim 1, characterized in that, The step of calculating the local precision of the preprocessed image based on the preprocessed image and the predicted image includes: The number of pixels labeled with preset labels is obtained in the preprocessed image and the predicted image, respectively. Calculate the intersection-over-union ratio (IoU) of the preprocessed image and the predicted image; The intersection-union ratio is used as the local precision.
5. The method according to any one of claims 1 to 4, characterized in that, The step of determining the data for training the image annotation model based on at least one of the overall image accuracy and the local accuracy includes: The preprocessed images with a global precision greater than or equal to the global precision threshold, and / or the preprocessed images with a local precision greater than or equal to the local precision threshold, are determined as data for training the image annotation model.
6. The method according to claim 5, characterized in that, The method further includes: For preprocessed images with overall image precision less than the overall image precision threshold, and / or preprocessed images with local precision less than the local precision threshold, re-annotation is performed.
7. An image data acquisition device, characterized in that, The device includes: The acquisition module is used to acquire multiple images to be reviewed, wherein the images to be reviewed are images obtained by manually annotating the initial images; A preprocessing module is used to preprocess the plurality of images to be reviewed to obtain preprocessed images. The preprocessing is used to remove images that have failed to be labeled and / or whose size is inconsistent with the initial image from the plurality of images to be reviewed. The calculation module is used to calculate at least one of the overall image accuracy and local accuracy for the preprocessed image; A filtering module is used to determine data for training an image annotation model based on at least one of the overall image accuracy and the local accuracy. Specifically, the calculation module is used for: Obtain a predicted image, which is an labeled image obtained by predicting the initial image using the image annotation model, and the predicted image has the same number of pixels as the preprocessed image; Based on the preprocessed image and the predicted image, calculate the overall accuracy and / or local accuracy of the preprocessed image; The computing module is also used for: Compare the corresponding pixels of the preprocessed image and the predicted image to obtain the number of pixels in the preprocessed image and the predicted image that have the same label. The ratio of the number of pixels with the same label to the number of pixels in the predicted image or the preprocessed image is used as the overall image accuracy.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Text image annotation system and method, computer equipment and storage medium
CN111898411A
Image scene recognition method and device based on artificial intelligence and electronic equipment
CN112699855A