Image processing method and electronic device
By performing feature extraction and clustering on the original image, class clusters are generated, and category labels are updated according to manual review and labeling operations, the problem of low efficiency and accuracy of image labeling data review is solved, and efficient and automated image labeling data processing is achieved.
Patent Information
- Application Number
- CN202411822033.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-13
AI Technical Summary
In the field of computer vision, the review efficiency and accuracy of image labeled data are low, making it difficult to meet the needs of large-scale data processing.
By obtaining the original image and its annotation data, cropping out the regional image, performing feature extraction and clustering processing, class clusters are generated, and category labels are updated according to manual review and labeling operations to improve the quality of image annotation data.
It improves the automation of image labeling data review tasks, reduces the workload of manual review, improves the efficiency of processing large amounts of image data, reduces costs, and improves the quality and accuracy of image labeling data.
Smart Images

Figure CN119992035A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and in particular to an image processing method and an electronic device. Background Art
[0002] In the field of computer vision, more and more application scenarios require high-quality image annotation data. In order to ensure the quality of image annotation data, in actual applications, a large number of professionals are often required to conduct multiple rounds of review of image annotation data. The review efficiency and accuracy are low, which makes it difficult to meet the needs of large-scale data processing. Summary of the invention
[0003] Multiple aspects of the present application provide an image processing method and an electronic device to improve the quality of image annotation data and ensure the review accuracy and review efficiency of the image annotation data.
[0004] An embodiment of the present application provides an image processing method, including: obtaining multiple original images and annotation data of each original image, the annotation data including the category label and position information of at least one regional image in the original image; cropping at least one regional image from the original image according to the position information of at least one regional image in the original image; inputting the regional image into a feature extraction network for feature extraction to obtain a feature vector of the regional image; and performing clustering processing based on the feature vectors of the multiple regional images using a clustering algorithm to obtain at least two clusters, each cluster including at least two regional images; in response to a manual review instruction for the cluster, displaying the regional images in the cluster; and in response to a manual annotation operation for an abnormal regional image manually screened out from the cluster, updating the category label of the abnormal regional image so that the abnormal regional image becomes a normal regional image in the cluster.
[0005] An embodiment of the present application also provides an electronic device, including: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to execute the steps in the image processing method.
[0006] In this embodiment, multiple original images and annotation data of each original image are obtained, and the annotation data includes the category label and position information of at least one regional image in the original image; at least one regional image is cropped from the original image according to the position information of at least one regional image in the original image; the regional image is input into the feature extraction network for feature extraction to obtain the feature vector of the regional image; and clustering is performed based on the feature vectors of multiple regional images using a clustering algorithm to obtain at least two clusters, each cluster including at least two regional images; in response to the manual review instruction of the cluster, the regional images in the cluster are displayed; and in response to the manual annotation operation of the abnormal regional image manually screened from the cluster, the category label of the abnormal regional image is updated so that the abnormal regional image becomes a normal regional image in the cluster. Thus, the automation degree of the image annotation data review task is improved, the workload of manually reviewing images is greatly reduced, the efficiency of processing a large amount of image data is improved, the labor cost of image annotation data review is reduced, the quality of image annotation data is improved, and the review accuracy and efficiency of image annotation data are guaranteed; it can also flexibly respond to different image recognition scenarios and requirements, and has good adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0008] Figure 1 A flowchart of an image processing method provided in an embodiment of the present application;
[0009] Figure 2 A flowchart of another image processing method provided in an embodiment of the present application;
[0010] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0011] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0012] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the access relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B may be singular or plural. In the text description of the present application, the character " / " generally indicates that the previous and next associated objects are in an "or" relationship. In addition, in the embodiments of the present application, "first", "second", "third", etc. are only used to distinguish the contents of different objects and have no other special meanings.
[0013] The following specific embodiments are used to describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following is a detailed description of the technical solutions provided by each embodiment of the present application in conjunction with the accompanying drawings.
[0014] Figure 1 A flowchart of an image processing method provided in an embodiment of the present application. Figure 1 , the method may include the following steps:
[0015] 101. Acquire multiple original images and annotation data of each original image, where the annotation data includes a category label and position information of at least one region image in the original image.
[0016] In practical applications, the original image can be annotated manually or automatically. There may be one or more target objects in the original image. If the original image is a road image, the target objects in the road image include, but are not limited to, vehicles, roads, buildings, people, etc. If the original image is an industrial product image, the target objects in the original image include, but are not limited to, various quality defects such as bulges, dents or cracks. It is understandable that the target objects in the original image vary depending on the application scenario, and there is no limitation on this.
[0017] In this embodiment, the regional image in the original image can be understood as the image area where the target object is located in the original image. For example, if the original image is a road image, at least one regional image in the original image includes but is not limited to: a vehicle image, a road image, a building image or a person image.
[0018] In practical applications, when annotating an original image, the category label and location information of at least one regional image in the original image may be determined, but the present invention is not limited thereto. The category label of the regional image represents the category information of the target object in the regional image, and the location information of at least one regional image may refer to the image location information of at least one regional image in the original image.
[0019] 102. Crop at least one regional image from the original image according to position information of at least one regional image in the original image.
[0020] It can be understood that cropping the region image from the original image can reduce interference information in feature extraction of the region image and improve the accuracy of feature extraction of the region image.
[0021] In practical applications, at least one regional image is located in the original image according to the position information of at least one regional image in the original image, and the located at least one regional image is cropped out from the original image.
[0022] In practical applications, a detection frame may be located in the original image according to the position information of the at least one regional image, and the at least one regional image may be cropped from the original image according to the detection frame of the at least one regional image.
[0023] In practical applications, there is no restriction on the shape of the detection frame. Further optionally, according to the position information of at least one regional image in the original image, the implementation method of cutting out at least one regional image from the original image includes: determining the shape information of the detection frame according to application requirements; locating the detection frame of at least one regional image in the original image according to the position information of at least one regional image in the original image and the shape information of the detection frame; and cutting out at least one regional image from the original image according to the detection frame of at least one regional image.
[0024] It can be understood that determining the shape information of the detection frame according to application requirements can make the processed image more suitable for downstream application scenarios and improve the quality of the processed image.
[0025] In actual applications, different application scenarios have different application scenario requirements. Optionally, the implementation method of determining the shape information of the detection frame according to the application requirements includes: if the application requirements are related to the target detection task or the text recognition task, then the shape information of the detection frame is determined to be a rectangle; if the application requirements are related to the semantic segmentation task, then the shape information of the detection frame is determined to be an external graphic. Among them, the external graphic includes, but is not limited to: an external rectangle, an outer diameter polygon, etc., and the shape of the external graphic is related to the shape of the target object in the regional image.
[0026] 103. Input the regional image into the feature extraction network for feature extraction to obtain a feature vector of the regional image.
[0027] In practical applications, the feature extraction network can be any neural network with image feature extraction function, and there is no restriction on this. For example, the feature extraction network is a convolutional neural network trained based on training data sets in different fields.
[0028] Further optionally, in order to capture the key information in the image and improve the richness and accuracy of the feature representation, a classification model can be obtained, and the last fully connected layer in the classification model can be deleted to obtain a feature extraction network; wherein the classification model is trained according to training data sets in different fields, and the classification model is used to extract features of the input image, obtain the feature vector of the image, and use the fully connected layer to classify the image based on the feature vector of the image.
[0029] It is understandable that the classification model trained according to the training data sets in different fields has good generalization performance, robustness and adaptability.
[0030] In practical applications, the classification model can be any residual network with classification processing function. Further optionally, the classification model is a Resnet18 model, which is a residual network for image classification.
[0031] In practical applications, the feature extraction network obtained based on the classification model can extract features from high-dimensional images and obtain a feature vector of 512 dimensions.
[0032] Further optionally, in order to improve the accuracy of feature extraction, the regional image is input into the feature extraction network for feature extraction, and the feature vector of the regional image is obtained by: if the image format of the regional image is not in RGB format, the regional image is converted into a regional image in RGB format; the image size of the regional image in RGB format is adjusted to a preset size; the regional image in RGB format of the preset size is input into the feature extraction network for feature extraction, and the feature vector of the regional image is obtained.
[0033] It can be understood that through format conversion and size adjustment, the consistency and standardization of the input image of the feature extraction network can be ensured, and the quality of feature extraction of the feature extraction network can be improved.
[0034] Specifically, RGB format refers to the red, green, blue (Red, Green, Blue) format, and images in RGB format are color images. In practical applications, the image format of the regional image may not be RGB format, but may be a grayscale image. Therefore, the grayscale image needs to be converted to a color image. Converting the non-RGB regional image to RGB format ensures that all input images are in the same format. The unified image format enables the feature extraction network to process all input images consistently, avoiding processing errors or performance degradation caused by inconsistent formats.
[0035] Specifically, the RGB format regional images are resized to a preset size to ensure that all input images have the same resolution. The standardized image size enables the feature extraction network to process the input images more efficiently, reducing processing errors or performance degradation caused by inconsistent sizes. In actual applications, the preset size can be set as needed without restriction.
[0036] 104. Perform clustering processing based on feature vectors of multiple regional images using a clustering algorithm to obtain at least two clusters, each cluster including at least two regional images.
[0037] Specifically, by using a clustering algorithm to perform clustering based on the feature vectors of multiple region images, similar region images can be classified into the same cluster, thereby achieving classification of the region images. In practical applications, at least two clusters can be stored for subsequent viewing.
[0038] 105. In response to a manual review instruction for the cluster, display the regional image in the cluster; and in response to a manual labeling operation on the abnormal region image manually screened out from the cluster, update the category label of the abnormal region image to make the abnormal region image become a normal region image in the cluster.
[0039] In actual applications, the user can trigger a manual review instruction for any cluster on the interactive interface. In response to the manual review instruction for the cluster, the regional images in the cluster are displayed for the user to view. The user selects the abnormal region images from the multiple region images included in the cluster based on professional knowledge. The multiple region images included in the cluster can be divided into normal region images and abnormal region images. The category labels of each normal region image are the same, and the category labels of the abnormal region image and the normal region image are different. The user can re-label the category label of the abnormal region image, that is, label the category label of the abnormal region image as the category label of the normal region image, thereby making the abnormal region image a normal region image in the cluster, and completing the review of the category labels of each region image in the cluster.
[0040] The technical solution provided in the embodiments of the present application improves the degree of automation of image annotation data review tasks, significantly reduces the workload required for manual image review, improves the efficiency of processing large amounts of image data, reduces the labor cost of image annotation data review, improves the quality of image annotation data, and ensures the accuracy and efficiency of review of image annotation data; it can also flexibly respond to different image recognition scenarios and needs and has good adaptability.
[0041] In some optional embodiments, after clustering is performed based on feature vectors of multiple region images using a clustering algorithm to obtain at least two clusters, it is possible to: identify the abnormal region images in each cluster and add the identified abnormal region images to the abnormal cluster; display the abnormal region images in the abnormal cluster in response to a manual review instruction for the abnormal cluster; and update the category label of the abnormal region image in response to a manual annotation operation for the abnormal region image to make the abnormal region image a normal region image.
[0042] In practical applications, in order to further improve the review accuracy and efficiency of image annotation data, it is also possible to identify abnormal area images in each cluster and add the identified abnormal area images to the abnormal cluster. Subsequent users can specifically review the area images in the abnormal cluster.
[0043] Further optionally, in order to accurately identify the abnormal region image in each cluster, the cluster center of each cluster can be determined; the region image in each cluster whose distance from the cluster center does not meet the preset distance requirement is determined as the abnormal region image in each cluster. The preset distance requirement can be flexibly set as needed.
[0044] Further optionally, in order to accurately identify the abnormal region image in each cluster, it is also possible to: based on each first region image in the multiple first region images included in each cluster, according to the feature vector of the first region image and the feature vector of each second region image, respectively calculate the feature similarity between the first region image and each second region image, the first region image is any one in the cluster, the second region image is any one in the cluster, and the first region image is different from the second region image; according to the feature similarity between each first region image and each second region image, determine the average feature similarity of the cluster; for each first region image, determine the number of abnormal feature similarities that are less than the average feature similarity among the multiple feature similarities corresponding to the first region image; if the number of abnormal feature similarities of the first region image is greater than the preset number threshold, then determine that the first region image is an abnormal region image in the cluster. Wherein, the number threshold is set as needed, for example, 20.
[0045] For example, a cluster includes 100 regional images. For each regional image, the feature similarity between the regional image and the other 99 regional images is calculated. According to the feature similarity between each regional image and the other 99 regional images, the average feature similarity corresponding to the cluster is calculated. For each regional image, the number of feature similarities less than the average feature similarity is determined among the 99 feature similarities corresponding to the regional image. For example, the regional image has 50 feature similarities less than the average feature similarity. Taking the number threshold set to 20 as an example, since the regional image has 50 feature similarities less than the average feature similarity, the regional image is an abnormal regional image.
[0046] In this embodiment, the review of abnormal region images in the abnormal cluster can improve the review accuracy and review efficiency. In practical applications, the user can trigger a manual review instruction for the abnormal cluster on the interactive interface, and in response to the manual review instruction for the abnormal cluster, the abnormal region images in the abnormal cluster are displayed for the user to view. The user can re-label the category label of the abnormal region image, that is, label the category label of the abnormal region image as the category label of the normal region image, thereby making the abnormal region image a normal region image, and completing the review of the category label of each abnormal region image in the abnormal cluster.
[0047] In order to better understand the technical solution of this application, Figure 2 Introduce a scenario implementation example. In various application scenarios such as defect detection, text recognition, and vehicle violation detection, high-quality image annotation data is required. For details, see Figure 2 First, multiple original images and annotation data of the original images are collected. Then, according to the annotation data of the original image, the original image is cropped to crop the original image into multiple regional images, each of which has a category label. Then, each regional image is input into a feature extraction network for feature extraction to obtain a feature vector of each regional image; then, clustering is performed according to the feature vector of each regional image to obtain multiple clusters, and the multiple clusters include cluster 1, cluster 2, and cluster 3. In practical applications, users can correct the category labels of abnormal regional images in cluster 1, cluster 2, and cluster 3. In practical applications, abnormal regional images in cluster 1, cluster 2, and cluster 3 can also be concentrated into abnormal clusters, and users can correct the category labels of abnormal regional images in abnormal clusters.
[0048] It should be noted that the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 to 105 can be device A; for another example, the execution subject of steps 101 and 102 can be device A, and the execution subject of steps 103 to 105 can be device B; and so on.
[0049] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel, and the sequence numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0050] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0051] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device includes: a memory 31 and a processor 32;
[0052] The memory 31 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.
[0053] The memory 31 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0054] The processor 32 is coupled to the memory 31 and is used to execute the computer program in the memory 31, so as to: obtain multiple original images and annotation data of each original image, the annotation data including the category label and position information of at least one regional image in the original image; crop at least one regional image from the original image according to the position information of at least one regional image in the original image; input the regional image into a feature extraction network for feature extraction to obtain a feature vector of the regional image; and perform clustering processing based on the feature vectors of the multiple regional images using a clustering algorithm to obtain at least two clusters, each cluster including at least two regional images; in response to a manual review instruction for the cluster, display the regional images in the cluster; and in response to a manual annotation operation for an abnormal regional image manually screened out from the cluster, update the category label of the abnormal regional image so that the abnormal regional image becomes a normal regional image in the cluster.
[0055] Optionally, the processor 32 is used to: obtain a classification model, and delete the last fully connected layer in the classification model to obtain a feature extraction network; wherein the classification model is trained based on training data sets in different fields, the classification model is used to extract features of the input image, obtain the feature vector of the image, and use the fully connected layer to classify the image based on the feature vector of the image.
[0056] Optionally, the classification model is a Resnet18 model, which is a residual network for image classification.
[0057] Optionally, when the processor 32 inputs the regional image into a feature extraction network for feature extraction and obtains a feature vector of the regional image, it is specifically used for: if the image format of the regional image is not in RGB format, converting the regional image into a regional image in RGB format; adjusting the image size of the regional image in RGB format to a preset size; inputting the regional image in RGB format of a preset size into a feature extraction network for feature extraction to obtain a feature vector of the regional image.
[0058] Optionally, the processor 32 uses a clustering algorithm to perform clustering processing based on the feature vectors of multiple area images, and after obtaining at least two clusters, it is also used to: identify the abnormal area images in each cluster, and add the identified abnormal area images to the abnormal cluster; in response to a manual review instruction for the abnormal cluster, display the abnormal area images in the abnormal cluster; and in response to a manual annotation operation on the abnormal area image, update the category label of the abnormal area image to make the abnormal area image become a normal area image.
[0059] Optionally, when the processor 32 identifies the abnormal area image in each cluster, it is specifically used to: determine the cluster center of each cluster; and determine the area image in each cluster whose distance from the cluster center does not meet the preset distance requirement as the abnormal area image in each cluster.
[0060] Optionally, when the processor 32 identifies the abnormal region image in each cluster, it is specifically used to: based on each first region image among the multiple first region images included in each cluster, according to the feature vector of the first region image and the feature vector of each second region image, respectively calculate the feature similarity between the first region image and each second region image, the first region image is any one in the cluster, the second region image is any one in the cluster, and the first region image is different from the second region image; according to the feature similarity between each first region image and each second region image, determine the average feature similarity of the cluster; for each first region image, determine the number of abnormal feature similarities that are less than the average feature similarity among the multiple feature similarities corresponding to the first region image; if the number of abnormal feature similarities of the first region image is greater than a preset number threshold, determine that the first region image is an abnormal region image in the cluster.
[0061] Optionally, when the processor 32 crops out at least one regional image from the original image based on the position information of at least one regional image in the original image, it is specifically used to: determine the shape information of the detection frame based on application requirements; locate the detection frame of at least one regional image in the original image based on the position information of at least one regional image in the original image and the shape information of the detection frame; and crop out at least one regional image from the original image based on the detection frame of at least one regional image.
[0062] Optionally, when the processor 32 determines the shape information of the detection box according to application requirements, it is specifically used to: if the application requirements are related to a target detection task or a text recognition task, determine the shape information of the detection box as a rectangle; if the application requirements are related to a semantic segmentation task, determine the shape information of the detection box as an external graphic.
[0063] Optional, such as Figure 3 As shown, the electronic device also includes: a communication component 33, a display 34, a power component 35, an audio component 36 and other components. Figure 3 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 3 In addition, Figure 3 The components in the dashed box are optional components, not mandatory components, and the specific components depend on the product form of the electronic device. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT (Internet of Things) device, or a server device such as a conventional server, a cloud server, or a server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it can include Figure 3 If the electronic device of this embodiment is implemented as a conventional server, cloud server or server array or other server-side equipment, it may not include Figure 3 Components within the dashed box.
[0064] For the detailed implementation process of the processor executing each action, please refer to the relevant description in the aforementioned method embodiment or device embodiment, which will not be repeated here.
[0065] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement each step that can be executed by an electronic device in the above method embodiment.
[0066] Accordingly, an embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement each step that can be executed by an electronic device in the above method embodiment.
[0067] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi (Wireless Fidelity), 2G (2 Generation, 2 Generation), 3G (3 Generation, 3 Generation), 4G (4 Generation, 4 Generation) / LTE (long Term Evolution), 5G (5 Generation, 5 Generation) and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (Near Field Communication, NFC) module to facilitate short-range communication. For example, the NFC module can be based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0068] The above-mentioned display includes a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0069] The power supply assembly provides power to various components of the device where the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply assembly is located.
[0070] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal.
[0071] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0072] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0073] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0075] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interface, network interface and memory.
[0076] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0077] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, Phase Change RAM (PRAM), Static Random-Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random Access Memory (RAM), Read Only Memory (ROM), Electrically-Erasable Programmable Read-Only Memory (EEPROM), flash memory or other memory technology, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this article, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0078] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0079] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. An image processing method, characterized in that: include: Acquire multiple original images and annotation data of each original image, wherein the annotation data includes a category label and position information of at least one region image in the original image; According to the position information of at least one regional image in the original image, cutting out at least one regional image from the original image; Inputting the regional image into a feature extraction network for feature extraction to obtain a feature vector of the regional image; and performing clustering processing based on the feature vectors of multiple regional images using a clustering algorithm to obtain at least two clusters, each cluster including at least two regional images; In response to a manual review instruction for the cluster, a regional image in the cluster is displayed; and in response to a manual annotation operation on an abnormal region image manually screened out from the cluster, a category label of the abnormal region image is updated to make the abnormal region image become a normal region image in the cluster.
2. The method according to claim 1, characterized in that The method of acquiring the feature extraction network includes: Obtaining a classification model, and deleting the last fully connected layer in the classification model to obtain the feature extraction network; The classification model is trained based on training data sets in different fields, and is used to extract features from an input image, obtain a feature vector of the image, and classify the image based on the feature vector of the image using the fully connected layer.
3. The method according to claim 2, characterized in that The classification model is a Resnet18 model, which is a residual network used for image classification.
4. The method according to claim 2, characterized in that: Inputting the region image into a feature extraction network for feature extraction to obtain a feature vector of the region image, including: If the image format of the region image is not in RGB format, converting the region image into a region image in RGB format; Adjust the image size of the area image in RGB format to a preset size; The regional image in RGB format of a preset size is input into a feature extraction network for feature extraction to obtain a feature vector of the regional image.
5. The method according to claim 1, characterized in that: After performing clustering processing based on the feature vectors of the multiple regional images using a clustering algorithm to obtain at least two clusters, the method further includes: Identify the abnormal region images in each cluster, and add the identified abnormal region images to the abnormal cluster; In response to a manual review instruction for the abnormal cluster, an abnormal region image in the abnormal cluster is displayed; and in response to a manual annotation operation for the abnormal region image, a category label of the abnormal region image is updated to make the abnormal region image become a normal region image.
6. The method according to claim 5, characterized in that Identifying abnormal region images in each cluster includes: Determine the cluster center of each cluster; The regional images in each cluster whose distance from the cluster center does not meet the preset distance requirement are determined as abnormal regional images in each cluster.
7. The method according to claim 5, characterized in that Identifying abnormal region images in each cluster includes: Based on each first region image of a plurality of first region images included in each cluster, according to a feature vector of the first region image and a feature vector of each second region image, respectively calculating a feature similarity between the first region image and each second region image, the first region image is any one of the clusters, the second region image is any one of the clusters, and the first region image is different from the second region image; Determining an average feature similarity of the clusters according to the feature similarity between each first region image and each second region image; For each first region image, determining the number of abnormal feature similarities that are smaller than the average feature similarity among a plurality of feature similarities corresponding to the first region image; If the number of abnormal feature similarities of the first region image is greater than a preset number threshold, the first region image is determined to be an abnormal region image in the cluster.
8. The method according to claim 1, characterized in that: According to the position information of at least one regional image in the original image, cutting out at least one regional image from the original image; Determine the shape information of the detection frame according to application requirements; Locating a detection frame of the at least one regional image in the original image according to position information of the at least one regional image in the original image and shape information of the detection frame; At least one regional image is cropped from the original image according to the detection frame of the at least one regional image.
9. The method according to claim 8, characterized in that Determine the shape information of the detection frame according to application requirements, including: If the application requirement is related to an object detection task or a text recognition task, determining that the shape information of the detection frame is a rectangle; If the application requirement is related to a semantic segmentation task, the shape information of the detection box is determined to be an external graphic.
10. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer programs; The processor is coupled to the memory and configured to execute the computer program to perform the steps of the method according to any one of claims 1 to 9.