Use case collection method, apparatus, device, medium, and program product
By automating the collection of use cases for multi-label image classification algorithms, the problems of misjudgment and incorrect judgment in existing technologies have been solved, improving the efficiency and accuracy of algorithm optimization and ensuring safety in fields such as autonomous driving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2026-03-17
AI Technical Summary
Existing multi-label image classification algorithms suffer from misjudgments and incorrect judgments in fields such as autonomous driving, which affects safety. The algorithms need to be optimized, but there is a lack of effective use case collection methods.
A method for collecting use cases is provided. By inputting labeled images into the algorithm to be evaluated, images with inconsistent labels are obtained. Second images with the same content are used for label verification. Use cases for algorithm optimization are collected, including recognition module, acquisition module and generation module, and optimization use cases are automatically generated.
It improves the efficiency and accuracy of algorithm optimization, reduces the investment of human resources, provides more use cases that are conducive to algorithm optimization, and ensures the accuracy of multi-label recognition.
Smart Images

Figure CN115171107B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a use case collection method, apparatus, device, medium, and program product. Background Technology
[0002] With the advent of the mobile internet era, the amount of image data has exploded exponentially. To facilitate image classification, applications for image classification have been developed.
[0003] Generally, an image contains objects of more than one class, meaning image classification tasks are often multi-label image classification tasks. As applications for multi-label image classification become increasingly prevalent in daily life, the accuracy of these applications is becoming increasingly important. For example, in the field of autonomous driving, the classification results of applications on objects in images directly impact the safety of autonomous driving.
[0004] Algorithms applied to multi-label image classification will always have a certain probability of misjudging and incorrectly classifying targets in images. Therefore, many use cases are needed to optimize the algorithm. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and program product for collecting use cases. The technical solution is as follows:
[0006] According to one aspect of this application, a use case collection method is provided, the method comprising:
[0007] The first image with labels is input into the algorithm to be evaluated, and the algorithm outputs the first set of labels for the first image.
[0008] In response to the existence of N tags in the first tag set that are different from the tags in the labeled tag set of the first image, a second image with the same content as the first image is obtained;
[0009] The second image is input into the algorithm to be evaluated, and the algorithm outputs a second tag set for the second image.
[0010] In response to the second tag set containing at least one of the N tags, the second image is added to the use case set for algorithm optimization, where N is a positive integer.
[0011] According to another aspect of this application, a use case collection apparatus is provided, the apparatus comprising:
[0012] The recognition module is used to input the first image with labels into the algorithm to be evaluated, and the algorithm to be evaluated outputs the first set of labels for the first image.
[0013] The acquisition module is used to acquire a second image with the same content as the first image in response to the presence of N tags in the first tag set that are different from the tags in the labeled tag set of the first image;
[0014] The recognition module is used to input the second image into the algorithm to be evaluated, and the algorithm to be evaluated outputs a second tag set of the second image;
[0015] A generation module is configured to add the second image to a use case set for algorithm optimization, where N is a positive integer, in response to the second tag set containing at least one of the N tags.
[0016] According to another aspect of this application, a computer device is provided, the computer device including a processor and a memory connected to the processor, the memory storing program instructions, wherein the processor executes the program instructions to implement the use case collection method as provided in various aspects of this application.
[0017] According to another aspect of this application, a computer-readable storage medium is provided, wherein program instructions are stored therein, which, when executed by a processor, implement the use case collection method as provided in various aspects of this application.
[0018] According to another aspect of this application, a computer program product (or computer program) is provided, the computer program product (or computer program) including computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, causing the computer device to perform the methods provided in various alternative implementations of the above-described use case collection method.
[0019] According to another aspect of this application, a chip is provided that includes programmable logic circuitry and / or program instructions, which, when the chip is running, are used to implement the use case collection methods provided in various aspects of this application.
[0020] The beneficial effects of the technical solutions provided in this application embodiment may include:
[0021] In the above-described use case collection method, after the algorithm to be evaluated identifies the first image, a first tag set is obtained. If there are N tags in the first tag set that are different from the tags in the already labeled tag set of the first image, it can be considered that the algorithm to be evaluated has made a mistake in identifying the tags of the first image. In this case, the terminal obtains a second image with the same content as the first image, and the algorithm to be evaluated identifies the second image to obtain a second tag set. If there is at least one of the N tags in the second tag set, it is determined that the algorithm to be evaluated has also made a mistake in identifying at least one tag of the second image. Then, the second image is added to the use case set. Subsequently, for the accurate identification of the N tags, more use cases that are conducive to the optimization of the algorithm to be evaluated are provided. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A schematic diagram of a communication system provided in an exemplary embodiment of this application is shown;
[0024] Figure 2 A flowchart illustrating a use case collection method provided in an exemplary embodiment of this application is shown;
[0025] Figure 3 This illustration shows a schematic diagram of a first image provided in an exemplary embodiment of this application;
[0026] Figure 4 A flowchart of a use case collection method provided by another exemplary embodiment of this application is shown;
[0027] Figure 5 A flowchart of a use case collection method provided by another exemplary embodiment of this application is shown;
[0028] Figure 6 A flowchart of a use case collection method provided by another exemplary embodiment of this application is shown;
[0029] Figure 7 This illustration shows a feature analysis diagram of a use case provided by an exemplary embodiment of this application;
[0030] Figure 8 A block diagram of a use case collection apparatus provided in an exemplary embodiment of this application is shown;
[0031] Figure 9A schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0033] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0034] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0035] With the increasing prevalence of multi-label image classification applications in daily life, the accuracy of these applications is becoming increasingly important. For example, in the field of autonomous driving, the classification results of objects in images directly affect the safety of autonomous driving. Therefore, before setting up multi-label recognition algorithms in applications, it is necessary to test and optimize the algorithms to ensure their accuracy in multi-label recognition. The optimization process requires numerous use cases. This application proposes a use case collection method that can collect a sufficient number of targeted use cases to effectively optimize the algorithm to be evaluated. Detailed implementation steps are provided in the following embodiments.
[0036] The use case collection method provided in this application can be applied to computer systems. For example... Figure 1 The diagram shown is a block diagram of a computer system provided in an exemplary embodiment of this application. The computer system includes a computer device 120, a communication network 140, and a server 160.
[0037] Computer device 120 contains an operating system and applications. This operating system provides a runtime environment for testing the multi-label classifier. For example, the operating system could be Android, iOS, HermeticOS, Windows, or Linux.
[0038] For example, a multi-label classifier is used to classify images with multiple labels. For instance, an image containing "cat" and "dog" can be classified using a multi-label classifier, determining that the image's labels include "cat" and "dog." Multi-label classifiers can be applied to various applications to support multi-label classification tasks within those applications. For example, such applications may include instant messaging applications, news and information applications, lifestyle service applications, financial service applications, etc.
[0039] For example, the computer device 120 may include at least one of a smartphone, tablet computer, laptop computer, desktop computer, or notebook computer.
[0040] A communication network 140 exists between computer device 120 and server 160. This communication network 140 can be a wired network or a wireless network. For example, a wired network can be a metropolitan area network, a local area network, a fiber optic network, etc.; a wireless network can be a mobile communication network, a wireless Fidelity network (WiFi), etc.
[0041] Server 160 provides background services for the operation of the operating system and / or applications in computer device 120. For example, server 160 provides background services for the testing and optimization of a multi-label classifier in computer device 120. Server 160 may include at least one of the following: a single server 160, multiple servers 160, a cloud computing platform, and a virtualization center.
[0042] For example, the use case collection method provided in this application can be completed independently by the computer device 120; for instance, the data required for testing the multi-label classifier is stored in the local memory of the computer device 120, and the computer device 120 can independently test the multi-label classifier and collect use cases for algorithm optimization based on the test results.
[0043] Alternatively, the use case collection method provided in this application can be completed collaboratively by computer device 120 and server 160. For example, when executing the above use case collection method, computer device 120 needs to obtain the data required for testing the multi-label classifier from server 160. For instance, server 160 includes data storage server 1602 and search engine server 1604. Computer device 120 can obtain test samples from data storage server 1602 via communication network 140, use the test samples to test the multi-label classifier, and obtain test results. If the test results indicate that the labels of the test samples identified by the multi-label classifier are incorrect, use cases are obtained from search engine server 1604 based on the test samples, and the use cases that meet the requirements are stored in data storage server 1602 for later use. Computer device 120 then optimizes the multi-label classifier based on the above use cases.
[0044] Those skilled in the art will understand that the number of computer devices 120 in the above-described computer system can be more or less. For example, the computer device 120 in the above-described computer system can be only one, or dozens or hundreds, or more. The embodiments of this application do not limit the number and type of computer devices 120 in the computer system.
[0045] Figure 2 A flowchart of a use case collection method provided in an exemplary embodiment of this application is shown, which can be applied to Figure 1 In the computer device shown, the method includes:
[0046] Step 202: Input the first image with labels into the algorithm to be evaluated, and the algorithm to be evaluated will output the first label set of the first image.
[0047] For example, a first image labeled with a label may be stored in the memory of a computer device, or in a data storage server. When collecting use cases, the computer device retrieves the first image from the memory, or retrieves the first image from the data storage server via a communication network.
[0048] For example, the algorithm to be evaluated includes a multi-label classifier. This multi-label classifier is composed of at least one of a neural network, the K-nearest neighbor algorithm, a decision tree, a kernel method for vector output, and a boosting algorithm. The types of multi-label classifiers are not limited in the embodiments of this application. The algorithm to be evaluated is stored in the memory of a computer device, or in a data storage server. When collecting use cases, the computer device also retrieves the algorithm to be evaluated from the memory, or retrieves it from the data storage server via a communication network.
[0049] The computer device inputs the first image into the algorithm to be evaluated. The algorithm performs label recognition on the first image to obtain a first label set for the first image. The first label set includes at least one first label, that is, the first label set includes one, two, or more labels. For example, Figure 3 The images contained images of cats and dogs. The computer device input these images into the algorithm to be evaluated and identified the image's label set {cat, rabbit}.
[0050] Step 204: In response to the existence of N tags in the first tag set that are different from the tags in the labeled tag set of the first image, obtain the second image with the same content as the first image.
[0051] The first image carries a set of labeled tags, which includes at least one labeled tag of the first image. These labeled tags are the correct labels for the first image; for example, the labeled tags of the first image are obtained through manual annotation.
[0052] For example, the computer device compares the first tag set with the tags in the labeled tag set, and filters the tags from the first tag set that are different from the tags in the labeled tag set; in response to the existence of N tags in the first tag set that are different from the tags in the labeled tag set, it determines that the tags of the first image identified by the algorithm to be evaluated are incorrect, and obtains a second image with the same content as the first image.
[0053] For example, computer equipment uses the algorithm to be evaluated to... Figure 3 After labeling the images, a first label set {cat, rabbit} is obtained; the labeled label set for the first image is {cat, dog}; the computer device determines that the label "rabbit" in the first label set is different from all the labels in the labeled label set, and then obtains a second image with the same content as the first image. For example, "same content" means that the elements in the images have the same features; for instance, the first and second images have the same content if both the first and second images contain elements of dogs and cats; or, for example, the images of cats and dogs are displayed side-by-side in both the first and second images.
[0054] Optionally, the number of tags in the first tag set may be equal to or different from the number of tags in the labeled tag set.
[0055] For example, after determining N labels, the computer device also labels the first image with N more labels.
[0056] Step 206: Input the second image into the algorithm to be evaluated, and the algorithm to be evaluated will output the second label set of the second image.
[0057] The computer device inputs the second image into the algorithm to be evaluated, and identifies a second tag set for the second image. The second tag set includes at least one second tag. Optionally, the number of tags in the second tag set may be equal to or different from the number of tags in the first tag set.
[0058] Step 208: In response to the second tag set containing at least one of N tags, add the second image to the use case set for algorithm optimization, where N is a positive integer.
[0059] The computer device iterates through the second tags in the second tag set and determines that the second tag set contains at least one of the N tags. This indicates that the algorithm to be evaluated has made a recognition error in the tag recognition of the second image, and this recognition error is partially or completely the same as the recognition error of the N tags in the first image. Therefore, the computer device adds the second image to the use case set for algorithm optimization for the recognition of the N tags.
[0060] Optionally, in response to the second tag set containing at least one of N tags, the computer device adds the second image to the use case set corresponding to the at least one of the included tags.
[0061] For example, N tags can be divided into at least one set of error tags, each set of error tags including at least one error tag from the N tags, and the number of tags in different error tag sets may be the same or different; in response to the second tag set containing at least one error tag from the N tags, the computer device adds the second image to the use case set corresponding to the error tag set consisting of at least one error tag.
[0062] In other words, different error labels and combinations of error labels correspond to their own use case sets. Based on the different error labels identified in the second images, multiple second images are added to their respective use case sets, thus allowing for targeted optimization of the evaluation algorithm according to different error labels. For example, if N labels include "rabbit" and "tiger," and second image 1 contains the error label "rabbit," then second image 1 is added to the use case set corresponding to "rabbit"; second image 2 contains the error label "tiger," then second image 2 is added to the use case set corresponding to "tiger"; and second image 3 contains both the error labels "rabbit" and "tiger," then second image 3 is added to the use case sets corresponding to both "rabbit" and "tiger."
[0063] Optionally, in response to the second tag set having the same tags as the first tag set, the computer device adds the second image to the use case set corresponding to the N tags. Theoretically, the second image in the use case set corresponding to the N tags should have the same error tag as the first image.
[0064] For example, the first image may include at least two images. Different first images may share a common use case set, or each first image may have its own corresponding use case set. For instance, a computer device may add a second image to the use case set corresponding to the first image.
[0065] In summary, the use case collection method provided in this embodiment obtains a first tag set after the algorithm to be evaluated identifies the first image. If there are N tags in the first tag set that are different from the tags in the already labeled tag set of the first image, it can be considered that the algorithm to be evaluated has made a mistake in identifying the tags of the first image. In this case, the terminal obtains a second image with the same content as the first image, and identifies the second image through the algorithm to be evaluated to obtain a second tag set. If there is at least one of the N tags in the second tag set, it is determined that the algorithm to be evaluated has also made a mistake in identifying at least one tag of the second image. Then, the second image is added to the use case set. Subsequently, for the accurate identification of the N tags, more use cases that are conducive to the optimization of the algorithm to be evaluated are provided.
[0066] In some embodiments, the algorithm to be evaluated can also be tested using the first image, so that while the algorithm to be evaluated is being tested, the use case collection method provided in this embodiment is executed to complete the collection of use cases for algorithm optimization.
[0067] For example, the second image mentioned above can be obtained by a computer device searching the internet, as shown in the following manner. Figure 4 As shown in the Chinese embodiment. Figure 4 A flowchart of a use case collection method provided in an exemplary embodiment of this application is shown, which can be applied to Figure 1 In the computer device shown, the method includes:
[0068] Step 202: Input the first image with labels into the algorithm to be evaluated, and the algorithm to be evaluated will output the first label set of the first image.
[0069] Step 304: In response to the fact that N tags in the first tag set are different from the tags in the already labeled tag set, the first image is used as input data for web search to obtain the second image.
[0070] The computer device iterates through the first tags in the first tag set and determines that there are N tags in the first tag set that are different from the tags in the already labeled tag set. The computer then calls the network search engine, uses the first image as input data for the network search, and retrieves at least two second images.
[0071] Optionally, the computer device uses the first image as input data for a web search, retrieves G candidate images, where G is a positive integer greater than 1; and selects the first K candidate images from the G candidate images as the second image, where K is a positive integer less than or equal to G. For example, the computer device uses the first image as a search keyword, retrieves G candidate images arranged in sequence, and extracts the first K candidate images from these G candidate images as the second image. The value of K can be an empirical value set by the user. For example, K can be 50; the computer device selects the first 30 candidate images from the 50 candidate images as the second image.
[0072] In some embodiments, the computer device can also search for a second image with the same content as the first image based on tags in a set of labeled tags for the first image. For example, the computer device concatenates the tags in the set of labeled tags to generate search keywords; the computer device calls a web search engine to search for the search keywords, obtaining G candidate images; and extracts the top K candidate images from the G candidate images as the second image. For example, the set of labeled tags includes the tags {cat, dog, rabbit, snake}, the computer device concatenates cat, dog, rabbit, and snake to obtain the search keyword "cat and dog and rabbit and snake"; the search keyword "cat and dog and rabbit and snake" is input into a web search engine to obtain a second image containing images of cats, dogs, rabbits, and snakes.
[0073] For example, the aforementioned at least two second images may have been found by searching at least two web search engines.
[0074] Step 206: Input the second image into the algorithm to be evaluated, and the algorithm to be evaluated will output the second label set of the second image.
[0075] Step 208: In response to the second tag set containing at least one of N tags, add the second image to the use case set for algorithm optimization, where N is a positive integer.
[0076] In summary, the use case collection method provided in this embodiment can obtain a second image with the same content as the first image when the algorithm to be evaluated misidentifies the first image; and collect the second image for algorithm optimization when the algorithm to be evaluated also misidentifies the second image. This generates optimization use cases for the algorithm to be evaluated without manual intervention, saving manpower and improving the efficiency of generating optimization use cases. Secondly, the method can also select the top K candidate images as the second images after obtaining multiple candidate images. This is because the multiple candidate images provided by the search engine are arranged in descending order of similarity to the first image, thus ensuring that the elements included in the top K candidate images have a higher similarity to the elements in the first image.
[0077] Since the analysis of algorithm optimization direction requires many use cases for support, the use case collection method described in the above embodiments can be used multiple times to collect use cases during the use case collection process. For example, such as... Figure 5 The diagram illustrates a flowchart of a use case collection method provided in an exemplary embodiment of this application, which can be applied to... Figure 1 In the computer device shown, the method includes:
[0078] Step 202: Input the first image with labels into the algorithm to be evaluated, and the algorithm to be evaluated will output the first label set of the first image.
[0079] Step 204: In response to the existence of N tags in the first tag set that are different from the tags in the labeled tag set of the first image, obtain the second image with the same content as the first image.
[0080] For example, in this embodiment, step 204 can also be implemented by step 304.
[0081] Step 206: Input the second image (or the third image) into the algorithm to be evaluated, and the algorithm to be evaluated will output the second label set (or the third label set) of the second image (or the third image).
[0082] Step 208: In response to the second tag set (or the third tag set) containing at least one of N tags, add the second image (or the third image) to the use case set for algorithm optimization, where N is a positive integer.
[0083] Step 410: Count the number of use cases in the use case set.
[0084] For example, after acquiring multiple (i.e. at least two) second images, the computer device performs steps 206 to 208 as described above on the multiple second images. After traversing the multiple second images, it counts the number of use cases in the use case set corresponding to the first image.
[0085] When a first image corresponds to multiple use case sets, the computer device counts the total number of use cases across all use case sets. When multiple first images share a single use case set, the use cases for each first image are stored in association with the image's identifier. Based on the label of the current first image, the number of use cases corresponding to the current first image in the use case set is counted.
[0086] Step 412: In response to the number of use cases being less than the number threshold, obtain a third image with the same content as the first image. The third image is different from the second image.
[0087] For example, the above quantity threshold is preset, and the above quantity threshold is an empirical value.
[0088] Optionally, in response to the number of use cases being less than a threshold, the computer device uses the first image as input data for a web search and obtains the third image.
[0089] For example, if the third image and the second image are obtained using the same web search engine, the computer device uses the first image as input data for the web search and obtains G candidate images; the (K+1)th to the 2Kth images among the G candidate images are determined as the third image, where 2K is a positive integer less than or equal to G.
[0090] If the third image and the second image are obtained using different web search engines, the computer device uses the first image as input data for the web search and obtains G candidate images; the first K images among the G candidate images are determined as the third image.
[0091] Optionally, in response to the number of use cases being less than a threshold, the computer device performs augmentation processing on the first image to obtain an augmented image; the augmented image is then used as input data for a web search to obtain a third image. For example, the computer device may perform augmentation processing on the first image in at least one of the following ways: flipping, such as flipping the image vertically and / or horizontally; changing the colors, such as altering the image's brightness, contrast, saturation, and hue.
[0092] Then, returning to step 206, the third image can be regarded as the second image of another batch.
[0093] In summary, the use case collection method provided in this embodiment can continuously acquire optimized use cases, thereby providing sufficient use cases for the subsequent optimization process of the algorithm to be evaluated.
[0094] In other embodiments, when the number of use cases is sufficient, the computer device can also analyze the collected use cases to obtain reference information about the optimization direction of the algorithm to be evaluated, for example, such as... Figure 5 The method also includes the following steps:
[0095] Step 414: In response to the number of use cases being greater than or equal to the number threshold, analyze the similarity features between the use cases in the use case set and the first image.
[0096] The similarity features between the above use case and the first image include at least one of color similarity features, texture similarity features, and spatial similarity features.
[0097] For example, the computer device calculates the similarity of at least two colors between the use case and the first image in at least two color spaces; based on the weights configured in the at least two color spaces, it calculates a weighted value of the at least two color similarities to obtain color similarity features.
[0098] For example, the computer device uses a sliding window to segment the test case into blocks, obtaining M first block images; and uses a sliding window to segment the first image into blocks, obtaining M second block images; wherein, there is a one-to-one correspondence between the M first block images and the M second block images, and M is a positive integer greater than 1; the structural similarity between the corresponding first block images and the second block images is calculated to obtain M structural similarities; the average of the M structural similarities is calculated to obtain the average structural similarity between the test case and the first image, and the average structural similarity is used as a texture similarity feature.
[0099] For example, the computer device determines a first spatial relationship between targets included in the use case; and determines a second spatial relationship between targets included in the use case; and determines the spatial similarity between the use case and the first image based on the similarity between the first spatial relationship and the second spatial relationship, and uses the spatial similarity as a spatial similarity feature.
[0100] The computer device can also summarize the similar characteristics of multiple use cases in each use case set, thereby providing reference data for the optimization direction of the algorithm to be evaluated.
[0101] In summary, the use case analysis method provided in this embodiment can provide reference data for the optimization direction of the algorithm to be evaluated by analyzing the similarity features between use cases with error labels and the first image; in particular, when use cases are collected separately for different error labels, targeted reference data can be provided for the optimization direction of the algorithm to be evaluated.
[0102] In the embodiments of this application, a large number of labeled image resources (first images) are input into the algorithm to be evaluated. After the labeled images are processed by the algorithm, the corresponding labels of the labeled images are output. After comparing the output labels with the labeled labels, image resources with incorrect labels are identified, and these image resources are marked as bad cases and the incorrect labels are recorded. Then, the bad case resources are searched in a search engine to obtain a certain number of second images. The searched second images are fed back into the algorithm to be evaluated. If the same incorrect label is generated, they are archived into the same type of bad case. When a certain number of bad cases are accumulated, feature analysis of batch bad cases is performed to quickly determine the similar features of the images and indicate the direction of algorithm optimization; or the batch of bad cases can be used as material to supplement the training set to accelerate the algorithm iteration and optimization.
[0103] like Figure 6 The main implementation of this solution is as follows:
[0104] S1, Image Resources: Prepare the labeled test set resources to form the labeled result set. The URL is a Uniform Resource Locator.
[0105] Image Name Tag 1 Tag 2 …… Tag z Storage path Image 1 cat people …… Trees URL: / / xx Image 2 dog sheep …… grassland URL: / / yy ……
[0106] S2 inputs the test set resources into the algorithm to be evaluated. After processing by the algorithm, it outputs the labels predicted by the algorithm to be evaluated.
[0107] Image Name Tag 1 Tag 2 …… Tag z Image 1 dog people …… Trees Image 2 dog sheep …… grassland ……
[0108] S3, Label Error: Compare the output labels in S2 with the labels in S1 based on the image name.
[0109] S3.1 If the labels of the images are completely identical, the prediction is correct and no further action is taken.
[0110] S3.2 If the image label is incorrect, it is a bad case. For example, if the cat in image 1 is predicted as a dog, record the incorrect label (an image may have multiple incorrect labels, but this example only shows one incorrect label, z is a positive integer greater than 2).
[0111] Image Name Error label 1 Error label 2 …… Error label z Image 1 dog ……
[0112] S4. In a search engine (multiple search engines can be used simultaneously), use the image with the incorrect tag as the image source, search for images, and get the first K images in the returned results (in this example, image 1 can be used as the search source image to get similar images).
[0113] S5. The acquired K images are fed back into the algorithm to be evaluated, and the error labels of the output results of the K images are retrieved.
[0114] S5.1 If there is no label that matches the error label, it is not a bad case and no action is taken.
[0115] S5.2 If a label that matches the error label exists, it is considered a bad case of the same type (e.g., in this case, if the output label "dog" exists in the returned image, it is considered a bad case of the same type).
[0116] S6 archives bad cases, summarizing similar bad cases found in multiple engines for the same image into a single table.
[0117]
[0118]
[0119] S7. Once the badcase resources have accumulated to a predetermined number, perform common feature analysis on the badcase set. If the number is insufficient, continue image search to obtain images from the (K+1)th to the 2Kth images and perform steps S4 to S6. Repeat this process until sufficient resources are obtained (the loop can end after obtaining 0 similar badcases in two consecutive loops), and then extract badcase features.
[0120] S7.1, Feature Analysis: Input the badcase set into the algorithm to be evaluated, and extract features such as color similarity, texture similarity, shape similarity, spatial similarity, and semantic similarity from the image containing the misidentified object, based on the algorithm's prediction logic. Figure 7 As shown.
[0121] S7.1.1 Calculate the color feature similarity score to obtain the color similarity features.
[0122] The computer device calculates the color similarity features of at least two colors between the bad example (i.e., the use case) and the first image in at least two color spaces. Based on the weights configured in the at least two color spaces, a weighted value is calculated for the color similarity. For example, the color spaces include at least two of the RGB, HSV, and Lab color spaces. Here, R represents red, G represents green, and B represents blue; H represents hue, S represents saturation, V represents value, and L represents lightness. A positive a value represents red, a negative a value represents green, a positive b value represents yellow, and a negative b value represents blue.
[0123] For example, the color feature similarity score S-color can be calculated using the following formula:
[0124] S-color = W rgb ×S rgb +W hsv ×S hsv +W lab ×S lab ;
[0125] Among them, W rgb : RGB color space similarity weighting coefficient;
[0126] S rgb : RGB color space similarity, calculating the probability value of histogram similarity;
[0127] W hsv HSV color space similarity weighting coefficient;
[0128] S hsvHSV color space similarity, calculating the probability value of histogram similarity;
[0129] W lab :Lab color space similarity weighting coefficient;
[0130] S lab Lab color space similarity, calculating the probability value of histogram similarity;
[0131] S-color: Represents the weighted sum of the probabilities of similarity in multiple color spaces. Images with scores TOPP (i.e., the top p) or scores that meet a certain threshold Ts-color can continue to be processed in subsequent stages.
[0132] Weight setting instructions:
[0133] Si={S rgb S hsv S lab}, S max =max(S) rgb S hsv S lab ), Wi = {W rgb W hsv W lab}
[0134] A) Weights can be user-defined, Wi = weight ratio / sum of weight ratios;
[0135] B) Generally, RGB can be given higher priority, with a recommended weighting of 2:1:1;
[0136] C) If Si / S max If the weight of Si is less than 1 / 10, then the weight of Si can be set to 0, and the weight ratio can be distributed to the remaining color space according to the proportional relationship.
[0137] S7.1.2 Calculate the texture feature similarity score to obtain the texture similarity features.
[0138] The computer device uses a sliding window to segment the test case into blocks, obtaining M first block images; and uses a sliding window to segment the first image into blocks, obtaining M second block images; wherein, there is a one-to-one correspondence between the M first block images and the M second block images, and M is a positive integer greater than 1; the structural similarity between the corresponding first block images and the second block images is calculated to obtain M structural similarities; the average of the M structural similarities is calculated to obtain the average structural similarity between the test case and the first image, and the average structural similarity is used as a texture similarity feature.
[0139] For example, SSIM (Structure Similarity Index Measure) measures image similarity from three aspects: brightness, contrast, and structure. SSIM values range from [0,1], with larger values indicating less image distortion. Since the searched images do not have strong consistency, they need to be divided into blocks. A sliding window can be used to divide the image into blocks, with a total of M blocks. Considering the influence of window shape on the blocks, Gaussian weighting is used to calculate the mean, variance, and covariance of each window. Then, the structure similarity SSIM of the corresponding blocks is calculated. Finally, the average value is used as the measure of the structure similarity between the two images, i.e., the average structure similarity MSSIM.
[0140] For example, the sliding window size can be set flexibly; to balance performance and effect, a size of 11*11 or 13*13 (pixels) is recommended. The mean of the SSIM calculated for each sliding window is then normalized. The retrieved materials are sorted by MSSIM, and images with high structural similarity (TOPp) or those meeting a certain threshold (Tmssim) are selected for subsequent steps.
[0141] S7.1.3 Calculate spatial feature similarity to obtain spatial similar features.
[0142] The computer device determines a first spatial relationship between targets included in the use case; and determines a second spatial relationship between targets included in the use case; based on the similarity between the first spatial relationship and the second spatial relationship, it determines the spatial similarity (i.e., spatial feature similarity) between the use case and the first image, and uses the spatial similarity as a spatial similarity feature.
[0143] For example, spatial relationship statistics of targets are performed on the retrieved images. The main spatial relationships presented in this scheme are: connection / adjacency relationship R1, overlap / intersection relationship R2, containment / enclosure relationship R3, top / bottom / left / right relationship R4. Of course, some relationships such as rotation, reversal, scale change, and orientation of targets can also be considered; the pairwise relationships between core targets in the original image and the retrieved image are recorded.
[0144]
[0145]
[0146] Steps for spatial similarity feature statistics:
[0147] Target consistency: Determine whether targets are consistent; the higher the consistency, the higher the priority. Calculate the spatial relationship between targets; if the spatial relationship is consistent with the original image, the spatial similarity is increased by 1. For example, if targets A and B in the first image have a spatial relationship R1, the spatial similarity is increased by 1. The larger the sum of the values, the higher the similarity of the spatial relationship.
[0148] S7.2, based on the various features analyzed in S7.1, find the main logic of the algorithm's misidentification.
[0149] S7.3 also allows adding bad case sets directly to the training set to optimize the training set and achieve algorithm optimization through iterative optimization.
[0150] In summary, the use case collection method provided in this embodiment acquires a large number of similar resources through image search and feeds them back into the model to extract similar bad cases; through multiple iterations of backflow, it quickly obtains bad cases that meet the conditions; using the bad case set, it feeds back into the algorithm to be evaluated to obtain the judgment features of the error labels, thereby finding the main logic for similarity prediction; using the bad case set, it feeds back into the algorithm to be evaluated to obtain the judgment features of the error labels, thereby finding the main logic for similarity prediction. It should be noted that some steps are not shown in the figures.
[0151] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0152] Figure 8 A structural block diagram of a use case collection apparatus provided in an exemplary embodiment of this application is shown. This use case collection apparatus can be implemented as all or part of a computer device through software, hardware, or a combination of both. The apparatus includes:
[0153] The recognition module 502 is used to input the first image with labels into the algorithm to be evaluated, and the algorithm to be evaluated outputs the first label set of the first image;
[0154] The acquisition module 504 is used to acquire a second image with the same content as the first image in response to the presence of N tags in the first tag set that are different from the tags in the labeled tag set of the first image;
[0155] The recognition module 502 is used to input the second image into the algorithm to be evaluated, and the algorithm to be evaluated outputs a second tag set of the second image.
[0156] The generation module 506 is configured to add the second image to a use case set for algorithm optimization in response to the second tag set containing at least one of the N tags, where N is a positive integer.
[0157] In some embodiments, the generation module 506 is configured to add the second image to the use case set corresponding to the at least one of the N tags in response to the second tag set containing at least one of the N tags.
[0158] In some embodiments, the generation module 506 is configured to add the second image to the use case set corresponding to the N tags in response to the second tag set having the same tags as the first tag set.
[0159] In some embodiments, the acquisition module 504 is configured to, in response to the N tags being different from the tags in the labeled tag set, use the first image as input data for a web search to obtain the second image.
[0160] In some embodiments, the acquisition module 504 is used for:
[0161] Using the first image as input data for a web search, G candidate images are obtained, where G is a positive integer greater than 1;
[0162] The first K images out of the G candidate images are selected as the second image, where K is a positive integer less than or equal to G.
[0163] In some embodiments, the device further includes: a statistics module 508;
[0164] The statistics module 508 is used to count the number of use cases in the use case set;
[0165] The acquisition module 504 is configured to acquire a third image with the same content as the first image in response to the number of use cases being less than a quantity threshold, wherein the third image is different from the second image.
[0166] In some embodiments, the device further includes: an analysis module 510;
[0167] Analysis module 510 is used to analyze the similarity features between the use cases in the use case set and the first image in response to the number of use cases being greater than or equal to the number threshold.
[0168] In some embodiments, the similarity features include color similarity features;
[0169] Analysis module 510 is used to calculate the similarity between the use case and the first image in at least two color spaces; and to calculate the weighted value of the similarity between the at least two colors based on the weights configured in the at least two color spaces, thereby obtaining the color similarity feature.
[0170] In some embodiments, the similarity features include texture similarity features;
[0171] Analysis module 510 is used to segment the use case using a sliding window to obtain M first segmented images; and to segment the first image using the sliding window to obtain M second segmented images; wherein, the M first segmented images and the M second segmented images are in one-to-one correspondence, and M is a positive integer greater than 1; the structural similarity between the corresponding first segmented images and the second segmented images is calculated to obtain M structural similarities; the average value of the M structural similarities is calculated to obtain the average structural similarity between the use case and the first image, and the average structural similarity is used as the texture similarity feature.
[0172] In some embodiments, the similarity features include spatial similarity features;
[0173] Analysis module 510 is used to determine a first spatial relationship between targets included in the use case; and to determine a second spatial relationship between targets included in the use case; based on the similarity between the first spatial relationship and the second spatial relationship, to determine the spatial similarity between the use case and the first image, and to use the spatial similarity as the spatial similarity feature.
[0174] It should be noted that the use case collection device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the use case collection method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the use case collection device and the use case collection method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0175] Figure 9 A schematic diagram of a computer device provided in an exemplary embodiment of this application is shown. This computer device may be a terminal or a server executing the use case collection method provided in this application. Specifically:
[0176] Computer device 700 includes a central processing unit (CPU) 701, a system memory 704 including random access memory (RAM) 702 and read-only memory (ROM) 703, and a system bus 705 connecting the system memory 704 and the CPU 701. Computer device 700 also includes a basic input / output system (I / O system) 706 that facilitates information transfer between various devices within the computer, and a mass storage device 707 for storing the operating system 713, application programs 714, and other program modules 715.
[0177] The basic input / output system 706 includes a display 708 for displaying information and an input device 709 for user input, such as a mouse or keyboard. Both the display 708 and the input device 709 are connected to the central processing unit 701 via an input / output controller 710 connected to the system bus 705. The basic input / output system 706 may also include the input / output controller 710 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 710 also provides output to a display screen, printer, or other types of output devices.
[0178] Mass storage device 707 is connected to central processing unit 701 via a mass storage controller (not shown) connected to system bus 705. Mass storage device 707 and its associated computer-readable media provide non-volatile storage for computer device 700. That is, mass storage device 707 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.
[0179] Computer-readable media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile optical disc (DVD), or solid-state drives (SSD), other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Random access memory can include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 704 and the mass storage device 707 mentioned above can be collectively referred to as memory.
[0180] According to various embodiments of this application, the computer device 700 can also be connected to a remote computer on a network, such as the Internet, for operation. That is, the computer device 700 can be connected to a network 712 via a network interface unit 711 connected to the system bus 705, or the network interface unit 711 can be used to connect to other types of networks or remote computer systems (not shown).
[0181] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU to implement the use case collection method described above.
[0182] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the use case collection method described in the above embodiments.
[0183] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).
[0184] This application also provides a computer program product (or computer program) including computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the use case collection method described above in the various embodiments.
[0185] According to another aspect of this application, a chip is provided, the chip including programmable logic circuitry and / or program instructions, which, when the chip is running, are used to implement the use case collection methods described in the various embodiments above.
[0186] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0187] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0188] The above description is merely an exemplary embodiment that can be implemented in this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A use case collection method characterized by comprising: The method comprises: inputting a first picture labeled with a label into an algorithm to be evaluated, outputting a first label set of the first picture by the algorithm to be evaluated; in response to there being N labels in the first label set different from the labels in the labeled label set of the first picture, obtaining a second picture with the same content as the first picture; inputting the second picture into the algorithm to be evaluated, outputting a second label set of the second picture by the algorithm to be evaluated; in response to the second label set containing at least one of the N labels, adding the second picture to a use case set for algorithm optimization, wherein N is a positive integer.
2. The method of claim 1, wherein, The response to the second label set containing at least one of the N labels includes: in response to the second label set containing at least one of the N labels, adding the second picture to a use case set corresponding to the at least one label.
3. The method of claim 2, wherein, The response to the second label set containing at least one of the N labels includes: in response to the second label set being the same as the labels in the first label set, adding the second picture to a use case set corresponding to the N labels.
4. The method according to any one of claims 1 to 3, characterized in that, The response to there being N labels in the first label set different from the labels in the labeled label set of the first picture includes: in response to the N labels being different from the labels in the labeled label set, taking the first picture as input data for network search to search for the second picture.
5. The method of claim 4, wherein, The taking the first picture as input data for network search to search for the second picture includes: taking the first picture as input data for network search to search for G candidate pictures, wherein G is a positive integer greater than 1; determining the first K of the G candidate pictures as the second picture, wherein K is a positive integer less than or equal to G.
6. The method according to any one of claims 1 to 3, characterized in that, After the adding the second picture to the use case set for algorithm optimization, the method comprises: counting the number of use cases in the use case set; in response to the number of use cases being less than a number threshold, obtaining a third picture with the same content as the first picture, wherein the third picture is different from the second picture.
7. The method of claim 6, wherein, The method further comprises: in response to the number of use cases being greater than or equal to the number threshold, analyzing similar features between the use cases in the use case set and the first picture.
8. The method of claim 7, wherein, The similar features include color similarity features; The analysis of the similar features between the use cases in the use case set and the first picture includes: calculating at least two color similarities between the use cases and the first picture in at least two color spaces; based on the weights configured for the at least two color spaces, calculating weighted values of the at least two color similarities to obtain the color similarity features.
9. The method of claim 7, wherein, The similar features include texture similarity features; The analysis of the similar features between the use cases in the use case set and the first picture includes: The use case is divided into M first sub-block images by using a sliding window, and the first picture is divided into M second sub-block images by using the sliding window; wherein the M first sub-block images and the M second sub-block images correspond to each other one by one, and M is a positive integer greater than 1; The structural similarity between the corresponding first sub-block image and the second sub-block image is calculated to obtain M structural similarities; The average value of the M structural similarities is calculated to obtain the average structural similarity between the use case and the first picture, and the average structural similarity is taken as the texture similarity feature.
10. The method of claim 7, wherein, The similarity feature includes a spatial similarity feature; The analysis of the similarity feature between the use case in the use case set and the first picture includes: determining a first spatial relationship between the targets contained in the use case; and determining a second spatial relationship between the targets contained in the use case; Based on the similarity of the first spatial relationship and the second spatial relationship, the spatial similarity between the use case and the first picture is determined, and the spatial similarity is taken as the spatial similarity feature.
11. A use collection apparatus characterized by comprising: The device includes: An identification module for inputting a first picture labeled with a label into a to-be-evaluated algorithm, and outputting a first label set of the first picture by the to-be-evaluated algorithm; An acquisition module for acquiring a second picture with the same content as the first picture in response to the existence of N labels in the first label set being different from the labels in the labeled label set of the first picture; The identification module inputs the second picture into the to-be-evaluated algorithm, and outputs a second label set of the second picture by the to-be-evaluated algorithm; A generation module for adding the second picture to a use case set for algorithm optimization in response to the second label set containing at least one of the N labels, wherein N is a positive integer.
12. A computer device, comprising: The computer device includes a processor, a memory connected to the processor, and program instructions stored on the memory, wherein the processor executes the program instructions to implement the use case collection method of any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and the program instructions are executed by the processor to implement the use case collection method of any one of claims 1 to 10.
14. A computer program product, characterised in that, The computer program product includes computer instructions stored in a computer-readable storage medium; the processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to make the computer device execute the use case collection method of any one of claims 1 to 10.
Citation Information
Patent Citations
Image processing method and device, electronic apparatus, and storage medium
WO2019233421A1
Image multi-tag classification method and apparatus, computer device, and storage medium
WO2022121485A1