Optimization method of image recognition model and computer readable storage medium
By quantifying and comparing the image label evaluation scores of the image recognition model, and learning labels with better rationality, the classification rationality and accuracy of the image recognition model are improved, solving the problem of large discrepancies between image labels and visual interpretation in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-08-23
- Publication Date
- 2026-05-19
AI Technical Summary
Existing image recognition models cannot improve classification accuracy during optimization and iteration, resulting in a significant difference between image labels and human visual interpretation, thus reducing classification accuracy.
After the image recognition model completes image recognition, the evaluation score of each image label is obtained to quantify the classification rationality of each image label. Based on the evaluation score of the image label, the evaluation result of the corresponding image recognition model is obtained. By comparing the classification rationality of multiple models, the image labels with better rationality are learned to improve the classification rationality of the target model.
This improves the classification rationality of the image recognition model, making image labels closer to human visual interpretation, and enhancing the model's classification rationality and accuracy.
Smart Images

Figure CN117671314B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, and in particular relates to an optimization method for an image recognition model and a computer-readable storage medium. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, AI has been widely applied in fields such as computer vision, natural language processing, intelligent robots, deep learning, and data mining. Among these, image recognition technology is an important branch of AI. Through image recognition, computers can perform quantitative analysis on images and identify the features contained within them, thus replacing human visual interpretation.
[0003] Currently, after creating an image recognition model, it needs to be optimized through training, testing, and analysis. After multiple optimizations, the best-performing model can be selected, completing one iteration of the image recognition model. Through optimization and iteration, the classification accuracy of the image recognition model for different features can be continuously improved, bringing the classification accuracy closer to a preset target (e.g., 95%). However, optimization and iteration cannot improve the reasonableness of the image recognition model's classification of different features. Image recognition models match features in an image with image labels according to preset classification rules. If the image labels are set irrationally, or the matching mechanism in the preset classification rules is flawed, the reasonableness of the image recognition model's classification can easily decrease, resulting in a significant difference between the image labels output for the same features and human visual interpretation. Therefore, how to improve the reasonableness of image recognition model classification has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of this application provide an optimization method for an image recognition model and a computer-readable storage medium to solve the problem that existing algorithm optimization methods cannot improve the classification rationality of image recognition models.
[0005] The first aspect of this application provides a method for optimizing an image recognition model, comprising:
[0006] The images are input into i image recognition models respectively, and the image labels output by each image recognition model are obtained;
[0007] Based on the feature proportion of each image feature reflected by each image label in the corresponding image, the evaluation score of each image label corresponding to the corresponding image is obtained respectively;
[0008] Based on the evaluation scores of each image label corresponding to the same image recognition model, the evaluation result of the corresponding image recognition model is determined;
[0009] Based on the evaluation results of each of the image recognition models, the target image recognition model is optimized, wherein the target image recognition model is at least one of the i image recognition models, and i is greater than or equal to 2.
[0010] The first aspect of the present application provides an optimization method for an image recognition model. After the image recognition model completes image recognition, it obtains the evaluation score of each image label to quantify the classification rationality of each image label. Based on the evaluation score of the image label, it obtains the evaluation result of the corresponding image recognition model to quantify the classification rationality of each image recognition model. By comparing the classification rationality of multiple image recognition models, when the classification rationality of the target image recognition model is poor or there are image labels with poor classification rationality, it can learn the image labels with better classification rationality in each image recognition model, thereby improving the classification rationality of the target image recognition model and making the image labels output by the image recognition model increasingly closer to human visual interpretation.
[0011] A second aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the optimization method for the image recognition model provided in the first aspect of this application.
[0012] It is understandable that the beneficial effects of the second aspect mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application;
[0015] Figure 2 This is a schematic diagram of the architecture of the terminal device provided in the embodiments of this application;
[0016] Figure 3 This is a schematic diagram of the first process of the optimization method for the image recognition model provided in the embodiments of this application;
[0017] Figure 4 This is a logical diagram illustrating the optimization of an image recognition model provided in an embodiment of this application;
[0018] Figure 5 This is a schematic diagram of the second process of the optimization method for the image recognition model provided in the embodiments of this application;
[0019] Figure 6 This is a schematic diagram of the third process of the optimization method for the image recognition model provided in the embodiments of this application;
[0020] Figure 7 This is a schematic diagram of the fourth process of the optimization method for the image recognition model provided in the embodiments of this application;
[0021] Figure 8 This is a fifth flowchart illustrating the optimization method for the image recognition model provided in this application embodiment. Detailed Implementation
[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0023] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0024] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0026] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0027] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0028] In applications, optimization and iteration can continuously improve the classification accuracy of image recognition models for different features, bringing the accuracy closer to a preset target (e.g., 95%). However, optimization and iteration cannot improve the reasonableness of image recognition models in classifying different features. Image recognition models match features in an image with image labels according to preset classification rules. If the image labels are set unreasonably, or if the matching mechanism in the preset classification rules is flawed, the reasonableness of the image recognition model's classification can easily decrease, resulting in a significant difference between the image labels output for the same features and human visual interpretation. Therefore, improving the reasonableness of image recognition model classification has become an urgent problem to be solved.
[0029] To address the aforementioned technical problems, this application provides an optimization method for an image recognition model. After the image recognition model completes image recognition, it obtains an evaluation score for each image label to quantify the classification rationality of each label. Based on the evaluation scores of the image labels, it obtains the evaluation results of the corresponding image recognition model to quantify the classification rationality of each model. By comparing the classification rationality of multiple image recognition models, when the target image recognition model has poor classification rationality or contains image labels with poor classification rationality, it can learn from the image labels with better classification rationality among various image recognition models. This improves the classification rationality of the target image recognition model, making the image labels output by the image recognition model increasingly closer to human visual interpretation.
[0030] The image recognition model method provided in this application can be applied to terminal devices. Terminal devices can be mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. This application does not impose any limitations on the specific type of terminal device.
[0031] Figure 1 An exemplary structural diagram of terminal device 1 is shown. Terminal device 1 may include a processor 10, a memory 20, a power module 30, an audio module 40, a camera module 50, a sensor module 60, an input module 70, a display module 80, and a wireless communication module 90, etc. The audio module 40 may include a speaker 41 and a microphone 42, etc.; the camera module 50 may include a short-focus camera 51, a long-focus camera 52, and a flash 53, etc.; the sensor module 60 may include an infrared sensor 61, an accelerometer 62, a position sensor 63, a fingerprint sensor 64, and an iris sensor 65, etc.; the input module 70 may include a touch panel 71 and an external input unit 72, etc.; and the wireless communication module 90 may include wireless communication units such as Bluetooth, ZigBee, Optical Wireless, Wireless Local Area Network (WLAN), and Near Field Communication (NFC).
[0032] In applications, processor 10 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0033] In applications, memory 20 may be an internal storage unit of the terminal device in some embodiments, such as a hard disk or memory of the terminal device. In other embodiments, memory 20 may be an external storage device of the terminal device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., provided on the terminal device. Furthermore, memory 20 may include both internal and external storage units of the terminal device. Memory 20 is used to store operating system 21, application program 22, computer program 23, and data, etc. Memory 20 may also be used to temporarily store data that has been output or will be output. When processor 10 executes the computer program 23 stored on memory 20, it implements the steps in the following various positioning method embodiments.
[0034] In the application, when the image recognition model is used on the terminal device 1, the computer program 23 includes the image recognition model and stores it in the memory 20. The terminal device 1 can acquire the image to be recognized through the short-focus camera 52 or the long-focus camera 53, and run the image recognition model stored in the memory 20 through the processor 10 to recognize the image to be recognized and obtain the features contained in the image to be recognized.
[0035] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on terminal device 1. In other embodiments of this application, terminal device 1 may include more or fewer components than illustrated, or combine certain components, or different components, such as input / output devices, network access devices, etc. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0036] like Figure 2 The diagram shows the architecture of the operating system running on the terminal device 1 provided in this application embodiment. The layered architecture can divide the operating system into multiple layers, which communicate with each other through software interfaces. In some embodiments, the operating system can be divided into an application layer 100, a framework layer 200, a system runtime library layer 300, a hardware abstraction layer 400, and a Linux kernel layer 500 from top to bottom.
[0037] It should be noted that the type of operating system can be Android, a customized operating system based on Android, or other different types of operating systems. This application embodiment does not limit the specific type of operating system.
[0038] In application, the five layers of the operating system are explained below:
[0039] Application layer 100 may include built-in applications 120 and upper-layer applications 110 provided by third-party application providers. Applications in application layer 100 can interact directly with users to realize different functions provided by the application. For example, application layer 100 may include built-in applications such as email, telephone, calendar, camera, contacts and Bluetooth, as well as upper-layer applications such as map positioning, food delivery and video playback.
[0040] The framework layer 200 may include application programming interfaces (APIs) and programming frameworks. The APIs can be used to provide operating system developers with interfaces for developing applications, and can also be used to call corresponding basic services when applications in the application layer implement different functions. For example, the framework layer may include different types of APIs such as window managers, content providers, phone managers, location managers, and view systems.
[0041] The window manager is used to manage window programs. Specifically, it can be used to get the window size, and also to determine whether there is a status bar, whether the screen is locked, and whether the screen is captured.
[0042] The content provider is used to store and retrieve data and make the data accessible to applications. The data may include videos, images, audio, call logs, contacts, browsing history, and bookmarks.
[0043] The phone manager is used to provide communication functions for terminal device 1, such as managing call status, including answering and hanging up calls;
[0044] The location manager is used to obtain the location information of terminal device 1, which may include different types of location information such as satellite location information (obtained through the Global Positioning System), network location information (obtained through network positioning services), and fused location information (obtained through fused positioning services);
[0045] The view system provides visual controls, such as controls for displaying text and controls for displaying images, and is also used to build applications for displaying application layer 100. The view system can run multiple visual controls simultaneously, allowing the terminal device to display multiple views at the same time, for example, a view displaying text and a view displaying images simultaneously.
[0046] The system runtime library layer (Native) 300 may include a C / C++ library 310 and an Android runtime library 320. The C / C++ library 310 may include a drawing function library, a font engine, a rendering engine, a multimedia library, and a database engine. Specifically, the drawing function library may be OpenGL For Embedded Systems 3D (an open 3D graphics library for embedded systems); the font engine provides different fonts, specifically FreeType (a portable font engine); the rendering engine renders 2D or 3D graphics, specifically Skia Graphics Library (an engine for rendering 2D graphics); the multimedia library (Media Framework) supports playback, recording, and replay of different audio and video formats; and the database engine provides storage functionality for different types of databases, allowing different types of data to be stored in different databases as needed, or stored uniformly in a single database.
[0047] The Android runtime library 320 includes core libraries and the Android Runtime (ART). In Android 5.0 and later, the Dalvik virtual machine was replaced by ART. The core libraries provide most of the functionality of the Java language core libraries, allowing developers to write Android applications using Java. Compared to the Java Virtual Machine (JVM), the Dalvik virtual machine is specifically designed for mobile devices, allowing multiple instances of the virtual machine to run simultaneously within limited memory, and each Dalvik application executes as an independent Linux process. Independent processes prevent all programs from being shut down in the event of a virtual machine crash. ART, which replaces the Dalvik virtual machine, uses a different mechanism. In Dalvik, bytecode needs to be converted into machine code by a just-in-time (JIT) compiler every time an application runs, which slows down application performance. In the ART environment, the bytecode is pre-compiled into machine code during the first installation, making it a truly native application.
[0048] The Hardware Abstraction Layer (HAL) 400 is an interface layer located between the operating system kernel and the hardware circuitry. Its purpose is to abstract the hardware, hiding the platform-specific hardware interface details to protect hardware manufacturers' intellectual property. This provides the operating system with a virtual hardware platform, making it hardware-independent and portable across multiple platforms. From a software and hardware testing perspective, both software and hardware testing can be performed separately based on the HAL, allowing for parallel testing. Specifically, image recognition models can run on the HAL 400.
[0049] Linux kernel layer 500 can be used to provide system services for the operating system, including operating system security services, memory management, process management, network protocol stack and driver model, etc.
[0050] Understandable, Figure 2 The diagram shown is merely an example of an operating system architecture. The operating system architecture can also be four-layer or six-layer. The architecture of a four-layer operating system may include: an application layer, a framework layer, a system runtime library layer, and a Linux kernel layer. The architecture of a six-layer operating system may include: an application layer, a framework layer, a system runtime library layer, a hardware abstraction layer, a Linux kernel layer, and a hardware device layer. The method embodiments of this application do not limit the number of architecture layers or the specific architecture of the operating system.
[0051] like Figure 3 As shown, the image recognition model optimization method provided in this application embodiment includes the following steps S301 to S304:
[0052] Step S301: Input the images into i image recognition models respectively to obtain the image labels output by each image recognition model.
[0053] In applications, image recognition models can be built and trained based on neural network (NN) models. Specifically, image recognition models can be built and trained based on one or more different types of neural network models, such as convolutional neural networks (CNN), region-convolutional neural networks (R-CNN), fully convolutional networks (FCN), region-fullly convolutional networks (R-FCN), and feature pyramid networks (FPN). This application does not impose any restrictions on the specific network structure of the image recognition model.
[0054] In applications, the number of images input to the image recognition model can be one or multiple images. This embodiment of the application does not impose any limitation on the number of images used for evaluation. The images are input into i image recognition models respectively. For each image recognition model, by recognizing each input image, at least one image feature can be obtained for each image. An image label corresponding to each image feature is determined according to classification rules, and the image label corresponding to each image feature is output to achieve image recognition of the input images. When outputting image labels, the image labels can display the location of the corresponding image feature in the corresponding image, or the correspondence between image features and image labels can be output in text form. This embodiment of the application does not impose any limitation on the output method of image labels. The classification rules are determined according to the actual image recognition model selected, and this embodiment of the application does not impose any limitation on the specific matching mechanism of the classification rules.
[0055] For example, suppose an image is input into an image recognition model. The image recognition model obtains two image features in the image: a monitor and a capsule. According to the classification rules, it can determine that the image label corresponding to the monitor is "computer," and the image label corresponding to the capsule is "medicine." The image recognition model can output the image label of "computer" to the location of the monitor in the image, and the image label of "medicine" to the location of the medicine in the image. Alternatively, it can output the correspondence between the image feature "monitor" and the image label "computer" in text form, and the correspondence between the image feature "capsule" and the image label "medicine."
[0056] In one embodiment, before step S301, the method further includes:
[0057] Images are obtained based on preset image features.
[0058] In applications, images containing preset image features can be obtained by setting preset image features. Specifically, a preset image set and preset image features can be input into an object detection model, which can then filter the preset image set based on the preset image features to obtain the images mentioned above, all of which contain the preset image features. Alternatively, the images can be obtained from the Internet (which can include various open-source search engines and various applications) using a web crawler based on the preset image features.
[0059] In applications, associative image features can be obtained based on the semantic meaning of preset image features. The images can be obtained by inputting the preset image set, preset image features, and associative image feature values into the target detection model, and / or by crawling the image based on the preset image features and associative image features, which can improve the image diversity in the evaluation process.
[0060] In applications, multiple images can be obtained through multiple preset image features, and when these images are input into an image recognition model for optimization, the model can perform specific optimization on images containing the corresponding preset image features, thereby improving the rationality of the model's classification of the corresponding preset image features.
[0061] In one embodiment, step S301 is followed by:
[0062] Based on the semantic meaning of the image labels output by each image recognition model, multiple image labels with the same semantic meaning are integrated into the same image label to improve the consistency of the image labels output by each image recognition model.
[0063] In applications, for each image recognition model, the word meaning of each image label output by the image recognition model can be obtained. Specifically, the word meaning of each image label can be obtained through Natural Language Processing (NLP) algorithms, or through Natural Language Understanding (NLU) algorithms that further enhance word meaning analysis capabilities based on NLP algorithms, or through a pre-stored word meaning lookup table, which includes the word meaning of each image label output by the image recognition model.
[0064] In the application, after obtaining the meaning of each image tag, multiple image tags with the same meaning can be integrated into the same image tag. Specifically, multiple image tags with the same meaning can be filtered based on a thesaurus and integrated into the same image tag. The thesaurus includes multiple integrated tags, and each integrated tag has multiple corresponding meanings. By inputting each image tag and its corresponding meaning into the thesaurus, the thesaurus can convert the corresponding image tag into an integrated tag according to the meaning, thereby integrating multiple image tags with the same meaning into the same image tag, so as to improve the consistency of the image tags output by each image recognition model.
[0065] Step S302: Based on the feature proportion of each image label in the corresponding image, obtain the evaluation score of each image label corresponding to the corresponding image.
[0066] In the application, for each image tag, the feature percentage of the image features reflected by the image tag in the corresponding image can be obtained. Specifically, all image features in the corresponding image can be obtained, and the feature percentage of the image features reflected by the image tag in all the image features can be determined based on the percentage of the number of image features reflected by the image tag in all the image features (for example, assuming the corresponding image includes 10 image features and the image tag reflects 2 image features, then the percentage of the number of image features is 20%). The feature percentage of the image tag in the corresponding image can be equal to the percentage of the number of image features. Alternatively, the feature percentage of each image feature in the corresponding image can be obtained, and the feature percentage of the image tag can be obtained by summing the feature percentages of each image feature reflected by the image tag (for example, assuming the corresponding image includes 3 image features, the feature percentage of the first image feature is 20%, the feature percentage of the second image feature is 45%, the feature percentage of the third image feature is 35%, and the image tag reflects the first and second image features, then the feature percentage of the image tag is 65%). This application embodiment does not impose any restrictions on the calculation method of the feature percentage of the image tag.
[0067] In the application, for each image label, after obtaining the feature proportion of the image label, the evaluation score of the image label in the corresponding image can be obtained based on the feature proportion of the image label. Specifically, the evaluation score of an image label in a corresponding image can be determined based on a pre-stored evaluation score comparison table and the feature proportion of the image label. The evaluation score comparison table includes the correspondence between the feature proportion of the image label and the evaluation score. The correspondence can be a one-to-one correspondence between the feature proportion of the image label and the evaluation score (for example, the evaluation score is equal to a preset multiple of the feature proportion of the image label. Assuming the preset multiple is 100 times, and the feature proportion of the image label is 65%, then the evaluation score is 65 points). Alternatively, the correspondence can be a one-to-one correspondence between the range of the feature proportion of the image label and the evaluation score (for example, assuming there are five ranges of the feature proportion of the image label, the first range is [0, 20%), and the corresponding evaluation score is 1 point; the second range is [20%, 40%), and the corresponding evaluation score is 2 points; the third range is [40%, 60%), and the corresponding evaluation score is 3 points; the fourth range is [60%, 80%), and the corresponding evaluation score is 4 points; the fifth range is [80%, 100%), and the corresponding evaluation score is 5 points). This application does not impose any restrictions on the scoring method for image tags.
[0068] It should be noted that for each image recognition model, different input images can output the same image label. Therefore, when obtaining the evaluation score of the image label, it is necessary to match the evaluation score of the image label with the corresponding image. Alternatively, the evaluation score of each image label corresponding to the image can be obtained on an image-by-image basis.
[0069] Step S303: Based on the evaluation scores of each image label corresponding to the same image recognition model, determine the evaluation result of the corresponding image recognition model.
[0070] In applications, for each image recognition model, the evaluation result can be determined based on the evaluation scores of all image labels to quantify the performance of the image recognition model. By processing the evaluation scores of image labels using different quantification methods, various evaluation results can be obtained, including evaluation items such as average score, recall, correct label rate, or cumulative frequency. Specifically, the average score can be obtained by acquiring the evaluation scores of all image labels output by the image recognition model; the recall value can be obtained based on the number of images input to the image recognition model and the number of image labels with evaluation scores greater than a preset evaluation score. Specifically, the recall value can be obtained by dividing the number of image labels with evaluation scores greater than the preset evaluation score by the number of images, reflecting the number of correct labels corresponding to each image; the calculation methods for the correct label rate and cumulative frequency can be referred to the relevant descriptions in the following embodiments.
[0071] It is easy to understand that the more evaluation items included in the evaluation results of an image recognition model, the more comprehensive the performance quantification of the image recognition model can be. This application embodiment does not impose any restrictions on the evaluation items included in the evaluation results of the image recognition model.
[0072] Step S304: Optimize the target image recognition model based on the evaluation results of each image recognition model, wherein the target image recognition model is at least one of the i image recognition models, and i is greater than or equal to 2.
[0073] In the application, after obtaining the evaluation results of each image recognition model, the target image recognition model can be optimized for each evaluation item. The i image recognition models can include at least one target image recognition model and at least one optimized image recognition model. Specifically, when the average score of the target image recognition model is less than a preset average score, for each labeled image, at least one label to be modified corresponding to the labeled image is obtained. The label to be modified satisfies the following conditions: the image label output by the optimized image recognition model based on the labeled image contains the label to be modified; the image label output by the target image recognition model based on the labeled image contains the label to be modified; and the evaluation score of the label to be modified output by the optimized image recognition model is greater than the evaluation score of the label to be modified output by the target image recognition model.
[0074] In application, multiple optimized image recognition models are filtered based on the labels to be modified output by labeled images to obtain at least one target label with the highest evaluation score. The labels to be modified in the target image recognition model can be modified to the above target label to modify the image features reflected by the labels to improve the classification rationality of the image recognition model.
[0075] For example, suppose the label to be modified is "computer". In the target image recognition model, the image feature reflected by the label to be modified is "keyboard" and the image feature reflected by the target label is "monitor". In the labeled image, the evaluation score when the image feature is "monitor" is greater than the evaluation score when the image feature is "keyboard". Therefore, the image feature reflected by the label to be modified can be changed from "keyboard" to "monitor".
[0076] Figure 4 An exemplary schematic diagram of the logic for optimizing an image recognition model is shown, wherein, Figure 4 The example shown only illustrates that when i=2, the image recognition model includes a target image recognition model and an optimized image recognition model.
[0077] In application, the image recognition model optimization method provided in this application embodiment can obtain the evaluation score of each image label after the image recognition model completes image recognition, so as to quantify the classification rationality of each image label, and obtain the evaluation result of the corresponding image recognition model based on the evaluation score of the image label, so as to quantify the classification rationality of each image recognition model. By comparing the classification rationality of multiple image recognition models, when the classification rationality of the target image recognition model is poor, or there are image labels with poor classification rationality, the image labels with better classification rationality in each image recognition model can be learned, thereby improving the classification rationality of the target image recognition model and making the image labels output by the image recognition model increasingly closer to human visual interpretation.
[0078] like Figure 5 As shown, in one embodiment, based on Figure 3 The corresponding embodiment includes the following steps S501 to S505:
[0079] Step S501: Input the images into i image recognition models respectively to obtain the image labels output by each image recognition model.
[0080] In application, the optimization method provided in step S501 is the same as the optimization method provided in step S301 above, and will not be repeated here.
[0081] Step S502: Train the evaluation model based on the feature proportion of the image features reflected by multiple sample labels in the corresponding image and the reference score of the corresponding sample label.
[0082] In application, the network structure of the evaluation model can refer to the network structure of the image recognition model described above, and will not be repeated here. Multiple sample labels can be set in the evaluation model to be trained. By inputting the images into the evaluation model, the model outputs the training proportion and training score of the image features reflected by each sample label in the corresponding image. For each sample label in each image, a first loss function is established based on the first difference between the training proportion of the image features reflected by the sample label in the corresponding image and the feature proportion of the sample label in the corresponding image. A second loss function is established based on the second difference between the training score and the reference score of the sample label. A comprehensive loss function is generated based on the first and second loss functions. The evaluation model to be trained is then trained until the first difference is less than or equal to a preset first difference and the second difference is less than or equal to a preset second difference. Training stops then, and the trained evaluation model is output.
[0083] Step S503: Input each image label into the trained evaluation model, obtain the feature proportion of each image label in the corresponding image, and obtain the evaluation score of each image label corresponding to the corresponding image.
[0084] In application, for each image recognition model, the output image labels can be input into the trained evaluation model to obtain the evaluation scores of each image label corresponding to the corresponding image, which improves the accuracy and automation of obtaining evaluation scores.
[0085] Step S504: Based on the evaluation scores of each image label corresponding to the same image recognition model, determine the evaluation result of the corresponding image recognition model;
[0086] Step S505: Optimize the target image recognition model based on the evaluation results of each image recognition model, wherein the target image recognition model is at least one of the i image recognition models, and i is greater than or equal to 2.
[0087] In application, the optimization methods provided in steps S504 and S505 are the same as those provided in step S301 above, and will not be repeated here.
[0088] like Figure 6 As shown, in one embodiment, based on Figure 5 The corresponding embodiment includes the following steps S601 to S608:
[0089] Step S601: Input the images into i image recognition models respectively, and obtain the image labels output by each image recognition model;
[0090] Step S602: Train the evaluation model based on the feature proportion of the image features reflected by multiple sample labels in the corresponding image and the reference score of the corresponding sample label;
[0091] Step S603: Input each image label into the trained evaluation model, obtain the feature proportion of each image label in the corresponding image, and obtain the evaluation score of each image label corresponding to the corresponding image.
[0092] In application, the optimization methods provided in steps S601 to S603 are the same as those provided in steps S501 to S503 above, and will not be repeated here.
[0093] Step S604: Based on the evaluation scores of each image label, filter out the correct labels corresponding to the corresponding images. The evaluation score of the correct label is greater than or equal to the preset evaluation score.
[0094] In the application, it can be determined whether the evaluation score of each image label is greater than or equal to the preset evaluation score. When the evaluation score of an image label is greater than or equal to the preset evaluation score, the image label is marked as the correct label corresponding to the corresponding image.
[0095] For example, assuming the scoring method of 1 to 5 points is adopted in the above embodiments, the preset scoring value can be 2 points or 3 points. This application embodiment does not impose any restrictions on the specific size of the preset scoring value.
[0096] Step S605: For each image recognition model, based on the number of identical image labels output by the image recognition model and the number of identical image labels that are selected as correct labels, determine the correct label rate of the identical image labels output by the image recognition model.
[0097] In applications, for each image recognition model, the number of identical image labels output by the image recognition model and the number of identical image labels that are selected as correct labels can be obtained, thus determining the correct label rate of the identical image labels output by the image recognition model.
[0098] For example, suppose the input image recognition model contains 3 images, namely image 1, image 2, and image 3. In image 1 and image 2, the image recognition model outputs the image label "computer". Then the image recognition model outputs the same image label "computer", and the number of times the same image label "computer" is output is 2. Furthermore, the same image label "computer" is selected as the correct label in image 1 but not in image 2. Therefore, the correct label rate of the image recognition model outputting the same image label "computer" is 50%.
[0099] In applications, by obtaining the correct label rate of the same image label output by the image recognition model, the reasonableness of the classification of each image label in the corresponding image recognition model can be quantified. The higher the correct label rate, the better the reasonableness of the classification of the image label in the corresponding image recognition model.
[0100] It should be noted that if the number of identical image labels output by the image recognition model is 1, then when the image is selected as the correct label, the correct label rate for the corresponding identical image label is 100%; when the image is not selected as the correct label, the correct label rate for the corresponding identical image label is 0.
[0101] Step S606: For each image recognition model, based on the number of identical correct labels output by the image recognition model and the number of image labels output by the image recognition model, determine the frequency of occurrence of identical correct labels output by the image recognition model.
[0102] In applications, for each image recognition model, the number of identical correct labels output by the image recognition model and the number of image labels output by the image recognition model can be obtained to determine the frequency of identical correct labels output by the image recognition model.
[0103] For example, suppose the image recognition model outputs the same correct labels including "computer", "toothbrush" and "mobile phone", and the number of the same correct labels "computer" is 3, "toothbrush" is 6, "mobile phone" is 9, and the number of image labels output by the image recognition model is 30. Then the frequency of the same correct label "computer" is 10%, the frequency of the same correct label "toothbrush" is 20%, and the frequency of the same correct label "mobile phone" is 30%.
[0104] Step S607: Based on the evaluation results of each image recognition model, correct the image labels corresponding to the labeled images; wherein, after the labeled images are input into the target image recognition model, the target image recognition model may have missed recognition or misrecognition phenomena.
[0105] In application, the optimization method provided in step S607 can refer to the optimization method provided in step S304 above, and will not be repeated here. The difference is that the target image recognition model can also be optimized based on the correct label rate and occurrence frequency.
[0106] In the application, at least one missed label corresponding to the labeled image can be obtained based on the evaluation results of each image recognition model. The missed label meets the following conditions: the image label output by the target image recognition model based on the labeled image does not contain the missed label or its approximate label, and the image label output by the optimized image recognition model based on the labeled image contains the missed label and is determined to be the correct label. At least one missed label is added as the image label corresponding to the labeled image. Specifically, the approximate label of the missed label can be obtained according to a pre-stored approximate label lookup table. The approximate label lookup table includes approximate labels of at least one image label. When a missed label is obtained, it can be determined whether the missed label has an approximate label by looking up the table. If so, the approximate label of the missed label is obtained to further improve the recognition rate of the missed label.
[0107] In application, the missing identification tags can also meet the following conditions: the correct label rate of the missing identification tags is greater than or equal to the preset correct label rate, and the occurrence frequency of the missing identification tags is greater than or equal to the preset occurrence frequency (or, the occurrence number of the missing identification tags is greater than or equal to the preset occurrence number), so as to ensure that the classification of the image tags selected as missing identification tags is reasonable enough and the occurrence frequency is high enough.
[0108] Step S608: Based on the labeled images and the corrected image labels, perform secondary training on the target image recognition model.
[0109] In the application, after the image labels are corrected, the target image recognition model is retrained based on the corrected image labels and the corresponding labeled images to improve the classification rationality of the target image recognition model after retraining and reduce the phenomenon of missed recognition.
[0110] like Figure 7 As shown, in one embodiment, based on Figure 6 The corresponding embodiment includes the following steps S701 to S710:
[0111] Step S701: Input the images into i image recognition models respectively, and obtain the image labels output by each image recognition model;
[0112] Step S702: Train the evaluation model based on the feature proportion of the image features reflected by multiple sample labels in the corresponding image and the reference score of the corresponding sample label;
[0113] Step S703: Input each image label into the trained evaluation model, obtain the feature proportion of each image label in the corresponding image, and obtain the evaluation score of each image label corresponding to the corresponding image.
[0114] Step S704: Based on the evaluation scores of each image label, filter out the correct labels corresponding to the corresponding images. The evaluation score of the correct label is greater than or equal to the preset evaluation score.
[0115] Step S705: For each image recognition model, based on the number of identical image labels output by the image recognition model and the number of identical image labels that are selected as correct labels, determine the correct label rate of the identical image labels output by the image recognition model.
[0116] Step S706: For each image recognition model, based on the number of identical correct labels output by the image recognition model and the number of image labels output by the image recognition model, determine the frequency of occurrence of identical correct labels output by the image recognition model.
[0117] In application, the optimization methods provided in steps S701 to S706 are the same as those provided in steps S601 to S606 above, and will not be repeated here.
[0118] Step S707: For each image recognition model, based on the occurrence frequency of each correct label output by the image recognition model, obtain the cumulative frequency of the correct labels output by the image recognition model, and obtain the cumulative number of correct labels whose cumulative frequency is within a preset cumulative frequency range. The cumulative number is used to characterize the diversity of correct labels contained in the image recognition model.
[0119] In the application, for each image recognition model, the correct labels can be sorted in descending order according to their frequency of occurrence, and the frequency of occurrence of each correct label can be accumulated according to the index of the correct label. For example, assuming there are three correct labels, "mobile phone", "toothbrush" and "computer", the frequency of occurrence of "mobile phone" is 30%, the frequency of occurrence of "toothbrush" is 20%, and the frequency of occurrence of "computer" is 10%. After sorting the correct labels in descending order according to their frequency of occurrence, mobile phone is the first correct label, toothbrush is the second correct label, and computer is the third correct label, with a cumulative frequency of 60%.
[0120] In application, the cumulative number of correct labels falling within a preset cumulative frequency range can be obtained. For example, assuming the preset cumulative frequency range is [20%, 95%], the cumulative frequencies of the three correct labels "mobile phone," "toothbrush," and "computer" all fall within this range. Assuming the preset cumulative frequency range is [35%, 95%], the cumulative frequencies of the two correct labels "toothbrush" and "computer" fall within this range. By setting a preset cumulative frequency range, correct labels with lower frequencies can be filtered out, while those with higher frequencies can be removed, improving the reasonableness of the judgment regarding the diversity of correct labels included in the image recognition model.
[0121] Step S708: When the cumulative number of target image recognition models is less than or equal to the preset cumulative number, obtain at least one missed recognition label corresponding to the labeled image based on the evaluation results of each image recognition model.
[0122] In application, when the cumulative number of target image recognition models is less than or equal to the preset cumulative number, it indicates that the diversity of correct labels contained in the target image recognition model is too low. At least one missed label corresponding to the labeled image can be obtained based on the evaluation results of each image recognition model. The method for obtaining the missed label can refer to the relevant description in step S607 above, and will not be repeated here.
[0123] Step S709: Add at least one missed identification label as an image label corresponding to the labeled image.
[0124] Step S710: Based on the labeled images and the corrected image labels, perform secondary training on the target image recognition model.
[0125] In application, the optimization methods provided in steps S709 and S710 can be referred to the relevant descriptions in steps S607 and S608 above, and will not be repeated here.
[0126] In application, by obtaining the cumulative number of target image recognition models, the diversity of image labels output by the target image recognition models can be quantified based on the cumulative number, and missed labels can be obtained when the diversity of image labels is too low, thereby improving the flexibility of adding new image labels.
[0127] like Figure 8 As shown, in one embodiment, based on Figure 5 The corresponding embodiment includes the following steps S801 to S807:
[0128] Step S801: Input the images into i image recognition models respectively, and obtain the image labels output by each image recognition model;
[0129] Step S802: Train the evaluation model based on the feature proportion of the image features reflected by multiple sample labels in the corresponding image and the reference score of the corresponding sample label;
[0130] Step S803: Input each image label into the trained evaluation model, obtain the feature proportion of each image label in the corresponding image, and obtain the evaluation score of each image label corresponding to the corresponding image.
[0131] Step S804: Based on the evaluation scores of each image label corresponding to the same image recognition model, determine the evaluation result of the corresponding image recognition model.
[0132] In application, the optimization methods provided in steps S801 to S804 are the same as those provided in steps S501 to S504 above, and will not be repeated here.
[0133] Step S805: Based on the evaluation results of each image recognition model, obtain at least one misidentified label corresponding to the labeled image; the misidentified label meets the following conditions: the image label output by the optimized image recognition model based on the labeled image contains the misidentified label and the evaluation score is greater than or equal to the preset evaluation score; the image label output by the target image recognition model based on the labeled image does not contain the misidentified label, but contains an approximate label containing the misidentified label and the evaluation score of the approximate label is less than or equal to the preset evaluation score.
[0134] Step S806: Modify the similar labels of the misidentified labels to the misidentified labels;
[0135] Step S807: Based on the labeled images and the corrected image labels, perform secondary training on the target image recognition model.
[0136] In application, at least one misidentified label corresponding to the labeled image can be obtained based on the evaluation results of each image recognition model. The misidentified label meets the following conditions: the image label output by the optimized image recognition model based on the labeled image contains the misidentified label and its evaluation score is greater than or equal to the preset evaluation score, indicating that the recognition effect and classification rationality of the misidentified label are good. The image label output by the target image recognition model based on the labeled image does not contain the misidentified label, but contains an approximate label of the misidentified label, and the evaluation score of the approximate label is less than or equal to the preset evaluation score, indicating that the recognition effect and classification rationality of the approximate label of the misidentified label are poor. Specifically, the approximate label of the misidentified label can be obtained according to a pre-stored approximate label lookup table. The recognition method is the same as the method for recognizing approximate labels of missed labels, and will not be repeated here.
[0137] In applications, within the target image recognition model, approximate labels of misidentified labels can be modified to the misidentified labels. Based on the misidentified labels and their corresponding labeled images, the target image recognition model can be retrained to further improve the classification rationality of the retrained target image recognition model and reduce misidentification.
[0138] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0139] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the embodiments of the optimization methods for the various image recognition models described above.
[0140] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include at least: any entity or device capable of carrying the computer program code to the camera terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0141] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0142] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] In the embodiments provided in this application, it should be understood that the disclosed terminal devices and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0144] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. An optimization method for an image recognition model, characterized in that, include: The images are input into i image recognition models respectively, and the image labels output by each image recognition model are obtained; Based on the feature proportion of each image feature reflected by each image label in the corresponding image, the evaluation score of each image label corresponding to the corresponding image is obtained respectively; Based on the evaluation scores of each image label corresponding to the same image recognition model, the evaluation result of the corresponding image recognition model is determined; Based on the evaluation results of each of the image recognition models, the target image recognition model is optimized, wherein the target image recognition model is at least one of the i image recognition models, and i is greater than or equal to 2; The evaluation results include the correct label rate and the frequency of occurrence. Determining the evaluation result of the corresponding image recognition model based on the evaluation scores of each image label corresponding to the same image recognition model includes: Based on the evaluation scores of each image tag, the correct tag corresponding to the corresponding image is selected, and the evaluation score of the correct tag is greater than or equal to the preset evaluation score. For each image recognition model, the correct label rate of the image recognition model outputting the same image labels is determined based on the number of identical image labels output by the image recognition model and the number of identical image labels that are selected as correct labels. For each image recognition model, the frequency of occurrence of the same correct label output by the image recognition model is determined based on the number of identical correct labels output by the image recognition model and the number of image labels output by the image recognition model.
2. The optimization method as described in claim 1, characterized in that, Before obtaining the evaluation score of each image label corresponding to the corresponding image based on the feature proportion of each image label reflected in the corresponding image, the method further includes: The evaluation model is trained based on the feature proportion of the image features reflected by multiple sample labels in the corresponding image and the reference score of the corresponding sample label. The step of obtaining evaluation scores for each image label corresponding to a given image based on the feature proportion of each image label reflected in the corresponding image includes: Each image label is input into a trained evaluation model to obtain the feature proportion of the image features reflected by each image label in the corresponding image, so as to obtain the evaluation score of each image label corresponding to the corresponding image.
3. The optimization method as described in claim 1, characterized in that, The step of optimizing the target image recognition model based on the evaluation results of each of the image recognition models includes: Based on the evaluation results of each image recognition model, the image labels corresponding to the labeled images are corrected; wherein, after the labeled images are input into the target image recognition model, the target image recognition model may have missed recognition or misrecognition phenomena. The target image recognition model is trained a second time based on the labeled image and the corrected image label.
4. The optimization method as described in claim 3, characterized in that, The i image recognition models include at least one target image recognition model and at least one optimized image recognition model. The step of correcting the image labels corresponding to the labeled images based on the evaluation results of each of the image recognition models includes: Based on the evaluation results of each of the image recognition models, at least one missing label corresponding to the labeled image is obtained; the missing label satisfies the following conditions: the image label output by the target image recognition model based on the labeled image does not contain the missing label or a label approximate to the missing label, and the image label output by the optimized image recognition model based on the labeled image contains the missing label and is determined to be the correct label; The at least one missed identification label is added as an image label corresponding to the labeled image.
5. The optimization method as described in claim 4, characterized in that, The missing identification tag also meets the following conditions: the correct tag rate of the missing identification tag is greater than or equal to the preset correct tag rate, and the occurrence frequency of the missing identification tag is greater than or equal to the preset occurrence frequency.
6. The optimization method as described in claim 4, characterized in that, The evaluation results also include cumulative frequency, and the method further includes: For each image recognition model, based on the occurrence frequency of each correct label output by the image recognition model, the cumulative frequency of the correct label output by the image recognition model is obtained, and the cumulative number of correct labels located in a preset cumulative frequency range is obtained. The cumulative number is used to characterize the diversity of the correct labels included in the image recognition model. The step of obtaining at least one missed label corresponding to the labeled image based on the evaluation results of each of the image recognition models includes: When the cumulative number of correct labels output by the target image recognition model is less than or equal to a preset cumulative number, at least one missing label corresponding to the labeled image is obtained based on the evaluation results of each image recognition model.
7. The optimization method as described in claim 2, characterized in that, The i image recognition models include at least one target image recognition model and at least one optimized image recognition model. The optimization of the target image recognition model based on the evaluation results of each of the image recognition models includes: Based on the evaluation results of each image recognition model, at least one misidentified label corresponding to the labeled image is obtained; the misidentified label satisfies the following conditions: the image label output by the optimized image recognition model based on the labeled image contains the misidentified label and the evaluation score is greater than or equal to a preset evaluation score; the image label output by the target image recognition model based on the labeled image does not contain the misidentified label, but contains an approximate label of the misidentified label and the evaluation score of the approximate label is less than or equal to a preset evaluation score. Modify the approximate label of the misidentified label to the misidentified label; The target image recognition model is trained a second time based on the labeled image and the corrected image label.
8. The optimization method according to any one of claims 1 to 7, characterized in that, The method further includes: Based on the meaning of the image labels output by each of the image recognition models, multiple image labels with the same meaning are integrated into the same image label to improve the consistency of the image labels output by each of the image recognition models.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the optimization method as described in any one of claims 1 to 8.