Icon detection method and device, electronic equipment and storage medium
By establishing a first training model for detecting the icon bounding box and a second training model for extracting feature vectors, combined with a vector database, the problems of insufficient icon detection capabilities and low accuracy in the prior art are solved, and efficient and accurate detection of different icon categories are achieved.
Patent Information
- Application Number
- CN202510029772.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, when detecting icons, the detection capability is insufficient and the accuracy is low, especially when new icon categories are encountered.
By establishing two training models: the first training model is used to detect the bounding box of the icon, the second training model is used to extract the feature vectors of the image, and construct a vector database based on these feature vectors to achieve accurate detection of the icon.
It enhances the detection ability of different icon categories, improves the accuracy and efficiency of icon detection, and can effectively handle new icon categories.
Smart Images

Figure CN119992045A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an icon detection method, device, electronic device and storage medium. Background Art
[0002] With the popularization of information technology and smart devices, icon detection, as a computer vision technology, has become a key application technology in many fields. It mainly involves accurately identifying and locating icons from images or user interfaces. These icons may be application symbols, website navigation signs, or any small representative graphics. With the widespread use of mobile devices and smart wearable devices, and the continuous advancement of artificial intelligence technology, the application of icon detection will penetrate deeper into people's daily lives. For example, device control in smart homes, interactive interfaces of in-vehicle infotainment systems, and even medical health monitoring equipment may take advantage of the convenience brought by icon detection. At the same time, with the continuous optimization of machine learning models, especially deep learning, the speed and accuracy of icon detection are expected to be significantly improved. In the future, with the development of edge computing and the Internet of Things, icon detection technology may be integrated into various edge devices to achieve more efficient local processing and instant response.
[0003] At present, when using machine learning models to detect icons, it is often only possible to detect icon categories preset in the machine learning model. When new icon categories appear, the existing machine learning models are insufficient in their detection capabilities for the new icon categories, which can easily lead to inaccurate icon detection. Summary of the invention
[0004] The present application provides an icon detection method, device, electronic device and storage medium to solve the problems of insufficient detection capability and low accuracy in icon detection in the prior art.
[0005] In a first aspect, the present application provides an icon detection method, comprising:
[0006] Establishing a first training model and a second training model, wherein the first training model is used to detect a boundary box of an icon in an image to be detected, and the second training model is used to extract a feature vector of the image;
[0007] Determine a vector database according to the second training model, wherein the vector database stores feature vectors corresponding to each first icon sample in the first icon sample set;
[0008] When a target image to be detected is acquired, detecting a bounding box of an icon in the target image using the first training model to obtain a target bounding box;
[0009] Determine a target area of the target bounding box in the target image;
[0010] Extracting a feature vector of the image in the target area using the second training model to obtain a target feature vector;
[0011] An icon detection result corresponding to the icon in the target image is determined according to the target feature vector and the vector database.
[0012] In an optional embodiment, before executing the step of detecting the bounding box of the icon in the target image using the first training model to obtain the target bounding box when the target image to be detected is acquired, the method further includes:
[0013] When a second icon sample is obtained, the feature vector of the image in the second icon sample is extracted using the second training model to obtain a feature vector corresponding to the second icon sample, wherein the second icon sample is different from each of the first icon samples in the first icon sample;
[0014] Updating the vector database using the feature vector corresponding to the second icon sample to obtain the updated vector database;
[0015] The step of determining the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database includes:
[0016] An icon detection result corresponding to the icon in the target image is determined according to the target feature vector and the updated vector database.
[0017] In an optional implementation, determining the vector database according to the second training model includes:
[0018] Acquire the first icon sample set;
[0019] For each of the first icon samples in the first icon sample set, extracting a feature vector of the image in the first icon sample using the second training model to obtain a feature vector corresponding to the first icon sample;
[0020] A vector database is determined according to feature vectors corresponding to all the first icon samples in the first icon sample set.
[0021] In an optional implementation, determining the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database includes:
[0022] Determine the similarity between the target feature vector and the feature vectors corresponding to each of the first icon samples in the first icon sample set in the vector database to obtain a similarity set;
[0023] Determine a first target similarity from the similarity set, the first target similarity being greater than a preset similarity threshold, the preset similarity threshold representing a lower limit value of the first target similarity corresponding to a feature vector similar to the target feature vector existing in the vector database;
[0024] The first icon sample corresponding to the first target similarity is determined as an icon detection result corresponding to the icon in the target image.
[0025] In an optional implementation, the establishing of the first training model includes:
[0026] Obtain the first icon sample set and the Yolo-V5 model to be trained;
[0027] For each of the first icon samples in the first icon sample set, determine sample annotation information corresponding to the first icon sample to obtain a sample annotation information set, where the sample annotation information includes a bounding box of the first icon sample;
[0028] The Yolo-V5 model to be trained is trained according to the first icon sample set and the sample annotation information set to establish the first training model.
[0029] In an optional implementation, the detecting the bounding box of the icon in the target image by using the first training model to obtain the target bounding box includes:
[0030] Detecting the bounding box of the icon in the target image using the first training model to obtain each candidate bounding box in the target image and a confidence level corresponding to each candidate bounding box;
[0031] Determine, by using the first training model, from among the candidate bounding boxes in the target image, the candidate bounding boxes whose confidences are greater than a preset confidence threshold, wherein the preset confidence threshold represents a lower limit value of the confidence corresponding to the bounding box that meets the icon;
[0032] The candidate bounding box whose confidence is greater than the preset confidence threshold is determined as the target bounding box.
[0033] In an optional implementation, the establishing of the second training model includes:
[0034] Obtain a third icon sample set, positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set, and a Shufflenet-V2 model to be trained, wherein the positive icon samples are icon samples that are identical or similar to the third icon samples, and the negative icon samples are icon samples that are different from the third icon samples;
[0035] For each of the third icon samples in the third icon sample set, construct an icon sample triple according to the third icon sample, the positive icon sample corresponding to the third icon sample, and the negative icon sample corresponding to the third icon sample, so as to obtain an icon sample triple set;
[0036] The Shufflenet-V2 model to be trained is trained according to the icon sample triplet set to establish the second training model.
[0037] In a second aspect, the present application provides an icon detection device, comprising:
[0038] An establishment module, used to establish a first training model and a second training model, wherein the first training model is used to detect a boundary box of an icon in an image to be detected, and the second training model is used to extract a feature vector of the image;
[0039] A determination module, configured to determine a vector database according to the second training model, wherein the vector database stores feature vectors corresponding to each first icon sample in the first icon sample set;
[0040] A detection module, configured to detect a bounding box of an icon in a target image to be detected by using the first training model when the target image to be detected is acquired, so as to obtain a target bounding box;
[0041] The determination module is further used to determine a target area of the target bounding box in the target image;
[0042] The detection module is further used to extract the feature vector of the image in the target area by using the second training model to obtain a target feature vector;
[0043] The determination module is further used to determine the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database.
[0044] In a third aspect, the present application provides an electronic device, comprising: a processor and a memory, wherein the processor is used to execute an icon detection program stored in the memory to implement the icon detection method described above.
[0045] In a fourth aspect, the present application provides a storage medium storing one or more programs, wherein the one or more programs can be executed by one or more processors to implement the icon detection method as described above.
[0046] The above-mentioned technical scheme provided by the embodiment of the present application has the following advantages compared with the prior art. The icon detection method provided by the embodiment of the present application includes establishing a first training model and establishing a second training model. The first training model is used to detect the bounding box of the icon in the image to be detected, and the second training model is used to extract the feature vector of the image; according to the second training model, a vector database is determined, and the vector database stores the feature vectors corresponding to each first icon sample in the first icon sample set; when the target image to be detected is obtained, the bounding box of the icon in the target image is detected using the first training model to obtain the target bounding box; the target area of the target bounding box in the target image is determined; the feature vector of the image in the target area is extracted using the second training model to obtain the target feature vector; according to the target feature vector and the vector database, the icon detection result corresponding to the icon in the target image is determined. Through the above method, the present application establishes a first training model for detecting the boundary box of the icon in the image to be detected and a second training model for extracting the feature vector of the image, and determines a vector database for determining the icon detection result based on the established second training model. When the target image to be detected is obtained, the first training model is used to identify and locate the icon in the target image, and the feature vector of the image of the located icon is extracted using the second training model, so as to realize icon detection according to the extracted feature vector and the vector database, so that the present application enhances the detection capability of different icon categories through the combination of the first training model and the second training model, and improves the accuracy and efficiency of icon detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0049] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0050] Figure 1 A flowchart of an icon detection method provided in an embodiment of the present application;
[0051] Figure 2 A flowchart of another icon detection method provided in an embodiment of the present application;
[0052] Figure 3 A flowchart of another icon detection method provided in an embodiment of the present application;
[0053] Figure 4 A schematic diagram of a first training model and a second training model processing process provided in an embodiment of the present application;
[0054] Figure 5 A schematic diagram of the structure of an icon detection device provided in an embodiment of the present application;
[0055] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0056] In the above attached figure:
[0057] 501, establishment module; 502, determination module; 503, detection module;
[0058] 600, electronic device; 601, processor; 602, memory; 6021, operating system; 6022, application; 603, user interface; 604, network interface; 605, bus system. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0060] The disclosure below provides many different embodiments or examples to implement different structures of the present invention. In order to simplify the disclosure of the present invention, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present invention. In addition, the present invention can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0061] refer to Figure 1 , Figure 1 A flowchart of an icon detection method provided in an embodiment of the present application is shown below. An icon detection method provided in this embodiment includes the following steps:
[0062] S101: Establish a first training model and establish a second training model.
[0063] In this embodiment, the first training model is used to detect the bounding box of the icon in the image to be detected, and the second training model is used to extract the feature vector of the image. The first training model is obtained by training the Yolo-V5 model, and the second training model is obtained by training the Shufflenet-V2 model.
[0064] Among them, when the target image to be detected is obtained, the first training model can be used to identify and locate the icon in the target image to obtain the target bounding box of the icon in the target image, and the area where the target bounding box is located in the target image, and the second training model can be used to extract the feature vector of the image in the located area to obtain the target feature vector, so that the detection result of the icon in the target image can be determined according to the target feature vector, thereby improving the accuracy and efficiency of icon detection and enhancing the adaptability to different icon categories.
[0065] S102: Determine a vector database according to the second training model.
[0066] In this embodiment, the vector database stores feature vectors corresponding to each first icon sample in the first icon sample set. The vector database is used to determine the icon detection result in the image to be detected. Each first icon sample in the first icon sample set needs to cover various shapes and styles. Each first icon sample in the first icon sample set can be selected according to actual needs. In this embodiment, the specific form of each first icon sample in the first icon sample set is not limited.
[0067] Among them, after the second training model is established, since the second training model can extract the feature vector of the image, the second training model can be used to extract the feature vector of the image of each first icon sample in the first icon sample set to obtain the feature vector corresponding to each first icon sample in the first icon sample set, so as to establish a final vector database based on the feature vector corresponding to each first icon sample in the first icon sample set. After determining that the vector database is obtained, the image in the boundary box of the icon in the image detected by the first training model and the vector database can be used to realize the detection of the icon in the image.
[0068] S103: When a target image to be detected is obtained, a bounding box of an icon in the target image is detected using the first training model to obtain a target bounding box.
[0069] In this embodiment, the target image to be detected can be uploaded by the user based on the client. When the target image to be detected is obtained, the bounding box of the icon in the target image is detected using the first training model. When the bounding boxes of multiple icons are detected from the target image using the first training model, a bounding box that meets the conditions is selected from the bounding boxes of the multiple icons, and the bounding box that meets the conditions is determined as the target bounding box.
[0070] The target bounding box may be zero, one, or at least two. When the target bounding box is not determined, return to re-execute step S103. When one target bounding box or at least two bounding boxes are obtained, the subsequent step S104 may be executed to obtain the final icon detection result corresponding to the icon.
[0071] It should be noted that when the target video to be detected is obtained, it is necessary to perform frame processing on the target video to determine the video key frame from the target video, and determine the video key frame as the target image to be detected.
[0072] S104: Determine a target area of the target bounding box in the target image.
[0073] In this embodiment, after obtaining the target bounding box shaped like an icon in the target image according to step S103, the sub-image (ie, the target area) where the target bounding box shaped like an icon is located in the target image can be determined.
[0074] S105: Extracting a feature vector of the image in the target area using the second training model to obtain a target feature vector.
[0075] In this embodiment, after obtaining the target area of the target bounding box in the target image, the second training model can be used to extract the feature vector of the image in the target area, so as to retrieve the extracted target feature vector in the vector database, thereby obtaining the icon detection result corresponding to the icon in the target image, improving the detection capability of icons of different categories, and improving the accuracy of icon detection.
[0076] S106: Determine the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database.
[0077] In this embodiment, after obtaining the target feature vector, the target feature vector can be searched in the vector database to determine a feature vector similar to the target feature vector from the vector database. Since the feature vectors corresponding to the first icon samples in the first icon sample set are stored in the vector database, after determining the feature vector similar to the target feature vector from the vector database, the icon detection result corresponding to the icon in the target image can be determined based on the feature vector similar to the target feature vector, and the icon detection result is the first icon sample corresponding to the feature vector similar to the target feature vector.
[0078] The present embodiment provides an icon detection method, which establishes a first training model for detecting the boundary box of the icon in the image to be detected and a second training model for extracting the feature vector of the image, and determines a vector database for determining the icon detection result based on the established second training model. When the target image to be detected is obtained, the first training model is used to identify and locate the icon in the target image, and the second training model is used to extract the feature vector of the image of the located icon, so as to realize icon detection according to the extracted feature vector and the vector database, so that the present application enhances the detection capability of different icon categories through the combination of the first training model and the second training model, and improves the accuracy and efficiency of icon detection.
[0079] refer to Figure 2 , Figure 2 A flowchart of another icon detection method provided in an embodiment of the present application is shown below. An icon detection method provided in this embodiment includes the following steps:
[0080] S201: Establish a first training model and establish a second training model.
[0081] In this embodiment, the first training model is used to detect the boundary box of the icon in the image to be detected, and the second training model is used to extract the feature vector of the image.
[0082] In this embodiment, establishing the first training model in step S201 includes:
[0083] Obtain the first icon sample set and the Yolo-V5 model to be trained;
[0084] For each first icon sample in the first icon sample set, determine the sample annotation information corresponding to the first icon sample to obtain a sample annotation information set;
[0085] The Yolo-V5 model to be trained is trained according to the first icon sample set and the sample annotation information set to establish a first training model.
[0086] The sample annotation information includes the bounding box of the first icon sample. The first icon sample set can refer to the above description, which will not be described in detail in this embodiment.
[0087] In this embodiment, when labeling each first icon sample in the first icon sample set, the bounding box and category labeling can be performed manually. After the manual labeling, the sample labeling information corresponding to each first icon sample in the first icon sample set can be uploaded through the client, so that when the first training model is established, the sample labeling information corresponding to each first icon sample in the first icon sample set can be determined, thereby obtaining the sample labeling information set corresponding to the first icon sample set. After obtaining the sample labeling information set corresponding to the first icon sample set, each first icon sample in the first icon sample set can be preprocessed, such as adjusting the size of the first icon sample, normalizing, etc., to obtain the processed first icon sample set, thereby facilitating the training of the model. After the preprocessing of the first icon sample set is completed, the processed first icon sample set and the sample labeling information can be used to train the Yolo-V5 model to be trained. The goal of this stage is to enable the trained model to detect the bounding box of the icon in the image, rather than the preset category of a specific icon. During the training process of the Yolo-V5 model to be trained, if the model converges, the Yolo-V5 model to be trained is no longer trained, so that the model parameters are frozen to establish the first training model.
[0088] In this embodiment, the step S201 of establishing the second training model includes:
[0089] Obtain a third icon sample set, positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set, and a Shufflenet-V2 model to be trained;
[0090] For each third icon sample in the third icon sample set, construct an icon sample triple according to the third icon sample, the positive icon sample corresponding to the third icon sample, and the negative icon sample corresponding to the third icon sample to obtain an icon sample triple set;
[0091] According to the icon sample triplet set, the Shufflenet-V2 model to be trained is trained to establish a second training model.
[0092] Among them, the positive icon sample is an icon sample that is the same as or similar to the third icon sample, and the negative icon sample is an icon sample that is different from the third icon sample. When constructing the icon sample triplet, a preset loss function is used, and the preset loss function is a Triplet Loss loss function. The goal of the preset loss function is to minimize the distance between the third icon sample (i.e., the anchor sample) and the positive icon sample, and to maximize the distance between the third icon sample (i.e., the anchor sample) and the negative icon sample.
[0093] Specifically, the third icon sample set and the positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set can be manually selected. After the selection, the third icon sample set and the positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set are uploaded through the client, so as to obtain the third icon sample set and the positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set when establishing the second training model. After obtaining the third icon sample set and the positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set, the above-mentioned icon samples can be enhanced and preprocessed, and based on the processed icon samples, a plurality of icon sample triplets are constructed using a preset loss function to obtain an icon sample triple set. The Shufflenet-V2 model to be trained is trained using the icon sample triple set until the model converges to establish the second training model. In each iteration, an icon sample triple is randomly selected from the icon sample triple set, the feature representation is calculated by network forward propagation, and then the loss is calculated and the network weight is updated by back propagation. Finally, a linear layer is added to the last layer of the Shufflenet-V2 model to convert the high-dimensional feature map into a 512-dimensional one-dimensional vector. This 512-dimensional output vector is used to represent the features of the input image.
[0094] More specifically, the schematic diagram of using the first training model to detect the bounding box of the icon in the image to obtain the bounding box and the schematic diagram of using the second training model to extract the feature vector of the image in the bounding box to obtain the feature vector can be referred to Figure 4 shown.
[0095] S202: Determine a vector database according to the second training model.
[0096] In this embodiment, the vector database stores feature vectors corresponding to each first icon sample in the first icon sample set. In step S202, the vector database is determined according to the second training model, including:
[0097] Obtain a first icon sample set;
[0098] For each first icon sample in the first icon sample set, extract a feature vector of the image in the first icon sample using the second training model to obtain a feature vector corresponding to the first icon sample;
[0099] A vector database is determined according to feature vectors corresponding to all first icon samples in the first icon sample set.
[0100] Among them, the first icon sample set can refer to the above description, and this embodiment does not elaborate on the first icon sample set. In this embodiment, in order to determine the icon detection result corresponding to the icon in the image based on the vector database, the second training model is used to extract the feature vectors of the images in each first icon sample in the first icon sample set, so as to obtain the feature vectors corresponding to each first icon sample in the first icon sample set, and then use the feature vectors corresponding to each first icon sample in the first icon sample set as the vector database. When the feature vector of the icon-like shape in the image to be detected is obtained, the feature vector of the icon-like shape is detected in the vector database, so as to determine the icon detection result corresponding to the icon in the image to be detected.
[0101] S203: When the second icon sample is obtained, the feature vector of the image in the second icon sample is extracted using the second training model to obtain a feature vector corresponding to the second icon sample.
[0102] S204: updating the vector database using the feature vector corresponding to the second icon sample to obtain an updated vector database.
[0103] With respect to the above-mentioned steps S203 and S204, the second icon sample is different from each first icon sample in the first icon sample. When a new icon category is added, the traditional icon detection method needs to retrain or fine-tune the machine learning model, which affects the icon detection efficiency. In order to solve the above-mentioned problem and improve the icon detection efficiency in this embodiment, when a new icon category is added, it is only necessary to use the second training model to extract the feature vector of the image in the second icon sample (i.e., the icon of the newly added icon category), so as to obtain the feature vector corresponding to the second icon sample, and store the feature vector corresponding to the second icon sample in the vector database to complete the update of the vector database, thereby obtaining the updated vector database. After obtaining the updated vector database, the updated vector database can be used to realize the detection of the newly added icons, without the need for tedious data collection and model training processes. Only an icon sample of a newly added icon category needs to be provided to realize the detection of the newly added icons, thereby improving the efficiency of icon detection.
[0104] It should be noted that after the feature vector corresponding to the second icon sample is stored in the vector database, the second icon sample is synchronously updated to the first icon sample in the updated vector data.
[0105] S205: When the target image to be detected is acquired, the bounding box of the icon in the target image is detected using the first training model to obtain a target bounding box.
[0106] In this embodiment, in step S205, the first training model is used to detect the bounding box of the icon in the target image to obtain the target bounding box, including:
[0107] Detecting the bounding box of the icon in the target image using the first training model to obtain each candidate bounding box in the target image and the confidence corresponding to each candidate bounding box;
[0108] Determine, using the first training model, from among the candidate bounding boxes in the target image, a candidate bounding box whose confidence is greater than a preset confidence threshold;
[0109] The candidate bounding box whose confidence is greater than a preset confidence threshold is determined as the target bounding box.
[0110] The preset confidence threshold represents the lower limit of the confidence corresponding to the bounding box of the icon, and the preset confidence threshold can be set according to actual needs. In this embodiment, the specific value of the preset confidence threshold is not limited. In this embodiment, the preset confidence threshold is set, so that after the candidate bounding boxes of each icon are determined from the target image using the first training model, the candidate bounding boxes of each icon can be filtered using the preset confidence threshold to reduce the amount of subsequent data processing.
[0111] S206: Determine a target area of the target bounding box in the target image.
[0112] S207: Extracting a feature vector of the image in the target area using the second training model to obtain a target feature vector.
[0113] Regarding the above-mentioned step S206 and step S207, step S206 is consistent with the above-mentioned step S104, and step S207 is consistent with the above-mentioned step S105. For details, please refer to the above-mentioned step S104 and step S105, which will not be described in detail in this embodiment.
[0114] S208: Determine the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database.
[0115] In this embodiment, determining the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database in step S208 includes:
[0116] Determine the similarity between the target feature vector and the feature vectors corresponding to each first icon sample in the first icon sample set in the vector database to obtain a similarity set;
[0117] Determine a first target similarity from the similarity set, the first target similarity being greater than a preset similarity threshold;
[0118] The first icon sample corresponding to the first target similarity is determined as the icon detection result corresponding to the icon in the target image.
[0119] Among them, the preset similarity threshold represents the lower limit value of the first target similarity corresponding to the feature vector similar to the target feature vector in the vector database. The preset similarity threshold can be set according to actual needs. In this embodiment, the specific value of the preset similarity threshold is not limited. When determining the similarity between the target feature vector and the feature vector corresponding to the first icon sample, the cosine similarity between the target feature vector and the feature vector corresponding to the first icon sample can be determined. After obtaining the cosine similarity between the target feature vector and the feature vector corresponding to each first icon sample in the first icon sample set in the vector database, all cosine similarities greater than the preset similarity threshold are determined from all the determined cosine similarities, and the first icon samples corresponding to each cosine similarity greater than the preset similarity threshold are determined from the vector database. The first icon sample determined from the vector database is determined as the icon detection result corresponding to the icon in the target image, thereby realizing the detection of the icon.
[0120] It should be noted that after obtaining the updated vector database, the icon detection result corresponding to the icon in the target image is determined according to the target feature vector and the updated vector database. The vector data in the above can be replaced with the updated vector database to obtain the icon detection result corresponding to the icon in the target image.
[0121] The present embodiment provides an icon detection method, which establishes a first training model for detecting the boundary box of the icon in the image to be detected and a second training model for extracting the feature vector of the image, and determines a vector database for determining the icon detection result based on the established second training model. When the target image to be detected is obtained, the first training model is used to identify and locate the icon in the target image, and the second training model is used to extract the feature vector of the image of the located icon, so as to realize icon detection according to the extracted feature vector and the vector database, so that the present application enhances the detection capability of different icon categories through the combination of the first training model and the second training model, and improves the accuracy and efficiency of icon detection.
[0122] As an example, refer to Figure 3 , let's take a closer look at the entire icon detection process, as follows:
[0123] 1. Data preprocessing step: This step mainly preprocesses the input data source for subsequent input. There are three main types of input data sources: videos to be detected, images to be detected, and newly added icon images.
[0124] 1) For the input video to be detected, the data preprocessing steps are as follows:
[0125] Step 1: Cut the input video into frames in sequence;
[0126] Step 2: Perform differential calculation on the video frames to extract the video key frames;
[0127] Step 3: Scale the extracted video key frames to 224*224 and normalize them.
[0128] 2) For input images to be detected or newly added icon images, the data preprocessing steps are as follows:
[0129] The input image size is scaled to 224*224 and normalized.
[0130] 2. Icon candidate bounding box detection steps
[0131] This step mainly inputs the video key frame or the image to be detected after data preprocessing into the first training model to determine the target bounding box. The processing steps are as follows:
[0132] Step 1: Collect a large number of icons of different categories and mark them;
[0133] Step 2: Input the annotated icon in step 1 into the Yolo-V5 model for training. After the model training converges, freeze the model parameters to obtain the first training model.
[0134] Step 3: Input the image data obtained in the data preprocessing step into the first training model in step 2 for inference, and output all detected candidate bounding boxes;
[0135] Step 4: According to the preset confidence threshold, all candidate bounding boxes in step 3 are filtered and the target bounding boxes that meet the threshold condition are output;
[0136] Step 5: Correspond the target bounding box in step 4 to the pixel points of the original image, and output the target area (i.e., sub-image) corresponding to the target bounding box in the original image.
[0137] 3. Feature vector extraction steps
[0138] This step mainly inputs the sub-image corresponding to the target bounding box output in the icon candidate bounding box detection step into the second training model, and outputs a 512-dimensional feature vector containing high-dimensional information. The processing steps are as follows:
[0139] Step 1: Collect a large number of icon samples that are identical or similar to the icon sample as positive icon samples, and different icon samples as negative icon samples;
[0140] Step 2: Enhance the data collected in step 1. The enhancement methods include random cropping, mirror flipping, rotation, random brightness adjustment, random contrast adjustment, Gaussian blur, scaling, and random addition of interference lines.
[0141] Step 3: Based on the enhanced icon samples, positive icon samples and negative icon samples, the loss function Triplet Loss is used to construct multiple icon sample triplets. The Shufflenet-V2 model is iteratively trained using multiple icon sample triplets until the model converges. Finally, the model outputs a one-dimensional vector of 512 dimensions after the linear layer.
[0142] 4. Vector database update steps
[0143] This step mainly inputs the newly added icon images manually configured into the second training model, and dynamically adds the 512-dimensional feature vector of the high-dimensional information output by the model to the vector database to complete the update of the vector database for use in the retrieval stage. The processing steps are as follows:
[0144] Step 1: preprocess the manually configured icon image;
[0145] Step 2: Input the preprocessed icon image into the second training model for inference, and output a 512-dimensional feature vector;
[0146] Step 3: Store the feature vector corresponding to the configured icon image into the vector database to complete the update of the vector database.
[0147] 5. Retrieval steps
[0148] This step mainly searches the vector database for the feature vectors corresponding to the video frame to be detected or the image to be detected that has been inferred by the second training model, and outputs the retrieved icons that meet the similarity threshold condition as the result. The processing steps are as follows:
[0149] Step 1: Retrieve the feature vectors corresponding to the video frame to be detected or the image to be detected inferred by the second training model in the vector database, and calculate the similarity between the feature vectors by cosine similarity during the retrieval;
[0150] Step 2: According to the manually set similarity threshold, the similarity search results in step 1 are filtered, and the icons corresponding to the similarities that meet the threshold conditions are output.
[0151] refer to Figure 5 , Figure 5 A schematic diagram of the structure of an icon detection device provided in an embodiment of the present application. An icon detection device provided in an embodiment of the present application includes: an establishment module 501, a determination module 502 and a detection module 503. Among them, the establishment module 501 is used to establish a first training model and a second training model, the first training model is used to detect the bounding box of the icon in the image to be detected, and the second training model is used to extract the feature vector of the image; the determination module 502 is used to determine a vector database according to the second training model, and the vector database stores the feature vectors corresponding to each first icon sample in the first icon sample set; the detection module 503 is used to detect the bounding box of the icon in the target image using the first training model when the target image to be detected is obtained, so as to obtain the target bounding box; the determination module 502 is also used to determine the target area of the target bounding box in the target image; the detection module 503 is also used to extract the feature vector of the image in the target area using the second training model to obtain the target feature vector; the determination module 502 is also used to determine the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database.
[0152] The icon detection device provided in this embodiment further includes an updating module, which is used to:
[0153] When a second icon sample is obtained, the feature vector of the image in the second icon sample is extracted using the second training model to obtain a feature vector corresponding to the second icon sample, wherein the second icon sample is different from each of the first icon samples in the first icon sample;
[0154] The vector database is updated using the feature vector corresponding to the second icon sample to obtain the updated vector database
[0155] In this embodiment, the determination module 502 is further configured to:
[0156] An icon detection result corresponding to the icon in the target image is determined according to the target feature vector and the updated vector database.
[0157] In this embodiment, the determination module 502 is further configured to:
[0158] Acquire the first icon sample set;
[0159] For each of the first icon samples in the first icon sample set, extracting a feature vector of the image in the first icon sample using the second training model to obtain a feature vector corresponding to the first icon sample;
[0160] A vector database is determined according to feature vectors corresponding to all the first icon samples in the first icon sample set.
[0161] In this embodiment, the determination module 502 is further configured to:
[0162] Determine the similarity between the target feature vector and the feature vectors corresponding to each of the first icon samples in the first icon sample set in the quantity database to obtain a similarity set;
[0163] Determine a first target similarity from the similarity set, the first target similarity being greater than a preset similarity threshold, the preset similarity threshold representing a lower limit value of the first target similarity corresponding to a feature vector similar to the target feature vector existing in the vector database;
[0164] The first icon sample corresponding to the first target similarity is determined as an icon detection result corresponding to the icon in the target image.
[0165] In this embodiment, the establishing module 501 is further used for:
[0166] Obtain the first icon sample set and the Yolo-V5 model to be trained;
[0167] For each of the first icon samples in the first icon sample set, determine sample annotation information corresponding to the first icon sample to obtain a sample annotation information set, where the sample annotation information includes a bounding box of the first icon sample;
[0168] The Yolo-V5 model to be trained is trained according to the first icon sample set and the sample annotation information set to establish the first training model.
[0169] In this embodiment, the detection module 503 is further used for:
[0170] Detecting the bounding box of the icon in the target image using the first training model to obtain each candidate bounding box in the target image and a confidence level corresponding to each candidate bounding box;
[0171] Determine, by using the first training model, from among the candidate bounding boxes in the target image, the candidate bounding boxes whose confidences are greater than a preset confidence threshold, wherein the preset confidence threshold represents a lower limit value of the confidence corresponding to the bounding box that meets the icon;
[0172] The candidate bounding box whose confidence is greater than the preset confidence threshold is determined as the target bounding box.
[0173] In this embodiment, the establishing module 501 is further used for:
[0174] Obtain a third icon sample set, positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set, and a Shufflenet-V2 model to be trained, wherein the positive icon samples are icon samples that are identical or similar to the third icon samples, and the negative icon samples are icon samples that are different from the third icon samples;
[0175] For each of the third icon samples in the third icon sample set, construct an icon sample triple according to the third icon sample, the positive icon sample corresponding to the third icon sample, and the negative icon sample corresponding to the third icon sample, so as to obtain an icon sample triple set;
[0176] The Shufflenet-V2 model to be trained is trained according to the icon sample triplet set to establish the second training model.
[0177] The present embodiment provides an icon detection device, which establishes a first training model for detecting the boundary box of an icon in an image to be detected and a second training model for extracting the feature vector of the image, and determines a vector database for determining the icon detection result based on the established second training model. When a target image to be detected is acquired, the first training model is used to identify and locate the icon in the target image, and the feature vector of the image of the located icon is extracted using the second training model, so as to realize icon detection based on the extracted feature vector and the vector database, thereby enabling the present application to enhance the detection capability of different icon categories and improve the accuracy and efficiency of icon detection through the combination of the first training model and the second training model.
[0178] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application, Figure 6 The electronic device 600 shown includes: at least one processor 601, a memory 602, at least one network interface 604 and other user interfaces 603. The various components in the electronic device 600 are coupled together via a bus system 605. It is understood that the bus system 605 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 605 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 605 is not shown in FIG. Figure 6 Various buses are labeled as bus system 605 .
[0179] The user interface 603 may include a display, a keyboard, or a pointing device (eg, a mouse, a trackball, a touch pad, or a touch screen).
[0180] It can be understood that the memory 602 in the embodiment of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 602 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0181] In some implementations, the memory 602 stores the following elements, executable units or data structures, or a subset thereof, or an extended set thereof: an operating system 6021 and an application 6022 .
[0182] The operating system 6021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application 6022 includes various application programs, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. The program for implementing the method of the embodiment of the present invention can be included in the application 6022.
[0183] In the embodiment of the present invention, by calling the program or instruction stored in the memory 602, specifically, the program or instruction stored in the application 6022, the processor 601 is used to execute the method steps provided by each method embodiment.
[0184] The method disclosed in the above embodiment of the present invention can be applied to the processor 601, or implemented by the processor 601. The processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 601 or the instruction in the form of software. The above processor 601 can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as a hardware decoding processor to execute, or the hardware and software units in the decoding processor can be executed. The software unit can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 602, and the processor 601 reads the information in the memory 602 and completes the steps of the above method in combination with its hardware.
[0185] It is understood that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPDevice, DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application, or a combination thereof.
[0186] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0187] The electronic device provided in this embodiment may be Figure 6 The electronic device shown in FIG. 1 may perform the following steps: Figure 1 to Figure 3 All steps of the icon detection method in order to achieve Figure 1 to Figure 3 For details, please refer to the technical effect of the icon detection method shown in Figure 1 to Figure 3 For the sake of brevity, the relevant description is not repeated here.
[0188] The embodiment of the present invention also provides a storage medium (computer-readable storage medium). The storage medium here stores one or more programs. The storage medium may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a read-only memory, a flash memory, a hard disk or a solid-state drive; the memory may also include a combination of the above-mentioned types of memory.
[0189] When one or more programs in the storage medium can be executed by one or more processors, the icon detection method executed on the icon detection device side can be implemented.
[0190] The processor is used to execute the icon detection program stored in the memory to implement the following steps of the icon detection method executed on the icon detection device side.
[0191] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0192] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0193] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An icon detection method, characterized in that: include: Establishing a first training model and a second training model, wherein the first training model is used to detect a boundary box of an icon in an image to be detected, and the second training model is used to extract a feature vector of the image; Determine a vector database according to the second training model, wherein the vector database stores feature vectors corresponding to each first icon sample in the first icon sample set; When a target image to be detected is acquired, detecting a bounding box of an icon in the target image using the first training model to obtain a target bounding box; Determine a target area of the target bounding box in the target image; Extracting a feature vector of the image in the target area using the second training model to obtain a target feature vector; An icon detection result corresponding to the icon in the target image is determined according to the target feature vector and the vector database.
2. The method according to claim 1, characterized in that Before executing the step of detecting the bounding box of the icon in the target image by using the first training model to obtain the target bounding box when the target image to be detected is acquired, the method further includes: When a second icon sample is obtained, the feature vector of the image in the second icon sample is extracted using the second training model to obtain a feature vector corresponding to the second icon sample, wherein the second icon sample is different from each of the first icon samples in the first icon sample; Updating the vector database using the feature vector corresponding to the second icon sample to obtain the updated vector database; The step of determining the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database includes: An icon detection result corresponding to the icon in the target image is determined according to the target feature vector and the updated vector database.
3. The method according to claim 1, characterized in that Determining a vector database according to the second training model includes: Acquire the first icon sample set; For each of the first icon samples in the first icon sample set, extracting a feature vector of the image in the first icon sample using the second training model to obtain a feature vector corresponding to the first icon sample; A vector database is determined according to feature vectors corresponding to all the first icon samples in the first icon sample set.
4. The method according to claim 1, characterized in that The step of determining the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database includes: Determine the similarity between the target feature vector and the feature vectors corresponding to each of the first icon samples in the first icon sample set in the vector database to obtain a similarity set; Determine a first target similarity from the similarity set, the first target similarity being greater than a preset similarity threshold, the preset similarity threshold representing a lower limit value of the first target similarity corresponding to a feature vector similar to the target feature vector existing in the vector database; The first icon sample corresponding to the first target similarity is determined as an icon detection result corresponding to the icon in the target image.
5. The method according to claim 1, characterized in that The establishing of the first training model comprises: Obtain the first icon sample set and the Yolo-V5 model to be trained; For each of the first icon samples in the first icon sample set, determine sample annotation information corresponding to the first icon sample to obtain a sample annotation information set, where the sample annotation information includes a bounding box of the first icon sample; The Yolo-V5 model to be trained is trained according to the first icon sample set and the sample annotation information set to establish the first training model.
6. The method according to claim 5, characterized in that The detecting the bounding box of the icon in the target image by using the first training model to obtain the target bounding box includes: Detecting the bounding box of the icon in the target image using the first training model to obtain each candidate bounding box in the target image and a confidence level corresponding to each candidate bounding box; Determine, by using the first training model, from among the candidate bounding boxes in the target image, the candidate bounding boxes whose confidences are greater than a preset confidence threshold, wherein the preset confidence threshold represents a lower limit value of the confidence corresponding to the bounding box that meets the icon; The candidate bounding box whose confidence is greater than the preset confidence threshold is determined as the target bounding box.
7. The method according to claim 1, characterized in that The establishing of the second training model comprises: Obtain a third icon sample set, positive icon samples and negative icon samples corresponding to each third icon sample in the third icon sample set, and a Shufflenet-V2 model to be trained, wherein the positive icon samples are icon samples that are identical or similar to the third icon samples, and the negative icon samples are icon samples that are different from the third icon samples; For each of the third icon samples in the third icon sample set, construct an icon sample triple according to the third icon sample, the positive icon sample corresponding to the third icon sample, and the negative icon sample corresponding to the third icon sample, so as to obtain an icon sample triple set; The Shufflenet-V2 model to be trained is trained according to the icon sample triplet set to establish the second training model.
8. An icon detection device, characterized in that: include: An establishment module, used to establish a first training model and a second training model, wherein the first training model is used to detect a boundary box of an icon in an image to be detected, and the second training model is used to extract a feature vector of the image; A determination module, configured to determine a vector database according to the second training model, wherein the vector database stores feature vectors corresponding to each first icon sample in the first icon sample set; A detection module, configured to detect a bounding box of an icon in a target image to be detected by using the first training model when the target image to be detected is acquired, so as to obtain a target bounding box; The determination module is further used to determine a target area of the target bounding box in the target image; The detection module is further used to extract the feature vector of the image in the target area by using the second training model to obtain a target feature vector; The determination module is further used to determine the icon detection result corresponding to the icon in the target image according to the target feature vector and the vector database.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is used to execute an icon detection program stored in the memory to implement the icon detection method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the icon detection method according to any one of claims 1 to 7.
Citation Information
Cited By
High-quality cargo sample image database construction method
CN120763353A