A self-learning-based category detection method, apparatus, and device

By employing a self-learning category detection method, the system automatically identifies product categories in the merchandise cabinet using detection and feature extraction models. This solves the problem of automatic detection in existing technologies, enabling highly accurate unmanned management and expanding self-learning capabilities.

CN115861633BActive Publication Date: 2026-05-26HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2022-12-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technology cannot effectively and automatically detect the categories of goods already placed in the display cabinet, making it difficult for brand owners to manage product placement and harming their interests.

Method used

A self-learning-based category detection method is adopted. The location and features of the target object are determined by the first detection model and the first feature extraction model. The similarity judgment and open set score calculation are performed by combining the feature set. The category of the product is automatically detected, and the model and feature set are updated under the self-learning condition.

Benefits of technology

It enables automatic detection of product categories in the merchandise cabinet, improves detection accuracy, achieves unmanned management, saves human resources, and has self-learning capabilities, allowing it to automatically expand and update the feature set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115861633B_ABST
    Figure CN115861633B_ABST
Patent Text Reader

Abstract

This application provides a self-learning-based category detection method, apparatus, and device. The method includes: inputting a target image into a first detection model to obtain the location information of a target object in the target image; extracting a sub-image corresponding to the location information from the target image; inputting the sub-image into a first feature extraction model to obtain the target feature corresponding to the target object; determining the similarity between the target feature and each registered feature in the feature set; if the highest similarity is less than a first similarity threshold, determining the open set score value corresponding to the target object based on the target feature and each registered feature in the feature set; and determining the target category of the target object as the labeled category corresponding to the registered feature corresponding to the highest similarity, or determining the target category as an unregistered category, based on the open set score value. Through the technical solution of this application, the target category of a target object can be automatically detected, improving the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular to a self-learning-based category detection method, apparatus, and device. Background Technology

[0002] In shopping malls, supermarkets and other similar places, merchandise cabinets (also known as vending machines, which can take the form of freezers, beverage cabinets, display cases, vending machines, etc.) are usually deployed. Various types of goods can be placed in the merchandise cabinets, and buyers can automatically purchase the goods in the merchandise cabinets without the need for the seller's intervention.

[0003] In some cases, brand owners will provide merchandise counters to shopping malls, supermarkets and other venues. For this type of merchandise counter, a specific type of merchandise (i.e., the brand owner's merchandise) must be placed in the counter, and other types of merchandise should not be placed, or only a small number of other types of merchandise are allowed to be placed.

[0004] Due to the sheer number of shopping malls and supermarkets, coupled with personnel costs, brand owners find it difficult to physically inspect whether designated product types are displayed in their display cases, thus harming their interests. Therefore, if the categories of products already displayed in a display case could be automatically detected, the usage of these cases could be managed, ensuring that sellers place the designated product types as required, thus protecting the interests of brand owners. However, currently, there is no reasonable way to automatically detect the categories of products already displayed in a display case. Summary of the Invention

[0005] This application provides a self-learning-based category detection method, the method comprising:

[0006] The acquired target image is input into the first detection model to obtain the location information corresponding to the target object in the target image. A sub-image corresponding to the location information is extracted from the target image, and the sub-image is input into the first feature extraction model to obtain the target features corresponding to the target object.

[0007] Determine the similarity between the target feature and each registered feature in the configured feature set;

[0008] If the highest similarity is not less than the first similarity threshold, then the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity.

[0009] If the highest similarity is less than the first similarity threshold, then based on the target feature and each registered feature in the feature set, the open set score corresponding to the target object is determined;

[0010] Based on the open set score, the target category of the target object is determined to be the labeled category corresponding to the registered feature with the highest similarity, or the target category is determined to be an unregistered category.

[0011] This application provides a self-learning-based category detection device, the device comprising:

[0012] The acquisition module is used to input the target image into the first detection model to obtain the location information corresponding to the target object in the target image, extract the sub-image corresponding to the location information from the target image, and input the sub-image into the first feature extraction model to obtain the target features corresponding to the target object;

[0013] The determination module is used to determine the similarity between the target feature and each registered feature in the feature set; if the highest similarity is not less than a first similarity threshold, the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity; if the highest similarity is less than the first similarity threshold, the open set score corresponding to the target object is determined based on the target feature and each registered feature in the feature set; the target category of the target object is determined based on the open set score to be the labeled category corresponding to the registered feature corresponding to the highest similarity, or the target category is determined to be an unregistered category.

[0014] This application provides an electronic device, including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; wherein, the processor is configured to execute the machine-executable instructions to implement the self-learning-based category detection method of the above example.

[0015] As can be seen from the above technical solutions, in this embodiment, the target category of the target object can be automatically detected. If the target object is a product already displayed in a merchandise cabinet, the target category of the product already displayed in the merchandise cabinet can be automatically detected, resulting in higher detection accuracy. It can utilize intelligent algorithms and image processing technology to achieve automatic detection of the target category of the target object, realizing unmanned management, improving detection accuracy, and saving human resources. A recognition system with open set recognition capability is proposed, possessing self-learning automatic iterative optimization capability. It can automatically expand the calibrated categories, automatically update the registered features in the expanded feature set, and automatically update the first detection model and the first feature extraction model, achieving high self-learning accuracy. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this application.

[0017] Figure 1 This is a flowchart illustrating a self-learning-based category detection method in one embodiment of this application.

[0018] Figure 2 This is a schematic diagram of the structure of an object recognition system according to one embodiment of this application;

[0019] Figure 3 This is a schematic diagram of an object category identification process in one embodiment of this application;

[0020] Figure 4 This is a schematic diagram of an automatic training process in one embodiment of this application;

[0021] Figure 5 This is a schematic diagram of the structure of a self-learning-based category detection device according to one embodiment of this application;

[0022] Figure 6 This is a hardware structure diagram of an electronic device according to one embodiment of this application. Detailed Implementation

[0023] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” as used in this application and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.

[0024] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."

[0025] This application proposes a self-learning-based category detection method in its embodiments. See [link to relevant documentation]. Figure 1 As shown, the method may include:

[0026] Step 101: Input the acquired target image into the first detection model to obtain the location information of the target object in the target image, extract the sub-image corresponding to the location information from the target image, and input the sub-image into the first feature extraction model to obtain the target features corresponding to the target object.

[0027] Step 102: Determine the similarity between the target feature and each registered feature in the configured feature set.

[0028] Step 103: Select the highest similarity from all similarities and determine whether the highest similarity is not less than the first similarity threshold. If yes, proceed to step 104; otherwise, proceed to step 105.

[0029] Step 104: If the highest similarity is not less than the first similarity threshold, then the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity.

[0030] Step 105: If the highest similarity is less than the first similarity threshold, then based on the target feature and each registered feature in the feature set, determine the open set score value corresponding to the target object, and determine the target category of the target object as the labeled category corresponding to the registered feature corresponding to the highest similarity based on the open set score value, or determine the target category of the target object as the unregistered category based on the open set score value.

[0031] For example, determining the target category of a target object as the labeled category corresponding to the registered feature corresponding to the highest similarity based on the open set score, or determining the target category of a target object as an unregistered category based on the open set score, may include: if the open set score is not greater than a preset open set threshold, then determining the target category of the target object as the labeled category corresponding to the registered feature corresponding to the highest similarity; or, if the open set score is greater than the preset open set threshold, then determining the target category of the target object as an unregistered category.

[0032] For example, determining the open set score of the target object based on the target feature and each registered feature in the feature set may include: determining the distance between the target feature and the central feature of each labeled category; wherein the feature set may include registered features of multiple labeled categories, and each labeled category corresponds to multiple registered features, and the central feature of each labeled category may be obtained based on the multiple registered features of the labeled category; for each labeled category, determining the weight coefficient corresponding to the labeled category based on the distance corresponding to the labeled category and the distances corresponding to all labeled categories, and determining the feature difference score between the target feature and the central feature of the labeled category based on the distance corresponding to the labeled category and the weight coefficient; and determining the open set score based on the feature difference score corresponding to each labeled category.

[0033] In one possible implementation, if the self-learning triggering condition for category detection is met, sample images from the open set dataset can be input into the second detection model to obtain the location information corresponding to the sample objects in the sample images. A sub-image corresponding to this location information is then extracted from the sample images and input into the second feature extraction model to obtain the sample features corresponding to the sample objects. The open set dataset can include multiple sample images, each containing sample objects corresponding to unregistered categories, and is constructed based on target images corresponding to target objects of unregistered categories. The sample features corresponding to multiple sample objects are clustered to obtain an initial cluster set. The sample features in the initial cluster set are then filtered to obtain a target cluster set corresponding to the initial cluster set. The first detection model and the first feature extraction model are retrained based on the target cluster set to obtain retrained first detection model and first feature extraction model. And / or, a new category corresponding to the target cluster set is determined, the sample features in the target cluster set are updated in the feature set, and the new category is used as the labeled category corresponding to the sample features.

[0034] For example, filtering sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set may include, but is not limited to: filtering first-class sample features and / or second-class sample features from all sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set; wherein: the process of determining the first-class sample features may include, but is not limited to: for each sample feature in the initial cluster set, determining the similarity between the sample feature and each registered feature in the feature set; if the highest similarity is not less than a second similarity threshold, then the sample feature is determined as the first-class sample feature; the process of determining the second-class sample features may include, but is not limited to: for each sample feature in the initial cluster set, determining the local density corresponding to the sample feature; if the local density corresponding to the sample feature is less than a density threshold, then the sample feature is determined as the second-class sample feature.

[0035] For example, retraining the first detection model and the first feature extraction model based on the target cluster set to obtain the retrained first detection model and the first feature extraction model may include: selecting sample images from the open set dataset that correspond to the sample features in the target cluster set, inputting the sample images into the first detection model to obtain the location information corresponding to the sample objects in the sample images, extracting sub-images corresponding to the location information from the sample images, inputting the sub-images into the first feature extraction model to obtain the predicted features corresponding to the sample objects; determining a target loss value based on the sample features in the target cluster set and the predicted features, and adjusting the network parameters of the first detection model and the first feature extraction model based on the target loss value to obtain the retrained first detection model and the first feature extraction model.

[0036] For example, a test dataset can also be obtained, which may include multiple test images and the true category of the test object in each test image. Test features corresponding to the test object in each test image can be determined based on a retrained first detection model and a retrained first feature extraction model. Multiple candidate similarity thresholds and multiple candidate open set thresholds can be obtained. For each threshold combination, which may include a candidate similarity threshold and a candidate open set threshold, the predicted category of the test object can be determined based on the candidate similarity threshold, the candidate open set threshold, and the test features corresponding to the test object. The accuracy corresponding to the threshold combination is determined based on the predicted categories and true categories of all test objects. Based on this, and based on the accuracy corresponding to each threshold combination, a target threshold combination can be selected from all threshold combinations. The candidate similarity thresholds in the target threshold combination are updated to the first similarity threshold, and the candidate open set thresholds in the target threshold combination are updated to the preset open set threshold.

[0037] As can be seen from the above technical solutions, in this embodiment, the target category of the target object can be automatically detected. If the target object is a product already displayed in a merchandise cabinet, the target category of the product already displayed in the merchandise cabinet can be automatically detected, resulting in higher detection accuracy. It can utilize intelligent algorithms and image processing technology to achieve automatic detection of the target category of the target object, realizing unmanned management, improving detection accuracy, and saving human resources. A recognition system with open set recognition capability is proposed, possessing self-learning automatic iterative optimization capability. It can automatically expand the calibrated categories, automatically update the registered features in the expanded feature set, and automatically update the first detection model and the first feature extraction model, achieving high self-learning accuracy.

[0038] The technical solutions described above in the embodiments of this application will be explained below in conjunction with specific application scenarios.

[0039] Before introducing the technical solution of this embodiment, let's first introduce the concepts involved in this embodiment:

[0040] Object: The object in this embodiment can be a product object, such as a product object in a product cabinet (also called a display cabinet, which can take the form of a freezer, beverage cabinet, display case, vending machine, etc.), a product object on a shelf, or other types of objects. There are no restrictions on this.

[0041] Feature Set: The feature set, also known as the object registration base library, can include registration features for multiple categorized classes, with each categorized class corresponding to multiple registration features. For example, based on a configured list of object categories, which may include multiple object categories (such as product categories), images of each object category can be pre-collected (e.g., multiple images from different angles). For each object category, features of each image of that object category can be extracted, and these features can be stored in the feature set as registration features for that object category. The object category itself can then be used as the categorized class for these registration features.

[0042] For example, the feature set may include registered features a1, a2, and a3 for class a, which are features of multiple images of class a from different angles. The feature set may also include registered features b1, b2, and b3 for class b, and so on.

[0043] For example, feature sets support quickly adding, deleting, or associating with specified object categories. For instance, you can add a new registered feature to an existing category, or add a registered feature for a new category to the feature set. You can delete a registered feature of a category from the feature set, or delete all registered features of a category from the feature set. You can associate a registered feature in the feature set with an object category; that is, the object category becomes the labeling category of the registered feature.

[0044] For example, an image database corresponding to the feature set can be constructed. The image database is used to store multiple images of each object category from different angles. That is, each registered feature in the feature set corresponds to each image in the image database. In other words, there is a one-to-one correspondence between the registered feature and the image, indicating that the registered feature is a feature of the image. This embodiment does not limit how to extract the registered features from the image.

[0045] First detection model and second detection model: Both the first detection model and the second detection model are used to output the location information of objects in the image. For example, after the image is input to the first detection model or the second detection model, the first detection model or the second detection model can output the location information of objects in the image. When there are multiple objects in the image, the location information of multiple objects can be output.

[0046] For example, the position information corresponding to an object can be the object's coordinate frame in the image, such as a rectangular coordinate frame, a circular coordinate frame, an irregularly shaped coordinate frame, etc., or other types of position information. There are no restrictions on this, as long as it can reflect the object's position in the image. Taking a rectangular coordinate frame as an example, the position information can be the coordinates of the four vertices of the rectangular coordinate frame, or the coordinates of one vertex of the rectangular coordinate frame plus the length and height of the rectangular coordinate frame, or the coordinates of the top-left and bottom-right vertices of the rectangular coordinate frame, or the coordinates of the bottom-left and top-right vertices of the rectangular coordinate frame.

[0047] For example, the difference between the first detection model and the second detection model is that the first detection model is an online-deployed model used to perform category detection in real time. For instance, when category detection is implemented on a front-end device (such as a camera, smartphone, etc.), the first detection model is deployed on the front-end device, which captures images and performs category detection based on the first detection model. When category detection is implemented on a cloud server, the first detection model is deployed on the cloud server, the front-end device captures images, transmits the images to the cloud server, and the cloud server performs category detection based on the first detection model. The second detection model is an offline-deployed model and does not need to perform category detection in real time. For instance, when the second detection model is deployed on a cloud server, and the self-learning trigger condition is met, the cloud server can perform self-learning and optimization on the online-deployed first detection model based on the processing results of the second detection model to improve the performance of the first detection model. The self-learning and optimization process can run offline.

[0048] For example, the first detection model is trained on a small amount of sample images, while the second detection model is trained on a large amount of sample images. The second detection model has higher accuracy and better generalization performance than the first model, and its training effect is also better (e.g., the second detection model can adapt to the current scene more quickly). Because the second detection model is trained on a large amount of sample images, while the first detection model is trained on a small amount of sample images, the training time for the second detection model is longer, while the training time for the first detection model is shorter. In other words, the training of the second detection model takes longer than the training time for the first detection model. Furthermore, because the second detection model has higher accuracy than the first detection model, its inference time is longer, while the inference time for the first detection model is shorter. That is, the second detection model requires a longer inference process, while the first detection model requires a shorter inference process.

[0049] For example, the second detection model is a domain-specific large model, also known as a visual large model or expert model, used to detect bounding boxes in images. It possesses high recognition accuracy and good generalization performance in a specific task domain. The training method for the second detection model is as follows: The detection model is trained using a large amount of sample images. A large batch of labeled data (accumulated across different business processes) is then collected to fine-tune the detection model, resulting in the second detection model. This second detection model already has sufficiently high performance for data labeling during the self-learning process. When the amount of business data is large enough, further fine-tuning can be performed to obtain even better results.

[0050] First feature extraction model and second feature extraction model: Both the first feature extraction model and the second feature extraction model are used to output the features corresponding to objects in the image, that is, to extract the features of objects in the image. There are no restrictions on this feature extraction process. For example, after the image is input to the first feature extraction model or the second feature extraction model, the first feature extraction model or the second feature extraction model can output the features corresponding to objects in the image. When there are multiple objects in the image, the features corresponding to multiple objects can be output.

[0051] The difference between the first and second feature extraction models lies in their deployment methods. The first feature extraction model is deployed online for real-time category detection. For example, when performing category detection on a front-end device, the first feature extraction model is deployed there. The front-end device captures images and performs category detection based on the first feature extraction model. Similarly, when performing category detection on a cloud server, the first feature extraction model is deployed there. The front-end device captures images, transmits them to the cloud server, and the cloud server performs category detection based on the first feature extraction model. The second feature extraction model, on the other hand, is deployed offline and does not require real-time category detection. For example, when the second feature extraction model is deployed on a cloud server, and a self-learning trigger condition is met, the cloud server uses the processing results of the second feature extraction model to perform self-learning and optimization on the online-deployed first feature extraction model to improve its performance. This self-learning and optimization process can run offline.

[0052] The first feature extraction model is trained on a small number of sample images, while the second feature extraction model is trained on a large number of sample images. The second feature extraction model has higher accuracy and better generalization performance than the first. Its training effect is also better (e.g., it adapts to the current scene more quickly). However, because the second feature extraction model is trained on a large number of sample images, while the first model is trained on a small number, the training time for the second model is longer, while the training time for the first model is shorter. In other words, the second model takes longer to train, while the first model takes much shorter. Furthermore, because the second feature extraction model has higher accuracy, its inference time is longer, while the first model takes much shorter. This means the second model requires a longer inference process, while the first model completes its inference process much faster.

[0053] The second feature extraction model is a domain-specific large model, also known as a visual large model or expert model. It is used to extract features from images and possesses high recognition accuracy and good generalization performance in a specific task domain. The training method for the second feature extraction model is as follows: The model is trained using a large amount of sample images. Then, a large batch of labeled data (accumulated from different business processes) is collected to fine-tune the feature extraction model, resulting in the second feature extraction model. This second feature extraction model already has sufficiently high performance for data labeling during the self-learning process. When the amount of business data is large enough, further fine-tuning can be performed to obtain even better results.

[0054] This application proposes a category detection method with self-learning capabilities, which can be applied to object recognition systems (such as commodity object recognition systems). See also... Figure 2As shown, the object recognition system includes an object recognition subsystem, a data annotation subsystem, and an automatic training subsystem. The object recognition subsystem uses front-end devices (such as cameras, smartphones, etc.) to capture images of the target area to be recognized. These images are then input into the object recognition subsystem, which performs a series of operations including detection, modeling, comparison, and open-set recognition to identify the location and category information of all objects in the image. Unregistered objects are output as unregistered categories. The data annotation subsystem automatically collects all registered objects, automatically classifies them to create category profiles, expands the feature set by adding new categories, and automatically collects actual system operation data. This data is then processed by algorithms to automatically calibrate the operation data. The calibrated data is then sent to the automatic training subsystem for automatic training. After training, the first detection model and the first feature extraction model are updated for the object recognition subsystem.

[0055] In one possible implementation, for the object category recognition process (such as the product category recognition process) of the object recognition subsystem, see [link to relevant documentation]. Figure 3 As shown, the object category identification process may include:

[0056] Step 301: Obtain the target image of the target area to be identified.

[0057] For example, when class detection is performed on a front-end device (such as a camera, smartphone, etc.), the front-end device can capture a target image of the area to be identified. When class detection is performed on a cloud server, the front-end device can capture a target image of the area to be identified and transmit the target image to the cloud server, which then retrieves the target image of the area to be identified.

[0058] Step 302: Input the target image into the first detection model to obtain the location information of the target object in the target image. The location information of the target object can be the coordinate box corresponding to the target object.

[0059] For example, the target image may include at least one object, each referred to as a target object. After the target image is input into the first detection model, the first detection model can output the location information corresponding to each target object in the target image. The processing procedure of the first detection model is not limited. Since the processing procedure for each target object is the same, the processing procedure for one target object will be used as an example in the following embodiments.

[0060] Step 303: Based on the location information corresponding to the target object, extract the sub-image corresponding to the location information from the target image (i.e., the sub-image where the target object is located, and the sub-image includes the target object), and input the sub-image into the first feature extraction model to obtain the target features corresponding to the target object.

[0061] For example, the location information corresponding to the target object can be the coordinate frame corresponding to the target object, and a sub-image corresponding to the coordinate frame can be extracted from the target image, which may include the target object.

[0062] After obtaining the sub-image, it can be input into the first feature extraction model, which will extract the features corresponding to the sub-image. The feature extraction process of the first feature extraction model is not restricted, and the features corresponding to the sub-image are the target features corresponding to the target object.

[0063] Step 304: Determine the similarity between the target feature and each registered feature in the feature set.

[0064] For example, after obtaining the target features corresponding to the target object, the target features can be compared with each registered feature in the feature set to obtain the similarity between the target features and each registered feature in the feature set, such as cosine similarity or other types of similarity.

[0065] Step 305: Select the highest similarity from all similarities and determine whether the highest similarity is not less than the first similarity threshold. If yes, proceed to step 306; otherwise, proceed to step 307.

[0066] For example, after obtaining the similarity between the target feature and each registered feature in the feature set, they can be sorted in descending order of similarity or in ascending order of similarity. In this way, the highest similarity can be selected from all similarities.

[0067] The first similarity threshold can be a similarity threshold configured based on experience, or it can be a similarity threshold determined during the self-learning process. The determination process is described in subsequent embodiments and is not limited thereto.

[0068] Step 306: Determine the target category of the target object as the labeled category corresponding to the registered feature with the highest similarity. For example, assuming the similarity between the target feature and the registered feature A is the highest similarity, then the target category of the target object can be determined as the labeled category corresponding to the registered feature A.

[0069] For example, if the highest similarity is not less than the first similarity threshold, it means that the target category of the target object is a closed set category. That is, the target category of the target object needs to be determined from the labeled categories of all registered objects. In other words, the labeled category corresponding to the registered feature with the highest similarity is taken as the target category.

[0070] Step 307: Based on the target feature and each registered feature in the feature set, determine the open set score corresponding to the target object. If the open set score is not greater than a preset open set threshold, the target category of the target object is determined to be the labeled category corresponding to the registered feature with the highest similarity; or, if the open set score is greater than the preset open set threshold, the target category of the target object is determined to be an unregistered category.

[0071] For example, if the highest similarity score is less than the first similarity threshold, it indicates that the target category of the target object is an open-set category. This means the target category could be the registered object's labeled category or an unregistered category. Therefore, it is necessary to further determine whether the target category is the registered object's labeled category or an unregistered category. For instance, if the open-set score corresponding to the target object is not greater than a preset open-set threshold, then the target category of the target object needs to be determined from the labeled categories of all registered objects. In other words, the labeled category corresponding to the registered feature with the highest similarity score is taken as the target category. If the open-set score corresponding to the target object is greater than the preset open-set threshold, then the target category of the target object is an unregistered category.

[0072] The preset open set threshold can be an open set threshold configured based on experience, or it can be an open set threshold determined during the self-learning process. The determination process is described in subsequent embodiments and is not limited thereto.

[0073] In one possible implementation, the process of determining the open set fraction value may include:

[0074] Step 3071: Obtain the registration feature matrix, which includes the central feature of each labeled category. The feature set includes registration features for multiple labeled categories, with each labeled category corresponding to multiple registration features, and the central feature of each labeled category is obtained based on the multiple registration features of that labeled category.

[0075] For example, suppose the feature set includes registered features a1, a2, and a3 for category a, and registered features b1, b2, and b3 for category b. Therefore, the registered feature matrix can include the central features of category a and the central features of category b. The central features of category a can be obtained based on registered features a1, a2, and a3, and the central features of category b can be obtained based on registered features b1, b2, and b3.

[0076] For example, when obtaining the central feature of a labeling category based on multiple registered features of the labeling category, the central feature can be determined based on the average feature of all registered features, or the registered feature with the highest local density can be selected as the central feature from all registered features. Of course, the above methods are just examples and are not limited thereto. For example, any registered feature can be selected as the central feature from all registered features.

[0077] The process of selecting the registration feature with the highest local density from all registration features (i.e., all registration features of the labeled category) as the center feature may include, but is not limited to: for each registration feature, such as registration feature a1, calculating the distance between this registration feature and other registration features (i.e., every other registration feature besides this one, such as registration feature a2 and registration feature a3), and counting the number of distances less than a preset distance threshold (which can be configured empirically). This number is the local density of the registration feature. After obtaining the local density of each registration feature, the registration feature with the highest local density can be selected as the center feature.

[0078] Step 3072: Determine the distance between the target feature and the central feature of each labeled category.

[0079] For example, the registration feature matrix includes the central features of each labeled category, such as the central features of labeled category a, the central features of labeled category b, etc. It can determine the distance between the target feature and the central features of labeled category a, and the distance between the target feature and the central features of labeled category b.

[0080] Step 3073: For each labeled category, determine the weight coefficient corresponding to that labeled category based on the distance corresponding to that labeled category and the distances corresponding to all labeled categories. The distance corresponding to that labeled category can refer to the distance between the target feature and the central feature of that labeled category.

[0081] For example, based on the distance corresponding to each labeled category, the weight coefficient corresponding to the labeled category can be determined by the following formula. Of course, the following formula is just an example and is not a limitation.

[0082]

[0083] In the above formula, i represents the i-th labeled category, and the value of i ranges from 1 to N, where N represents the total number of labeled categories, and softmin(d) i Let represent the weight coefficient corresponding to the i-th labeled category, and ... This represents the sum of distances corresponding to all calibrated categories.

[0084] Step 3074: For each calibration category, determine the feature difference score between the target feature and the central feature of the calibration category based on the distance corresponding to the calibration category and the weight coefficient corresponding to the calibration category.

[0085] For example, the feature difference score between the target feature and the central feature of the labeled category can be determined using the following formula. Of course, the following formula is just an example and is not a limitation.

[0086] γ=d*(1-softmin(d))

[0087] For each labeled category, in the above formula, γ represents the feature difference score between the target feature and the central feature of that labeled category, d represents the distance corresponding to that labeled category, and softmin(d) represents the weight coefficient corresponding to that labeled category. In summary, we can obtain the feature difference score corresponding to each labeled category, that is, the feature difference score between the target feature and the central feature of each labeled category.

[0088] Step 3075: Determine the open set score based on the feature difference score corresponding to each labeled category. For example, based on the feature difference score corresponding to each labeled category, the minimum feature difference score can be used as the open set score, or the average feature difference score can be used as the open set score; there is no restriction on this.

[0089] In summary, we can obtain the open set score value corresponding to the target object, and then determine the target category of the target object based on the open set score value, such as the labeled category corresponding to the registered feature or the unregistered category.

[0090] This completes the object category recognition process of the object recognition subsystem. For each frame of target image, the above object category recognition process can be used to determine the target category of each target object in the target image. The target category can be the labeled category corresponding to a certain registered feature or an unregistered category.

[0091] In one possible implementation, if the target object in the target image belongs to an unregistered category, the target image can be stored in an open-set dataset as a sample image in the open-set dataset, so that self-learning of category detection can be achieved based on the open-set dataset in subsequent processes. Obviously, in the object category recognition process, multiple target images can be stored in an open-set dataset, and for ease of distinction, these target images in the open-set dataset can be called sample images.

[0092] In one possible implementation, if the self-learning trigger condition for category detection is met, the data annotation subsystem and the automatic training subsystem can be triggered. The data annotation subsystem then performs automatic labeling, and the automatic training subsystem automatically trains and updates the first detection model and the first feature extraction model. Specifically, the self-learning trigger condition for category detection is determined to be met when the user manually triggers self-learning, or when the number of sample images in the open set dataset reaches a preset threshold. These are merely examples of self-learning trigger conditions and are not intended to be limiting. For the data annotation and automatic training processes, see [link to relevant documentation]. Figure 4 As shown, the process may include:

[0093] Step 401: Input the sample images in the open set dataset into the second detection model to obtain the location information of the sample objects in the sample images. Extract the sub-image corresponding to the location information from the sample images and input the sub-image into the second feature extraction model to obtain the sample features corresponding to the sample objects.

[0094] For example, an open set dataset may include multiple sample images, and for each sample image, it may include at least one object. For each object, if the target category corresponding to the object is the labeled category corresponding to a certain registered feature, then the object is not a sample object; if the target category corresponding to the object is an unregistered category, then the object is a sample object. The data annotation and automatic training processes are processing processes for sample objects, that is, sample objects in sample images need to correspond to unregistered categories.

[0095] For example, after inputting the sample images (such as each sample image) in the open set dataset into the second detection model, the second detection model can output the location information corresponding to each object in the sample image. The processing procedure of the second detection model is not limited, and the location information corresponding to the sample object can be selected from the location information corresponding to each object, that is, the location information corresponding to the sample object is obtained.

[0096] For example, the location information corresponding to the sample object can be the coordinate box corresponding to the sample object. A sub-image corresponding to the coordinate box can be extracted from the sample image, and the sub-image can include the sample object. After obtaining the sub-image, it can be input into the second feature extraction model, which extracts the features corresponding to the sub-image. The feature extraction process of the second feature extraction model is not limited, and the features corresponding to the sub-image can be the sample features corresponding to the sample object.

[0097] Step 402: Cluster the sample features corresponding to multiple sample objects to obtain an initial cluster set.

[0098] For example, an open set dataset may include multiple sample images. Sample features corresponding to sample objects in each sample image can be obtained, thus yielding sample features corresponding to multiple sample objects. Based on these sample features, clustering can be performed to obtain at least one initial cluster set. For instance, K-means clustering, partitioning clustering, hierarchical clustering, density-based clustering, grid-based clustering, or model-based clustering algorithms can be used to cluster these sample features; no restriction is placed on the specific clustering algorithm used.

[0099] In summary, by clustering multiple sample features, an initial cluster set can be obtained, and each initial cluster set can include at least one sample feature. In one possible implementation, the number of sample features in the initial cluster set can be counted. If the number of sample features is less than a preset threshold (which can be configured empirically), this initial cluster set can be filtered out, meaning it will not participate in subsequent processes, thus eliminating initial cluster sets with too few features. If the number of sample features is not less than the preset threshold, this initial cluster set can be retained, meaning it will participate in subsequent processes. In another possible implementation, all initial cluster sets can be retained for participation in subsequent processes.

[0100] Step 403: For each initial cluster set, filter the sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set. For example, the first type of sample features and / or the second type of sample features can be filtered from all the sample features of the initial cluster set (i.e., the first type of sample features and / or the second type of sample features are removed) to obtain the target cluster set corresponding to the initial cluster set.

[0101] In one possible implementation, for each sample feature in the initial cluster set, the similarity between the sample feature and each registered feature in the feature set is determined; if the highest similarity is not less than the second similarity threshold, the sample feature is determined as a first-class sample feature, that is, the sample feature is considered to be a feature of the closed set category and is removed; otherwise, the sample feature is not determined as a first-class sample feature.

[0102] For example, a feature comparison can be performed between a sample feature and each registered feature in the feature set to obtain the similarity between the sample feature and each registered feature in the feature set, such as cosine similarity. After obtaining the similarity between the sample feature and each registered feature in the feature set, they are sorted in descending order of similarity or in ascending order of similarity, and the highest similarity is selected from all similarities. Then, it is determined whether the highest similarity is not less than a second similarity threshold. The second similarity threshold can be a similarity threshold configured based on experience; it can be greater than or less than the first similarity threshold, without restriction. If the highest similarity is not less than the second similarity threshold, the sample feature is identified as a first-class sample feature and filtered from the initial cluster set. If the highest similarity is less than the second similarity threshold, the sample feature is not identified as a first-class sample feature, that is, it is retained in the initial cluster set.

[0103] In one possible implementation, for each sample feature in the initial cluster set, the local density corresponding to the sample feature can be determined. If the local density corresponding to the sample feature is less than the density threshold (configured based on experience), then the sample feature is determined as a second type of sample feature. Otherwise, if the local density corresponding to the sample feature is not less than the density threshold, then the sample feature is not determined as a second type of sample feature.

[0104] For example, after obtaining the initial cluster set, there may be erroneous noise sample features (sample features that are not in the current category or sample features with abnormal coordinate frames). Therefore, the sample features in the initial cluster set can be denoised. During the denoising process, the sample features in the initial cluster set (i.e., second-type sample features) can be selected using the local density method. Local density can be defined as the number of sample features whose distance from the sample feature is less than a preset distance threshold (which can be configured empirically). The center of the initial cluster set is generally the sample feature with the highest local density. Sample features with lower local density are uncommon data forms, usually noise. Therefore, the second-type sample features in the initial cluster set can be selected using the local density method and treated as noise.

[0105] For example, for each sample feature in the initial cluster set, calculate the distance between that sample feature and other sample features (i.e., every sample feature in the initial cluster set other than that sample feature), and count the number of distances less than a preset distance threshold. This number is the local density of that sample feature.

[0106] After obtaining the local density of each sample feature, if the local density corresponding to the sample feature is less than the density threshold, then the sample feature is identified as a second-class sample feature and filtered out from the initial cluster set. If the local density corresponding to the sample feature is not less than the density threshold, then the sample feature is not identified as a second-class sample feature, that is, the sample feature is retained in the initial cluster set.

[0107] In one possible implementation, after filtering the first type of sample features and / or the second type of sample features from all sample features of the initial cluster set, the cluster set after filtering the sample features can be used as the target cluster set corresponding to the initial cluster set. Alternatively, the cluster set after filtering the sample features can be deduplicated, and the cluster set after deduplication can be used as the target cluster set corresponding to the initial cluster set.

[0108] The deduplication operation on the cluster set involves calculating the similarity between any two sample features in the cluster set, such as using a hash algorithm. If the similarity is greater than or equal to a preset threshold, either of the two sample features is retained; otherwise, both sample features are retained. After performing this operation on any two sample features in the cluster set, the deduplication of sample features is completed.

[0109] Step 404: Determine the new category corresponding to the target cluster set, update the sample features in the target cluster set to the feature set, and use the new category as the label category corresponding to the sample features.

[0110] For example, for each target cluster set, the sample features in that target cluster set can be updated to the feature set, meaning the sample features become the registered features in the feature set. A default category can be generated for the target cluster set as the new category, and this new category can be used as the labeling category corresponding to the sample features (i.e., the registered features). Alternatively, the user can configure a category for the target cluster set, and the user-configured category can be used as the new category, which can then be used as the labeling category corresponding to the sample features.

[0111] Step 405: Retrain the first detection model and the first feature extraction model based on the target cluster set to obtain the retrained first detection model and the first feature extraction model.

[0112] For example, for each sample feature in the target cluster set, since the sample feature is obtained based on the second detection model and the second feature extraction model, and the inference accuracy of the second detection model and the second feature extraction model is very high, the inference result can be directly used as the data annotation ground truth. Therefore, each sample feature in the target cluster set can be used as the data annotation ground truth. Based on this, sample images corresponding to the sample features (i.e., the sample features corresponding to the sample objects) in the target cluster set can be selected from the open set dataset, and the sample images can be input into the first detection model to obtain the location information corresponding to the sample object in the sample image. A sub-image corresponding to the location information can be extracted from the sample image, and the sub-image can be input into the first feature extraction model to obtain the predicted features corresponding to the sample object.

[0113] Then, the target loss value can be determined based on the sample features in the target cluster set and the predicted features corresponding to each sample object. For example, for each sample object, the loss value corresponding to that sample object can be determined based on the sample features and the predicted features corresponding to that sample object, such as determining the loss value based on the difference between the sample features and the predicted features. The target loss value can also be determined based on the loss value corresponding to each sample object, such as determining the target loss value based on the sum of the loss values ​​corresponding to each sample object.

[0114] After obtaining the target loss value, the network parameters of the first detection model and the first feature extraction model can be adjusted based on the target loss value to obtain the retrained first detection model and the first feature extraction model. For example, based on the target loss value, gradient descent or other methods can be used to adjust the network parameters of the first detection model and the first feature extraction model, with the goal of making the target loss value smaller and smaller, resulting in the adjusted detection model and the adjusted feature extraction model. If the adjusted detection model and the adjusted feature extraction model have converged, the adjusted detection model is used as the retrained first detection model, and the adjusted feature extraction model is used as the retrained first feature extraction model, completing the retraining process. If the adjusted detection model and / or the adjusted feature extraction model have not converged, the adjusted detection model is used as the first detection model, and the adjusted feature extraction model is used as the first feature extraction model, and the process returns to execute the operation of "inputting the sample image into the first detection model to obtain the location information corresponding to the sample object in the sample image, extracting the sub-image corresponding to the location information from the sample image, and inputting the sub-image into the first feature extraction model to obtain the predicted features corresponding to the sample object".

[0115] For example, for each target cluster set, a default category can be generated as a new category for that target cluster set, or the user can configure a category for that target cluster set and use the user-configured category as the new category. Based on this, when the sample image is input into the first detection model, the new category can also be used as the category of the sample image and input into the first detection model. In this way, the retrained first detection model and the first feature extraction model can also support the new category corresponding to the target cluster set, that is, they have the ability to identify the new category, and the new category becomes the label category of the registered object.

[0116] Step 406: Determine the first similarity threshold and the preset open set threshold for the retrained first detection model and the first feature extraction model. At this point, the first detection model, the first feature extraction model, the first similarity threshold, and the preset open set threshold have been successfully updated. Based on the updated first detection model, the first feature extraction model, the first similarity threshold, and the preset open set threshold, steps 301-307 can be executed.

[0117] In one possible implementation, when determining the first similarity threshold and the preset open set threshold, the first similarity threshold may be a similarity threshold configured based on experience, and the preset open set threshold may be an open set threshold configured based on experience. Alternatively, the first similarity threshold and the preset open set threshold may be determined using the following steps:

[0118] Step 4061: Obtain the test dataset, which may include multiple test images and the true category of the test object (the object in the test image is referred to as the test object) in each test image.

[0119] For example, multiple test images can be acquired, and the true category of the test object in each test image can be assigned, indicating that the category of the test object is the true category, which is an accurate and reliable category.

[0120] Step 4062: Obtain multiple candidate similarity thresholds and multiple candidate open set thresholds.

[0121] For example, it can obtain all possible candidate similarity thresholds, such as obtaining a configured similarity threshold range, and using all or part of the similarity thresholds within that range as candidate similarity thresholds, such as candidate similarity threshold c1, candidate similarity threshold c2, ..., and so on.

[0122] For example, it can obtain all possible candidate open set thresholds, such as obtaining the configured open set threshold range, and using all or part of the open set thresholds within the open set threshold range as candidate open set thresholds, such as candidate open set threshold d1, candidate open set threshold d2, candidate open set threshold d3, ..., and so on.

[0123] Step 4063: Construct multiple threshold combinations, each including a candidate similarity threshold and a candidate open set threshold. For example, threshold combination e1 includes candidate similarity threshold c1 and candidate open set threshold d1, threshold combination e2 includes candidate similarity threshold c1 and candidate open set threshold d2, and so on.

[0124] Step 4064: Determine the test features corresponding to the test object in each test image based on the retrained first detection model and the retrained first feature extraction model.

[0125] For example, for each test image, the test image can be input into the retrained first detection model to obtain the location information of the test object in the test image, and a sub-image corresponding to the location information can be extracted from the test image. The sub-image is then input into the retrained first feature extraction model to obtain the test features corresponding to the test object.

[0126] Step 4065: For each threshold combination, based on the candidate similarity threshold and candidate open set threshold in the threshold combination, and the test features corresponding to the test object, determine the predicted category of the test object.

[0127] For example, after obtaining the test features corresponding to the test object, the similarity between the test feature and each registered feature in the feature set is determined. The highest similarity is selected from all similarities, and it is determined whether the highest similarity is not less than the candidate similarity threshold in the threshold combination. If so, the predicted category is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity. If not, based on the test feature and each registered feature in the feature set, the open set score corresponding to the test object is determined. If the open set score is not greater than the candidate open set threshold in the threshold combination, the predicted category is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity; if the open set score is greater than the candidate open set threshold in the threshold combination, the predicted category is determined to be an unregistered category. In summary, the predicted category of the test object can be obtained.

[0128] Step 4066: Determine the accuracy corresponding to the threshold combination based on the predicted and true categories of all test objects. For example, for each test object, if the predicted category and the true category are the same, increment the accuracy corresponding to the threshold combination by 1; if the predicted category and the true category are different, do not increment the accuracy corresponding to the threshold combination, i.e., keep it unchanged. After performing the above processing on the test objects in all test images, the accuracy corresponding to the threshold combination can be obtained.

[0129] Step 4067: Based on the accuracy corresponding to each threshold combination, select the target threshold combination from all threshold combinations. For example, the threshold combination with the highest accuracy can be used as the target threshold combination.

[0130] Step 4068: Update the candidate similarity threshold in the target threshold combination to the first similarity threshold, and update the candidate open set threshold in the target threshold combination to the preset open set threshold.

[0131] In summary, this embodiment can automatically trigger the update process of the first similarity threshold and the preset open set threshold. For example, a statistical threshold update method can be used to calculate the accuracy for each of all threshold combinations (candidate similarity threshold and candidate open set threshold), and then update the first similarity threshold and the preset open set threshold. To improve calculation speed, the similarity threshold and the open set threshold can be approximated, reducing the number of accuracy calculation rounds. After calculating for all threshold combinations, the threshold combination with the highest accuracy on the test data is selected as the first similarity threshold and the preset open set threshold, thus improving the open set recognition capability.

[0132] In summary, the self-learning process in this embodiment includes three automatic iterative optimization processes: automatic iterative optimization of the feature set (step 404), automatic iterative optimization of the first detection model and the first feature extraction model (step 405), and automatic iterative optimization of the first similarity threshold and the preset open set threshold (step 406). These three processes interact to gradually improve the recognition accuracy and capability.

[0133] As can be seen from the above technical solutions, in this embodiment, the target category of the target object can be automatically detected. If the target object is a product already displayed in a shelf, the target category of the product already displayed in the shelf can be automatically detected, resulting in higher detection accuracy. It can utilize intelligent algorithms and image processing technology to achieve automatic detection of the target category of the target object, realizing unmanned management, improving detection accuracy, and saving human resources. A recognition system with open set recognition capability is proposed, possessing self-learning automatic iterative optimization capability. It can automatically expand the calibrated categories, automatically update the registered features in the expanded feature set, and automatically update the first detection model and the first feature extraction model, achieving high self-learning accuracy. The proposed object recognition system has open set recognition capability, which is more in line with actual application scenarios. It can automatically iteratively expand recognition capabilities, has a wide update range, low human intervention cost, and good self-learning effect in the self-learning process.

[0134] Based on the same concept as the methods described above, this application proposes a self-learning-based category detection device, see [link to relevant documentation]. Figure 5 The diagram shown is a structural schematic of the device, which may include:

[0135] The acquisition module 51 is used to input the target image into the first detection model to obtain the location information corresponding to the target object in the target image, extract the sub-image corresponding to the location information from the target image, and input the sub-image into the first feature extraction model to obtain the target feature corresponding to the target object; the determination module 52 is used to determine the similarity between the target feature and each registered feature in the feature set; if the highest similarity is not less than the first similarity threshold, the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity; if the highest similarity is less than the first similarity threshold, the open set score value corresponding to the target object is determined based on the target feature and each registered feature in the feature set; the target category of the target object is determined based on the open set score value to be the labeled category corresponding to the registered feature corresponding to the highest similarity, or the target category is determined to be an unregistered category.

[0136] For example, the determining module 52 determines the target category of the target object as the labeled category corresponding to the registered feature corresponding to the highest similarity based on the open set score value, or, when determining the target category as an unregistered category, it is specifically used to: if the open set score value is not greater than a preset open set threshold, then determine the target category of the target object as the labeled category corresponding to the registered feature corresponding to the highest similarity; or, if the open set score value is greater than the preset open set threshold, then determine the target category as an unregistered category.

[0137] For example, when determining the open set score of the target object based on the target feature and each registered feature in the feature set, the determining module 52 is specifically used to: determine the distance between the target feature and the central feature of each calibration category; the feature set includes registered features of multiple calibration categories, and each calibration category corresponds to multiple registered features, and the central feature of each calibration category is obtained based on the multiple registered features of that calibration category; for each calibration category, determine the weight coefficient corresponding to that calibration category based on the distance corresponding to that calibration category and the distances corresponding to all calibration categories, determine the feature difference score between the target feature and the central feature of that calibration category based on the distance corresponding to that calibration category and the weight coefficient; and determine the open set score based on the feature difference score corresponding to each calibration category.

[0138] For example, the apparatus further includes: a processing module, configured to, if a self-learning trigger condition for category detection is met, input sample images from an open set dataset into a second detection model to obtain location information corresponding to sample objects in the sample images; extract sub-images corresponding to the location information from the sample images; input the sub-images into a second feature extraction model to obtain sample features corresponding to the sample objects; wherein, the open set dataset includes multiple sample images, and the sample objects in each sample image correspond to unregistered categories, and the open set dataset is constructed based on target images corresponding to target objects of unregistered categories; cluster the sample features corresponding to multiple sample objects to obtain an initial cluster set; filter the sample features in the initial cluster set to obtain a target cluster set corresponding to the initial cluster set; retrain the first detection model and the first feature extraction model based on the target cluster set to obtain a retrained first detection model and a first feature extraction model; and / or, determine a new category corresponding to the target cluster set, update the sample features in the target cluster set to the feature set, and use the new category as the labeling category corresponding to the sample features.

[0139] For example, when the processing module filters the sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set, it specifically performs the following steps: filtering first-class sample features and / or second-class sample features from all sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set; wherein: the process of determining the first-class sample features includes: for each sample feature in the initial cluster set, determining the similarity between the sample feature and each registered feature in the feature set; if the highest similarity is not less than a second similarity threshold, then the sample feature is determined as the first-class sample feature; the process of determining the second-class sample features includes: for each sample feature in the initial cluster set, determining the local density corresponding to the sample feature; if the local density corresponding to the sample feature is less than a density threshold, then the sample feature is determined as the second-class sample feature.

[0140] For example, when the processing module retrains the first detection model and the first feature extraction model based on the target cluster set to obtain the retrained first detection model and the first feature extraction model, it specifically performs the following steps: selecting sample images from the open set dataset that correspond to the sample features in the target cluster set, and inputting the sample images into the first detection model to obtain the location information corresponding to the sample objects in the sample images; extracting sub-images corresponding to the location information from the sample images, and inputting the sub-images into the first feature extraction model to obtain the predicted features corresponding to the sample objects; determining a target loss value based on the sample features in the target cluster set and the predicted features, and adjusting the network parameters of the first detection model and the first feature extraction model based on the target loss value to obtain the retrained first detection model and the first feature extraction model.

[0141] For example, the processing module is further configured to: acquire a test dataset, the test dataset including multiple test images and the true category of the test object in each test image; determine the test features corresponding to the test object in each test image based on the retrained first detection model and the retrained first feature extraction model; acquire multiple candidate similarity thresholds and multiple candidate open set thresholds; for each threshold combination, the threshold combination including a candidate similarity threshold and a candidate open set threshold, determine the predicted category of the test object based on the candidate similarity threshold, the candidate open set threshold and the test features corresponding to the test object; determine the accuracy corresponding to the threshold combination based on the predicted categories and true categories of all test objects; and select a target threshold combination from all threshold combinations based on the accuracy corresponding to each threshold combination, update the candidate similarity threshold in the target threshold combination to the first similarity threshold, and update the candidate open set threshold in the target threshold combination to the preset open set threshold.

[0142] Based on the same application concept as the above method, this application proposes an electronic device, see [link to application]. Figure 6 As shown, the electronic device includes a processor 61 and a machine-readable storage medium 62, the machine-readable storage medium 62 storing machine-executable instructions that can be executed by the processor 61; the processor 61 is used to execute the machine-executable instructions to implement the self-learning-based category detection method disclosed in the above example of this application.

[0143] Based on the same concept as the above method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the self-learning-based category detection method disclosed in the above examples of this application.

[0144] The aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0145] The systems, devices, modules, or units described in the above embodiments can be implemented by a computer or entity, or by a product with a certain function. A typical implementation device is a computer, which can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0146] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0147] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0148] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0149] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0150] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0151] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A self-learning-based category detection method, characterized in that, The method includes: The acquired target image is input into the first detection model to obtain the location information corresponding to the target object in the target image. A sub-image corresponding to the location information is extracted from the target image, and the sub-image is input into the first feature extraction model to obtain the target features corresponding to the target object. Determine the similarity between the target feature and each registered feature in the configured feature set; If the highest similarity is not less than the first similarity threshold, then the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity. If the highest similarity is less than the first similarity threshold, then based on the target feature and each registered feature in the feature set, the open set score corresponding to the target object is determined; Based on the open set score, the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity, or the target category is determined to be an unregistered category; If the self-learning triggering condition for category detection is met, the method further includes: The sample images in the open set dataset are input into the second detection model to obtain the location information corresponding to the sample objects in the sample images. A sub-image corresponding to the location information is extracted from the sample images, and the sub-image is input into the second feature extraction model to obtain the sample features corresponding to the sample objects. The open set dataset includes multiple sample images, and the sample objects in each sample image correspond to unregistered categories. The open set dataset is constructed based on the target images corresponding to the target objects of the unregistered categories. Clustering is performed on the sample features corresponding to multiple sample objects to obtain an initial cluster set; the sample features in the initial cluster set are then filtered to obtain the target cluster set corresponding to the initial cluster set; The first detection model and the first feature extraction model are retrained based on the target cluster set to obtain the retrained first detection model and the first feature extraction model; and / or, a new category corresponding to the target cluster set is determined, the sample features in the target cluster set are updated to the feature set, and the new category is used as the labeling category corresponding to the sample features.

2. The method according to claim 1, characterized in that, The step of determining the target category of the target object based on the open set score as the labeled category corresponding to the registered feature corresponding to the highest similarity, or determining the target category as an unregistered category, includes: If the open set score is not greater than a preset open set threshold, then the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity; or, If the open set score is greater than the preset open set threshold, then the target category is determined to be an unregistered category.

3. The method according to claim 1, characterized in that, The step of determining the open set score corresponding to the target object based on the target feature and each registered feature in the feature set includes: Determine the distance between the target feature and the central feature of each labeled category; wherein the feature set includes registered features of multiple labeled categories, and each labeled category corresponds to multiple registered features, and the central feature of each labeled category is obtained based on the multiple registered features of that labeled category; For each labeled category, a weight coefficient corresponding to the labeled category is determined based on the distance corresponding to the labeled category and the distances corresponding to all labeled categories. A feature difference score between the target feature and the central feature of the labeled category is determined based on the distance corresponding to the labeled category and the weight coefficient. The open set score is determined based on the feature difference score corresponding to each calibrated category.

4. The method according to claim 1, characterized in that, The step of filtering sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set includes: Filter the first type of sample features and / or the second type of sample features from all sample features of the initial cluster set to obtain the target cluster set corresponding to the initial cluster set; wherein: The process of determining the first type of sample features includes: for each sample feature in the initial cluster set, determining the similarity between the sample feature and each registered feature in the feature set; if the highest similarity is not less than the second similarity threshold, then the sample feature is determined as the first type of sample feature; The process of determining the second type of sample features includes: for each sample feature in the initial cluster set, determining the local density corresponding to the sample feature; if the local density corresponding to the sample feature is less than the density threshold, then the sample feature is determined as the second type of sample feature.

5. The method according to claim 1, characterized in that, The step of retraining the first detection model and the first feature extraction model based on the target cluster set to obtain the retrained first detection model and the first feature extraction model includes: Select a sample image from the open set dataset that corresponds to the sample features in the target cluster set, and input the sample image into the first detection model to obtain the location information of the sample object in the sample image. Extract a sub-image from the sample image that corresponds to the location information, and input the sub-image into the first feature extraction model to obtain the predicted features corresponding to the sample object. The target loss value is determined based on the sample features in the target cluster set and the predicted feature, and the network parameters of the first detection model and the first feature extraction model are adjusted based on the target loss value to obtain the retrained first detection model and the first feature extraction model.

6. The method according to claim 2, characterized in that, After retraining the first detection model and the first feature extraction model based on the target cluster set to obtain the retrained first detection model and the first feature extraction model, the method further includes: Obtain a test dataset, which includes multiple test images and the true category of the test object in each test image; determine the test features corresponding to the test object in each test image based on the retrained first detection model and the retrained first feature extraction model; Multiple candidate similarity thresholds and multiple candidate open set thresholds are obtained; for each threshold combination, which includes a candidate similarity threshold and a candidate open set threshold, the predicted category of the test object is determined based on the candidate similarity threshold, the candidate open set threshold, and the test features corresponding to the test object; the accuracy corresponding to the threshold combination is determined based on the predicted categories and true categories of all test objects. Based on the accuracy corresponding to each threshold combination, a target threshold combination is selected from all threshold combinations. The candidate similarity threshold in the target threshold combination is updated to the first similarity threshold, and the candidate open set threshold in the target threshold combination is updated to the preset open set threshold.

7. A category detection device based on self-learning, characterized in that, The device includes: The acquisition module is used to input the target image into the first detection model to obtain the location information corresponding to the target object in the target image, extract the sub-image corresponding to the location information from the target image, and input the sub-image into the first feature extraction model to obtain the target features corresponding to the target object; The determination module is used to determine the similarity between the target feature and each registered feature in the feature set; if the highest similarity is not less than a first similarity threshold, the target category of the target object is determined to be the labeled category corresponding to the registered feature corresponding to the highest similarity; if the highest similarity is less than the first similarity threshold, the open set score corresponding to the target object is determined based on the target feature and each registered feature in the feature set; the target category of the target object is determined based on the open set score to be the labeled category corresponding to the registered feature corresponding to the highest similarity, or the target category is determined to be an unregistered category. The device further includes: a processing module, configured to, if the self-learning triggering condition for category detection is met, input sample images from the open set dataset into a second detection model to obtain location information corresponding to sample objects in the sample images; extract sub-images corresponding to the location information from the sample images; input the sub-images into a second feature extraction model to obtain sample features corresponding to the sample objects; wherein the open set dataset includes multiple sample images, and the sample objects in each sample image correspond to unregistered categories, and the open set dataset is constructed based on target images corresponding to target objects of unregistered categories; cluster the sample features corresponding to multiple sample objects to obtain an initial cluster set; filter the sample features in the initial cluster set to obtain a target cluster set corresponding to the initial cluster set; retrain the first detection model and the first feature extraction model based on the target cluster set to obtain a retrained first detection model and a first feature extraction model; and / or determine a new category corresponding to the target cluster set, update the sample features in the target cluster set to the feature set, and use the new category as the label category corresponding to the sample features.

8. The apparatus according to claim 7, Its features are, in, The determining module determines the target category of the target object as the labeled category corresponding to the registered feature corresponding to the highest similarity based on the open set score value, or, when determining the target category as an unregistered category, it is specifically used as follows: if the open set score value is not greater than a preset open set threshold, then the target category of the target object is determined as the labeled category corresponding to the registered feature corresponding to the highest similarity; or, if the open set score value is greater than the preset open set threshold, then the target category is determined as an unregistered category. Specifically, when the determining module determines the open set score of the target object based on the target feature and each registered feature in the feature set, it is used to: determine the distance between the target feature and the central feature of each labeled category; wherein the feature set includes registered features of multiple labeled categories, and each labeled category corresponds to multiple registered features, and the central feature of each labeled category is obtained based on the multiple registered features of that labeled category; for each labeled category, determine the weight coefficient corresponding to that labeled category based on the distance corresponding to that labeled category and the distances corresponding to all labeled categories, determine the feature difference score between the target feature and the central feature of that labeled category based on the distance corresponding to that labeled category and the weight coefficient; and determine the open set score based on the feature difference score corresponding to each labeled category. Specifically, when the processing module filters the sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set, it filters first-type sample features and / or second-type sample features from all sample features in the initial cluster set to obtain the target cluster set corresponding to the initial cluster set. The process for determining the first-type sample features includes: for each sample feature in the initial cluster set, determining the similarity between the sample feature and each registered feature in the feature set; if the highest similarity is not less than a second similarity threshold, then the sample feature is determined as a first-type sample feature. The process for determining the second-type sample features includes: for each sample feature in the initial cluster set, determining the local density corresponding to the sample feature; if the local density corresponding to the sample feature is less than a density threshold, then the sample feature is determined as a second-type sample feature. Specifically, when the processing module retrains the first detection model and the first feature extraction model based on the target cluster set to obtain the retrained first detection model and the first feature extraction model, it performs the following steps: selecting sample images from the open set dataset that correspond to the sample features in the target cluster set, inputting the sample images into the first detection model to obtain the location information corresponding to the sample objects in the sample images, extracting sub-images corresponding to the location information from the sample images, inputting the sub-images into the first feature extraction model to obtain the predicted features corresponding to the sample objects; determining a target loss value based on the sample features in the target cluster set and the predicted features, and adjusting the network parameters of the first detection model and the first feature extraction model based on the target loss value to obtain the retrained first detection model and the first feature extraction model. The processing module is further configured to: acquire a test dataset, which includes multiple test images and the true category of the test object in each test image; determine the test features corresponding to the test object in each test image based on the retrained first detection model and the retrained first feature extraction model; acquire multiple candidate similarity thresholds and multiple candidate open set thresholds; for each threshold combination, which includes a candidate similarity threshold and a candidate open set threshold, determine the predicted category of the test object based on the candidate similarity threshold, the candidate open set threshold, and the test features corresponding to the test object; determine the accuracy corresponding to the threshold combination based on the predicted categories and true categories of all test objects; and select a target threshold combination from all threshold combinations based on the accuracy corresponding to each threshold combination, update the candidate similarity threshold in the target threshold combination to the first similarity threshold, and update the candidate open set threshold in the target threshold combination to the preset open set threshold.

9. An electronic device, characterized in that, include: A processor and a machine-readable storage medium storing machine-executable instructions that can be executed by the processor; wherein the processor is configured to execute the machine-executable instructions to implement the method of any one of claims 1-6.