Man-machine collaborative machine visual perception method and device based on multi-granularity learning
Through the human-machine collaborative machine vision perception method based on multi-grain learning, compact intra-class features and scattered inter-class features are generated, and combined with the active learning module to select decision boundary samples for annotation, the problem of expensive data annotation and inconsistent distribution in open scenarios is solved, and the accuracy and robustness of robot visual perception is improved.
Patent Information
- Application Number
- CN202510349538.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
Existing robot visual perception methods have the problem of large-scale data annotation and inconsistent data distribution in open scenarios, which makes it difficult to accurately identify task-related categories, affecting the recognition ability and robustness of intelligent robots.
Using a human-machine collaborative machine vision perception method based on multi-grain learning, compact in-class features and scattered inter-class features are generated through the feature center of mass module. Combined with the multi-grain size module and the active learning module, the feature similarity and boundary distance of the sample are calculated, the samples located at the decision boundary are selected for annotation, and the training set is updated to improve the classification ability of the model.
The accuracy of the model's identification of known categories is improved, and the decision-making ability and robustness of intelligent robots in complex environments is enhanced, especially in the real world, which is improved by 3%.
Smart Images

Figure CN120279315A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent robot visual perception in machine learning, and particularly to a human-machine collaborative machine vision perception method and device based on multi-granularity learning. Background Art
[0003] The visual perception system is a core component of an intelligent robot. This system enables the intelligent robot to obtain spatial information in a complex real-world scenario through a visual sensor, thereby accurately depicting the surrounding environment. With the aid of the visual perception system, the intelligent robot will possess the ability to observe the scenario in real time, which can significantly improve its intelligence level. Visual perception technology utilizes machine learning to achieve the perception of scene targets and objects. Currently, existing robot visual perception systems are mainly applied to closed-set environments, that is, the training set and test set samples are consistent in terms of category and distribution.
[0004] Existing robot visual perception methods based on deep learning obtain environmental data through a visual sensor, and after learning these data through a deep neural network, they judge the current environmental state. Although the performance of existing methods has been continuously breaking through, the following problems still exist in the practical application of existing methods:
[0005] (1) Labeled data is expensive. The improvement of the performance of existing robot visual perception methods based on deep learning depends on large datasets with manually labeled tags. However, it is very expensive and time-consuming to label large-scale data with high-quality annotations. Therefore, learning with limited labeled data in the real world is a major challenge.
[0006] (2) The data distribution in the open scenario is inconsistent. Existing closed-set robot visual perception methods cannot accurately distinguish task-irrelevant categories from task-relevant categories in an open scenario, but tend to select them as annotation objects because there is more uncertainty. The inability to accurately identify task-irrelevant categories will result in a waste of the annotation budget.
[0007] Therefore, in the open scenario of the real world, how to solve the problems of expensive large-scale data annotation in the field of visual perception and accurately identify task-relevant categories in the scenario is the key to enhancing the target recognition ability of intelligent robots. Summary of the Invention
[0008] The present invention provides a human - machine collaborative machine vision perception method and device based on multi - granularity learning. The present invention designs an active learning method based on a multi - granularity feature tree, which constrains the active learning process to find sample annotations that are similar to the target category and located on the category boundary. The annotated samples are used to update the model, enabling the model to accurately discover and learn the sample features that are similar to the target category and easily confused in the open set under a small - scale network structure, improving the accuracy of the model in recognizing known classes and the ability to detect unknown classes in the real world, further enhancing the detection ability of intelligent robots for known and unknown classes in the real world, enabling intelligent robots to make more accurate and intelligent decisions in complex environments, and improving the robustness of intelligent robots. The details are described below:
[0009] In a first aspect, a human - machine collaborative machine vision perception method based on multi - granularity learning, the method includes:
[0010] Obtain the features of the training set samples through the feature extractor in the detector module, and calculate the feature centroid by the feature centroid module;
[0011] Calculate the closed - set entropy score for the unlabeled set samples through the binary classifier in the detector module, and obtain the multi - granularity feature tree of the unlabeled set samples through the multi - granularity module;
[0012] Calculate the feature similarity score and the boundary distance score through the feature centroid module and the multi - granularity feature tree, calculate the active query strategy through the active learning module using the boundary distance score, the feature similarity score, and the closed - set entropy score, obtain the samples in the unlabeled set that are most similar to the known classes and located on the decision boundary through the active query strategy and annotate them, and update the training set after annotation; the robot recognizes the environment based on the updated training set.
[0013] Among them, the feature centroid module generates the feature centroids of all known classes through clustering, and makes the intra - class features of known classes compact and the inter - class features dispersed through the sample large - margin loss.
[0014] Among them, the detector module uses the binary classifier inside to learn the boundary between target categories with the sample features, updates the positive decision boundary and the nearest negative decision boundary of each sample through the binary cross - entropy loss function, and makes the unknown class samples in the training set obtain a high closed - set entropy through the closed - set entropy loss function; the detector module obtains the sample features of each sample in the unlabeled set, and provides the sample features to the internal binary classifier, the active learning module, and the multi - granularity module, and the binary classifier calculates the closed - set entropy score of the samples.
[0015] Among them, the multi-granularity module obtains the multi-granularity feature tree of the unlabeled set samples. The multi-granularity feature tree calculates the similarity score between the sample feature and the feature centroid through cosine similarity, and calculates the boundary distance score between the sample feature and the feature centroid through the ratio of the two nearest Euclidean distances in the multi-granularity feature tree.
[0016] In a second aspect, a human-machine collaborative machine vision perception device based on multi-granularity learning, the device includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to enable the device to execute the method described in any one of the first aspect.
[0017] In a third aspect, a computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is enabled to execute the method described in any one of the first aspect.
[0018] The beneficial effects of the technical solution provided by the present invention are as follows:
[0019] 1. The feature centroid module proposed by the present invention obtains the feature centroid according to the training set samples. The feature centroid has compact intra-class features and dispersed inter-class features, making it easier to distinguish between different categories; the multi-granularity module obtains the multi-granularity feature tree according to the unlabeled sample features, and preliminarily divides the boundary between in-distribution and out-of-distribution samples;
[0020] 2. The active learning module calculates the similarity and boundary distance of the samples in the multi-granularity feature tree to the target category centroid. The feature similarity helps the model determine whether the unlabeled sample belongs to the target category, and the boundary distance helps the model determine the information content of the unlabeled sample, improving the accuracy of the robot vision perception method in distinguishing known classes; the binary classifier uses hard negative classifier sampling to help the model learn a more accurate category boundary, improving the problem in the prior art that the recognition of easily confused target categories is inaccurate;
[0021] 3. The active learning strategy used in the visual perception technology proposed by the present invention pays more attention to the samples at the category boundary in the dataset compared with the active learning strategies of the existing technologies, thereby improving the classification ability of the model and making the robot more intelligent after being applied;
[0022] 4. The visual perception technology proposed by the present invention has increased the accuracy of distinguishing known category objects in the environment by up to 3% compared with the existing technologies. Description of the Drawings
[0023] Figure 1 It is a flowchart of a human-machine collaborative machine vision perception method based on multi-granularity learning;
[0024] Figure 2It is an architecture diagram of human-machine collaborative machine vision perception based on multi-granularity learning;
[0025] Figure 3 It is a structural schematic diagram of a human-machine collaborative machine vision perception device based on multi-granularity learning.
[0026] In the accompanying drawings, the list of components with each label is as follows:
[0027] 210: Sample;
[0028] 220: Detector module;
[0029] 230: Feature centroid module;
[0030] 240: Active learning module;
[0031] 250: Multi-granularity module;
[0032] 260: Target model;
[0033] 310: Processor; 320: Memory;
[0034] 330: Sensor; 340: Bus. Detailed implementation manners
[0035] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0036] Embodiment 1
[0037] To solve the problems existing in the background technology, most of the existing robot vision perception methods need to satisfy that the distributions of the training and test sets and the class distributions are consistent. However, in real scenarios, it is very difficult to achieve this constraint. For related technologies such as open sets and active learning, there are also problems such as inaccurate discrimination of difficult samples at the class boundary and insufficient correlation between the queried samples and the target class.
[0038] An embodiment of the present invention proposes a human - machine collaborative machine vision perception method based on multi - granularity learning. The features of the training set samples are obtained through the feature extractor inside the detector module 220. The feature centroid is calculated through the feature centroid module 230 for the sample features. The closed - set entropy score of the unlabeled set samples is calculated through the binary classifier inside the detector module 220 for the unlabeled set sample features. The multi - granularity feature tree of the unlabeled set samples is obtained through the multi - granularity module 250. The feature similarity score and the boundary distance score are calculated through the feature centroid and the multi - granularity feature tree of the unlabeled set samples. The active query strategy is calculated through the active learning module 240 for the boundary distance score, the feature similarity score, and the closed - set entropy score. Using the samples obtained by querying with the active learning module 240 to train the model can improve the classification performance of the robot vision perception model for known classes in the real scene, further improve the robustness of the intelligent robot, and enable the intelligent robot to make more accurate and reasonable decisions in a complex real - world environment.
[0039] See Figure 1 , the method includes the following steps:
[0040] 101: Generate feature centroids;
[0041] Among them, the feature extractor inside the detector module 220 obtains the sample features of each sample 210 in the training set and provides the sample features to the feature centroid module 230. The feature centroid module 230 generates the feature centroids of all known classes through clustering and makes the intra - class features of known classes compact and the inter - class features dispersed through the sample large - margin loss.
[0042] 102: Sample closed - set entropy constraint;
[0043] Among them, the feature extractor inside the detector module 220 obtains the training set sample features and provides the sample features to the detector module 220. The detector module 220 learns the boundaries between target classes using the sample features through the internal binary classifier, updates the positive decision boundary and the nearest negative decision boundary of each sample through the binary cross - entropy loss function, and makes the unknown class samples in the training set obtain high closed - set entropy through the closed - set entropy loss function to reduce the interference of unknown class samples.
[0044] 103: Active learning strategy;
[0045] Among them, the feature extractor inside the detector module 220 obtains the sample features of each sample in the unlabeled set, and provides the sample features to the internal binary classifier, active learning module 240, and multi-granularity module 250. The binary classifier calculates the closed-set entropy score of the sample; the multi-granularity module 250 obtains the multi-granularity feature tree of the unlabeled set samples. The multi-granularity feature tree calculates the similarity score between the sample feature and the feature centroid through cosine similarity, and calculates the boundary distance score between the sample feature and the feature centroid through the ratio of the two nearest Euclidean distances in the multi-granularity feature tree. The active learning module 240 combines the similarity score, boundary distance score, and closed-set entropy score to obtain an active query strategy, and obtains the samples in the unlabeled set that are most similar to the known classes and located on the decision boundary through the active query strategy for annotation. After annotation, the training set is updated.
[0046] 104: Robot environment recognition.
[0047] Among them, the visual perception sensor 330 installed on the robot obtains the surrounding environment information, and transmits the information to the memory 320 through the bus 350. The memory 320 transmits the information to the processor 310 according to the bus 350. The processor 310 identifies the most valuable known samples in the environment and filters the unknown samples according to the method proposed in the embodiment of the present invention. The identified samples will be transmitted to the memory 320, and after being processed by the processor 310, the robot is trained to improve the robot's perception ability of the known class samples in the environment and the robustness of the unknown class samples.
[0048] In summary, through the above steps 101-step 104, the embodiment of the present invention realizes calculating the boundary distance score, feature similarity score, and closed-set entropy score of the unlabeled samples to the known class feature centroid based on the binary classifier and multi-granularity module, so that the model can learn the difficult samples obtained by querying according to the active learning module, improve the classification performance of the robot vision perception model for the known classes in the real scene, further improve the robustness of the intelligent robot, enable the intelligent robot to make more accurate and reasonable decisions in the complex real-world environment, and improve the safety of the robot.
[0049] Embodiment 2
[0050] The following combines Figure 2 , calculation formulas, and tables to further introduce the solution in Embodiment 1, as described in detail below:
[0051] I. Model framework
[0052] The model framework consists of a detector module 220, a feature centroid module 230, an active learning module 240, a multi-granularity module 250, and a target model 260, as Figure 1 shown. The following components are the core of the method:
[0053] 1. Detector module 220
[0054] In the embodiment of the present invention, the sample features are obtained through the feature extractor inside the detector module 220. The binary classifier in the detector module 220 further uses the sample features to learn the boundaries between target categories and assigns high closed-set entropy to unknown class samples, enhancing the robot's recognition ability for known class samples and robustness to unknown class samples.
[0055] 2. Feature centroid module 230
[0056] In the embodiment of the present invention, the feature centroids of target categories are obtained through the feature centroid module 230. The feature centroids judge whether the unlabeled samples are located on the decision boundaries of target categories in the active learning module 240 and judge the feature similarity between the unlabeled samples and the feature centroids, facilitating the robot to discover samples that are similar to known classes and located on the decision boundaries.
[0057] 3. Active learning module 240
[0058] The method proposed in the embodiment of the present invention combines the feature similarity score, boundary distance score, and closed-set entropy score through the active learning module 240, and can obtain known class samples at the classification boundary from the unlabeled sample set. By labeling the samples and updating the model, the classification ability of the robot for known class samples in real-world scenarios is enhanced.
[0059] 4. Multi-granularity module 250
[0060] The method proposed in the embodiment of the present invention combines the multi-granularity feature tree in the multi-granularity module 250 with the target category feature centroids in the feature centroid module 230 to obtain the feature similarity score and boundary distance score, facilitating the preliminary division of samples in the unlabeled set and further providing guarantee for the active learning module 240 to select samples.
[0061] II. Data division, sampling, and construction of multi-granularity structure
[0062] 1. Data set
[0063] The visual perception method uses the CIFAR10, CIFAR100, and TinyImageNet robot vision perception data sets. The specific details are as follows:
[0064] (1) CIFAR10 data set. It consists of 60,000 32×32 color images of 10 different categories, with 5,000 training images and 1,000 test images for each category. These categories include common items such as airplanes, cars, and cats. In the experiment, 4 categories are selected as known categories.
[0065] (2) CIFAR100 dataset. It consists of 60,000 32×32 color images in 100 different categories, with 500 training images and 100 test images for each category. These categories include common items and portraits in life, such as bridges, mushrooms, and men. In the experiment, 40 categories are selected as known categories.
[0066] (3) TinyImageNet dataset. It is a large dataset containing 100,000 training images and 20,000 test images in 200 categories. Each category has 500 training images, 50 validation images, and 50 test images.
[0067] III. Feature extractor inside the detector module 220
[0068] The visual perception method uses a Residual Network (ResNet) as the feature extractor. For all datasets in the experiments proposed in the embodiments of the present invention, the ResNet18 network is used to extract features.
[0069] Among them, the ResNet18 network is well-known to those skilled in the art, and the embodiments of the present invention will not elaborate on it.
[0070] IV. Feature centroid module 230
[0071] In the robot visual perception method, the feature centroid module 230 calculates the feature centroid of the target category according to the sample features obtained by the feature extractor inside the detector module 220. By alternately updating the sample features and the feature centroid, the distance between each sample feature and the corresponding category centroid is minimized, and the distance to other centroids is maximized, where the category boundary is calculated by the large margin loss; during training, the internal classes of the same category and samples of different categories are sampled to form a mini-batch. The sampled samples are grouped according to the categories they belong to, and the feature centroid of each group is updated by the sample features of this batch.
[0072] The calculation method of the large margin loss function of the sample is as follows:
[0073] L LM (z, {c i}) = max(0, ∑ i=y ‖z - c i ‖ - ∑ i≠y ‖z - c i ‖ + m) (1)
[0074] Among them, z is the sample feature; {c i} represents the feature centroid of category i; y represents the category of the sample; m represents the margin threshold. The calculation method of the centroid update process is as follows:
[0075]
[0076] Among them, I(·) represents the indicator function, B represents the batch size, and {c i} represents the feature centroid of class i.
[0077] V. Detector Module 220
[0078] In the robot vision perception method, the detector module 220 is shown as follows. The detector module 220 consists of a feature extractor and K binary classifiers, where K is the number of target classes. The i-th binary classifier uses the samples of the i-th class in the known classes as positive samples and the samples of other known classes as negative samples. Since updating the boundaries for other known classes each time will cause confusion in the boundaries learned by the classifier, the vision perception method uses hard negative classifier sampling and only updates the nearest class boundary for each binary classifier, which can ensure learning accurate class boundaries. The binary cross-entropy loss function L bce is calculated as follows:
[0079]
[0080] Among them, (x i , y i ) represents the sample and its label, n l represents the number of samples in the known classes, D L represents the set of samples in the known classes for active query, p i represents the probability that the sample is classified as a positive class by the i-th binary classifier, and the calculation method is:
[0081] p i = σ(G i (f)) (4)
[0082] Among them, σ represents the Softmax function, f = F(x) represents the feature extractor. G i represents the i-th binary classifier.
[0083] To ensure that the binary classifier outputs a high known entropy for unknown class samples, the closed-set entropy loss of the unknown class samples is obtained from the sample features based on the binary classifier, and the calculation method is:
[0084]
[0085] Among them, K represents the number of known class categories, and n open represents the number of unknown class samples.
[0086] VI. Active Learning Module 240
[0087] The visual perception method obtains the multi-granularity feature tree of unlabeled samples according to the multi-granularity module 250, calculates the boundary distance score and feature similarity score from the sample features to the feature centroid according to the multi-granularity feature tree, obtains the closed-set entropy of the unlabeled samples according to the binary classifier inside the detector module 220, obtains the active query strategy according to the boundary distance score, feature similarity score and closed-set entropy score of the unlabeled samples, and selects samples from the unlabeled set according to the active query strategy to be labeled by an expert; the labeled samples are used to update the model.
[0088] The calculation method of the closed-set entropy of unlabeled samples is as follows:
[0089]
[0090] Among them, x represents the sample, K represents the number of known class categories, and H i represents the entropy of the i-th binary classifier, and the calculation method is:
[0091] H i (x) = -p i ·log(p i ) - (1 - p i )·log(1 - p i ) (7)
[0092] Among them, p i represents the probability that the sample is classified as the positive class by the i-th binary classifier.
[0093] The calculation method of the boundary distance of unlabeled samples is as follows:
[0094]
[0095] Among them, x represents the sample, respectively represent the first two closest distances between the sample features and the centroids of the known class features, and the distance uses the Euclidean distance.
[0096] The calculation method of the feature similarity score of unlabeled samples is as follows:
[0097]
[0098] Among them, c i represents the centroid of the existing class features, z represents the unlabeled sample features, <·,·> represents the dot product, and ‖·‖ represents the norm.
[0099] The calculation method of the active learning strategy is as follows:
[0100] U(x) = U sim ×U dist -U info (11)
[0101] VII. Target Model
[0102] In the robot vision perception method, the target model uses the feature extractor inside the detector module 220 to extract the features of the training set samples, and calculates the cross-entropy loss for the sample features. The calculation method of the cross-entropy loss function is as follows:
[0103]
[0104] where (x i , y i ) represents the sample and its label, n l represents the number of known class samples, and f(·) represents the feature extractor.
[0105] VIII. Training Strategy
[0106] The training set is denoted as The number of samples in the training set is represented by N L . In the training stage, the samples are sent into the detector module 220 and the target model simultaneously. In the detector module 220, the samples first pass through the internal feature extractor, and the extracted sample features are input into the feature centroid module 230 and the binary classifier inside the detector module 220. Using the extracted sample labels and features, the feature centroid c i of each target class is obtained. The unlabeled set is denoted as The number of samples in the unlabeled set is represented by N U . After using the feature extractor inside the detector module 220 to extract the sample features, the sample features are sent into the multi-granularity module 250, and the boundary distance U dist and feature similarity U sim of the samples are calculated using the multi-granularity feature tree and the feature centroid. The known entropy U info of the samples is calculated using the binary classifier. By combining the boundary distance, feature similarity, and known entropy of the unlabeled set samples, an active query strategy is obtained. According to the active query strategy, unlabeled samples are selected, and the labeled samples are used to update the detector module 220 and the target model. In the detector module 220, the large margin loss L LM , binary cross-entropy loss L bce , and closed-set entropy loss L em are calculated. In the target model, the samples pass through the internal feature extractor, and the extracted sample features are input into the target model to calculate the cross-entropy loss L ce . The total loss L is iteratively optimized through backpropagation. The following Algorithm 1 summarizes the working process of the training process applied to this invention.
[0107]
[0108]
[0109] In summary, the embodiments of the present invention realize the construction of the feature centroid of the target category through the above parts, constrain the target categories to form clear boundaries, and form an active learning strategy based on the boundaries and sample features, so that the model can accurately query and learn the most valuable samples for the task in the data set under a small-scale network structure, improve the model's ability to distinguish known classes and the model's ability to recognize unknown classes in real scenes, and further improve the intelligent robot's ability to recognize target category objects in the real world, so that the intelligent robot can make more accurate and reasonable decisions in complex environments, thereby enhancing the robustness of the intelligent robot.
[0110] Example 3
[0111] In combination with Table 1, Table 2, and Table 3, specific experimental data are used to verify the feasibility of the schemes in Examples 1 and 2, as described below:
[0112] First, the details of the experiment, using the ResNet18 network as the backbone network of the visual perception method. In terms of optimization, the gradient descent optimizer is used, the momentum value is set to 0.9, the initial learning rate is set to 0.01, the batch size is set to 128, and a total of 100 iterations are performed. In all experiments, 10 cycles of active sampling are performed, and 1500 samples are queried for annotation in each cycle. Other parameters are set by default.
[0113] For each data set, some randomly selected categories are considered known and the rest are unknown, and the mismatch rate is used for category selection. The mismatch rate is defined as Where |K| is the number of known categories and |U| is the number of unknown categories. For CIFAR-10, CIFAR-100, and TinyImageNet, 1%, 8%, and 8% of samples are randomly sampled from known categories to initialize the labeled datasets, respectively. Each experiment is performed four times, and the known / unknown category division is the same for all compared methods.
[0114] The second is the robot visual perception method compared with the visual perception method, including: Random, Uncertainty, Certainty, Coreset, BALD, LfOSA, EOAL. The specific details are as follows:
[0115] (1) Random: Randomly select samples from the unlabeled set for labeling.
[0116] (2) Uncertainty: Samples with the largest uncertainty in neural network prediction entropy are selected from the unlabeled set for labeling.
[0117] (3) Certainty: Select samples with the highest certainty of neural network prediction from the unlabeled set for labeling.
[0118] (4) Coreset: Measure the diversity of samples in the unlabeled set through clustering methods, and select the most representative samples for annotation.
[0119] (5) BALD: Use Dropout Bayesian inference approximation as an active learning strategy.
[0120] (6) LfOSA: Use a Gaussian mixture model to reject open-set samples and achieve higher known-class purity in the query set.
[0121] (7) EOAL: Propose an entropy-based open-set active learning framework that uses the distributions of known and unknown classes as an active learning strategy to select high-information samples in the unlabeled set.
[0122] In addition, the evaluation metric in the multi-granularity inference method is: Acc. Acc (Accuracy) is one of the commonly used evaluation metrics in the field of machine learning and represents the accuracy rate. It is a metric that measures the proportion of samples predicted correctly by the classification model on the test set, especially suitable for classification problems with relatively balanced class distributions.
[0123] For the classification task, the visual perception method needs to identify samples from classes that were not learned during the training of each dataset. In each split of the classification experiment, a ResNet18 network with an initial learning rate of 0.01 is trained, with a fixed batch size of 128, and a multi-step learning rate scheduler is used. The model is iterated 100 times, and the learning rate is adjusted during the 60th iteration of training, with a decay factor set to 0.1. The optimizer is the stochastic gradient descent optimizer, with a Newton momentum of 0.9 and a weight decay of 5e-4. For the hyperparameters, the scaling parameter λ of the total loss function is set to 0.1, and the margin threshold m in the large margin loss function is set to 10.
[0124] Tables 1, 2, and 3 show the Acc using the average value of four trials under different mismatch rates on the public dataset, comprehensively reflecting the performance of the model provided by the embodiments of the present invention. The results prove that the robot visual perception method exceeds the existing technologies in terms of the accuracy performance of identifying known classes.
[0125] Table 1 Comparison table of methods for known-class classification performance under the Acc evaluation metric and a mismatch rate of 20%
[0126]
[0127] Table 2 Comparison table of methods for known-class classification performance under the Acc evaluation metric and a mismatch rate of 30%
[0128]
[0129] Table 3 Comparison Table of Known Class Classification Performance under the Acc Evaluation Index and a Mismatch Rate of 40%
[0130]
[0131] Among them, the best performance is shown in bold.
[0132] Example 4
[0133] A human - machine collaborative machine vision perception device based on multi - granularity learning. Refer to Figure 2 , the device includes: a vision perception sensor, a processor, a memory, and a bus. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method described above.
[0134] Obtain the features of the training set samples through the feature extractor in the detector module, and calculate the feature centroid by the feature centroid module for the sample features;
[0135] Calculate the closed - set entropy score for the unlabeled set sample features through the binary classifier in the detector module, and obtain the multi - granularity feature tree of the unlabeled set samples through the multi - granularity module;
[0136] Calculate the feature similarity score and the boundary distance score through the feature centroid module and the multi - granularity feature tree, calculate the active query strategy through the active learning module using the boundary distance score, the feature similarity score, and the closed - set entropy score, obtain the samples in the unlabeled set that are most similar to the known classes and located on the decision boundary through the active query strategy and label them, and update the training set after labeling; the robot identifies the environment based on the updated training set.
[0137] Among them, the feature centroid module generates the feature centroids of all known classes through clustering, and makes the intra - class features of the known classes compact and the inter - class features dispersed through the sample large - margin loss.
[0138] Among them, the detector module uses the binary classifier inside to learn the boundaries between target classes with the sample features, updates the positive decision boundary and the nearest negative decision boundary of each sample through the binary cross - entropy loss function, and makes the unknown class samples in the training set obtain a high closed - set entropy through the closed - set entropy loss function; the detector module obtains the sample features of each sample in the unlabeled set and provides the sample features to the internal binary classifier, active learning module, and multi - granularity module, and the binary classifier calculates the closed - set entropy score of the sample.
[0139] Among them, the multi - granularity module obtains the multi - granularity feature tree of the unlabeled set samples. The multi - granularity feature tree calculates the similarity score between the sample features and the feature centroid through cosine similarity, and calculates the boundary distance score between the sample features and the feature centroid through the ratio of the two nearest Euclidean distances in the multi - granularity feature tree.
[0140] It should be noted here that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments, and the embodiments of the present invention will not be elaborated herein.
[0141] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as a computer, a single-chip microcomputer, a microcontroller, etc. The execution subject of the above-mentioned sensor needs to support sampling of components during the operation of industrial equipment, such as a bearing vibration sensor. In specific implementation, the embodiments of the present invention do not limit the execution subject, and it can be selected according to the needs in actual applications.
[0142] In specific implementation, the embodiments of the present invention do not limit the execution subject, and it can be selected according to the needs in actual applications. Data signals are transmitted between the sensor, the memory, and the processor through a bus, and the embodiments of the present invention will not be elaborated herein.
[0143] Embodiment 5
[0144] Based on the same inventive concept, the embodiments of the present invention also provide a computer-readable storage medium. The storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the method steps in the above embodiments.
[0145] The computer-readable storage medium includes, but is not limited to, flash memory, hard disk, solid-state drive, etc.
[0146] It should be noted here that the description of the readable storage medium in the above embodiments corresponds to the method description in the embodiments, and the embodiments of the present invention will not be elaborated herein.
[0147] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.
[0148] The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium or a semiconductor medium, etc.
[0149] Except for special instructions, the embodiments of the present invention do not limit the models of each device, and any device that can perform the above functions can be used.
[0150] Those skilled in the art can understand that the attached drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.
[0151] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A human-machine collaborative machine vision perception method based on multi-granularity learning, characterized in that, The method includes: Obtaining the features of the training set samples through the feature extractor in the detector module, and calculating the feature centroid of the sample features through the feature centroid module; Calculating the closed-set entropy score of the unlabeled set samples through the binary classifier in the detector module, and obtaining the multi-granularity feature tree of the unlabeled set samples through the multi-granularity module; Calculating the feature similarity score and the boundary distance score through the feature centroid module and the multi-granularity feature tree, calculating the active query strategy by the active learning module based on the boundary distance score, the feature similarity score, and the closed-set entropy score, obtaining the samples in the unlabeled set that are most similar to the known classes and located on the decision boundary through the active query strategy and labeling them, and updating the training set after labeling; The robot identifies the environment based on the updated training set.
2. The human-machine collaborative machine vision perception method based on multi-granularity learning according to claim 1, wherein The feature centroid module generates the feature centroids of all known classes through clustering, and makes the intra-class features of the known classes compact and the inter-class features dispersed through the sample large margin loss.
3. A human-machine collaborative machine vision perception method based on multi-granularity learning according to claim 1, characterized in that, The detector module uses the binary classifier inside to learn the boundary between target classes using the sample features, updates the positive decision boundary and the nearest negative decision boundary of each sample through the binary cross-entropy loss function, and makes the unknown class samples in the training set obtain a high closed-set entropy through the closed-set entropy loss function; The detector module obtains the sample features of each sample in the unlabeled set and provides the sample features to the internal binary classifier, the active learning module, and the multi-granularity module, and the binary classifier calculates the closed-set entropy score of the sample.
4. A human-machine collaborative machine vision perception method based on multi-granularity learning according to claim 1, characterized in that The multi-granularity module obtains the multi-granularity feature tree of the unlabeled set samples. The multi-granularity feature tree calculates the similarity score between the sample features and the feature centroid through cosine similarity, and calculates the boundary distance score between the sample features and the feature centroid through the ratio of the two nearest Euclidean distances in the multi-granularity feature tree.
5. A human-machine collaborative machine vision perception device based on multi-granularity learning, characterized in that, The device includes: a processor and a memory. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method described in any one of claims 1-4.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by the processor, the processor executes the method described in any one of claims 1-4.