Zero-Shot Object Detection Method Based on Repository Robust Region Feature Synthesizer
By introducing intra-class semantic divergence and inter-class structure maintenance modules into the object detection network, and using the repository optimization feature generator, the problem of insufficient accuracy of existing methods in complex detection scenarios is solved, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202310254266.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-16
AI Technical Summary
The existing zero-sample object detection method is not effective in complex and open detection scenarios, and fails to effectively synthesize intra-class diversity and inter-class discriminable visual features, resulting in insufficient detection accuracy.
Build an object detection network, add in-class semantic divergence modules and inter-class structure maintenance modules, optimize feature generators through repository optimization, and use cross-modal contrast enhancement loss training generators and discriminators to synthesize in-class diversity, inter-class discrimination and cross-domain discrimination visual features.
Improves the accuracy and robustness of zero-sample object detection, especially in complex and generalized zero-sample detection scenarios.
Smart Images

Figure CN116416619B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a zero-shot object detection method based on a repository robust region feature synthesizer. Background Art
[0002] With the rapid development of Internet technology, data containing visual information such as videos and images has grown exponentially. The need to extract and represent visual objects with semantic information from massive visual data, and then realize the perception and understanding of visual content, is becoming increasingly strong. However, object detection algorithms based on deep learning require a large amount of training data. Therefore, using zero-shot object detection algorithms for visual perception has great application value.
[0003] Existing zero-shot object detection methods, such as Gtnet: Generative transfer network for zero-shot object detection proposed by Shizhen Zhao et al. in 2020 and Synthesizing the Unseen for Zero-shot Object Detection proposed by Nasir Hayat et al. in 2020. These methods first use a generative model to synthesize visual features of unseen classes, and then use the synthesized visual features to train an unseen class classifier in a supervised manner to achieve the detection of unseen classes. Although these methods have achieved good results, they mainly follow the ideas proposed in the zero-shot classification framework and do not consider the high intra-class diversity, simultaneous occurrence of complex categories, and open test scenarios in real object detection scenarios. Therefore, the synthesized visual features may perform well in less complex classification scenarios, but are not sufficient to obtain satisfactory results in actual complex detection scenarios. Summary of the Invention
[0004] To overcome the deficiencies of the prior art, the present invention provides a zero-shot object detection method based on a repository robust region feature synthesizer. A target detection network model is constructed, and semantic descriptors and random noise are input. In the first stage of training, a feature generator is trained on visible class visual features. To obtain intra-class diverse and inter-class distinguishable features, an intra-class semantic divergence module and an inter-class structure preservation module are added to constrain the learning process. After training, this feature generator is used to synthesize invisible class visual features; in the second stage of training, a repository is constructed to store real visible class storage units and the invisible class storage units synthesized in the first stage. The cross-modal contrast enhancement loss based on this repository enables the feature synthesizer to be optimized on both visible class and invisible class data simultaneously. The generator can synthesize intra-class diverse, inter-class distinguishable, and cross-domain distinguishable visual features for training a robust object detector, and finally obtain accurate zero-shot object detection results.
[0005] A zero-shot object detection method based on a repository robust region feature synthesizer, characterized by the following steps:
[0006] Step 1: Construct a target detection network model, including two parts: a target detection component and an invisible classifier learning component. Among them, the target detection component adopts the Faster r-cnn target detection network structure for object detection. In this network, there is a region proposal network for extracting region features; the invisible classifier learning component includes a generative adversarial network and an invisible class classifier network. The generative adversarial network includes a generator network and a discriminator network for synthesizing invisible class region features. The invisible class classifier network is trained using the invisible class region features synthesized by the generative adversarial network as input to obtain trained complete invisible class classifier weight parameters, and then used to update the target detection component; the network parameters are initialized using the uniform distribution initialization method;
[0007] Step 2: Divide the PASCAL VOC 2007+2012 image dataset into two sub-datasets: visible classes and invisible classes;
[0008] Step 3: Input the images in the visible class sub-dataset, as well as their corresponding labels and bounding box coordinates, into the target detection component for network training. When training, the cross-entropy loss function is used as the network's loss function, and the stochastic gradient descent algorithm is used to optimize the network parameters;
[0009] Step 4: Input the images in the visible class sub-dataset into the target detection component trained in Step 3, and use the region proposal network therein to extract the visible class region features F s ;
[0010] Step 5: Input the random noise z~N(0,1) and the visible class semantic descriptor ws ∈W s The generator network input into the generative adversarial network takes the real region feature f of the images in the visible class sub-dataset s ∈F s and the synthetic region feature input into the discriminator network of the generative adversarial network, and then the two networks are alternately iteratively trained separately to obtain the trained generator network and discriminator network; where, W s represents the visible class semantic descriptor, generated using the FastText model, represents the set of synthetic region features, composed of all region features synthesized by the generation network;
[0011] During training, the adaptive moment estimation algorithm is used to optimize the network parameters, and the loss function of the network is set as follows:
[0012]
[0013] where, G represents the generator network, D represents the discriminator network, represents the WGAN loss, represents the classification loss, represents the intra-class semantic divergence loss, represents the inter-class structure preservation loss, η is the balance coefficient of the classification loss term, set η = 0.01, λ1 is the balance coefficient of the intra-class semantic divergence loss term, set λ1 = 0.001, λ2 is the balance coefficient of the inter-class structure preservation loss term, set λ2 = 0.001;
[0014] The inter-class structure preservation loss is calculated according to the following formula:
[0015]
[0016] where, E[·] represents the expectation, represents the synthetic positive example region feature, represents the synthetic negative example region feature, a positive example refers to the region feature corresponding to the random noise input within the sphere with the random noise input z corresponding to the region feature as the center of the sphere and a radius of r, and a negative example refers to the region feature corresponding to the random noise input outside the sphere, set r = 10 -6 , N is the number of negative examples, set N = 10, τ1 is the temperature coefficient one, set τ1 = 0.1;
[0017] The inter-class structure preservation loss is calculated according to the following formula:
[0018]
[0019] Among them, g + represents the positive example region feature. The positive example refers to the region feature with the same class label as . Φ = {g j} represents the region feature set, including the real region feature and the synthesized region feature gj represents the j-th region feature, and τ2 is the temperature coefficient two, with τ2 set to 0.1;
[0020] Step 6: Input the random noise z ∼ N(0, 1) and the semantic descriptor w of the unseen class u ∈W u into the generator network trained in Step 5 to synthesize the region features of the unseen class All the region features of the unseen class form the region feature set of the unseen class Among them, W u represents the semantic descriptor of the unseen class and is generated using the FastText model;
[0021] Step 7: Calculate the initial memory unit of the asynchronous repository, including the visible class memory unit and the unseen class memory unit The calculation formulas are as follows:
[0022]
[0023]
[0024] Among them, represents the set of real region features of the k-th visible class, and f i s represents the real region feature of the i-th visible class, represents the set contains the number of features, represents the set of region features of the synthesized k-th unseen class, represents the synthesized region feature of the i-th unseen class, represents the set contains the number of features;
[0025] Step 8: Input the random noise z ∼ N(0, 1), the semantic descriptor w of the visible class s ∈W s and the semantic descriptor w of the unseen class u ∈W u into the generator network of the generative adversarial network. The input of the discriminator network is the same as that in Step 5, and then the two networks are alternately iteratively trained separately again to obtain the trained generator network and discriminator network;
[0026] When training, set the loss function of the network as follows:
[0027]
[0028] Among them, is the cross-modal contrast enhancement loss corresponding to the repository, λ3 is the balance coefficient of the cross-modal contrast enhancement loss term, and set λ3 = 0.001;
[0029] The aforementioned cross-modal contrast enhancement loss The calculation formula is as follows:
[0030]
[0031] Among them, represents the synthesized visible or invisible class region feature; u + represents the positive example memory unit, and the positive example refers to the memory unit with the same class label as the region feature ; K s represents the number of visible class memory units, and set K s = 16; represents the k-th visible class memory unit; K u represents the number of visible class memory units, and set K u = 4; represents the k-th invisible class memory unit; τ3 is the temperature coefficient three, and set τ3 = 0.05;
[0032] During training, use the following formula to update the invisible class memory unit in the repository:
[0033]
[0034] Among them, ← represents the update operation, m is the momentum update factor, and set m = 0.0001; is the set of synthesized invisible class region features;
[0035] Step 9: Input the random noise z ∼ N(0, 1) and the invisible class semantic descriptor w u ∈ W u into the generator network trained in Step 8 to synthesize the invisible class region feature All invisible class region features form the invisible class region feature set
[0036] Step 10: The invisible class region feature set obtained in Step 9 Input it into the invisible class classifier network, train this network, and obtain the trained invisible class classifier network; during training, set the loss function of the network as the cross-entropy loss function, and use the adaptive moment estimation algorithm to optimize the network parameters; then replace the corresponding classifier network parameters in the target detection component with the weight parameters of the invisible class classifier network;
[0037] Step 11: Input the zero-shot test image set to be processed into the target detection component updated in Step 10, and the regional features output by the candidate region network therein constitute the candidate region feature set F;
[0038] Step 12: Use the formula to calculate and obtain the final classification score p i of the i-th candidate region feature f i in the set F; where, is the basic classifier score of the candidate region feature f i , which is obtained by inputting f i into the basic classifier in the target detection component updated in Step 10, β is the balance coefficient, and set β = 1.0, is the class-level classification prediction score of the candidate region feature f i , and the calculation formula is as follows:
[0039]
[0040] where, u is the updated memory unit in the repository, which is obtained after the memory unit in the repository is updated in Step 8, u i represents the memory unit with the same class as the i-th candidate feature f i , u k represents the memory unit corresponding to the k-th class in the memory bank, and K is the total number of memory unit classes.
[0041] The beneficial effects of the present invention are as follows: Since the intra-class semantic divergence module and the inter-class structure preservation module constructed optimize the intra-class and inter-class relationships of the synthetic features, and by further adding the repository to the optimization process of the region feature synthesizer, the cross-modal contrast enhancement loss enables the feature synthesizer to be optimized on both visible class and invisible class data at the same time, so that the feature synthesizer can finally synthesize region visual features with intra-class diversity, inter-class distinguishability, and cross-domain distinguishability; meanwhile, the updated repository memory unit is used to calculate the class-level classification score, which can further improve the final detection performance. Compared with the existing zero-shot object detection methods, the present invention has higher detection accuracy and better robustness when dealing with more challenging generalized zero-shot object detection scenarios. Description of the Drawings
[0042] Figure 1It is the flowchart of the zero-shot object detection method based on the repository robust region feature synthesizer of the present invention;
[0043] Figure 2 It is the experimental result diagram of the method of the present invention. Detailed implementation manners
[0044] The present invention will be further described below in conjunction with the drawings and embodiments. The present invention includes but is not limited to the following embodiments.
[0045] The present invention provides a zero-shot object detection method based on a repository robust region feature synthesizer, which combines a repository into a region feature generator model to synthesize visually distinctive features with inter-class diversity, inter-class distinguishability, and cross-domain distinguishability to complete the zero-shot object detection task. As Figure 1 shown, the specific implementation process is as follows:
[0046] Step 1: Construct a target detection network model, including two parts: a target detection component and an unseen classifier learning component. Among them, the target detection component adopts the Faster r-cnn target detection network structure for target detection. In this network, there is a region proposal network for extracting region features. The Faster r-cnn target detection network structure is described in the work "Faster r-cnn: Towards real-time object detection with region proposal networks" by Ren, Shaoqing, etc. in 2015.
[0047] The unseen classifier learning component uses the network architecture used by Nasir Hayat et al. in their 2020 work "Synthesizing the Unseen for Zero-shot Object Detection", including a generative adversarial network and an unseen class classifier network. The generative adversarial network includes a generator network and a discriminator network for synthesizing unseen class region features. The unseen class classifier network is trained using the unseen class region features synthesized by the generative adversarial network as input to obtain the trained weight parameters of the unseen class classifier, and then used to update the target detection component.
[0048] The network parameters are initialized using the uniform distribution initialization method.
[0049] Step 2: Divide the PASCAL VOC 2007+2012 image dataset into two sub-datasets: visible classes and invisible classes. The dataset is sourced from http: / / host.robots.ox.ac.uk / pascal / VOC / . This dataset contains a total of 20 classes of objects. Among them, the four classes of "car", "dog", "sofa", and "train" are invisible classes, and the remaining 16 classes are visible classes.
[0050] Step 3: Input the images in the visible-class sub-dataset, along with their corresponding labels and bounding box coordinates, into the object detection component for network training. During training, use the cross-entropy loss function as the network's loss function and the stochastic gradient descent algorithm for network parameter optimization.
[0051] Step 4: Input the images in the visible-class sub-dataset into the object detection component trained in Step 3, and use the region proposal network therein to extract the visible-class region features F s .
[0052] Step 5: Input the random noise z ∼ N(0, 1) and the visible-class semantic descriptor w s ∈W s into the generator network in the generative adversarial network. Input the real region features f s ∈F s of the images in the visible-class sub-dataset and the synthetic region features into the discriminator network in the generative adversarial network. Then, perform separate alternating iterative training on the two networks to obtain the trained generator network and discriminator network. Among them, W s represents the visible-class semantic descriptor, which is generated using the FastText model, represents the set of synthetic region features, which consists of all the region features synthesized by the generator network. The FastText model is described in the work "Advances in pre-training distributed word representations" by Tomas Mikolov et al. in 2017.
[0053] During training, use the adaptive moment estimation algorithm to optimize the network parameters. The network's loss function is set as follows:
[0054]
[0055] Among them, G represents the generator network, D represents the discriminator network, represents the WGAN loss, represents the classification loss, represents the intra-class semantic divergence loss, It represents the inter-class structure preservation loss. η is the balance coefficient of the classification loss term, and η is set to 0.01. λ1 is the balance coefficient of the intra-class semantic divergence loss term, and λ1 is set to 0.001. λ2 is the balance coefficient of the inter-class structure preservation loss term, and λ2 is set to 0.001.
[0056] The WGAN loss and the classification loss can be specifically referred to the work of Nasir Hayat et al. in 2020, "Synthesizing the Unseen for Zero-shot Object Detection".
[0057] Inter-class structure preservation loss It is calculated according to the following formula:
[0058]
[0059] Among them, E[·] represents the expectation. represents the synthetic positive example region feature. represents the synthetic negative example region feature. A positive example refers to the region feature corresponding to the random noise input within the sphere with the random noise input z corresponding to the region feature as the center and a radius of r. A negative example refers to the region feature corresponding to the random noise input outside the sphere. Set r = 10. , N is the number of negative examples, and N is set to 10. τ1 is the first temperature coefficient, and τ1 is set to 0.1. -6
[0060] Inter-class structure preservation loss It is calculated according to the following formula:
[0061]
[0062] Among them, g + represents the positive example region feature. A positive example refers to the region feature with the same class label as . Φ = {g j} represents the region feature set, including the real region feature the synthetic region feature and the background region feature f b , g j represents the j-th region feature. τ2 is the second temperature coefficient, and τ2 is set to 0.1.
[0063] Step 6: Input the random noise z ∼ N(0, 1) and the unseen class semantic descriptor w u ∈W u into the generator network trained in Step 5 to synthesize the unseen class region features All the unseen class region features form the unseen class region feature set Among them, W u represents the invisible class semantic descriptor, which is generated using the FastText model.
[0064] Step 7: Calculate the initial memory unit of the asynchronous repository, including the visible class memory unit and the invisible class memory unit The calculation formulas are as follows:
[0065]
[0066]
[0067] Among them, represents the set of true region features of the k-th visible class, and f i s represents the true region feature of the i-th visible class, represents the set contains the number of features, represents the set of region features of the synthesized k-th invisible class, represents the synthesized region feature of the i-th invisible class, represents the set contains the number of features.
[0068] Step 8: Input the random noise z ∼ N(0, 1), the visible class semantic descriptor w s ∈ W s and the invisible class semantic descriptor w u ∈ W u into the generator network of the generative adversarial network. The input of the discriminator network is the same as in Step 5, and then the two networks are alternately iteratively trained separately to obtain the trained generator network and discriminator network.
[0069] When training, set the loss function of the network as follows:
[0070]
[0071] Among them, is the cross-modal contrast enhancement loss corresponding to the repository, and λ3 is the balance coefficient of the cross-modal contrast enhancement loss term. Set λ3 = 0.001.
[0072] Cross-modal contrast enhancement loss The calculation formula is as follows:
[0073]
[0074] Among them, represents the synthesized region feature of the visible class or invisible class; u +Denote the positive example memory unit, and the positive example represents the memory unit with the same class label as the regional feature ; K s Denote the number of visible class memory units, set K s = 16; Denote the k-th visible class memory unit; K u Denote the number of visible class memory units, set K u = 4; Denote the k-th invisible class memory unit; τ3 is the temperature coefficient three, set τ3 = 0.05.
[0075] During training, the following formula is used to update the invisible class memory units in the repository :
[0076]
[0077] where, ← represents the update operation, m is the momentum update factor, set m = 0.0001; is the synthetic set of invisible class regional features.
[0078] Step 9: Input the random noise z ∼ N(0, 1) and the invisible class semantic descriptor w u ∈ W u into the generator network trained in Step 8 to synthesize the invisible class regional features All the invisible class regional features constitute the set of invisible class regional features
[0079] Step 10: Input the set of invisible class regional features obtained in Step 9 into the invisible class classifier network, train this network to obtain the trained invisible class classifier network; during training, set the loss function of the network as the cross-entropy loss function, and use the adaptive moment estimation algorithm to optimize the network parameters; then replace the corresponding classifier network parameters in the target detection component with the weight parameters of the invisible class classifier network.
[0080] Step 11: Input the zero-shot test image set to be processed into the target detection component updated in Step 10, and the regional features output by the candidate region network therein constitute the candidate region feature set F.
[0081] Step 12: Use the formula to calculate the final classification score p i of the i-th candidate region feature f i in the set F; where, is the basic classifier score of the candidate region feature f i , by inputting f iObtained from the basic classifier in the updated target detection component in input step 10, where β is the balance coefficient and β = 1.0 is set. For the candidate region feature f i The class-level classification prediction score, and the calculation formula is as follows:
[0082]
[0083] Among them, u is the updated memory unit in the repository, obtained after the memory unit of the repository is updated in step 8. u i Represents the memory unit with the same class as the i-th candidate feature f i And u k Represents the memory unit corresponding to the k-th class in the memory bank, and K is the total number of memory unit classes.
[0084] Figure 2 The figure shows the final result image obtained by processing the zero-shot test image set of the PASCAL VOC dataset using the method of the present invention. In the figure, the boxes represent the bounding boxes formed by the detected target coordinates, and Figure 2 It can be seen that the method of the present invention has achieved good zero-shot detection results under both the traditional zero-shot setting conditions, that is, the test images containing only invisible class targets (shown in the first row), and the generalized zero-shot conditions, that is, the test images containing both visible and invisible class targets (shown in the second row), and at the same time, correct classification and positioning have been achieved.
Claims
1. A zero-shot object detection method based on a repository robust region feature synthesizer, characterized in that The steps are as follows: Step 1: Construct a target detection network model, which includes two parts: a target detection component and an invisible classifier learning component. Among them, the target detection component adopts the Faster r-cnn target detection network structure for target detection. In this network, there is a region proposal network for extracting region features. The invisible classifier learning component includes a generative adversarial network and an invisible class classifier network. The generative adversarial network includes a generator network and a discriminator network for synthesizing invisible class region features. The invisible class classifier network is trained using the invisible class region features synthesized by the generative adversarial network as input to obtain the trained weight parameters of the invisible class classifier, and then used to update the target detection component. The network parameters are initialized using the uniform distribution initialization method; Step 2: Divide the PASCAL VOC 2007+2012 image dataset into two sub-datasets: visible classes and invisible classes; Step 3: Input the images in the visible class sub-dataset, as well as their corresponding labels and bounding box coordinates, into the target detection component for network training. During training, the cross-entropy loss function is used as the network's loss function, and the stochastic gradient descent algorithm is used to optimize the network parameters; Step 4: Input the images in the visible class sub-dataset into the object detection component trained in Step 3, and use the region proposal network therein to extract the visible class region feature F s ; Step 5: Input the random noise z ∼ N(0, 1) and the visible class semantic descriptor w s ∈ W s into the generator network in the generative adversarial network, and input the real region features f of the images in the visible class sub-dataset s ∈ F s and the synthetic region features into the discriminator network in the generative adversarial network, and then alternately train the two networks separately and iteratively to obtain the trained generator network and discriminator network; where W s represents the visible class semantic descriptor, generated using the FastText model, represents the set of synthetic region features, composed of all the region features synthesized by the generation network; During training, the adaptive moment estimation algorithm is used to optimize the network parameters. The network's loss function is set as follows: Among them, G represents the generator network, and D represents the discriminator network. represents the WGAN loss, represents the classification loss, represents the intra-class semantic divergence loss, represents the inter-class structure preservation loss. η is the balance coefficient of the classification loss term, and η = 0.01 is set. λ1 is the balance coefficient of the intra-class semantic divergence loss term, and λ1 = 0.001 is set. λ2 is the balance coefficient of the inter-class structure preservation loss term, and λ2 = 0.001 is set. The inter-class structure preservation loss is calculated according to the following formula: where E[·] represents expectation, represents the synthesized positive example region feature, represents the synthesized negative example region feature. A positive example refers to the region feature corresponding to the random noise input within the sphere with the random noise input z corresponding to the region feature as the center of the sphere and a radius of r. A negative example refers to the region feature corresponding to the random noise input outside the sphere. Set r = 10 -6 , N is the number of negative examples, set N = 10, τ1 is the temperature coefficient one, set τ1 = 0.1; The inter-class structure preservation loss is calculated according to the following formula: Among them, g + represents the positive sample region feature. A positive sample refers to the region feature with the same class label as . Φ = {g j} represents the set of region features, including the true region feature f s , the synthesized region feature and the background region feature f b . g j represents the j-th region feature, and τ2 is the temperature coefficient two, with τ2 set to 0.1; Step 6: Input the random noise z ∼ N(0, 1) and the semantic descriptor w of the unseen class u ∈ W u into the generator network trained in Step 5 to synthesize the feature of the unseen class region All the features of the unseen class regions form the set of features of the unseen class regions where W u represents the semantic descriptor of the unseen class and is generated using the FastText model; Step 7: Calculate the initialized memory cells of the asynchronous repository, including the visible class memory cells and the invisible class memory cells The calculation formulas are as follows: Among them, represents the set of true regional features of the k-th visible class, f i s represents the true regional feature of the i-th visible class, represents the set contains the number of features, represents the set of regional features of the k-th invisible class synthesized, represents the regional feature of the i-th invisible class synthesized, represents the set contains the number of features; Step 8: Input the random noise z ∼ N(0, 1), the visible class semantic descriptor w s ∈ W s and the invisible class semantic descriptor w u ∈ W u into the generator network of the generative adversarial network. The input of the discriminator network is the same as that in Step 5, and then the two networks are alternately iteratively trained separately again to obtain the trained generator network and discriminator network; During training, the network's loss function is set as follows: Among them, is the cross-modal contrast enhancement loss corresponding to the repository, λ3 is the balance coefficient of the cross-modal contrast enhancement loss term, and λ3 is set to 0.001; The cross-modal contrastive enhancement loss has the following calculation formula: Among them, represents the synthesized visible or invisible class region feature; u + represents the positive example memory unit, and the positive example represents the memory unit with the same class label as the region feature ; K s represents the number of visible class memory units, and set K s = 16; represents the k-th visible class memory unit; K u represents the number of visible class memory units, and set K u = 4; represents the k-th invisible class memory unit; τ3 is the temperature coefficient three, and set τ3 = 0.05; During training, the invisible class memory units in the repository are updated using the following formula as follows: Among them, ← represents the update operation, m is the momentum update factor, and m is set to 0.0001; is the synthetic invisible class region feature set; Step 9: Input the random noise z ∼ N(0, 1) and the semantic descriptor w of the invisible class u ∈ W u into the generator network trained in Step 8 to synthesize the feature of the invisible class region All the features of the invisible class regions form a set of features of the invisible class regions Step 10: Input the set of invisible class region features obtained in Step 9 into the invisible class classifier network, train the network, and obtain a trained invisible class classifier network; when training, set the loss function of the network as the cross-entropy loss function, and use the adaptive moment estimation algorithm to optimize the network parameters; then replace the corresponding classifier network parameters in the object detection component with the weight parameters of the invisible class classifier network; Step 11: Input the zero-shot test image set to be processed into the target detection component updated in Step 10. The region features output by the candidate region network therein constitute the candidate region feature set F; Step 12: Using the formula calculate the final classification score p of the i-th candidate region feature f in set F i ; where, i is the basic classifier score of the candidate region feature f i , obtained by inputting f i into the basic classifier in the object detection component updated in Step 10, β is the balance coefficient, and β = 1.0 is set, is the class-level classification prediction score of the candidate region feature f i , and the calculation formula is as follows: Among them, u is the updated memory unit in the repository, obtained after the memory unit in the repository is updated in step 8, u i represents the memory unit with the same category as the i-th candidate feature f i u k represents the memory unit corresponding to the k-th category in the memory bank, and K is the total number of memory unit categories.
Citation Information
Patent Citations
Non-contact multi-target behavior recognition method
CN114358103A
Associated memory apparatus
GB201906894D0