Fresh goods long tail identification learning method based on perception manifold volume information guidance
By introducing an iterative fine-tuning method of perceived manifold volume information guidance into the fresh product image recognition model, the problem of uneven recognition capabilities under long-tail distribution data is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510704596.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The fresh product image recognition model has an uneven recognition ability under the long-tail distribution data, resulting in insufficient recognition ability for categories with small samples.
It adopts an iterative fine-tuning method based on perceived manifold volume information guidance, by building a teacher network and a student network, using the teacher network to calculate manifold volumes of different categories, and adjust the classification weight of the student network accordingly, and perform block knowledge distillation training to balance the model's recognition ability of different categories.
Effectively balance the model's ability to identify different categories of goods, improves the recognition accuracy and robustness of long-tail distribution data, and improves the recognition performance of fresh goods in unmanned retail scenarios.
Smart Images

Figure CN120236147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of long-tail image recognition, and particularly to a long-tail recognition learning method for fresh food products guided by the volume information of the perceptual manifold. Background Art
[0002] In the modern retail industry, especially in unmanned retail stores, image recognition technology is widely used in automatic commodity recognition and settlement. However, in practical applications, in an open environment, the image recognition of fresh food products faces a severe challenge: the long-tail characteristic of data distribution. The long-tail distribution refers to the phenomenon that in a dataset, a small number of categories have a large number of samples, while the number of samples in the majority of categories is scarce. This imbalance in data distribution is particularly prominent in fresh food products because the sales situations of fresh food products vary greatly, the image data of best-selling products is abundant, while the image data of slow-selling or newly launched products is extremely limited.
[0003] The long-tail distribution problem has a significant impact on the training of image recognition models. Due to the imbalance in the amount of data, traditional image recognition models often tend to be biased towards categories with a larger number of samples during the training process, resulting in insufficient recognition ability for categories with a smaller number of samples. This bias not only reduces the overall recognition accuracy of the model but also limits its application value in actual scenarios. Although various long-tail recognition learning methods have been proposed currently, these methods still have limitations when dealing with fresh food product data. For example, existing long-tail learning methods often have difficulty effectively balancing the recognition performance between different categories, especially in the case of extremely unbalanced sample numbers.
[0004] In addition, the particularity of fresh food products also increases the recognition difficulty. The appearance characteristics of fresh food products may change due to factors such as freshness, lighting conditions, and placement methods, which requires the image recognition model to have stronger robustness and adaptability. However, existing long-tail learning methods often cannot fully consider the diversity and complexity of the data when dealing with these complex situations, resulting in poor performance of the model in actual applications.
[0005] In view of this, the present application is proposed. Summary of the Invention
[0006] The present invention provides a long-tail recognition learning method for fresh food products guided by the volume information of the perceptual manifold, which can at least partially improve the above problems.
[0007] To achieve the above object, the present invention adopts the following technical solutions: A long-tail recognition learning method for fresh food products guided by the volume information of the perceptual manifold, which includes: Based on the traditional long-tail recognition method, train a preset intelligent goods image recognition model to obtain a student network and a teacher network, and calculate the perceptual manifold volume of different categories based on the teacher network to obtain the required perceptual manifold volume information; Use the perceptual manifold volume information as a weight factor to adjust the logits output of the student network based on this weight factor, and calculate the logit re-weighted classification loss; According to the perceptual manifold volume information, perform block processing on the softmax layers of the student network and the teacher network, and calculate the block knowledge distillation loss based on the logit within each block; Combine the logit re-weighted classification loss and the block knowledge distillation loss based on the logit to obtain a total loss function, and fine-tune the student network according to the total loss function; Repeat the above fine-tuning steps until the fine-tuning of the preset number of cycles is completed. Copy the student network to replace the original teacher network and enter a new round of iterative fine-tuning until the obtained student network meets the preset standard.
[0008] In summary, in the actual scenario, the long-tail distribution of fresh produce data leads to a huge difference in the number of samples, making the traditional model tend to the categories with more samples during training and ignoring the categories with fewer samples. To solve this problem, the long-tail recognition learning method for fresh produce guided by perceptual manifold volume information is proposed. After the traditional model training is completed, an iterative optimization strategy based on perceptual manifold volume is introduced. By constructing a teacher network and a student network, the manifold volume of different categories is calculated using the teacher network, and the classification weights of the student network are adjusted accordingly. At the same time, the network is processed in blocks based on the manifold volume information, and knowledge distillation training is implemented between the blocks. Through periodic iterative updates, this method can effectively balance the model's recognition ability for different categories of goods, enabling the model to exhibit higher accuracy and robustness under long-tail distribution data, providing a better solution for fresh produce recognition in the unmanned retail scenario.
[0009] Specifically, after training the model with the traditional long-tail recognition learning method, the fresh produce long-tail recognition learning method guided by the perceptual manifold volume information innovatively adds an iterative fine-tuning stage guided by the perceptual manifold volume information. In this stage, the method sets the trained model as the student network of the existing long-tail recognition learning method, and duplicates it as the teacher network. The teacher network is used to calculate the perceptual manifold volumes of different categories, and based on this, the classification loss of logit calculation and re-weighting is adjusted to fine-tune the student network. At the same time, according to the manifold volume information, the softmax layers of the teacher and student networks are partitioned, and block knowledge distillation training is implemented for the student network between the corresponding partitions. By setting the number of iterative steps, the teacher network is periodically updated and the volume is recalculated to continuously iteratively fine-tune the student network, so that the perceptual manifold volumes of all categories are evenly distributed in the representation space learned by the student network, and the student network can balance learning and recognize various fresh produce categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of the fresh produce long-tail recognition learning method guided by the perceptual manifold volume information provided by an embodiment of the present invention; Figure 2 is an overall process framework diagram of the fresh produce long-tail recognition learning method guided by the perceptual manifold volume information provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the warm-up stage provided by an embodiment of the present invention; Figure 4 is a schematic diagram of the iterative fine-tuning stage provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0012] In real - world retail fresh - produce datasets, the severe long - tail distribution phenomenon is quite common. Some product categories have a large number of samples, while those with poor sales have significantly insufficient sample quantities. This situation leads to a serious imbalance in the training set when training an intelligent product image recognition system. Although there are currently various long - tail recognition learning methods, this problem still has not been effectively solved. To truly overcome the long - tail problem in the field of retail fresh - produce, the present invention proposes adding an iterative fine - tuning stage guided by the perceived manifold volume information after completing model training using traditional long - tail recognition learning methods. It has been found through research that the models trained by current long - tail recognition learning methods have a significant bias in the representation space for all categories towards the head categories (i.e., product categories with a large number of samples). From the perspective of manifold performance, the representations of head categories occupy a larger representation area, meaning that their perceived manifold volume is much larger than that of the tail categories (i.e., product categories with few samples).
[0013] Based on this, referring to Figure 1 、 Figure 2 shown, the first embodiment of the present invention discloses a long - tail recognition learning method for fresh - produce guided by perceived manifold volume information, which can be executed by a long - tail recognition learning device for fresh - produce guided by perceived manifold volume information (hereinafter referred to as the learning device), specifically, executed by one or more processors within the learning device to implement the following method: Please refer to Figure 3 、 Figure 4 , S1, Train a preset intelligent product image recognition model based on traditional long - tail recognition methods to obtain a student network and a teacher network, and calculate the perceived manifold volume of different categories based on the teacher network to obtain the final required perceived manifold volume information; Specifically, step S1 includes: Training a preset intelligent product image recognition model on a preset long - tail fresh - produce dataset using traditional long - tail recognition methods, and defining the trained intelligent product image recognition model as the student network; Copy the student network, and define the copy of the student network as the teacher network.
[0014] In this embodiment, first, a preset long-tailed fresh food product dataset is selected. This dataset contains images of various fresh food products, and the class distribution of these images exhibits typical long-tailed characteristics, that is, a small number of classes have a large number of samples, while the majority of classes have a small number of samples. Then, a traditional long-tailed recognition method is used to train a preset intelligent product image recognition model, which is usually a deep learning model. The traditional long-tailed recognition method used here can be any known deep learning method that can handle class imbalance problems, such as resampling techniques, class weight adjustment, etc. By training on the long-tailed fresh food product dataset, the model can learn the features of different fresh food product images and initially form the recognition ability for various classes of products. When the training is completed, the obtained intelligent product image recognition model is defined as the student network. The student network already has a certain recognition ability at this time, but its performance may be affected by the long-tailed distribution, and its recognition ability for the tail classes with fewer samples is relatively weak.
[0015] Next, in order to be able to optimize the student network using the perceptual manifold volume information in the follow-up, it is necessary to copy the trained student network to obtain a copy, and this copy is defined as the teacher network. The teacher network and the student network have the same structure and parameters in the initial stage, but they will play different roles in the subsequent iterative fine-tuning stage. The main role of the teacher network is to act as a guide and use the perceptual manifold volume information calculated by it to guide the optimization process of the student network, while the student network is continuously fine-tuned under the guidance of the teacher network to improve its recognition performance for long-tailed distribution data.
[0016] Among them, Figure 4 the snowflake in indicates that the model is in a frozen state and the parameters are fixed and not updated; Figure 4 the fire in indicates that the model is in an updated state and the parameters need to be optimized.
[0017] Preferably, based on the teacher network, calculate the perceptual manifold volume of different classes to obtain the final required perceptual manifold volume information, specifically: According to the teacher network, calculate the perceptual manifold volume of different classes, and its calculation formula is: , where is the perceptual manifold volume of class , is the representation matrix of class , is the average value of all representations of class , is the total number of all product classes, is the determinant function, indicating calculating the determinant value of the matrix; Among them, when the class When the number of samples of a class is 1, the class perceptual manifold volume ; Normalize the perceptual manifold volume corresponding to each class to obtain the required perceptual manifold volume information. The formula is: , where is the sum of the volumes of all classes.
[0018] In this embodiment, after obtaining the teacher network, the next key step is to calculate the perceptual manifold volumes of different classes based on the teacher network. The perceptual manifold volume is an important concept that reflects the distribution of different classes in the feature space. Specifically, for each class, the teacher network calculates the volume of its perceptual manifold. The calculation of the perceptual manifold volume can be achieved by analyzing the features extracted by the teacher network. For example, the determinant value of the covariance matrix of all sample features within the class can be calculated to measure the volume occupied by the class in the feature space. In this way, the perceptual manifold volume information of each class can be obtained, which will provide key guiding basis for the subsequent iterative fine-tuning stage. This formula measures the distribution range of the class in the feature space by calculating the square root of the determinant value of the covariance matrix of the sample features within the class, so as to obtain the perceptual manifold volume.
[0019] Specifically, when the number of samples of a certain class is only 1, in order to ensure the stability of the calculation and avoid numerical problems, it is stipulated that the perceptual manifold volume of this class is a preset small positive value. For example: when the class has 1 sample, the class perceptual manifold volume , can be ensured to be greater than zero by adding a small value . This processing method not only avoids calculation errors caused by too few samples, but also provides reasonable initial values for the subsequent optimization process.
[0020] After calculating the perceptual manifold volumes of each class, in order to enable fair comparison of these volume information between different classes and be used as weight factors in the subsequent optimization process, normalize the perceptual manifold volumes of each class. Through normalization, the perceptual manifold volume of each class is converted into a relative value, and the sum of these relative values is 1, thus ensuring that in the subsequent optimization process, the volume information of different classes can participate in the adjustment of the model in a balanced manner.
[0021] S2. Use the perceptual manifold volume information as a weight factor to adjust the logits output of the student network based on this weight factor, and calculate the logit and then weigh the classification loss; Specifically, step S2 includes: The formula for logit reweighted classification loss is: , where , is the student network, is the product image, is the corresponding class label, is a hyperparameter, is the set of all product classes, is the weight factor calculated using the normalized popularity volume, is the student network model 's logit output at the position of class label at this position, is the student network model 's logit output at the position of class label at this position.
[0022] In this embodiment, the calculated perceptual manifold volume information is used as a weight factor to adjust the logits output of the student network. Here, logits refer to the raw output of the student network before the softmax layer, and these output values reflect the model's original prediction confidence for each class. By introducing the perceptual manifold volume information as a weight factor, these logits can be adjusted so that the model pays more attention to the tail classes with fewer samples when calculating the classification loss, thereby reducing the model's over - bias towards the head classes. The introduction of the weight factor is an important innovation point of the present invention. It measures the "importance" of each class through the perceptual manifold volume information and adjusts the classification loss function accordingly. Specifically, for the tail classes with a smaller perceptual manifold volume, their weight factors will be relatively larger, so more attention will be given when calculating the classification loss; while for the head classes with a larger perceptual manifold volume, their weight factors will be relatively smaller, thus appropriately reducing their weight in the classification loss. This weight adjustment mechanism based on the perceptual manifold volume can effectively balance the model's recognition ability for different classes of products, especially improving the recognition performance for the tail classes.
[0023] Through this calculation method of logit reweighted classification loss, the student network can learn the features of various categories of goods more evenly during the training process, rather than simply being biased towards the head categories with a large number of samples. This improvement not only helps to improve the overall recognition accuracy of the model for long-tailed distribution data, but also enhances the robustness and generalization ability of the model in practical applications. For example, in the unmanned retail scenario, even if the sample quantity of some fresh goods is small, the model can accurately identify these goods, thereby improving the accuracy and reliability of the automatic settlement system. In addition, the logit reweighted classification loss guided by the perceived manifold volume information can be combined with other traditional long-tailed learning methods to further improve the performance of the model.
[0024] S3. According to the perceived manifold volume information, perform block processing on the softmax layers of the student network and the teacher network, and calculate the block knowledge distillation loss based on logits within each block; Specifically, step S3 includes: dividing the perceived manifold volume information of different categories into blocks so that the manifold volume of each category within each block is relatively consistent, and marked as , where is a preset constant, represents the th block, is the first block, is the second block, is the th block, represents the set of categories belonging to the th block, is the set of categories belonging to the th block, is the set of categories belonging to the th block, is the set of categories belonging to the th block; According to this block result, perform corresponding segmentation processing on the softmax layers of the student network and the teacher network, and calculate the corresponding block knowledge distillation loss based on logits on each block. The formula is: , where is the teacher network, is the temperature parameter, is to calculate softmax for logits only on the categories included, is the value at the category y position, is the number of categories included in
[0025] It further includes: based on a preset addition rule, judging whether to combine the loss function of traditional long-tail learning according to the current recognition situation to obtain a total loss function, where the loss function of traditional long-tail learning is a regularization term.
[0026] In this embodiment, different categories are segmented into several blocks according to the perceptual manifold volume information, so that the perceptual manifold volumes of each category within each block are relatively consistent. The purpose of this segmentation strategy is to ensure that the categories within each block have a similar feature space distribution during the knowledge distillation process, thereby improving the effect of knowledge distillation. For example, assume that the categories are divided into K blocks according to the perceptual manifold volume. The category sets within each block are respectively . This segmentation method not only helps to simplify the calculation, but also ensures that the categories within each block have similar distribution characteristics in the feature space, thus providing more favorable conditions for subsequent knowledge distillation.
[0027] Subsequently, corresponding segmentation processing is performed on the softmax layers of the student network and the teacher network according to the above segmentation results to ensure that the manifold volumes of the categories within each block are relatively consistent or the volume size differences are small. Specifically, for each block, the softmax function is calculated only on the categories included in that block. The purpose of this operation is to perform knowledge distillation independently within each block and avoid interference between the categories of different blocks. Within each block, the block knowledge distillation loss based on logits is calculated. By calculating the knowledge distillation loss independently within each block, the student network can more accurately learn the knowledge of the teacher network within that block, thereby improving the recognition ability of each category, especially the recognition performance of the tail categories.
[0028] In addition, the present invention also provides a flexible mechanism that can judge whether to combine the traditional long-tail learning loss function according to the current recognition situation. Specifically, based on a preset addition rule, for example, when the optimization effect of the block knowledge distillation loss is not obvious, the traditional long-tail learning loss function can be selectively added as a regularization term to the total loss function. This combination method not only makes full use of the advantages of the perceptual manifold volume information, but also retains the effectiveness of the traditional long-tail learning method, thereby further improving the performance of the model. Simply put, at this stage, the loss function used in traditional long-tail learning is used as an option according to the actual effect.
[0029] S4. Combine the logits and then weigh the classification loss and the block knowledge distillation loss based on logits to obtain a total loss function, and fine-tune the student network according to the total loss function; Specifically, step S4 includes: The formula of the total loss function is: , where 、 and are both hyperparameters, is an optional item, and the loss function of unified long-tail learning is adopted.
[0030] In this embodiment, the total loss function is constructed by re-weighting the classification loss and the block knowledge distillation loss based on logits. In addition, as an optional item, the traditional long-tail learning loss function can be selected to be added to the total loss function as needed. Through this comprehensive optimization method, the student network can be constrained by both the re-weighted classification loss based on logits and the block knowledge distillation loss during the training process. The re-weighted classification loss based on logits adjusts the classification weights by perceiving the manifold volume information, enabling the model to pay more attention to the tail categories with fewer samples; while the block knowledge distillation loss transfers knowledge to the student network under the guidance of the teacher network, further optimizing the classification performance of the student network. This dual-constraint mechanism can not only effectively balance the model's recognition ability for different categories of goods, but also improve the robustness and generalization ability of the model under complex long-tail distribution data.
[0031] Specifically, the student network is fine-tuned according to the total loss function. The fine-tuning process is an iterative optimization process, and the parameters of the student network are updated through the backpropagation algorithm, making the value of the total loss function gradually decrease. In each iteration, the parameters of the student network are adjusted according to the gradient of the total loss function, thereby gradually optimizing the performance of the model. By setting appropriate iteration times and learning rates, the student network can gradually learn a more balanced category recognition ability during the fine-tuning process, especially in the case of extremely unbalanced sample numbers, it can significantly improve the recognition accuracy of the tail categories. It should be noted that the teacher network is fixed and not updated.
[0032] In addition, by introducing the traditional long-tail learning loss function as a regularization term, the student network can further benefit from the advantages of traditional methods during the fine-tuning process. This combination method not only makes full use of the innovation of the manifold volume information perception and knowledge distillation technology, but also retains the effectiveness of traditional long-tail learning methods, enabling the model to more flexibly adjust the optimization strategy when facing complex long-tail distribution data.
[0033] The core purpose of the iterative fine-tuning guided by the manifold volume information perception is to promote a more balanced distribution of the manifold volumes of each category in the representation space, ultimately achieving the ideal effect of balanced learning.
[0034] S5. Repeat the above fine-tuning steps until the fine-tuning of the preset number of cycles is completed, copy the student network to replace the original teacher network, and enter a new round of iterative fine-tuning until the obtained student network meets the preset standard.
[0035] Specifically, step S5 includes: continuing to fine-tune the student network according to the total loss function. When the fine-tuning for K cycles is completed, the student network is copied to replace the original teacher network, obtaining a new teacher network, where K is the preset number of iteration steps; Recalculate the perceptual manifold volumes of different categories according to the new teacher network and perform a new round of iterative fine-tuning; Enter a new round of iterative fine-tuning until the obtained student network meets the preset standard.
[0036] In this embodiment, after the fine-tuning of the student network based on the total loss function is completed, the fine-tuning continues according to the preset number of iteration steps K. In each fine-tuning cycle, the student network updates its parameters according to the total loss function, gradually optimizing its ability to recognize different categories of goods. When the fine-tuning for K cycles is completed, the current student network is copied, and the copied network is used to replace the original teacher network, thus obtaining a new teacher network. The purpose of this replacement operation is to introduce the latest model parameters and knowledge, enabling the subsequent iterative fine-tuning to be guided based on a more optimized teacher network.
[0037] After obtaining the new teacher network, recalculate the perceptual manifold volumes of different categories according to the new teacher network. This recalculation process is necessary because as the student network is continuously optimized, its ability to represent data will also change, thereby affecting the distribution of the perceptual manifold volumes. By recalculating the perceptual manifold volumes, it can be ensured that the subsequent fine-tuning process can be optimized based on the latest and more accurate volume information, further improving the performance of the student network.
[0038] Subsequently, enter a new round of iterative fine-tuning. In the new iteration process, the student network continues to be fine-tuned according to the total loss function, and at the same time, adjusts the classification loss and knowledge distillation loss using the perceptual manifold volume information calculated by the new teacher network. This process will be repeated continuously. After each completion of the fine-tuning for K cycles, the student network will be copied to replace the teacher network, and the perceptual manifold volumes will be recalculated to start the next round of iterative fine-tuning. Through this periodic iterative update mechanism, the student network can gradually learn a more balanced category recognition ability. Especially when dealing with long-tailed distribution data with extremely unbalanced sample numbers, it can significantly improve the recognition performance of the tail categories.
[0039] The beneficial effect of this process is that through periodic iterative fine-tuning and the update of the teacher network, the student network can continuously absorb the latest knowledge and information, thereby gradually optimizing its ability to identify data with long-tailed distributions. Each iterative fine-tuning is based on the latest teacher network and the information of the perceptual manifold volume, which enables the student network to adjust its parameters more accurately to adapt to the complexity of the long-tailed distribution data. In addition, by setting preset criteria, such as the recognition accuracy of the model reaching a certain threshold or the value of the loss function no longer decreasing significantly, it can be ensured that the student network stops iterating after reaching the optimal performance, avoiding the overfitting problem caused by overtraining.
[0040] Finally, through the periodic iterative fine-tuning of this step, the student network can gradually reach the preset performance criteria and become a model with high recognition accuracy and strong robustness under long-tailed distribution data. This process not only solves the limitations of traditional long-tailed learning methods in processing fresh produce data but also provides a more efficient and accurate solution for the identification of fresh produce in the unmanned retail scenario. Through this continuously optimized strategy, the student network can better adapt to the long-tailed distribution problem in practical applications, thus significantly improving the intelligence level and user experience of the unmanned retail system.
[0041] In this embodiment, the long-tailed recognition learning method for fresh produce guided by the perceptual manifold volume information is guided by the perceptual manifold volume information, which can effectively enhance the robustness of the goods intelligent recognition system when facing long-tailed distributions in an open environment, and further improve the intelligence level of the fresh produce automatic settlement system.
[0042] In summary, the long-tailed recognition learning method for fresh produce guided by the perceptual manifold volume information introduces the perceptual manifold volume information as the key guiding factor for optimization on the basis of traditional long-tailed recognition learning methods. By calculating the perceptual manifold volumes of different classes and using them as weight factors to adjust the classification loss function, the model can pay more attention to the tail classes with fewer samples. This innovation not only balances the model's ability to identify different classes of goods but also significantly improves the recognition accuracy of the tail classes, solving the problem that the model in traditional methods tends to the head classes.
[0043] Secondly, through the collaborative optimization of the teacher network and the student network, the performance of the model is further improved. The teacher network provides the optimization direction and goal for the student network by calculating the perceptual manifold volume information. The student network then fine-tunes by re-weighting the classification loss and the block knowledge distillation loss through logits under the guidance of the teacher network. This optimization strategy based on knowledge distillation enables the student network to inherit the excellent features of the teacher network while avoiding the overfitting problem, further improving the generalization ability of the model.
[0044] In addition, a periodic iterative fine-tuning mechanism is introduced. In each iteration cycle, the student network is optimized according to the total loss function. After completing the fine-tuning for the preset number of steps, the student network is copied to replace the teacher network, the perceptual manifold volume is recalculated, and a new round of iterative fine-tuning is started. This mechanism not only ensures that the model can be optimized based on the latest information at each stage, but also introduces new knowledge and features by continuously updating the teacher network, enabling the student network to gradually improve its ability to recognize long-tail distribution data.
[0045] Generally speaking, the method for learning long-tail recognition of fresh food products guided by perceptual manifold volume information not only effectively solves the long-tail distribution problem, but also significantly improves the accuracy and robustness of fresh food product image recognition. In the unmanned retail scenario, this method can significantly improve the efficiency and accuracy of the automatic settlement system, reducing customer dissatisfaction and increased operating costs caused by recognition errors. In addition, this method also has good scalability and adaptability, and can be flexibly adjusted according to different data sets and application scenarios, providing strong technical support for the development of the unmanned retail industry.
[0046] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A long-tail recognition learning method for fresh goods guided by the volume information of the perceptual manifold, characterized in that Including: Based on the traditional long-tail recognition method, train a preset intelligent goods image recognition model to obtain a student network and a teacher network, and calculate the perceptual manifold volume of different categories based on the teacher network to obtain the required perceptual manifold volume information; Use the perceptual manifold volume information as a weight factor to adjust the logits output of the student network, and calculate the logit reweighted classification loss; According to the perceptual manifold volume information, perform block processing on the softmax layers of the student network and the teacher network, and calculate the block knowledge distillation loss based on logits within each block; Combine the logit reweighted classification loss and the block knowledge distillation loss based on logits to obtain a total loss function, and fine-tune the student network according to the total loss function; Repeat the above fine-tuning steps until the fine-tuning of the preset number of cycles is completed, copy the student network to replace the original teacher network, and enter a new round of iterative fine-tuning until the obtained student network meets the preset standard.
2. The method for identifying and learning the long tail of fresh goods guided by the volume information of the perceptual manifold according to claim 1, wherein Based on the traditional long-tail recognition method, train a preset intelligent goods image recognition model to obtain a student network and a teacher network. Specifically: Use the traditional long-tail recognition method to train the preset intelligent goods image recognition model on the preset long-tail fresh food dataset, and define the trained intelligent goods image recognition model as the student network; Copy the student network, and define the copy of the student network as the teacher network.
3. The method for identifying and learning the long tail of fresh goods guided by the volume information of the perceptual manifold according to claim 1, wherein Calculate the perceptual manifold volume of different categories based on the teacher network to obtain the required perceptual manifold volume information. Specifically: Calculate the perceived popularity volume of different categories according to the teacher network, and its calculation formula is: , where is the perceived manifold volume of category , is the representation matrix of category , is the average value of all representations of category , is the number of all product categories, is the determinant function, indicating to calculate the determinant value of the matrix; Among them, when the number of samples of class is 1, the perceptual manifold volume of class is ; Normalize the perceived popularity volume corresponding to each category to obtain the perceived manifold volume information required finally, and its formula is: , where is the total volume of all categories.
4. The method for identifying and learning the long tail of fresh goods guided by the volume information of the perceptual manifold according to claim 3, wherein The formula for the logit reweighted classification loss is as follows: , where , is the student network, is the product image, is the corresponding class label, is a hyperparameter, is the set of all product classes, is the weight factor calculated using the normalized popularity volume, is the student network model the value of the logit output of at this position of the class label, is the student network model the value of the logit output of at this position of the class label.
5. The method for identifying and learning the long tail of fresh goods guided by the volume information of the perceptual manifold according to claim 4, wherein According to the perceptual manifold volume information, perform block processing on the softmax layers of the student network and the teacher network, and calculate the block knowledge distillation loss based on logits within each block. Specifically: Segment the volume information of the perceptual manifolds of different categories into blocks so that the manifold volumes of each category within each block are relatively consistent, and label them as , where is a preset constant, represents the th block, is the first block, is the second block, is the th block, represents the set of categories belonging to the th block, is the set of categories belonging to the th block, is the set of categories belonging to the th block, is the set of categories belonging to the th block; According to this chunking result, the softmax layers of the student network and the teacher network are segmented accordingly, and the corresponding chunk-based knowledge distillation loss based on logits is calculated on each chunk. The formula is as follows: , where is the teacher network, is the temperature parameter, is to calculate softmax on logits only on the classes included in , is the value at the class y position, is the number of classes included in 6. The method for identifying and learning the long tail of fresh goods guided by the volume information of the perceptual manifold according to claim 5, wherein Also including: Based on the preset addition rule, judge whether to combine the loss function of traditional long-tail learning according to the current recognition situation to obtain a total loss function, where the loss function of traditional long-tail learning is a regularization term.
7. The method for identifying and learning the long tail of fresh goods guided by the volume information of the perceptual manifold according to claim 6, wherein The formula of the total loss function is as follows: , where , and are all hyperparameters, is an optional item, and the loss function of unified long-tail learning is adopted.
8. The method for identifying and learning the long tail of fresh goods guided by the volume information of the perceptual manifold according to claim 1, characterized in that Repeat the above fine-tuning steps until the fine-tuning of the preset number of cycles is completed, copy the student network to replace the original teacher network, and enter a new round of iterative fine-tuning until the obtained student network meets the preset standard. Specifically: Continue to fine-tune the student network according to the total loss function. When the fine-tuning of K cycles is completed, copy the student network to replace the original teacher network to obtain a new teacher network, where K is the preset number of iterative steps; Recalculate the perceptual manifold volume of different categories according to the new teacher network and perform a new round of iterative fine-tuning; Enter a new round of iterative fine-tuning until the obtained student network meets the preset standard.
Citation Information
Patent Citations
Small sample image classification method based on manifold learning and high-order graph neural network
CN113052263A
Long tail distribution visual classification method based on sample perception distillation
CN115995018A
Long-tail commodity recommendation method based on accurate symmetric positive definite manifold learning
CN119809766A
System and method for efficient machine learning
US20230419170A1