Real object identification method and device, medium and equipment

By constructing the target image pyramid and extracting dense features, combined with the neural network model optimized by nested wolf pack algorithm, the problem of low recognition accuracy of complex products is solved, and a more efficient recognition effect is achieved.

CN120047664APending Publication Date: 2025-05-27GUANGDONG POWER GRID MATERIALS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510008999.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has low accuracy when identifying complex products, making it difficult to deal with the detection of counterfeit and shoddy products.

Method used

By scaling multiple initial images of the product to be identified at different scales, a target image pyramid is constructed, dense features are extracted from each layer, and combined them to generate a target comprehensive feature map, and input it to a neural network model optimized based on the nested wolf pack algorithm for identification.

Benefits of technology

It improves the detection accuracy and efficiency of complex products, can more effectively identify multiple distinctive features of complex products, and enhances the ability to prevent counterfeit and shoddy products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047664A_ABST
    Figure CN120047664A_ABST
Patent Text Reader

Abstract

The invention discloses a real object identification method and device, a medium and equipment, and the method comprises the steps: carrying out the zooming of a plurality of initial images of a to-be-identified product at different scales, and obtaining a target image pyramid of the to-be-identified product; performing dense feature extraction on each layer of image in the target image pyramid to obtain dense features of each layer of the target image pyramid; overlapping the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the to-be-identified product; the comprehensive feature map of the to-be-recognized product is input into the trained target detection model, the recognition result of the to-be-recognized product is obtained, the target detection model is a neural network model optimized based on the nested wolf pack algorithm, the detection accuracy and efficiency of the target detection model are high, and the recognition accuracy is high. The comprehensive features comprise features of different scales of the to-be-identified product, and the comprehensive features are used as input data of the target detection model, so that the accuracy and efficiency of object identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of product identification, and particularly to a physical object identification method, device, medium and equipment. Background Art

[0002] In modern manufacturing and logistics industries, physical object identification and detection methods, as an important quality control means, have significantly improved the operation efficiency of warehousing and logistics.

[0003] Currently, physical object identification methods mainly rely on the combination of manual visual inspection and automated equipment assistance. Manual visual inspection depends on the experience and judgment of quality inspection personnel, and has a high accuracy rate for identifying products with obvious features. For products with unclear features, it is easily affected by subjective factors, resulting in inconsistencies and errors in the identification and detection results. In addition, the efficiency of manual visual inspection is low, and it is difficult to meet the requirements of large-scale and rapid outbound and inbound. Although automated detection equipment has improved the detection speed, there are still limitations in detection accuracy and flexibility when detecting complex products. With the continuous emergence of counterfeit and shoddy products and the continuous upgrading of counterfeiting means, the deficiencies of automated detection equipment in detection accuracy and efficiency make it difficult for enterprises to comprehensively and effectively prevent counterfeit and shoddy products.

[0004] Therefore, it is particularly urgent to develop a physical object identification method that can effectively improve the detection accuracy of complex products. Summary of the Invention

[0005] In view of this, the present invention provides a physical object identification method, device, medium and equipment, mainly aiming to solve the problem of low identification accuracy of existing identification methods for complex products.

[0006] According to one aspect of the present application, a physical object identification method is provided, and the method includes:

[0007] Performing scaling of multiple initial images of the product to be identified at different scales to obtain a target image pyramid of the product to be identified;

[0008] Performing dense feature extraction on each layer of the images in the target image pyramid respectively to obtain dense features of each layer of the target image pyramid;

[0009] Stacking the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be identified;

[0010] Inputting the comprehensive feature map of the product to be identified into a trained target detection model to obtain an identification result of the product to be identified, where the target detection model is a neural network model optimized based on the nested wolf pack algorithm.

[0011] Optionally, performing dense feature extraction on each layer of the target image pyramid respectively to obtain the dense features of each layer of the target image pyramid includes:

[0012] Performing local feature extraction on each layer of the target image pyramid based on a feature extractor to obtain the key features of each layer of the target image pyramid;

[0013] Performing interpolation processing on the key features of each layer of the target image pyramid to obtain the dense features of each layer of the target image pyramid.

[0014] Optionally, superimposing the dense features of each layer of the target image pyramid to obtain the target comprehensive feature map of the product to be recognized includes:

[0015] Determining the weights corresponding to the dense features of each layer in the target image pyramid, and performing weighted processing on the dense features based on the dense features of each layer of the target image pyramid and their corresponding weights to obtain the target comprehensive feature map of the product to be recognized.

[0016] Optionally, obtaining a trained target detection model by using the following method includes:

[0017] Performing scaling of multiple initial images of the training product at different scales to obtain a training image pyramid of the training product;

[0018] Performing local feature extraction on each layer of the training image pyramid based on a feature extractor to obtain the key features of each layer of the training image pyramid;

[0019] Performing interpolation processing on the key features of each layer of the training image pyramid to obtain the dense features of each layer of the training image pyramid;

[0020] Performing weighted processing on the dense features based on the dense features of each layer of the training image pyramid and their corresponding weights to obtain a training comprehensive feature map of the training product;

[0021] Based on the training comprehensive feature map of the training product, determining the recognition result of the training product, inputting the training comprehensive feature map and the recognition result of the training product into an initial detection model, and optimizing the initial detection model based on the nested wolf pack algorithm to obtain a trained target detection model.

[0022] Optionally, optimizing the initial detection model based on the nested wolf pack algorithm includes:

[0023] Take the loss function of the initial detection model as the fitness function of the artificial wolf pack, and take the weights and biases of the initial detection model as the positions of each artificial wolf in the artificial wolf pack, and perform initialization processing on the artificial wolf pack;

[0024] Calculate the fitness function value of each artificial wolf, and determine the lead wolf, scout wolf, and fierce wolf according to the fitness function value;

[0025] Based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the fierce wolf's siege behavior, update the lead wolf, and delete a number of artificial wolves according to the prey distribution rule, and generate a number of new artificial wolves;

[0026] Recalculate the fitness function value of each artificial wolf, determine the new lead wolf, scout wolf, and fierce wolf according to the new fitness function value, and again update the lead wolf based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the fierce wolf's siege behavior, and delete a number of artificial wolves according to the prey distribution rule, and generate a number of new artificial wolves until the preset stop iteration condition is met. Based on the position of the lead wolf at this time, obtain the trained target detection model.

[0027] Optionally, before scaling the multiple initial images of the training product at different scales, the physical object recognition method further includes:

[0028] Obtain multiple original images of the training product at multiple perspectives and under various lighting conditions, and perform preprocessing on each original image;

[0029] Add random noise and occlusion to each preprocessed original image to obtain multiple initial images of the training product.

[0030] Optionally, the loss function of the initial detection model is a label smoothing regularization loss function.

[0031] According to another aspect of the present application, there is provided a physical object recognition device, including:

[0032] A target image acquisition module, configured to scale multiple initial images of the product to be recognized at different scales to obtain a target image pyramid of the product to be recognized;

[0033] A dense feature extraction module, configured to perform dense feature extraction on each layer of the images in the target image pyramid to obtain the dense features of each layer of the target image pyramid;

[0034] A comprehensive feature acquisition module, configured to stack the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be recognized;

[0035] An identification result acquisition module, configured to input the comprehensive feature map of the product to be identified into a trained target detection model to obtain the identification result of the product to be identified, where the target detection model is a neural network model optimized based on the nested wolf pack algorithm.

[0036] Optionally, the dense feature extraction module is further configured to: perform local feature extraction on each layer of the image in the target image pyramid based on a feature extractor to obtain the key feature of each layer of the target image pyramid; perform interpolation processing on the key feature of each layer of the target image pyramid to obtain the dense feature of each layer of the target image pyramid.

[0037] Optionally, the comprehensive feature acquisition module is further configured to: determine the weight corresponding to the dense feature of each layer in the target image pyramid, and perform weighted processing on the dense feature based on the dense feature of each layer in the target image pyramid and its corresponding weight to obtain the target comprehensive feature map of the product to be identified.

[0038] Optionally, the physical object identification device further includes:

[0039] A model training module, configured to scale multiple initial images of the training product at different scales to obtain a training image pyramid of the training product; perform local feature extraction on each layer of the image in the training image pyramid based on a feature extractor to obtain the key feature of each layer of the training image pyramid; perform interpolation processing on the key feature of each layer of the training image pyramid to obtain the dense feature of each layer of the training image pyramid; perform weighted processing on the dense feature based on the dense feature of each layer in the training image pyramid and its corresponding weight to obtain a training comprehensive feature map of the training product; determine the identification result of the training product based on the training comprehensive feature map of the training product, input the training comprehensive feature map and the identification result of the training product into an initial detection model, and optimize the initial detection model based on the nested wolf pack algorithm to obtain a trained target detection model.

[0040] Optionally, the model training module is further configured to: use the loss function of the initial detection model as the fitness function of the artificial wolf pack, use the weights and biases of the initial detection model as the positions of each artificial wolf in the artificial wolf pack, and perform initialization processing on the artificial wolf pack; calculate the fitness function values of each artificial wolf, and determine the lead wolf, scout wolf, and fierce wolf according to the fitness function values; update the lead wolf based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the fierce wolf's siege behavior, and delete several artificial wolves according to the prey distribution rule, and generate several new artificial wolves; recalculate the fitness function values of each artificial wolf, determine the new lead wolf, scout wolf, and fierce wolf according to the new fitness function values, and update the lead wolf again based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the fierce wolf's siege behavior, and delete several artificial wolves according to the prey distribution rule, and generate several new artificial wolves until the preset stop iteration condition is met, and obtain the trained target detection model based on the position of the lead wolf at this time.

[0041] Optionally, the physical object recognition device further includes:

[0042] A preprocessing module, configured to obtain a plurality of original images of the training product under multiple perspectives and various lighting conditions, and perform preprocessing on each original image; add random noise and occlusion to each preprocessed original image to obtain a plurality of initial images of the training product.

[0043] Optionally, the loss function of the initial detection model is a label smoothing regularization loss function.

[0044] According to another aspect of the present application, there is provided a storage medium storing at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the physical object recognition method described above.

[0045] According to another aspect of the present application, there is provided a computer device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0046] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the physical object recognition method described above.

[0047] By virtue of the above technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:

[0048] A physical object recognition method, device, medium and equipment provided by the present application scale multiple initial images of a product to be recognized at different scales, combine the scaled images to obtain a target image pyramid of the product to be recognized, extract dense features of each layer of the target image pyramid, combine the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map, input the comprehensive feature map of the product to be recognized into an object detection model trained based on the wolf pack algorithm. The object detection model trained by nesting the wolf pack algorithm has high recognition accuracy and efficiency, and outputs accurate recognition results. Since the comprehensive features include features of different scales of the product to be recognized, reflecting various significant features of complex products, using the comprehensive features as the input data of the object detection model also improves the accuracy of physical object recognition.

[0049] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0051] Figure 1 The flowchart of a physical object recognition method provided by an embodiment of the present application is shown;

[0052] Figure 2 Another flowchart of a physical object recognition method provided by an embodiment of the present application is shown;

[0053] Figure 3 The block diagram of a physical object recognition device provided by an embodiment of the present application is shown;

[0054] Figure 4 The structural diagram of a computer device provided by an embodiment of the present invention is shown.

[0055] Among them,

[0056] Figure 3 In: 302 - Target image acquisition module; 304 - Dense feature extraction module; 306 - Comprehensive feature acquisition module; 308 - Recognition result acquisition module;

[0057] Figure 4 In: 402 - Processor; 404 - Communication interface; 406 - Memory; 408 - Communication bus; 410 - Program. Detailed implementation manners

[0058] In the following, the present invention will be described in detail with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0059] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific implementation manners, structures, features and their effects of the application according to the present invention. In the following description, different "one embodiment" or "embodiments" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0060] In view of the problem of low accuracy of the current physical object recognition method, the embodiments of the present application provide a physical object recognition method, as Figure 1 shown, the method includes:

[0061] 102: Scale multiple initial images of the product to be recognized at different scales to obtain a target image pyramid of the product to be recognized;

[0062] 104: Respectively perform dense feature extraction on each layer of the images in the target image pyramid to obtain the dense features of each layer of the target image pyramid;

[0063] 106: Overlay the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be recognized;

[0064] 108: Input the comprehensive feature map of the product to be recognized into the trained target detection model to obtain the recognition result of the product to be recognized, wherein the target detection model is a neural network model optimized based on the nested wolf pack algorithm.

[0065] In this embodiment, it is necessary to obtain multiple images of the product to be recognized, and these images are usually taken at different angles and under different lighting conditions. Scale these images at different scales, and then combine the scaled images hierarchically to obtain an image pyramid. An image pyramid is a multi-resolution representation method, which represents the original image by constructing a set of gradually reduced or enlarged images. Common types of image pyramids include Gaussian pyramids and Laplacian pyramids. A Gaussian pyramid is constructed by continuously downsampling an image (usually reducing it by half). The image in each layer of the pyramid is a blurred version of the previous layer image. A Laplacian pyramid is constructed by performing high-pass filtering on each layer of the image in the Gaussian pyramid. The image of each layer represents the difference between this layer and the next layer.

[0066] Then, dense feature extraction is performed on each layer of the image pyramid to obtain features of the product to be recognized at different scales. The features of each layer are combined to obtain a comprehensive feature map of the product to be recognized. This comprehensive feature recognition map is used as the input data of the trained object detection model. Since the comprehensive features include the features of the product to be recognized at different scales and the object detection model is optimized based on the wolf pack algorithm, the object detection model can accurately output the recognition result of the product to be recognized.

[0067] This application provides a physical object recognition method. Compared with the prior art, multiple initial images of the product to be recognized are scaled at different scales, and the scaled images are combined to obtain the target image pyramid of the product to be recognized. Dense features of each layer of the target image pyramid are extracted, and the dense features of each layer of the target image pyramid are combined to obtain a target comprehensive feature map. The comprehensive feature map of the product to be recognized is input into the object detection model trained based on the wolf pack algorithm. The object detection model trained by the nested wolf pack algorithm has high recognition accuracy and efficiency and can output accurate recognition results. Since the comprehensive features include the features of the product to be recognized at different scales and reflect various significant features of complex products, using the comprehensive features as the input data of the object detection model also improves the accuracy of physical object recognition.

[0068] In one embodiment, dense feature extraction is respectively performed on each layer of the target image pyramid to obtain the dense features of each layer of the target image pyramid, including:

[0069] Based on the feature extractor, local feature extraction is performed on each layer of the target image pyramid to obtain the key features of each layer of the target image pyramid;

[0070] Interpolation processing is performed on the key features of each layer of the target image pyramid to obtain the dense features of each layer of the target image pyramid.

[0071] In this embodiment, multiple images containing the same scene or target need to be obtained. These images are usually taken at different angles and under different lighting conditions. These images are scaled at different scales and then combined hierarchically to obtain an image pyramid. Then, a feature point extraction algorithm (such as SIFT, SURF, ORB, etc.) is used to extract features in each layer of the image pyramid to obtain the key features of each layer. These features are the pixels in the image that have obvious features (such as corner points, places with drastic texture changes, etc.). Finally, the sparse key features are converted into dense features through interpolation or other methods, and interpolation can be used to enhance the features.

[0072] In one embodiment, the dense features of each layer of the target image pyramid are superimposed to obtain the target comprehensive feature map of the product to be recognized, including:

[0073] Determine the weights corresponding to the dense features of each layer in the target image pyramid. Based on the dense features of each layer in the target image pyramid and their corresponding weights, perform weighted processing on the dense features to obtain the target comprehensive feature map of the product to be recognized.

[0074] In this embodiment, the dense features of different levels are combined into a multi-scale feature. The combination method can be simple splicing, or weighted or other more complex fusion methods. The comprehensive feature can represent the features of the product to be recognized at different scales and different levels. Performing physical object recognition based on the comprehensive feature can improve the recognition accuracy.

[0075] In one embodiment, as Figure 2 shown, the following method is used to obtain the trained object detection model, including:

[0076] 202: Obtain multiple original images of the training product under multiple perspectives and various lighting conditions, and preprocess each original image;

[0077] 204: Add random noise and occlusion to each preprocessed original image to obtain multiple initial images of the training product;

[0078] 206: Scale the multiple initial images of the training product at different scales to obtain the training image pyramid of the training product;

[0079] 208: Based on the feature extractor, perform local feature extraction on each layer of the image in the training image pyramid to obtain the key features of each layer of the training image pyramid;

[0080] 210: Perform interpolation processing on the key features of each layer of the training image pyramid to obtain the dense features of each layer of the training image pyramid;

[0081] 212: Based on the dense features of each layer in the training image pyramid and their corresponding weights, perform weighted processing on the dense features to obtain the training comprehensive feature map of the training product;

[0082] 214: Based on the training comprehensive feature map of the training product, determine the recognition result of the training product. Input the training comprehensive feature map and the recognition result of the training product into the initial detection model, and optimize the initial detection model based on the nested wolf pack algorithm to obtain the trained object detection model.

[0083] Specifically, collect a set of training images with clear labels and annotate them (these images should cover different object types, different perspectives, and different lighting conditions). Annotation can be done using common label management systems such as Annotator. Next, divide the annotated dataset into a training set and a test set (use various types of datasets such as PASCAL VOC, COCO, etc. to ensure that the model can exhibit good generalization ability in various situations).

[0084] Next, preprocessing steps will be adopted to augment the training data. This will include rotating, flipping, and cropping the images so that the model can learn different perspectives and poses. In addition, random noise and occlusions are added to each training image to help the model better handle uncertainties in practical applications.

[0085] First, combine multiple training images in a pyramid manner to obtain an image pyramid, then extract the dense features of each layer of the image pyramid, and then combine the dense features of each layer together to obtain a comprehensive feature. Divide the dataset into a training set and a test set, and use the nested wolf pack algorithm to train the model. During the training process, techniques such as cross-validation are used to ensure the generalization ability of the model.

[0086] Use the stacked dense feature (DensePoint Cloud) as the input feature, and use a method called "multi-scale feature extraction" to generate these embeddings. The following method can also be adopted, which involves two key steps: (1) Use traditional descriptor extractors such as SIFT, SURF, or ORB to find the local features, i.e., key features, of each pixel point, and combine these local features in a pyramid manner to produce a more global representation and obtain a comprehensive feature.

[0087] The advantage of the stacked dense feature is that it can make full use of the inherent semantic information of the object and improve the accuracy and robustness of recognition. The stacked dense feature can also be obtained through the following methods:

[0088] ① Methods based on deep neural networks: These methods usually use convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to extract feature points on the object surface. The main advantage of these methods is that they can perform fast parameter updates without losing too much information, thus adapting to changing environments.

[0089] ② Strategies based on divide-and-conquer algorithms: These methods divide the object surface into multiple regions and then process each region separately. The advantage of this method is that it can reduce the computational complexity while maintaining high recognition accuracy.

[0090] ③Multi-task learning-based method: This method attempts to extract information about each part of an object by sharing underlying features. The advantage of this method is that it can improve the generalization ability of the model and avoid overfitting.

[0091] In one embodiment, optimizing the initial detection model based on the nested wolf pack algorithm includes:

[0092] Taking the loss function of the initial detection model as the fitness function of the artificial wolf pack, and taking the weights and biases of the initial detection model as the positions of each artificial wolf in the artificial wolf pack, and performing initialization processing on the artificial wolf pack;

[0093] Calculating the fitness function value of each artificial wolf, and determining the lead wolf, scout wolf, and aggressive wolf according to the fitness function value;

[0094] Updating the lead wolf based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the aggressive wolf's besieging behavior, and deleting a number of artificial wolves according to the prey distribution rule, and generating a number of new artificial wolves;

[0095] Recalculating the fitness function value of each artificial wolf, determining the new lead wolf, scout wolf, and aggressive wolf according to the new fitness function value, and again updating the lead wolf based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the aggressive wolf's besieging behavior, and deleting a number of artificial wolves according to the prey distribution rule, and generating a number of new artificial wolves until the preset stop iteration condition is met, and obtaining the trained object detection model based on the position of the lead wolf at this time.

[0096] In this implementation, taking the loss function of the initial detection model as the fitness function of the artificial wolf pack, and taking the weights and biases of the initial detection model as the positions of each artificial wolf in the artificial wolf pack, the position of each wolf is a solution, and the optimal solution is obtained.

[0097] First, perform initialization processing on the artificial wolf pack, that is, randomly generate the positions of each artificial wolf, calculate the fitness function value of each artificial wolf according to the position of each artificial wolf, determine the lead wolf, scout wolf, and aggressive wolf according to the fitness function value, take the one with the largest fitness function value as the lead wolf, take a preset number of artificial wolves with relatively large fitness function values as scout wolves, and the remaining artificial wolves as aggressive wolves. The scout wolves are responsible for wandering and searching for prey in the solution space, that is, looking for better solutions. When a scout wolf finds that the fitness function value at a certain position is greater than the fitness function value of the lead wolf, it will update the position of the lead wolf, and the lead wolf initiates a summoning behavior. If the scout wolf does not find a better position during the wandering process, it will continue to wander until the maximum wandering times are reached, and at this time, the lead wolf will issue a summons at the original position.

[0098] Next, the fierce wolf will respond to the call of the alpha wolf and rush towards the alpha wolf with a large step length. During the rush, if the fierce wolf finds a fitness function value greater than the alpha wolf, then the fierce wolf will replace the alpha wolf and become the new alpha wolf, and continue to call. If the fitness function value of the fierce wolf is not as good as the alpha wolf, then it will continue to rush until it enters the siege range.

[0099] After entering the siege range, the fierce wolf will join forces with the scout wolf to capture the prey (the alpha wolf is considered the prey at this time). During the siege, if the fitness function value of other artificial wolves is greater than that of the alpha wolf, the alpha wolf's position will be updated again. This process will continue until the prey is captured, that is, the optimal solution is found.

[0100] In addition, there is an important update mechanism in the wolf pack algorithm, which is "survival of the strong". Under this mechanism, artificial wolves with smaller fitness function values ​​will be eliminated from the wolf pack, and new artificial wolves will be randomly generated in the solution space to achieve the update of the wolf pack. This process helps to maintain the diversity of individual wolves and prevent the algorithm from falling into a local optimal solution.

[0101] In summary, the process of updating the wolf pack's position in the wolf pack algorithm according to the "survival of the strong" rule is achieved through the wandering search of the scout wolf, the response and siege behavior of the fierce wolf, and the update mechanism of the wolf pack. This process not only reflects the group intelligence and collaboration ability of the wolf pack, but also ensures that the algorithm can continuously evolve to a better solution.

[0102] In one embodiment, in terms of model training, a regularization strategy called “Label Smoothing Regularization Loss Function” is used. This strategy can avoid model overfitting by balancing the weights between hard labels and soft labels.

[0103] The "label smoothing regularization loss function" is also called the label smoothing L1 / L2 loss. It is mainly used to solve classification problems, especially when dealing with data with uncertain labels.

[0104] The main idea of ​​this loss function is to use label smoothing technology to better balance model performance and generalization ability during the learning process. By introducing a penalty term, the model can minimize the value of this penalty term during the training process, thereby achieving better generalization effect.

[0105] Specifically, the goal of the label smoothing L1 / L2 loss is to minimize the following loss function:

[0106] L(y, \hat{y}) = -\sum\left[\begin{cases}\frac{(y_i - \hat{y}_i)^2}{(N - 1)} & \text{if} y_i = \hat{y}_i \\ \log(\frac{1}{N}) \cdot y_i & \text{else}\end{cases}\right]^T+\frac{\alpha}{N}

[0107] Among them, \(y\) is the true label, \(\hat{y}\) is the predicted label, \(\alpha\) is the penalty coefficient, and \(N\) is the number of all samples in the dataset.

[0108] When there is no correlation between the labels, it can be assumed that the label corresponding to each sample will not be much different from its neighbors. Therefore, the penalty term can take a relatively large value. When there is a strong correlation between the labels, it is hoped that the model can learn this correlation. At this time, the penalty term can be set to a relatively small value to encourage the model to learn a stronger correlation.

[0109] After the model training is completed, its performance on the test set will be evaluated (the trained model will be deployed to a real-time application named "Online Object Detection System". In this system, users can upload pictures, and the system will automatically detect and classify objects and display the results to users in real time). If the model performs well, the model can be continued to be used for actual applications. Otherwise, consider adjusting the network structure and parameters to improve its performance until the model performs well.

[0110] Furthermore, as an implementation of the method shown above Figure 1 shown, an embodiment of the present invention provides an object recognition device, as Figure 3 shown, the device includes:

[0111] A target image acquisition module 302, configured to scale multiple initial images of the product to be recognized at different scales to obtain a target image pyramid of the product to be recognized;

[0112] A dense feature extraction module 304, configured to perform dense feature extraction on each layer of the image in the target image pyramid to obtain dense features of each layer of the target image pyramid;

[0113] A comprehensive feature acquisition module 306, configured to stack the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be recognized;

[0114] An identification result acquisition module 308, configured to input the comprehensive feature map of the product to be recognized into the trained target detection model to obtain the identification result of the product to be recognized, where the target detection model is a neural network model optimized based on the nested wolf pack algorithm.

[0115] The present application provides a physical object recognition device. Compared with the prior art, multiple initial images of a product to be recognized are scaled at different scales, and the scaled images are combined to obtain a target image pyramid of the product to be recognized. Dense features of each layer of the target image pyramid are extracted, and the dense features of each layer of the target image pyramid are combined to obtain a target comprehensive feature map. The comprehensive feature map of the product to be recognized is input into a target detection model trained based on the wolf pack algorithm. The target detection model trained by the nested wolf pack algorithm has a high recognition accuracy and efficiency and outputs accurate recognition results. Since the comprehensive feature includes features of different scales of the product to be recognized and reflects various significant features of complex products, using the comprehensive feature as the input data of the target detection model also improves the accuracy of physical object recognition.

[0116] In one embodiment, the dense feature extraction module is further configured to: perform local feature extraction on each layer of the image in the target image pyramid based on a feature extractor to obtain key features of each layer of the target image pyramid; perform interpolation processing on the key features of each layer of the target image pyramid to obtain dense features of each layer of the target image pyramid.

[0117] In one embodiment, the comprehensive feature acquisition module is further configured to: determine weights corresponding to the dense features of each layer in the target image pyramid, and perform weighted processing on the dense features based on the dense features of each layer of the target image pyramid and their corresponding weights to obtain a target comprehensive feature map of the product to be recognized.

[0118] In one embodiment, the physical object recognition device further includes:

[0119] A model training module, configured to scale multiple initial images of a training product at different scales to obtain a training image pyramid of the training product; perform local feature extraction on each layer of the image in the training image pyramid based on a feature extractor to obtain key features of each layer of the training image pyramid; perform interpolation processing on the key features of each layer of the training image pyramid to obtain dense features of each layer of the training image pyramid; perform weighted processing on the dense features based on the dense features of each layer of the training image pyramid and their corresponding weights to obtain a training comprehensive feature map of the training product; determine the recognition result of the training product based on the training comprehensive feature map of the training product, input the training comprehensive feature map and the recognition result of the training product into an initial detection model, and optimize the initial detection model based on the nested wolf pack algorithm to obtain a trained target detection model.

[0120] In one embodiment, the model training module is further configured to: use the loss function of the initial detection model as the fitness function of the artificial wolf pack, use the weights and biases of the initial detection model as the positions of each artificial wolf in the artificial wolf pack, and perform initialization processing on the artificial wolf pack; calculate the fitness function values of each artificial wolf, and determine the lead wolf, scout wolf, and fierce wolf according to the fitness function values; update the lead wolf based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the fierce wolf's siege behavior, and delete a number of artificial wolves according to the prey distribution rule, and generate a number of new artificial wolves; recalculate the fitness function values of each artificial wolf, determine the new lead wolf, scout wolf, and fierce wolf according to the new fitness function values, and again update the lead wolf based on the scout wolf's wandering behavior, the lead wolf's summoning behavior, and the fierce wolf's siege behavior, and delete a number of artificial wolves according to the prey distribution rule, and generate a number of new artificial wolves until the preset stop iteration condition is satisfied, and obtain the trained target detection model based on the position of the lead wolf at this time.

[0121] In one embodiment, the physical object recognition device further includes:

[0122] A preprocessing module, configured to obtain a plurality of original images of the training product under multiple perspectives and various lighting conditions, and perform preprocessing on each original image; add random noise and occlusion to each preprocessed original image to obtain a plurality of initial images of the training product.

[0123] In one embodiment, the loss function of the initial detection model is a label smoothing regularization loss function.

[0124] According to an embodiment of the present invention, there is provided a storage medium storing at least one executable instruction, and the computer executable instruction can execute the physical object recognition method in any of the above method embodiments.

[0125] Figure 4 FIG. shows a schematic structural diagram of a computer device according to an embodiment of the present invention, and the specific implementation of the computer device in the specific embodiments of the present invention is not limited.

[0126] As Figure 4 shown, the computer device may include: a processor 402, a communication interface 404, a memory 406, and a communication bus 408.

[0127] Wherein: the processor 402, the communication interface 404, and the memory 406 communicate with each other through the communication bus 408.

[0128] The communication interface 404 is used to communicate with network elements of other devices such as clients or other servers.

[0129] A processor 402 is configured to execute a program 410, and specifically, can execute relevant steps in the above embodiments of the physical object recognition method.

[0130] Specifically, the program 410 may include program code, and the program code includes computer operation instructions.

[0131] The processor 402 may be a central processing unit (CPU), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The computer device includes one or more processors, which may be of the same type, such as one or more CPUs; or may be of different types, such as one or more CPUs and one or more ASICs.

[0132] A memory 406 is used to store the program 410. The memory 406 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0133] The program 410 is specifically configured to cause the processor 402 to perform the following operations:

[0134] Scale multiple initial images of the product to be recognized at different scales to obtain a target image pyramid of the product to be recognized;

[0135] Extract dense features from each layer of the images in the target image pyramid to obtain dense features of each layer of the target image pyramid;

[0136] Overlay the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be recognized;

[0137] Input the comprehensive feature map of the product to be recognized into the trained target detection model to obtain the recognition result of the product to be recognized, where the target detection model is a neural network model optimized based on the nested wolf pack algorithm.

[0138] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. In one embodiment, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0139] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present application within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.

Claims

1. A physical object identification method, characterized in that: include: Scaling the multiple initial images of the product to be identified at different scales to obtain a target image pyramid of the product to be identified; Performing dense feature extraction on each layer of the target image pyramid to obtain dense features of each layer of the target image pyramid; Overlaying the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be identified; The comprehensive feature map of the product to be identified is input into the trained target detection model to obtain the identification result of the product to be identified, wherein the target detection model is a neural network model optimized based on the nested wolf pack algorithm.

2. The object identification method according to claim 1, characterized in that: The step of extracting dense features from each layer of the target image pyramid to obtain dense features of each layer of the target image pyramid comprises: Extracting local features of each layer of the target image pyramid based on a feature extractor to obtain key features of each layer of the target image pyramid; Interpolation processing is performed on the key features of each layer of the target image pyramid to obtain dense features of each layer of the target image pyramid.

3. The object identification method according to claim 2, characterized in that: The method of superimposing the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be identified includes: The weight corresponding to each layer of dense features in the target image pyramid is determined, and based on each layer of dense features in the target image pyramid and the corresponding weights, the dense features are weighted to obtain a target comprehensive feature map of the product to be identified.

4. The object identification method according to claim 1, characterized in that: The trained target detection model is obtained using the following method, including: Scaling a plurality of initial images of a training product at different scales to obtain a training image pyramid of the training product; Extracting local features of each layer of the image in the training image pyramid based on a feature extractor to obtain key features of each layer of the training image pyramid; Performing interpolation processing on the key features of each layer of the training image pyramid to obtain dense features of each layer of the training image pyramid; Based on the dense features of each layer of the training image pyramid and their corresponding weights, weighted processing is performed on the dense features to obtain a training comprehensive feature map of the training product; Based on the training comprehensive feature graph of the training product, the recognition result of the training product is determined, the training comprehensive feature graph and the recognition result of the training product are input into the initial detection model, and the initial detection model is optimized based on the nested wolf pack algorithm to obtain a trained target detection model.

5. The object identification method according to claim 4, characterized in that: The optimizing the initial detection model based on the nested wolf pack algorithm includes: Using the loss function of the initial detection model as the fitness function of the artificial wolf pack, using the weight and bias of the initial detection model as the position of each artificial wolf in the artificial wolf pack, and initializing the artificial wolf pack; Calculating the fitness function value of each artificial wolf, and determining the leader wolf, the scout wolf and the fierce wolf according to the fitness function value; Based on the wandering behavior of the scout wolf, the calling behavior of the alpha wolf and the siege behavior of the fierce wolf, the alpha wolf is updated, and some artificial wolves are deleted according to the prey distribution rules to generate some new artificial wolves; The fitness function value of each artificial wolf is recalculated, and the new alpha wolf, scout wolf and fierce wolf are determined based on the new fitness function value. The alpha wolf is updated again based on the scout wolf's wandering behavior, the alpha wolf's calling behavior and the fierce wolf's siege behavior. According to the prey distribution rule, several artificial wolves are deleted and several new artificial wolves are generated until the preset stop iteration condition is met. Based on the position of the alpha wolf at this time, the trained target detection model is obtained.

6. The object identification method according to claim 4, characterized in that: Before scaling the multiple initial images of the training product at different scales, the physical object recognition method further includes: Obtain multiple original images of the training product under multiple viewing angles and multiple lighting conditions, and pre-process each original image; Random noise and occlusion are added to each preprocessed original image to obtain multiple initial images of the training product.

7. The object identification method according to any one of claims 4 to 6, characterized in that: The loss function of the initial detection model is a label smoothing regularization loss function.

8. A physical object recognition device, characterized in that: include: A target image acquisition module, used for scaling the multiple initial images of the product to be identified at different scales to obtain a target image pyramid of the product to be identified; A dense feature extraction module, used for performing dense feature extraction on each layer of the target image pyramid to obtain dense features of each layer of the target image pyramid; A comprehensive feature acquisition module, used for superimposing the dense features of each layer of the target image pyramid to obtain a target comprehensive feature map of the product to be identified; The recognition result acquisition module is used to input the comprehensive feature map of the product to be identified into the trained target detection model to obtain the recognition result of the product to be identified, wherein the target detection model is a neural network model optimized based on the nested wolf pack algorithm.

9. A storage medium storing at least one executable instruction, wherein the executable instruction enables a processor to execute an operation corresponding to the physical object identification method as described in any one of claims 1 to 7.

10. A computer device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the physical object identification method according to any one of claims 1-7.