Automatic hardware identification system and method

By using a feature extraction method combining cascade encoder units and feature pyramid networks in the metal tool recognition system, integrating attention mechanism and feature fusion, the problem of insufficient accuracy and robustness of metal tool recognition in the prior art is solved, and efficient identification of metal tools with similar appearance is achieved.

CN120107648APending Publication Date: 2025-06-06XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510033604.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

It is difficult for existing metal tools to accurately distinguish between similar and subtle differences in appearance, especially under the same model of metal tools produced by different manufacturers and diversified storage methods, the recognition accuracy and robustness are insufficient.

Method used

A feature extraction method combined with a cascade encoder unit and a feature pyramid network is adopted. The channel attention mechanism and spatial attention mechanism of the deep separation convolution layer are integrated into the channel attention mechanism and the spatial attention mechanism, and the multi-level features of the metal are extracted, and the global and local features are obtained through global average pooling and uniform grid division, and the prototype feature library and comparison learning module are identified.

Benefits of technology

It significantly improves the recognition accuracy of similar-shaped metals, enhances the ability to distinguish the same model of metals from different manufacturers, improves the robustness and accuracy of identification, and reduces the misidentification rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107648A_ABST
    Figure CN120107648A_ABST
Patent Text Reader

Abstract

The invention provides a hardware fitting automatic identification technology system and method, and the system employs a cascaded depth separable convolutional network and a feature pyramid network to extract the multi-scale features of a hardware fitting, and improves the feature extraction capability through integrating a channel attention mechanism and a space attention mechanism. The system combines global features and local features, carries out global average pooling on a multi-scale feature map to obtain global features, and carries out grid division on a target coding feature map to obtain local region features. Through a pre-stored prototype feature library, similarities between the global features and the local features of the to-be-identified hardware fitting and the prototype features of each category are calculated respectively, and a final similarity score is obtained through multi-level fusion, so that accurate identification of the hardware fitting category is realized. The technical scheme can effectively improve the accuracy of hardware fitting recognition, and is especially suitable for hardware fitting recognition scenes with high appearance similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision processing technology, and in particular to a hardware automatic identification system and method. Background Art

[0002] With the development of intelligent storage management of power equipment, hardware, as a material with a large storage scale, is particularly important for its automatic identification and classification management. At present, hardware identification technology mainly uses traditional computer vision methods and basic deep learning models for feature extraction and classification. These methods can achieve basic identification functions of hardware under ideal conditions. Some advanced deep learning models also introduce technologies such as attention mechanism and multi-scale feature extraction to improve recognition performance.

[0003] However, existing technologies still face many challenges in practical applications. First, there are many types of hardware and their appearances are very similar, so it is difficult for traditional methods to accurately distinguish subtle differences. Second, hardware of the same model produced by different manufacturers has subtle changes in appearance and size, which increases the difficulty of identification. In addition, hardware storage methods are diverse, including different packaging forms of standard hardware and bulk storage of non-standard hardware, which further increases the complexity of automated identification.

[0004] Therefore, there is an urgent need to develop an intelligent recognition system that can effectively handle the identification of subtle differences in hardware. Summary of the invention

[0005] The present application provides a hardware automatic identification system and method to achieve accurate separation and efficient sorting of power hardware.

[0006] The present application provides a hardware automatic identification system, comprising:

[0007] An image acquisition module, used to obtain image data of the hardware to be identified, preprocess the image data, and output a standardized image after preprocessing, wherein the preprocessing includes size standardization and data enhancement processing;

[0008] A feature extraction network, comprising a cascaded encoder unit and a feature pyramid network, wherein the cascaded encoder unit comprises a plurality of cascaded encoders, wherein the input of the first encoder is the preprocessed normalized image, and each encoder unit comprises two depthwise separable convolutional layers and a maximum pooling layer in sequence, wherein the depthwise separable convolutional layers are integrated with a channel attention mechanism and a spatial attention mechanism, and are used to extract multi-level features of hardware to be identified; the feature pyramid network is used to obtain the encoding feature map output by each encoder, and fuse the encoding feature map to generate a multi-scale feature map of the hardware to be identified;

[0009] A feature processing unit is used to receive the multi-scale feature map and the target coding feature map output by the last encoder; perform a global average pooling operation on the multi-scale feature map to extract the global features of the hardware to be identified; and perform a grid-even division on the target coding feature map to obtain multiple local area features;

[0010] The prototype feature library is used to store the features of each hardware category learned during the training phase. The features of each hardware category include a global prototype feature and multiple local area prototype features.

[0011] The comparative learning module is used to calculate the similarity between the global features of the hardware to be identified and the global prototype features of each category in the prototype feature library to obtain the global similarity score of each category; for multiple local area features of the hardware to be identified, the similarity between them and the corresponding local prototype features of each category in the prototype feature library is calculated respectively to obtain multiple local similarity scores of each category; the multiple local similarity scores of each category are processed to obtain the comprehensive local similarity score of the category;

[0012] The classification recognition module is used to process the global similarity score and the comprehensive local similarity score of each category to obtain the final similarity score of the category; sort the final similarity scores of all categories and select a specified number of categories with the highest scores as candidate categories; and select categories with final similarity scores higher than a preset threshold from the candidate categories as recognition results.

[0013] The beneficial effects of the present application mainly include: (1) By integrating the channel attention mechanism and the spatial attention mechanism in the deep separable convolutional layer and combining the multi-scale feature fusion of the feature pyramid network, the system can simultaneously focus on the overall structure and local detail features of the hardware, significantly improving the recognition accuracy of hardware with similar appearance, especially when dealing with the same model of hardware produced by different manufacturers, showing stronger differentiation ability. (2) A recognition strategy combining global features and local features is adopted, in which the global features are obtained by global average pooling of multi-scale feature maps, and the local features are obtained by grid-uniform division of the target encoding feature map. This dual feature extraction mechanism enables the system to characterize the hardware from multiple dimensions, effectively improving the robustness of recognition. (3) The prototype feature library is introduced to store the category features learned in the training phase, and a similarity calculation mechanism at both global and local levels is designed. Through multi-level processing and fusion of similarity scores, the system can more accurately capture the difference features between hardware categories and reduce the misrecognition rate. (4) In the classification decision stage, a strategy combining candidate category screening and threshold judgment is adopted, which not only ensures the recognition efficiency but also provides a guarantee mechanism for recognition reliability, making the system more practical and reliable in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic diagram of a hardware automatic identification system provided in the first embodiment of the present application.

[0015] Figure 2 This is a flow chart of a method for automatic identification of hardware provided in the second embodiment of the present application. DETAILED DESCRIPTION

[0016] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.

[0017] The first embodiment of the present application provides a hardware automatic identification system. Figure 1 , which is a schematic diagram of the first embodiment of the present application. Figure 1 A hardware automatic identification system is provided in detail for the first embodiment of the present application.

[0018] The hardware automatic identification system comprises an image acquisition module 101 , a feature extraction network 102 , a feature processing unit 103 , a prototype feature library 104 , a comparison learning module 105 and a classification identification module 106 .

[0019] The image acquisition module 101 is used to obtain image data of the hardware to be identified, preprocess the image data, and output a standardized image after preprocessing, wherein the preprocessing includes size standardization and data enhancement processing.

[0020] The image acquisition module 101 is the first key component of the hardware automatic identification system. It is mainly responsible for acquiring the original image data of the hardware to be identified and performing necessary preprocessing operations. This module uses industrial-grade image acquisition equipment, which usually includes a high-resolution industrial camera and a standard light source system to ensure the quality of the acquired image. The installation position of the image acquisition device is usually fixed directly above the hardware conveyor belt or the inspection station, maintaining a suitable shooting distance and angle so that the hardware can be fully presented in the image.

[0021] After acquiring the original image, the module will immediately perform preprocessing operations. The first step of preprocessing is size normalization, which will adjust the input images of different sizes to a preset standard size, such as 224×224 pixels or 448×448 pixels. The choice of size needs to strike a balance between maintaining image details and computational efficiency. Bilinear interpolation algorithms are usually used for image scaling to ensure image quality.

[0022] Data enhancement is the second important step of preprocessing, and its purpose is to improve the robustness of the system. Specific data enhancement operations include: random horizontal flip (probability is 0.5), random rotation (range between ±10 degrees), random brightness adjustment (range between 0.8-1.2), random contrast adjustment (range between 0.8-1.2), etc. These enhancement operations can help the system better cope with various changes in actual scenes.

[0023] In addition, preprocessing also includes image normalization operations, that is, scaling pixel values ​​to the [0,1] interval or performing standardization (minus the mean and divided by the standard deviation). This step is crucial for the stable training and reasoning of subsequent deep learning networks. The preprocessed standardized image will be used as the input of the feature extraction network, laying the foundation for subsequent feature extraction and recognition tasks.

[0024] The output of the image acquisition module is a standardized image after preprocessing. The pixel distribution of these images is more uniform, and the features are more prominent, which is conducive to the subsequent network processing. In order to ensure processing efficiency, the entire preprocessing process is executed in parallel on the GPU, and a single image can usually be processed in milliseconds. The module also performs a quality check on the processed image to ensure that it meets the requirements of subsequent processing. If the image quality is found to be unqualified (for example, blurry, too dark, or severely occluded), the system will issue a warning in time and require the image to be re-captured.

[0025] Furthermore, the image acquisition module includes a quality assessment unit, which is used to perform quality assessment on the preprocessed standardized image, specifically including: detecting the clarity, brightness uniformity and hardware integrity of the image. When the quality assessment result is lower than the preset standard, the image re-acquisition mechanism is triggered and the image acquisition parameters are adjusted.

[0026] The role of the quality assessment unit is to conduct a comprehensive assessment of the image clarity, brightness uniformity and hardware integrity based on the pre-processed standardized image data. If the assessment result does not meet the preset standard, the system will automatically start the image re-capture mechanism and adjust the image acquisition parameters according to the specific situation to optimize the acquisition conditions.

[0027] When evaluating the clarity of an image, the quality assessment unit determines its clarity by analyzing the edge information of the image. Edge information reflects the distribution of high-frequency components in the image and can be extracted through a dedicated edge detection algorithm. For images with higher clarity, the edge features are usually more obvious and concentrated, while the edge features of blurred images will show more scattered or irregular changes. In addition, the quality assessment unit will combine the local gradient changes of the image to assist in the judgment and ensure the accuracy of the clarity detection.

[0028] The detection of brightness uniformity is achieved by analyzing the brightness distribution of each area in the image. The quality assessment unit divides the image into multiple areas, calculates the average brightness value of each area, and then compares the brightness differences between these areas. If the brightness distribution is too uneven, such as some areas are too bright or too dark, it may affect the reliability of subsequent recognition results. Therefore, when it is detected that the brightness uniformity does not meet the expected standards, the quality assessment unit will feedback to the image re-capture mechanism to adjust the acquisition parameters, such as exposure time or light source position, to improve brightness uniformity.

[0029] The purpose of the hardware integrity test is to ensure that the target area of ​​the hardware in the image is not missing or blocked. The quality assessment unit uses the target detection algorithm to identify the position and shape of the hardware in the image, and compares the recognition result with the hardware template stored in the system. If a hardware area is detected to be obviously missing, such as part of the hardware is not captured due to improper shooting angle, or part of the hardware is blocked due to environmental interference, the system will determine that the hardware integrity does not meet the standard and trigger the re-collection process.

[0030] When the quality assessment unit completes the detection of clarity, brightness uniformity and hardware integrity, it will synthesize the evaluation results and compare them with the preset image quality standards. If the comprehensive evaluation result is lower than the standard, the system will automatically trigger the image re-capture mechanism. The image re-capture mechanism includes two parts: adjusting the image acquisition parameters and re-capturing the image. Adjusting the image acquisition parameters is to optimize the acquisition conditions, such as by changing the camera's exposure time, aperture size or focal length to ensure that the newly acquired image better meets the quality requirements; re-capturing the image is to obtain a replacement for the current image that does not meet the quality standards to ensure the reliability of subsequent processing and recognition.

[0031] The quality assessment unit can be implemented by hardware, such as an embedded image processing chip or a dedicated sensor module, or by software, such as an image processing algorithm integrated into an embedded image processing program or running on a general-purpose computing device. Regardless of the implementation method, the working principle and processing flow of the quality assessment unit are the same.

[0032] The feature extraction network 102 includes a cascaded encoder unit and a feature pyramid network, wherein the cascaded encoder unit includes a plurality of cascaded encoders, wherein the input of the first encoder is the preprocessed normalized image, and each encoder unit includes two depthwise separable convolutional layers and a maximum pooling layer in sequence, wherein the depthwise separable convolutional layers integrate a channel attention mechanism and a spatial attention mechanism to extract multi-level features of hardware to be identified; the feature pyramid network is used to obtain the encoded feature map output by each encoder, and fuse the encoded feature map to generate a multi-scale feature map of the hardware to be identified.

[0033] The feature extraction network 102 is the core processing unit of the system, which consists of two major parts: a cascade encoder unit and a feature pyramid network.

[0034] The cascade encoder unit adopts a multi-stage cascade architecture design, which specifically includes five encoder units. The first encoder directly receives the preprocessed standardized image as input, and each subsequent encoder takes the output of the previous encoder as input, forming a deep cascade structure. The internal structure of each encoder unit is the same, which includes two depth-separable convolution layers and a maximum pooling layer in sequence. This design not only ensures the depth of feature extraction, but also greatly reduces the amount of computation through depth-separable convolution.

[0035] In the design of the depthwise separable convolutional layer, 3×3 depthwise convolution is first used to extract features for each input channel independently, and then 1×1 point convolution is used to achieve information fusion between channels. This decomposition method not only significantly reduces the number of parameters and computational complexity, but also can more effectively extract spatial features and channel features. Specifically, for the case where the input feature map size is H×W×C, the depthwise convolution uses a step size of 1 and a padding method of SAME to keep the feature map size unchanged; the point convolution is responsible for adjusting the number of channels. In different encoder units, the number of output channels is 64, 128, 256, 512, and 1024, respectively.

[0036] In order to enhance the feature extraction capability of the network, a channel attention mechanism and a spatial attention mechanism are also integrated in each depth-separable convolutional layer. The channel attention mechanism adopts a squeeze-and-excitation structure. It first obtains channel descriptors through global average pooling, then learns the correlation between channels through two fully connected layers, and finally generates channel weights through a sigmoid function. The spatial attention mechanism helps the network focus on important spatial areas by learning the weight distribution in the spatial dimension. The outputs of these two attention mechanisms are combined with the original feature map through a multiplication operation, thereby enhancing meaningful feature responses.

[0037] The feature pyramid network is designed to solve the problem of multi-scale feature representation. The network first collects the encoded feature maps output by each encoder, which have a decreasing spatial resolution but increasing number of channels. Then, features of different scales are fused through top-down paths and lateral connections. Specifically, the high-level feature map is first adjusted to the same spatial size as the low-level feature map by upsampling, and then pixel-wise additively fused with the corresponding low-level feature map. This process starts from the top layer and proceeds step by step downward, eventually generating a set of multi-scale feature maps that contain both high-level semantic information and low-level detail features.

[0038] During the network training process, a stochastic gradient descent optimizer with momentum was used, with the initial learning rate set to 0.001 and a factor of 0.1 used to decay the learning rate every 30 cycles. To prevent overfitting, a batch normalization layer and a dropout layer were added after each depth-separable convolution layer, and the dropout rate was set to 0.3. Through this carefully designed network structure and training strategy, the feature extraction network can effectively capture the multi-level features of the hardware, providing strong feature support for subsequent recognition tasks.

[0039] Furthermore, the cascade encoder unit in the feature extraction network includes five encoders, the number of output channels of each encoder is 64, 128, 256, 512 and 1024 respectively, the depthwise separable convolution layer adopts a 3×3 convolution kernel, and the maximum pooling layer adopts a 2×2 pooling window with a step size of 2.

[0040] The feature extraction network aims to extract multi-level features of hardware from the preprocessed standardized images to support subsequent recognition tasks. The cascaded encoder unit is the core component of the feature extraction network, which is responsible for dimensionality reduction and feature extraction of the input image layer by layer to generate feature representations with different semantic levels. The cascaded encoder unit consists of five encoders cascaded in sequence, each of which consists of a set of depthwise separable convolutional layers and a maximum pooling layer. The encoder is designed in a hierarchical manner, and the number of feature channels in each layer increases successively to capture higher-level semantic information.

[0041] In the first encoder, the input is a preprocessed normalized image, which is processed by an initial depth-wise separable convolutional layer to output 64 feature channels. These feature channels represent the basic feature information of the image, such as edges, textures, etc. Subsequently, the feature map is reduced in dimension through a 2×2 max pooling layer. The max pooling operation slides a 2×2 window on the input feature map, moving 2 pixels at a time, retaining the maximum value in each window, thereby effectively reducing the size of the feature map while retaining significant features.

[0042] Starting from the second encoder, the input of each encoder is the output of the previous encoder. In the second encoder, the number of feature channels increases from 64 to 128, indicating that higher-level features are extracted. In each encoder, the depthwise separable convolution layer uses a 3×3 convolution kernel. This convolution kernel design helps to extract local features while maintaining computational efficiency, and its depthwise separable property allows separate operations in the spatial and channel dimensions, significantly reducing computational complexity. In addition, the maximum pooling operation of each layer is the same as the first layer, using a 2×2 pooling window with a step size of 2.

[0043] As the number of cascade levels increases, the number of feature channels output by each encoder is 64, 128, 256, 512, and 1024, respectively, indicating that the network gradually aggregates more semantic information. In this design, the low-level encoder focuses on local features, such as the edges and textures of the hardware, while the high-level encoder focuses more on global features, such as the overall shape and structure of the hardware. Through this layer-by-layer enhancement method, the feature extraction network can generate a multi-level feature representation containing rich semantic information, providing high-quality input for the subsequent feature pyramid network and classification recognition module.

[0044] This embodiment optimizes the parameter selection of the depthwise separable convolution layer and the maximum pooling layer. The depthwise separable convolution layer reduces the amount of computation by separating the convolution operations in the spatial and channel dimensions, allowing the system to remain efficient when processing high-resolution inputs. The 2×2 window and step size settings of the maximum pooling layer ensure that the feature map size is scaled down proportionally, preventing information loss while reducing the complexity of subsequent calculations. The design of the five-layer encoder balances the depth of feature extraction and computational efficiency, allowing the network to effectively extract complex multi-level features while meeting the needs of real-time recognition.

[0045] Through the above design, the feature extraction network can generate high-quality multi-scale features while maintaining efficient calculation, providing a solid foundation for the accurate identification of hardware.

[0046] Furthermore, the feature pyramid network fuses the output features of encoders at each level through a top-down feature fusion path, specifically including: upsampling the high-level features so that they have the same spatial resolution as the low-level features, and then additively fusing them with the corresponding low-level features, and sequentially processing them to obtain multi-scale feature maps with different resolutions.

[0047] The core function of the feature pyramid network is to integrate the features output from each level of encoder to generate a feature map with multi-scale information, so as to adapt to the different scales and complex shapes of hardware targets. The top-down feature fusion path is an effective multi-scale feature processing method. Its design purpose is to combine the semantic information of high-level features with the spatial details of low-level features, and provide feature representations with both global and local information for the subsequent recognition process.

[0048] In the feature fusion process, the feature pyramid network first receives the output features from encoders at all levels. The high-level output features of the encoder have a higher semantic level but lower spatial resolution, while the low-level output features have a higher spatial resolution but relatively less semantic information. In order to achieve effective fusion of the two, the feature pyramid network upsamples the high-level features so that their spatial resolution matches the corresponding low-level features.

[0049] The upsampling operation is implemented through an interpolation algorithm, such as bilinear interpolation or nearest neighbor interpolation. After upsampling, the size of the high-level features is expanded to the same resolution as the low-level features to ensure alignment in the spatial dimension. Subsequently, the feature pyramid network performs pixel-by-pixel additive fusion of the upsampled high-level features with the corresponding low-level features. Additive fusion is a simple and efficient operation that superimposes two features by pixel-by-pixel addition, so that the fused features retain both the semantic information of the high-level features and the spatial details of the low-level features.

[0050] This fusion process is performed layer by layer, starting from the encoder features of the highest layer and moving to lower layers in sequence. For example, after the features of the highest layer are upsampled, they are fused with the features of the next higher layer to generate a first-level multi-scale feature map. Then, the fusion result is further processed with the low-level features of the next layer to generate a second-level multi-scale feature map. And so on, until the encoder features of all layers are involved in the fusion.

[0051] Finally, the feature pyramid network generates a set of multi-scale feature maps with different resolutions. These feature maps contain different levels of semantic information and spatial details, and can adapt to the diversity of hardware targets in size, shape and complexity. Each level of feature map can be used alone for subsequent hardware classification and positioning, or it can be further fused or weighted to generate a more comprehensive feature representation.

[0052] This top-down feature fusion path has significant technical advantages. On the one hand, by upsampling and pixel-by-pixel addition fusion, information loss is effectively avoided, ensuring that high-level semantic information can be fully transmitted to low-level features; on the other hand, the fused multi-scale feature map can take into account both global and local features, making the hardware recognition system more robust when dealing with complex backgrounds, target occlusion and scale changes.

[0053] The feature processing unit 103 is used to receive the multi-scale feature map and the target coding feature map output by the last encoder; perform a global average pooling operation on the multi-scale feature map to extract the global features of the hardware to be identified; and evenly divide the target coding feature map into a grid to obtain multiple local area features.

[0054] As a key processing link of the system, the feature processing unit 103 undertakes the important task of converting the features output by the feature extraction network into a more compact and effective feature representation. The unit adopts a dual-path parallel processing architecture to process global features and local features respectively, thereby achieving comprehensive capture of hardware features.

[0055] In the global feature processing path, the feature processing unit first receives the multi-scale feature maps output by the feature pyramid network. These feature maps contain feature information of different spatial resolutions. After being fused by the feature pyramid network, they already have rich semantic information and detail features. For these multi-scale feature maps, the unit uses a global average pooling operation to process them. Specifically, for each feature map, the average value of all its spatial positions is calculated, and the two-dimensional feature map is compressed into a one-dimensional feature vector. This pooling operation not only greatly reduces the dimension of the feature, but also can extract a global feature representation with spatial invariance.

[0056] In the local feature processing path, the feature processing unit receives the target encoded feature map output by the last encoder. The output of the last encoder is chosen because this feature map has the richest semantic information while still retaining sufficient spatial resolution for local feature extraction. For this feature map, the unit adopts a grid-uniform partitioning strategy to divide it into N×N local regions of equal size, where N is usually 4 or 6. This uniform partitioning ensures that all parts of the hardware receive equal attention and that important local information is not missed. For each local region, the feature representation of the region is extracted through an average pooling operation, resulting in N×N local region features.

[0057] In order to ensure the effectiveness of the features, the feature processing unit also adopts some optimization strategies when extracting features. For example, before performing global average pooling, the feature map will be normalized so that the features of different channels have similar numerical ranges. When performing grid division, a certain overlap area will be reserved between adjacent areas, which can avoid the problem of feature breakage caused by strict division. At the same time, L2 normalization will be performed on the features extracted from each local area to enhance the discriminability of the features.

[0058] The global features and local features output by the feature processing unit will be passed to the subsequent contrastive learning module for processing. The dimension of the global feature is determined by the number of channels of the last encoder, usually 1024 dimensions, while the local feature is N×N feature vectors, each of which has a dimension of 1024. This feature representation method that contains both global information and local details provides comprehensive and effective feature support for subsequent hardware identification.

[0059] Furthermore, the grid in the feature processing unit is evenly divided into 4×4 divisions, with 25% overlap between adjacent regions, and each local region obtains a fixed-dimensional feature representation through an adaptive pooling operation.

[0060] The uniform grid division of the feature processing unit is a method of decomposing the target feature map into multiple local regions, aiming to extract local feature representations from the feature map, thereby enhancing the system's ability to capture detail information. In the present invention, the grid division method adopts a 4×4 uniform division, which means that the feature map is divided into 16 local regions. This division method allows each local region to cover a part of the target feature map while retaining the overall structure and distribution information.

[0061] In the specific implementation, in order to avoid the boundary of a single area being too rigid and information loss, the grid division introduces a 25% overlap area between adjacent areas. The design of the overlapping area allows each local area to not only contain its own core features, but also capture some feature information of the adjacent area, thereby increasing the connection and consistency between the areas. This processing method effectively improves the continuity and stability of feature extraction, especially in the feature extraction of the edges or key parts of the hardware.

[0062] After each local area is determined, in order to ensure the uniformity and efficiency of subsequent processing, the feature representation of the local area is converted into a fixed-dimensional feature through an adaptive pooling operation. Adaptive pooling is an operation that dynamically adjusts the size of the pooling window according to the target dimension, and can convert input features of any size into a predefined fixed size. In this embodiment, the goal of adaptive pooling is to generate a feature representation of consistent size for each local area, regardless of how the resolution of the original feature map changes. This method not only simplifies subsequent calculations, but also provides better compatibility for inputs of different resolutions.

[0063] Specifically, the adaptive pooling operation extracts features by calculating the average or maximum value within the local region. This extraction method can retain the global information or key feature points in the region, ensuring that the feature representation of each region is both robust and can fully express the local characteristics. Finally, the fixed-dimensional feature representations of all local regions will be integrated and input into the contrastive learning module for further similarity calculation and classification recognition.

[0064] This division and processing method has important technical significance in the task of automatic identification of hardware. On the one hand, the 4×4 grid division provides sufficient resolution, allowing the system to capture the subtle features of hardware; on the other hand, the 25% overlap area increases the continuity and robustness of feature extraction, reducing the risk of feature loss due to feature boundary cutting. At the same time, the introduction of adaptive pooling operations ensures the uniformity of local features, providing a solid foundation for subsequent analysis and calculation.

[0065] The prototype feature library 104 is used to store the features of each hardware category learned during the training phase. The features of each hardware category include a global prototype feature and multiple local area prototype features.

[0066] The prototype feature library 104 is the core component of the system for storing and managing the feature representation of hardware. Its design concept is derived from the prototype learning theory, and it supports subsequent recognition tasks by maintaining the standard feature representation of each hardware category. The feature library is obtained through the deep learning process during the training phase of the system and provides a stable and reliable feature reference during the actual application phase.

[0067] During the training phase, the construction process of the prototype feature library is carried out in an iterative optimization manner. The system first collects a large number of hardware training samples with category annotations. These samples need to cover various common hardware types, different shooting angles, and various practical application scenarios. For each category of training samples, the system extracts its global features and local features through the feature extraction network and feature processing unit. Then, through feature aggregation, the global features of all samples in the same category are weighted averaged to obtain the global prototype features of the category. Similarly, for each predefined local area, the system also weighted averages all sample features in the area to obtain the corresponding local prototype features.

[0068] The feature storage in the prototype feature library is organized using an efficient data structure. For each hardware category, the library stores a global prototype feature vector of a fixed dimension, which usually matches the output dimension of the last layer of the feature extraction network, with a typical value of 1024 dimensions. At the same time, multiple local prototype feature vectors corresponding to the grid division are also stored. If a 4×4 grid division is used, 16 local prototype feature vectors will be stored for each category. These feature vectors are all normalized by the L2 norm to ensure good numerical stability in subsequent similarity calculations.

[0069] In order to ensure the representativeness and discriminability of the prototype features, a quality assessment mechanism is also introduced in the process of building the feature library. The system will calculate the closeness of intra-class features and the separation of inter-class features, and use these indicators to evaluate the quality of the prototype features. If the quality of the prototype features of a certain category does not meet the standards, the system will trigger the feature relearning process of that category. In addition, the feature library also supports an incremental update mechanism. When a new hardware category is added, the features of the new category can be added to the library without affecting the features of the existing category.

[0070] In actual applications, the prototype feature library is loaded into memory to support fast feature retrieval. The library is implemented using an efficient data structure and indexing mechanism to ensure that the required prototype features can be quickly accessed when performing feature similarity calculations. At the same time, the system also implements feature compression and caching mechanisms to reduce storage overhead and improve access speed while ensuring recognition accuracy.

[0071] Through this design, the prototype feature library not only provides a stable and reliable feature reference, but also has good scalability and maintainability, and can meet various requirements for hardware identification in practical applications. The feature representation and organizational structure of the prototype feature library provide a solid foundation for the subsequent comparative learning module, enabling the system to accurately identify various types of hardware.

[0072] Furthermore, the prototype feature library includes a feature update mechanism, which is used to trigger the online update process of the prototype features when a specified number of high-confidence recognition samples are accumulated during the actual application of the system, and fine-tune the prototype features through momentum update to adapt to slight changes in the appearance of the hardware.

[0073] In the actual application of the hardware automatic identification system, the appearance of the hardware may change slightly due to the manufacturing process, wear and tear, or environmental conditions. If the prototype feature library in the system cannot be updated in time to reflect these changes, the recognition accuracy may decrease. In order to solve this problem, the present invention designs an online prototype feature update mechanism based on high-confidence samples, and fine-tunes the prototype features through momentum update, so that the system can dynamically adapt to changes in the appearance of the hardware and maintain high recognition accuracy.

[0074] The core logic of the feature update mechanism of the prototype feature library is to efficiently utilize the high-confidence samples accumulated during the actual recognition process. After each recognition task is completed, the system will screen the samples according to the confidence of the recognition result. Confidence is an important indicator to measure the reliability of the recognition result. Its value is calculated by the classification recognition module and reflects the similarity score between the hardware to be recognized and each category in the prototype feature library. For samples whose confidence exceeds the preset threshold, the system considers that their recognition results are reliable and can be used as a reference for feature update.

[0075] When the system accumulates a specified number of high-confidence samples, the online update process of the prototype features will be triggered. In the update process, the features of these high-confidence samples are first extracted and compared with the prototype features of the corresponding category in the current prototype feature library. The update adopts the momentum update method, which achieves fine-tuning by weighted averaging between the current prototype features and the new sample features, gradually introducing the feature information of the new samples without completely covering the original features. The weight distribution of the momentum update is usually dynamically adjusted according to the confidence of the new sample. High-confidence samples have a greater impact on the prototype features, while low-confidence samples have a smaller impact.

[0076] The advantage of this momentum update mechanism is that its update process is progressive and robust, which can effectively adapt to the gradual changes in the appearance of hardware and avoid the deviation of prototype features due to a small number of abnormal samples. In addition, the online update feature of this mechanism allows the system to complete the dynamic adjustment of the feature library without suspending operation, which is suitable for recognition scenarios with high real-time requirements.

[0077] In terms of specific implementation, the feature update mechanism of the prototype feature library can be implemented through software programming and integrated into the control module of the system. The update logic can be automatically triggered after the recognition task is completed, without manual intervention. The computational complexity of the entire update process is low, and the consumption of system resources is controllable, so it is suitable for embedded devices or scenarios with limited resources.

[0078] Through this design, the effectiveness of the prototype feature library can be dynamically maintained during the actual application of the system, allowing the hardware automatic identification system to operate stably for a long time, while adapting to subtle changes in the appearance of hardware, significantly improving recognition accuracy and robustness.

[0079] The comparative learning module 105 is used to calculate the similarity between the global features of the hardware to be identified and the global prototype features of each category in the prototype feature library to obtain the global similarity score of each category; for multiple local area features of the hardware to be identified, respectively calculate their similarity with the corresponding local prototype features of each category in the prototype feature library to obtain multiple local similarity scores of each category; and process the multiple local similarity scores of each category to obtain a comprehensive local similarity score of the category.

[0080] The contrast learning module 105 is the core discrimination unit of the system. Its main task is to evaluate the matching degree between the hardware to be identified and each standard category through accurate feature similarity calculation. This module adopts the strategy of dual comparison of global features and local features, and realizes accurate identification of hardware through multi-level similarity calculation.

[0081] In the global feature similarity calculation, the module first starts with the global features of the hardware to be identified. This global feature is a 1024-dimensional vector obtained by the feature processing unit through global average pooling. The module will calculate the similarity between this feature vector and the global prototype features of each category stored in the prototype feature library. The similarity calculation uses the cosine similarity method, that is, the dot product operation of the two feature vectors is performed and then divided by their respective L2 norms. This calculation method can not only effectively measure the similarity between feature vectors, but also has good robustness to the scale changes of feature vectors. Through this calculation process, the system obtains the global similarity score for each category.

[0082] In the local feature similarity calculation link, the module needs to handle more complex calculation tasks. For each local area feature of the hardware to be identified (assuming that 16 local features are obtained by using 4×4 grid division), the module needs to calculate the similarity between it and the local prototype feature of the corresponding position of the corresponding category in the prototype feature library. This position-corresponding matching strategy ensures that the comparison of local features is performed in the same semantic area, so that local differences can be captured more accurately. Cosine similarity is also used for calculation. For each category, the system will obtain a local similarity score equal to the number of grids.

[0083] In order to integrate multiple local similarity scores into a meaningful score, the contrastive learning module adopts a weighted average strategy. Specifically, each local region is first assigned an importance weight, which can be learned through the attention mechanism or pre-set according to the discriminability of the region. Then, the similarity scores of each local region are multiplied by the corresponding weights and summed, and finally normalized to obtain the comprehensive local similarity score of each category.

[0084] It is worth noting that in the actual calculation process, the module adopts a number of optimization measures to improve the calculation efficiency. For example, when performing batch processing, the features of multiple samples to be identified are organized into a matrix form, and matrix operations are used to accelerate the similarity calculation. At the same time, the module also implements a feature cache mechanism to cache the calculation results of features that appear repeatedly in a short period of time to avoid repeated calculations.

[0085] The global similarity score and comprehensive local similarity score output by the module will be passed to the classification and recognition module for final decision-making. This multi-level comparison strategy combining global and local features enables the system to focus on the overall shape and local details of the hardware at the same time, thereby improving the accuracy and robustness of recognition.

[0086] Furthermore, the contrastive learning module is specifically used for:

[0087] Use the following formula 1 to calculate the global similarity score of each category:

[0088]

[0089] Among them, S(c) represents the global similarity score of category c, which is used to measure the matching degree between the global feature F of the hardware to be identified and the global prototype feature of category c.

[0090] σ represents a nonlinear activation function, which is used to control the range of the similarity score. The nonlinear activation function includes a Sigmoid function or a Tanh function.

[0091] N represents the number of global prototype features corresponding to category c, which is generated by the system during the training phase and stored in the prototype feature library.

[0092] ω i Indicates dynamic adjustment of weights, see Formula 3 for details.

[0093] Cos(F, P(c, i)) represents the cosine similarity between the global feature F of the hardware to be identified and the i-th global prototype feature P(c, i), which is used to evaluate the directional similarity between the two. The value of cosine similarity is between -1 and 1, where 1 means they are exactly the same and -1 means they are completely opposite.

[0094] α represents the balance coefficient of global feature normalization, which is used to control the influence ratio of cosine similarity and feature distribution difference on the final score. The recommended value is 0.1 to 0.5, which is adjusted according to the experiment to ensure that the feature difference item does not excessively affect the total score.

[0095] F represents the global feature of the hardware to be identified, which is extracted by the feature processing unit;

[0096] P(c, i ) represents the i-th global prototype feature of category c, obtained from the prototype feature library;

[0097] P(c) is the weighted average of all global prototype features of category c;

[0098] μ c is the center vector of category c; ∈ is a small positive number used to avoid the denominator being zero;

[0099] ||F-μ c || represents the center vector μ of the global feature F of the hardware to be identified and the category c c The Euclidean distance reflects the degree of deviation between the hardware to be identified and the distribution center of category c.

[0100] ||P (c) -μ c || represents the weighted average P of the global prototype features of category c (c) and the category center vector μ cThe Euclidean distance is used to measure the distribution tightness of category prototype features.

[0101] Among them, the weighted average P(c) of all global prototype features of category c is calculated using the following formula 2:

[0102]

[0103] Where N represents the number of global prototype features corresponding to category c;

[0104] P(c, i ) represents the i-th global prototype feature of category c;

[0105] ω i With ω in formula 1 i The same means dynamic adjustment of weights, specifically the importance weight of the i-th global prototype feature of category c, which is determined by the distance between the feature and the category center. The smaller the distance, the greater the weight.

[0106] ω i is calculated according to the following formula 3:

[0107] ω i =exp(-λ·||P (c,i) -μ c ||)(3)

[0108] Among them, λ is an adjustment factor used to control the attenuation effect of distance on weight. The recommended value is 0.1 to 1.0, which is set according to the experimental distribution of category characteristics.

[0109] P(c, i) represents the i-th global prototype feature of category c; μc Represents the center vector of category c, calculated according to the following formula 4:

[0110]

[0111] Where N represents the number of global prototype features corresponding to category c.

[0112] Furthermore, the contrastive learning module is specifically used for:

[0113] According to the following formula 5, the comprehensive local similarity score is calculated:

[0114]

[0115] Among them, S( k ) represents the comprehensive local similarity score of category k, which is significant in providing a unified similarity evaluation value for category k, which is used to measure whether the hardware to be identified belongs to this category.

[0116] φ represents a normalization function of the comprehensive score, and the normalization function representing the comprehensive score includes Sigmoid or Softmax, which is used to limit the range of the final score;

[0117] U represents the number of local region features, which is the number of regions divided by the grid, indicating how many local regions the feature map is divided into. For example, if the feature map is divided into a 4×4 grid, then U=16.

[0118] η u To dynamically adjust the weight, it is used to measure the importance of each local area feature. Formula 6 calculates η u When the modulus length of the local feature ||L u ||Exponential processing is used to highlight the feature contribution of the significant area. The denominator contains the exponential sum of the feature modulus of all local areas to ensure weight normalization. The adjustment factor γ controls the distribution sensitivity of the weight. Its value is usually adjusted according to experimental results, and the recommended initial value is 1. When γ is large, the weight of the significant feature will increase significantly; when γ is small, the weight distribution is more uniform.

[0119] L u is the u-th local area feature of the hardware to be identified; Q u , k is the u-th local prototype feature corresponding to category k in the prototype feature library; ρ is the asymmetric similarity calculation function; ψ is the nonlinear feature distance transformation function; ξ is the weight coefficient used to balance similarity and distance, the recommended value is 0.5, and the value range is 0.1 to 1.

[0120] The dynamic adjustment weight ηu is calculated according to the following formula 6:

[0121]

[0122] Among them, L u is the u-th local area feature of the hardware to be identified; L v is the vth local area feature of the hardware to be identified; U represents the number of local area features; γ is a hyperparameter for adjusting the weight distribution of regional features, which is an adjustment factor for controlling the sensitivity of weight distribution. The recommended value is 1 and the range is 0.5 to 2;

[0123] The asymmetric similarity calculation function ρ uses the following formula 7:

[0124]

[0125] in, <L u , Q u , k> means for L u and Q u , k performs vector dot product operation;

[0126] Asymmetric similarity calculation function ρ(L u , Q u,k ) defines the local area feature L_u of the hardware to be identified and the u-th local prototype feature Q of category k u,k Formula 7 calculates the asymmetric normalized ratio of the dot product result and the characteristic modulus length. <L u , Q u , k> measures the directional similarity of two eigenvectors, while the asymmetric processing of modulus normalization uses ||Lu|| 0.5 and||Q u,k || 1.5 The purpose is to highlight the dominant role of prototype features on local similarity, while weakening the interference of large local modulus length in the features to be identified.

[0127] The nonlinear feature distance transformation function ψ adopts the following formula 8:

[0128] ψ(x)=log(1+x)(8)

[0129] Among them, x is the input variable.

[0130] The nonlinear feature distance transformation function ψ(x) is defined using Formula 8, which transforms the Euclidean distance square ||L u -Q u,k || 2 Convert to logarithmic form. The purpose of this transformation is to smooth the impact of distance values ​​and avoid excessive negative impact of regional features with larger distances on the similarity score. The function log(1+x) can retain high sensitivity in a small distance range while growing smoothly in a large distance range.

[0131] The classification recognition module 106 is used to process the global similarity score and the comprehensive local similarity score of each category to obtain the final similarity score of the category; sort the final similarity scores of all categories and select a specified number of categories with the highest scores as candidate categories; and select categories with final similarity scores higher than a preset threshold from the candidate categories as recognition results.

[0132] The classification recognition module 106 is the decision-making core of the system, responsible for integrating all similarity information and making the final classification judgment. This module adopts a multi-stage processing strategy, and ensures the accuracy and reliability of the recognition results through scientific fusion methods and strict screening mechanisms.

[0133] In the feature fusion stage, the classification and recognition module first needs to process two types of similarity scores from the contrastive learning module: the global similarity score and the comprehensive local similarity score. These two types of scores reflect the degree of match between the hardware to be identified and each category at different levels. The module uses an adaptive weighted approach to fuse these two types of scores, where the weight coefficient is determined by performance evaluation on the validation set. In specific implementation, for each category, its global similarity score and comprehensive local similarity score are multiplied by the corresponding weight coefficient (usually the global feature weight is slightly higher, such as 0.6, while the local feature weight is 0.4), and then the weighted results are added to obtain the final similarity score of the category. This weighting strategy takes into account the overall characteristics of the hardware while not ignoring the importance of local details.

[0134] After obtaining the final similarity scores of all categories, the module will sort the scores. The sorting uses a quick sort algorithm to sort all categories from high to low according to the final similarity scores. Then, the system will select a specified number of categories with the highest scores as candidate categories. This number is a configurable parameter, and in practical applications it is usually set to 3 to 5. This range can achieve a good balance between ensuring recognition accuracy and computational efficiency.

[0135] The final category determination stage adopts a threshold-based judgment mechanism. The system compares the final similarity score of the candidate category with the preset threshold. The setting of this threshold is crucial and needs to be determined through a large number of experiments during the system training phase. It is usually set between 0.75 and 0.85. Candidate categories below this threshold will be filtered out, and categories with scores exceeding the threshold will be determined as the final recognition results. This threshold screening mechanism effectively reduces the risk of misidentification.

[0136] In practical applications, the classification and recognition module also implements some optimization strategies. For example, when processing similarity scores, numerical stability optimization measures are adopted to avoid numerical overflow or precision loss. At the same time, the module also supports batch processing, which can make classification decisions for multiple hardware to be identified at the same time, significantly improving processing efficiency. In addition, the module also integrates the result visualization function, which can output the confidence distribution map of the recognition, helping users to understand the reliability of the recognition results more intuitively.

[0137] Through this multi-stage processing strategy, the classification and recognition module can accurately identify various types of hardware, even hardware with high similarity in appearance. The design of the module fully considers the actual application needs, ensuring the accuracy of recognition while also ensuring the practicality and reliability of the system.

[0138] Furthermore, the classification and recognition module adopts an adaptive threshold mechanism, which is specifically used to:

[0139] Dynamically adjust the similarity threshold for category determination based on the imaging quality of the hardware to be identified and the reliability of feature extraction;

[0140] When the final similarity score difference of multiple candidate categories is less than the preset difference threshold, they are marked as fuzzy categories to be manually confirmed.

[0141] The classification and recognition module is a key component in the hardware automatic recognition system that outputs the recognition results. Its core task is to make category determination based on the matching between the features of the hardware to be identified and the prototype feature library. In order to adapt to the differences that may occur in the hardware imaging quality and feature extraction process, this embodiment introduces an adaptive threshold mechanism in the classification and recognition module, so that the similarity threshold of category determination can be dynamically adjusted according to real-time conditions, thereby improving the reliability and flexibility of recognition.

[0142] The first part of the mechanism is to dynamically adjust the similarity threshold based on the imaging quality of the hardware to be identified and the reliability of feature extraction. In practical applications, imaging quality may be affected by many factors, such as lighting conditions, camera parameters, or external interference. The classification and recognition module calculates the imaging quality score by reading the image quality assessment results, such as clarity, brightness uniformity, and hardware integrity, and adjusts the similarity threshold for category determination based on this. When the imaging quality is high, the system can appropriately increase the similarity threshold to enhance the accuracy of recognition; when the imaging quality is low, the similarity threshold will be lowered to tolerate more uncertainty and avoid an increase in the misjudgment rate due to an excessively high threshold.

[0143] The reliability of feature extraction is also an important basis for dynamic adjustment of the threshold. In the feature extraction network, if the hierarchical distribution consistency of the multi-scale feature map is high, or the matching degree between the global feature and the local feature is good, it means that the reliability of the feature extraction process is high, and the system can appropriately increase the threshold. On the contrary, if there are significant distribution differences in the feature map, or there are large fluctuations in the local feature matching results, the system will lower the threshold to adapt to the unstable feature extraction results.

[0144] The second part of the adaptive threshold mechanism is to deal with the fuzzy judgment of candidate categories. When the classification recognition module generates a list of candidate categories based on the final similarity scores, it may encounter situations where the scores of multiple categories are very close. To avoid misjudgment, the classification recognition module compares the final similarity score differences of the candidate categories. If the difference between the highest score and the second highest score is less than the preset difference threshold, the system marks these candidate categories as fuzzy categories and outputs a prompt for manual confirmation. This design can not only reduce the system's misjudgment rate, but also retain the flexibility of decision-making in uncertain scenarios.

[0145] In terms of implementation, the adaptive threshold mechanism can be implemented through software programs and tightly integrated with the classification and recognition module. The dynamic adjustment logic of the threshold is based on the image quality assessment results and the internal state of feature extraction, which can be completed through real-time calculation, while the determination of fuzzy categories depends on the sorting and difference comparison of similarity scores, with simple logic and low computational complexity.

[0146] By introducing the adaptive threshold mechanism, the adaptability and reliability of the hardware automatic identification system are significantly enhanced. Whether it is dealing with input images of different qualities or processing the judgment of fuzzy categories under complex backgrounds, this mechanism can provide flexible response strategies to ensure the stable performance of the system in a variety of application scenarios.

[0147] The following is the reference implementation code of the system:

[0148]

[0149]

[0150]

[0151]

[0152]

[0153] The steps of training the hardware automatic identification system can be divided into several main stages: data preparation, model training, feature extraction optimization, prototype feature library construction and verification optimization. The following is the training process for implementing this system:

[0154] First, a large number of datasets covering a variety of hardware categories need to be acquired through the image acquisition module. These data include images of hardware from different perspectives, lighting conditions, and backgrounds. The acquired images will be preprocessed, including size standardization and data augmentation. Size standardization adjusts all images to a uniform size, while data augmentation expands the diversity of the dataset through rotation, flipping, random cropping, etc., and enhances the robustness of the model.

[0155] During the training phase, the cascade encoder units and feature pyramid network of the feature extraction network will be optimized first. After inputting the standardized image, the cascade encoder extracts multi-level features layer by layer, captures detail features through depthwise separable convolution, and improves the quality of feature extraction through channel attention and spatial attention mechanisms. The feature pyramid network fuses these features to generate multi-scale feature maps to capture the semantic information of the hardware at different scales.

[0156] Subsequently, the feature processing unit receives the extracted multi-scale feature map and further calculates the global features and local features. The global features are obtained through global average pooling and are used to describe the overall characteristics of the hardware; the local features are generated through meshing operations and cover multiple areas of the feature map. During the training process, it is necessary to ensure that these features can fully distinguish between different categories.

[0157] The construction of the prototype feature library is based on the aggregation of features for each category. In the training phase, high-confidence training samples are used to gradually generate global prototype features and local prototype features for each category through momentum updates. These prototype features need to be dynamically updated to adapt to possible minor changes in the appearance of hardware.

[0158] The core of the training is the contrastive learning module. The system calculates the similarity between the global features of the hardware to be identified and the global prototype features in the prototype feature library, and also calculates the similarity between the local features and the corresponding local prototype features. The similarity is calculated using methods such as cosine similarity and distance smoothing, and the results are used to generate global similarity scores and comprehensive local similarity scores. A loss function needs to be defined to maximize the similarity of the correct category and minimize the similarity of the wrong category. Commonly used methods include cross entropy loss or contrast loss.

[0159] Finally, the classification recognition module combines the global similarity score and the comprehensive local similarity score to obtain the final category score. During the training process, the candidate categories are screened out by adjusting the threshold, and the fuzzy categories are marked according to the confidence of the model prediction results for manual verification. In order to improve the performance of the model, the balance parameters in the similarity calculation are dynamically adjusted during the training phase to optimize the classification ability of the system.

[0160] After training is completed, the validation data set will be used to evaluate the model performance and quantify the system effect through indicators such as confusion matrix, accuracy, recall rate and F1 score. By repeatedly optimizing the above steps, it can be ensured that the system can demonstrate efficient hardware identification capabilities in practical applications.

[0161] In the above embodiment, a hardware automatic identification system is provided. Correspondingly, the present application also provides a hardware automatic identification method. Figure 2 , which is a flow chart of an embodiment of an automatic sorting method for electric hardware based on differential separation of the present application. Since this embodiment, i.e., the second embodiment, is basically similar to the first embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the first embodiment. The method embodiment described below is only illustrative.

[0162] A second embodiment of the present application provides a method for automatically identifying hardware, including:

[0163] Step S201: acquiring image data of hardware to be identified, and preprocessing the image data to obtain a preprocessed standardized image, wherein the preprocessing includes size standardization and data enhancement processing;

[0164] Step S202: extracting features through a feature extraction network, specifically including processing the preprocessed standardized image through a cascaded encoder unit, wherein the cascaded encoder unit includes a plurality of cascaded encoders, each encoder unit sequentially includes two depth-separable convolutional layers and a maximum pooling layer, wherein the depth-separable convolutional layer integrates a channel attention mechanism and a spatial attention mechanism, and obtains an encoding feature map output by each encoder; fusing the encoding feature map through a feature pyramid network to generate a multi-scale feature map of the hardware to be identified;

[0165] Step S203: performing feature processing, specifically including performing a global average pooling operation on the multi-scale feature map to extract the global features of the hardware to be identified; performing grid uniform division on the target encoding feature map output by the last encoder to obtain multiple local area features;

[0166] Step S204: performing feature comparison, specifically including calculating the similarity between the global features of the hardware to be identified and the global prototype features of each hardware category stored in advance, to obtain the global similarity score of each category; for multiple local area features of the hardware to be identified, respectively calculating the similarity between them and the corresponding local prototype features of each hardware category stored in advance, to obtain multiple local similarity scores of each category; processing the multiple local similarity scores of each category, to obtain the comprehensive local similarity score of the category;

[0167] Step S205: performing classification recognition, specifically including, for each category, processing its global similarity score and the comprehensive local similarity score to obtain a final similarity score for the category; sorting the final similarity scores of all categories, selecting a specified number of categories with the highest scores as candidate categories; and selecting categories with final similarity scores higher than a preset threshold from the candidate categories as recognition results.

[0168] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

Claims

1. A hardware automatic identification system, characterized in that: include: An image acquisition module, used to obtain image data of the hardware to be identified, preprocess the image data, and output a standardized image after preprocessing, wherein the preprocessing includes size standardization and data enhancement processing; A feature extraction network, comprising a cascaded encoder unit and a feature pyramid network, wherein the cascaded encoder unit comprises a plurality of cascaded encoders, wherein the input of the first encoder is the preprocessed normalized image, and each encoder unit comprises two depthwise separable convolutional layers and a maximum pooling layer in sequence, wherein the depthwise separable convolutional layers are integrated with a channel attention mechanism and a spatial attention mechanism, and are used to extract multi-level features of hardware to be identified; the feature pyramid network is used to obtain the encoding feature map output by each encoder, and fuse the encoding feature map to generate a multi-scale feature map of the hardware to be identified; A feature processing unit is used to receive the multi-scale feature map and the target coding feature map output by the last encoder; perform a global average pooling operation on the multi-scale feature map to extract the global features of the hardware to be identified; and perform a grid-even division on the target coding feature map to obtain multiple local area features; The prototype feature library is used to store the features of each hardware category learned during the training phase. The features of each hardware category include a global prototype feature and multiple local area prototype features. The comparative learning module is used to calculate the similarity between the global features of the hardware to be identified and the global prototype features of each category in the prototype feature library to obtain the global similarity score of each category; for multiple local area features of the hardware to be identified, the similarity between them and the corresponding local prototype features of each category in the prototype feature library is calculated respectively to obtain multiple local similarity scores of each category; the multiple local similarity scores of each category are processed to obtain the comprehensive local similarity score of the category; The classification recognition module is used to process the global similarity score and the comprehensive local similarity score of each category to obtain the final similarity score of the category; sort the final similarity scores of all categories and select a specified number of categories with the highest scores as candidate categories; and select categories with final similarity scores higher than a preset threshold from the candidate categories as recognition results.

2. The hardware automatic identification system according to claim 1, characterized in that: The image acquisition module includes a quality assessment unit, which is used to perform quality assessment on the preprocessed standardized image, specifically including: detecting the clarity, brightness uniformity and hardware integrity of the image. When the quality assessment result is lower than the preset standard, the image re-acquisition mechanism is triggered and the image acquisition parameters are adjusted.

3. The hardware automatic identification system according to claim 1, characterized in that: The cascade encoder unit in the feature extraction network includes five encoders, the number of output channels of each encoder is 64, 128, 256, 512 and 1024 respectively, the depth-separable convolution layer adopts a 3×3 convolution kernel, and the maximum pooling layer adopts a 2×2 pooling window with a step size of 2.

4. The hardware automatic identification system according to claim 1, characterized in that: The feature pyramid network fuses the output features of encoders at each level through a top-down feature fusion path, specifically including: upsampling high-level features so that they have the same spatial resolution as low-level features, and then additively fusing them with corresponding low-level features, and sequentially processing them to obtain multi-scale feature maps with different resolutions.

5. The hardware automatic identification system according to claim 1, characterized in that: The grid in the feature processing unit is evenly divided into 4×4 divisions, with 25% overlap between adjacent regions, and each local region obtains a fixed-dimensional feature representation through an adaptive pooling operation.

6. The hardware automatic identification system according to claim 1, characterized in that: The prototype feature library includes a feature update mechanism, which is used to trigger the online update process of the prototype features when a specified number of high-confidence recognition samples are accumulated during the actual application of the system, and fine-tune the prototype features through momentum update to adapt to slight changes in the appearance of the hardware.

7. The hardware automatic identification system according to claim 1, characterized in that: The classification and recognition module adopts an adaptive threshold mechanism, which is specifically used for: Dynamically adjust the similarity threshold for category determination based on the imaging quality of the hardware to be identified and the reliability of feature extraction; When the final similarity score difference of multiple candidate categories is less than the preset difference threshold, they are marked as fuzzy categories to be manually confirmed.

8. The hardware automatic identification system according to claim 1, characterized in that: The contrastive learning module is specifically used for: Use the following formula 1 to calculate the global similarity score of each category: Among them, S (c) represents the global similarity score of category c; σ represents a nonlinear activation function, which is used to control the range of the similarity score. The nonlinear activation function includes a Sigmoid or Tanh function; N represents the number of global prototype features corresponding to category c; ω i Indicates dynamic adjustment of weights; cos(F, P (c,i) ) represents the global feature F of the hardware to be identified and the i-th global prototype feature P (c,i) α represents the balance coefficient of global feature normalization; F represents the global feature of the hardware to be identified, which is extracted by the feature processing unit; P (c,i) represents the i-th global prototype feature of category c, obtained from the prototype feature library; P (c) is the weighted average of all global prototype features of category c; μ c is the center vector of category c; ∈ is a small positive number used to avoid the denominator being zero; Among them, the weighted average P of all global prototype features of category c (c) The following formula 2 is used for calculation: Where N represents the number of global prototype features corresponding to category c; P (c,i) represents the i-th global prototype feature of category c; ω i With ω in formula 1 i The same means that the weight is adjusted dynamically, which is calculated according to the following formula 3: oh i =exp(-λ·||P (c,i) -m c ||) (3) Where λ is the adjustment factor; P (c,i) represents the i-th global prototype feature of category c; μ c Represents the center vector of category c, calculated according to the following formula 4: Where N represents the number of global prototype features corresponding to category c.

9. The hardware automatic identification system according to claim 1, characterized in that: The contrastive learning module is specifically used for: According to the following formula 5, the comprehensive local similarity score is calculated: Among them, S (k) represents the comprehensive local similarity score of category k; φ represents the normalization function of the comprehensive score, and the normalization function representing the comprehensive score includes Sigmoid or Softmax, which is used to limit the final score range; U represents the number of local area features, which is the number of areas divided by the grid; η u To dynamically adjust the weight; L u is the u-th local area feature of the hardware to be identified; Q u,k is the u-th local prototype feature corresponding to category k in the prototype feature library; ρ is the asymmetric similarity calculation function; ψ is the nonlinear feature distance transformation function; ξ is the weight coefficient used to balance similarity and distance; Among them, the dynamic adjustment weight η u , calculated according to the following formula 6: Among them, L u is the uth local area feature of the hardware to be identified; l v is the vth local area feature of the hardware to be identified; U represents the number of local area features; γ is a hyperparameter for adjusting the weight distribution of regional features; The asymmetric similarity calculation function ρ uses the following formula 7: in, <L u , Q u,k 〉 indicates that for L u and Q u,k Perform vector dot product operations; The nonlinear feature distance transformation function ψ adopts the following formula 8: ψ(x)=log(1+x)(8) Among them, x is the input variable.

10. A method for automatically identifying hardware, characterized in that: include: Acquire image data of the hardware to be identified, and preprocess the image data to obtain a preprocessed standardized image, wherein the preprocessing includes size standardization and data enhancement processing; Extracting features through a feature extraction network, specifically including processing the preprocessed standardized image through a cascade encoder unit, wherein the cascade encoder unit includes a plurality of cascaded encoders, each encoder unit sequentially includes two depth-separable convolutional layers and a maximum pooling layer, wherein the depth-separable convolutional layer integrates a channel attention mechanism and a spatial attention mechanism, and obtains an encoding feature map output by each encoder; fusing the encoding feature map through a feature pyramid network to generate a multi-scale feature map of the hardware to be identified; Perform feature processing, specifically including performing a global average pooling operation on the multi-scale feature map to extract the global features of the hardware to be identified; and evenly dividing the target encoding feature map output by the last encoder into a grid to obtain multiple local area features; Perform feature comparison, specifically including calculating the similarity between the global features of the hardware to be identified and the global prototype features of each hardware category stored in advance, to obtain the global similarity score of each category; for multiple local area features of the hardware to be identified, respectively calculating the similarity between them and the corresponding local prototype features of each hardware category stored in advance, to obtain multiple local similarity scores of each category; processing the multiple local similarity scores of each category, to obtain the comprehensive local similarity score of the category; Classification recognition is performed, specifically including, for each category, processing its global similarity score and the comprehensive local similarity score to obtain a final similarity score for the category; sorting the final similarity scores of all categories, selecting a specified number of categories with the highest scores as candidate categories; and selecting categories with final similarity scores higher than a preset threshold from the candidate categories as recognition results.

Citation Information

Cited By

  • Universe unattended intelligent monitoring method based on multi-source data fusion

    CN120932182A

  • Global unattended intelligent monitoring method based on multi-source data fusion

    CN120932182B

  • Target tracking and counting method and system for electric power fittings

    CN121033106A

  • A power fitting target tracking counting method and system

    CN121033106B