A spare part identification method and system based on a visual AI model

CN122841909APending Publication Date: 2026-09-29CHINA TOBACCO GUIZHOU IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510360577.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明的目的在于解决现有技术中依赖人工识别备件所导致的效率低、准确性差和人力成本高等问题

Benefits of technology

[0045]采用上述技术方案,可以在用户访问页面时提前加载图像内容,减少用户等待时间,提升系统的响应速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841909A_ABST
    Figure CN122841909A_ABST
Patent Text Reader

Abstract

The application discloses a spare part recognition method based on a visual AI model, and comprises the following steps: collecting spare part images; manually labeling the spare part images to generate a labeled data set; constructing a SpareNet-Pro model for spare part recognition based on a YOLO model, retaining a backbone network and a classification head of the YOLO model, and adding a multi-scale feature fusion network; performing enhancement processing on the labeled data set to generate an enhanced sample set; training the SpareNet-Pro model by using the enhanced sample set; and recognizing images of spare parts to be recognized by using the trained SpareNet-Pro model to obtain the categories and confidence of the spare parts to be recognized. The application can effectively improve the searching efficiency of spare parts in the equipment maintenance and repair process, greatly improve the query speed and accuracy, improve the equipment maintenance efficiency and maintenance quality, and reduce the labor cost. The application also provides a spare part recognition system based on a visual AI model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning, and in particular to a spare parts identification method and system based on a visual AI model. Background Technology

[0002] With the continuous improvement of industrial automation and intelligence, spare parts management and equipment maintenance have become crucial aspects of enterprise operation. Spare parts are diverse and complex in model, and quick and accurate identification of spare parts is a key factor in ensuring normal equipment operation and timely maintenance.

[0003] Currently, most companies still rely on traditional manual identification and manual input queries when identifying spare parts. This is not only inefficient but also prone to errors, leading to inaccurate identification results. In such cases, temporary measures are often required, which affects the efficiency and accuracy of spare parts management. Summary of the Invention

[0004] The purpose of this invention is to solve the problems of low efficiency, poor accuracy, and high labor costs caused by relying on manual identification of spare parts in existing technologies. This invention provides a spare parts identification method and system based on a visual AI model. Through visual AI models and deep learning technologies, it achieves rapid and accurate identification of spare parts, effectively improving the efficiency of finding parts during equipment maintenance and repair, greatly increasing query speed and accuracy, improving equipment maintenance efficiency and quality, and reducing labor costs.

[0005] To address the aforementioned technical problems, embodiments of the present invention disclose a spare parts identification method based on a visual AI model, comprising the following steps:

[0006] Acquire images of spare parts;

[0007] The spare parts images were annotated using a manual annotation method to generate an annotated dataset;

[0008] A SpareNet-Pro model for spare parts identification is constructed. The SpareNet-Pro model is an improvement on the YOLO model, retaining the backbone network and classification head of the YOLO model and adding a multi-scale feature fusion network.

[0009] The labeled dataset is augmented to generate an augmented sample set, simulating the actual working conditions of spare parts when they are installed on the equipment.

[0010] The SpareNet-Pro model was trained using an augmented sample set;

[0011] The trained SpareNet-Pro model is used to identify images of spare parts to be identified, and the category and confidence level of the spare parts to be identified are obtained.

[0012] By employing the above technical solution, the spare parts identification method based on the SpareNet-Pro model can significantly improve the efficiency and accuracy of spare parts identification. By augmenting the labeled dataset to simulate various scenarios of spare parts in real-world working environments, the diversity of training samples is increased, improving the model's robustness and generalization ability. Training with the augmented sample set makes the training process more efficient while avoiding overfitting.

[0013] Optionally, a SpareNet-Pro model for spare parts identification is constructed, including:

[0014] The backbone network serves as the feature extraction network for the SpareNet-Pro model, used to extract low-level, mid-level, and deep features from the images input to the SpareNet-Pro model.

[0015] The multi-scale feature fusion network fuses low-level, mid-level, and deep features through upsampling and feature concatenation.

[0016] The classification head is used for image classification tasks. It outputs a spare part category prediction by connecting a fully connected layer after global average pooling.

[0017] By employing the above technical solution, and by retaining the backbone network and classification head of the YOLO model, combined with a multi-scale feature fusion network, the model can effectively extract multi-level features from the input image, improving its ability to identify spare parts under different scales and complex backgrounds. Simultaneously, multi-scale feature fusion makes the model more robust, enabling it to handle spare parts identification tasks at different scales and enhancing its adaptability in complex environments.

[0018] Optionally, the backbone network uses the C2f module to extract low-level, mid-level, and deep features from the images input to the SpareNet-Pro model.

[0019] By adopting the above technical solution, the C2f module helps the model learn more detailed information in the image at multiple levels, thereby enhancing the model's recognition performance.

[0020] Optionally, the YOLO model is the YOLOv8 model.

[0021] Optionally, the augmentation process involves performing standard data augmentation on each data sample in the labeled dataset during each training round, and determining, based on random values, to use at most two of the following: advanced data augmentation and occlusion, to generate an augmented sample set.

[0022] By adopting the above technical solution, each sample is randomly augmented during each round of training to simulate various situations that spare parts may encounter in actual work, thereby increasing the diversity of training samples, improving the generalization ability of the model, and avoiding overfitting of the model to a single mode during training.

[0023] Optional advanced data augmentation options include Mosaic augmentation, CutMix augmentation, or Copy-Paste augmentation.

[0024] Optionally, the SpareNet-Pro model can be trained using an augmented sample set, including:

[0025] The augmented sample set is divided into a training set, a validation set, and a test set;

[0026] Model training was performed using the AdamW optimizer.

[0027] Save the optimal model weights.

[0028] By adopting the above technical solution, the enhanced sample set can be divided into training set, validation set and test set, which can effectively evaluate the performance of the model and ensure the quality of the training set. Using the AdamW optimizer for training can effectively avoid overfitting and optimize the model weights, while saving the optimal model weights for subsequent applications.

[0029] Optionally, model training also includes weighting the loss of the augmented sample set using a loss function. If occlusion is detected in a data sample during augmentation, the loss of the data sample is multiplied by a loss amplification factor. The loss amplification factor is used to increase the loss weight of the data sample to improve the SpareNet-Pro model's ability to recognize partially occluded objects.

[0030] By adopting the above technical solution, a weighted loss function is introduced to process samples containing occlusion with weight, so that the model can pay more attention to these more difficult samples during training, thereby improving the model's ability to identify partially occluded spare parts.

[0031] Optionally, the initial value of the loss amplification factor is 1.5, and the loss amplification factor is adjusted during model training based on the error of the augmented samples with occlusion on the validation set.

[0032] By adopting the above technical solution, the loss amplification factor is adjusted according to the error of occluded samples on the validation set, which further optimizes the training process of the model, enabling the model to better adapt to different degrees of occlusion and improve its accuracy and robustness.

[0033] The present invention also discloses a spare parts identification system based on a visual AI model, used to implement the above-described identification method. The identification system includes:

[0034] The front-end system is developed based on the Flask framework. The front-end system includes the front-end UI interface, which is used for user interface interaction.

[0035] The backend system, developed based on the Flask framework, is used for data processing and model invocation.

[0036] The database uses MySQL to store the data required by the spare parts identification system.

[0037] The mobile application, developed based on the Android platform, includes a mobile UI interface. The mobile application is used to upload images of spare parts to be identified and to receive the category and confidence level of the spare parts to be identified.

[0038] The backend system communicates with the frontend system through a web interface, communicates with the mobile application through a mobile API interface, and communicates with the database through SQL statements and a database connection library.

[0039] By adopting the above technical solution, a complete spare parts identification system is formed through a front-end system, a back-end system, a database, and a mobile application. This system enables the uploading, processing, and identification of spare parts images, providing a convenient user experience. The front-end and back-end systems, developed using the Flask framework, and the use of MySQL for data storage, provide efficient and stable data storage and processing capabilities. Furthermore, the mobile application enhances the user experience.

[0040] Optionally, the database includes a spare parts information table, and may also include at least one of an identification result table, a user information table, and a backend management table. The spare parts information table is used to store at least one of the spare parts' category, number, URL, and image. The identification result table is used to record the identification results of each uploaded image during the identification process. The user information table is used to store the user's login information and usage records. The backend management table is used by the administrator to manage user permissions.

[0041] By adopting the above technical solution, we can ensure the orderly storage and management of data to support the efficient operation of the system. At the same time, the database design has good scalability, and the table structure can be further added or modified as needed to adapt to different business requirements.

[0042] Optionally, the front-end system and the mobile application upload images of the spare parts to be identified to the back-end system through the front-end UI and the mobile UI, respectively. The back-end system receives the images of the spare parts to be identified and calls the trained SpareNet-Pro model to perform image recognition. The recognition results are uploaded to the database, the front-end system, and the mobile application. The recognition results include confidence scores, as well as at least one of the spare parts' category, number, URL, and image.

[0043] Using the above technical solution, users can upload spare parts images and receive recognition results through the front end or mobile device, achieving an efficient operating experience.

[0044] Optionally, the front-end system can preload images by embedding spare parts images into the front-end UI.

[0045] By adopting the above technical solution, image content can be preloaded when users access the page, reducing user waiting time and improving system response speed. Attached Figure Description

[0046] Figure 1 The diagram shows a flowchart of a spare parts identification method based on a visual AI model according to an embodiment of the present invention. Detailed Implementation

[0047] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention is presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to these embodiments. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a deep understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0048] It should be noted that in this specification, similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0050] To address the problems of low efficiency, error-proneness, and high labor costs associated with traditional manual identification and manual input queries in the spare parts identification process, which make it difficult to meet the demand for fast and accurate spare parts identification, this invention discloses a spare parts identification method and system based on a visual AI model. By using visual AI models and deep learning technologies, it achieves fast and accurate identification of spare parts, effectively improving the efficiency of finding parts during equipment maintenance and repair, greatly improving query speed and accuracy, enhancing equipment maintenance efficiency and quality, reducing labor costs, and ensuring normal equipment operation and timely maintenance.

[0051] This application discloses a spare parts identification method based on a visual AI model, such as... Figure 1 As shown, the identification method includes:

[0052] S101: Acquire images of spare parts;

[0053] S102: Use manual annotation to annotate spare parts images and generate an annotated dataset;

[0054] S103: Construct a SpareNet-Pro model for spare parts identification. The SpareNet-Pro model is an improvement on the YOLO model, retaining the backbone network and classification head of the YOLO model and adding a multi-scale feature fusion network.

[0055] S104: Enhance the labeled dataset to generate an enhanced sample set to simulate the actual working conditions of spare parts when they are installed on the equipment;

[0056] S105: Train the SpareNet-Pro model using an enhanced sample set;

[0057] S106: Use the trained SpareNet-Pro model to identify the image of the spare part to be identified, and obtain the category and confidence level of the spare part to be identified.

[0058] By employing the aforementioned technical solution, and utilizing visual AI models and deep learning technologies, rapid and accurate identification of spare parts can be achieved. This effectively improves the efficiency of parts retrieval during equipment maintenance and repair, significantly enhancing query speed and accuracy, improving equipment maintenance efficiency and quality, reducing labor costs, and ensuring normal equipment operation and timely repair. Specifically, high-quality image acquisition and accurate manual annotation provide high-quality image and label data for subsequent model training. By improving the YOLO model and combining it with the optimized design of a multi-scale feature fusion network, a highly efficient, accurate, and robust SpareNet-Pro model for spare parts identification is constructed, capable of handling spare parts identification tasks at different scales and in complex backgrounds. The labeled dataset is augmented to simulate various changes in the real environment, enhancing the diversity of the dataset and the model's generalization ability. The SpareNet-Pro model is trained using the augmented sample set, enabling it to learn useful features from the augmented dataset and improve its recognition capabilities. Finally, the trained model accurately identifies spare parts from the images to be identified and returns the results to the user, providing an accurate and reliable spare parts identification service.

[0059] Hereinafter, steps S101 to S106 will be specifically described according to embodiments of this application.

[0060] It should be noted beforehand that in this embodiment, the YOLOv8 (You Only Look Once Version 8) model is used as a specific implementation of the YOLO model. YOLOv8 is a version of the YOLO object detection algorithm released in 2023. The core idea of ​​the YOLO algorithm is to transform the object detection task into a regression problem, predicting the category and location of multiple objects simultaneously in a single neural network. YOLOv8 differs from traditional two-stage object detection methods. In traditional two-stage object detection methods such as Faster R-CNN, the detection process is usually divided into a generation stage and a classification / regression stage. The generation stage uses a region proposal network to generate a series of candidate regions that may contain objects from the input image, while the classification / regression stage feeds the candidate regions extracted in the first stage into a classification network for further processing. YOLOv8, however, uses single-stage detection, eliminating the need for a candidate region generation process and directly predicting the category and location based on the input image, thus improving detection speed. The category and location are determined by the classification head and detection head, respectively. Furthermore, YOLOv8 relies on CNNs (Convolutional Neural Networks) for feature extraction, enabling efficient processing of image data.

[0061] The backbone network is the foundation of any object detection model. Its role is to extract features from the input image. The YOLOv8 backbone network typically consists of multiple convolutional layers, activation functions, pooling layers, etc., with the aim of extracting low-level, mid-level, and deep features from the input image. By extracting image features layer by layer, the backbone network can learn the spatial and semantic information in the image, helping subsequent object detection layers to accurately predict the location and category of the target.

[0062] Multi-scale fusion networks are a type of network structure used to improve the model's ability to detect targets at different scales. By fusing features from different levels, they can capture information about both large and small targets simultaneously.

[0063] The classification head is a crucial component of object detection models. Its task is to map the features extracted from the backbone network to different object categories, classifying the objects and outputting the category probability distribution for each detection box. In YOLOv8, the classification head typically includes fully connected layers or convolutional layers, responsible for mapping the features of each detection box to the corresponding object category.

[0064] In step S101, a suitable image acquisition device needs to be selected to acquire images of the spare parts. A high-resolution industrial camera or digital camera can be used. To ensure image clarity and consistency, good lighting conditions should be maintained during acquisition, and images should be taken from multiple angles. The shooting distance should ensure that the key features of the spare parts are clearly displayed. All acquired images should be saved in a standard format, and the number of images should be sufficiently diverse to avoid overfitting of the model due to an excessive proportion of one type of data.

[0065] For example, in this embodiment of the application, the resolution of the image acquisition device is 3840×2160 pixels, artificial light source is used to provide uniform illumination, the device position is fixed by a tripod, about 50cm directly above the spare parts, the image format is JPG, the acquired spare parts images include 210 images of brand new spare parts in the warehouse, 210 images of machine spare parts on site, and 210 images of images that have been artificially blurred after acquisition, for a total of 630 images, which are classified and organized to ensure the balance of data and improve the generalization ability and robustness of the model.

[0066] It should be noted that in some other implementations, since spare parts are installed on the equipment during use, it is not convenient to acquire images. Therefore, image acquisition can also be carried out only in the spare parts warehouse, and only images of brand-new spare parts can be acquired. The actual working conditions of the spare parts when they are installed on the equipment can be simulated during subsequent data augmentation.

[0067] In step S102, for the acquired spare parts images, each image is manually labeled to create labeled data, generating a labeled dataset. Specifically, the labeling process includes precisely selecting each spare part in the image and assigning an appropriate category label to each spare part. The labelers need to accurately assign categories to the spare parts based on their appearance, shape, and function, such as motors, sensors, valves, etc., and assign each spare part a unique identifier. Labeling can be done using rectangular or polygonal boxes to select the location of the spare part in the image, ensuring that each selected area completely covers the key parts of the spare part. Finally, after completing the labeling of all images, a labeled dataset is generated. This labeled dataset will be used as the basis for training the SpareNet-Pro model after subsequent augmentation processing, ensuring that the model can learn the accurate features of the spare parts and achieve accurate spare part recognition in practical applications. For example, Adobe Photoshop can be used for labeling.

[0068] In step S103, regarding the construction of the model, this embodiment of the application optimizes and improves the YOLOv8 model. Considering the characteristics of spare parts image acquisition and the many details on the surface of spare parts, and combined with the actual situation in the workshop, the YOLOv8 model is simplified. The backbone network and classification head of the YOLOv8 model are retained, the modules related to the detection head of the YOLOv8 model are removed, and the improved multi-scale fusion network is added to construct the SpareNet-Pro model for spare parts identification.

[0069] Specifically, the backbone of the SpareNet-Pro model employs the feature extraction module from YOLOv8, responsible for extracting low-level, mid-level, and deep features from the input image. The backbone extracts image features progressively through multiple convolutional layers, with layers connected via convolution, activation functions, and pooling operations. The specific structure includes multiple C2f modules, each containing a series of convolutional operations, effectively extracting multi-level information from the image. For example, in this embodiment, the first layer reduces the input 224×224×3 image size to 112×112×32 through convolutional operations, extracting low-level features for downsampling; the second layer further downsamples the feature map to 56×56×64, achieving preliminary feature fusion, and then extracts low-level features through the C2f module. As the backbone deepens, the spatial resolution of the features gradually decreases, while the number of channels gradually increases, thus achieving more complex feature extraction. Ultimately, the deepest feature extracted by the backbone network is 14×14×256, which is used for subsequent multi-scale feature fusion.

[0070] The multi-scale feature fusion network (Neck) of the SpareNet-Pro model fuses low-level, mid-level, and deep features extracted from the backbone network through upsampling and feature concatenation. The Neck utilizes features at different scales to fuse global semantic information and local details, thereby improving the model's robustness to incomplete targets, size variations, and partial occlusion. For example, in this embodiment, the Neck first upsamples the 14×14×256 deep features to 28×28 and concatenates them with the 28×28×128 mid-level features to obtain a mid-scale feature map. Then, a 1×1 convolution operation is used to reduce the dimensionality of the concatenated features, resulting in a 128-channel mid-scale fused feature map. This feature map is then upsampled again to 56×56 and concatenated with the 56×56×64 low-level features to obtain higher-resolution features. Finally, a 1×1 convolution is used again to reduce the dimensionality of the fused features, ultimately outputting a 64-channel feature map, which is then processed in the classification head.

[0071] The SpareNet-Pro model's classification head is responsible for transforming the feature maps extracted by the model into the final class prediction results. The classification head includes a global average pooling layer, a fully connected layer, and an activation function. For example, in this embodiment, the 56×56×64 feature map is first compressed into a 1×1×64 global feature vector through a global average pooling layer, preserving the global semantic information of the image. Then, after passing through the ReLU activation function, features are further extracted through a fully connected layer, and a Dropout operation is added to prevent overfitting. Finally, through another fully connected layer, the model outputs the predicted class. The output of the classification head is a vector of size num_classes, representing the predicted probability of each class.

[0072] In step S104, the labeled dataset is augmented to generate an augmented sample set. Depending on training needs, various data augmentation methods can be used, such as Mosaic augmentation, CutMix augmentation, Copy-Paste augmentation, random cropping and scaling, color perturbation, noise addition, and occlusion, to simulate the actual working conditions of spare parts installed on equipment. During each training round, standard data augmentation can be performed on each data sample in the labeled dataset, and at most two methods—advanced data augmentation and occlusion—can be selected based on random values ​​to generate an augmented sample set. This effectively improves the model's generalization ability, especially when dealing with complex spare part images in real-world applications, those with partial occlusion, or those with significant environmental changes. Advanced data augmentation includes Mosaic augmentation, CutMix augmentation, or Copy-Paste augmentation.

[0073] During the enhancement process, different enhancement methods are applied to the images in each epoch of training to improve the diversity of the training data. The selection of enhancement methods is usually random, and the application of a certain enhancement technique is determined based on a certain probability. For example, in the embodiments of this application, the collected spare parts images may differ significantly from the actual situation, and there may be many types of spare parts with many surface details. Therefore, for each data sample in the labeled dataset, a random value is first generated using the random.random() function. If this random value is less than 0.3, the current data sample is occluded, and the occlusion flag occlusion_flag is 1, indicating that occlusion is applied; otherwise, occlusion_flag is 0, indicating that occlusion is not applied. Next, a parameter, such as `aug_choice`, is defined to determine whether to use Mosaic enhancement, CutMix enhancement, and Copy-Paste enhancement. The `random.random()` function is used again to generate a random value and assign it to `aug_choice`. If `aug_choice < 0.15`, Mosaic enhancement is used; if 0.15 ≤ `aug_choice < 0.3`, CutMix enhancement is used; if 0.3 ≤ `aug_choice < 0.45`, Copy-Paste enhancement is used; and if `aug_choice ≥ 0.45`, advanced data augmentation is not used. Finally, standard data augmentation is applied to the image, including random cropping, scaling, color perturbation, and noise reduction.

[0074] Based on the above conditions for generating random values, the following four scenarios may occur: If the first value generated by `random.random()` is ≥ 0.3 and `aug_choice` is ≥ 0.45, then neither occlusion nor advanced data augmentation will be applied; only standard data augmentation will be used. If the first value generated by `random.random()` is ≤ 0.3 and `aug_choice` is ≥ 0.45, then advanced data augmentation will not be used; only standard data augmentation will be used, and occlusion will be applied. If the first value generated by `random.random()` is ≥ 0.3 and `aug_choice` < 0.45, then... No occlusion is applied; only standard data augmentation and advanced data augmentation are used, and one of the advanced data augmentation methods (Mosaic augmentation, CutMix augmentation, and Copy-Paste augmentation) is selected based on the value of aug_choice. If the value generated by random.random() for the first time is ≤0.3 and aug_choice<0.45, then occlusion, standard data augmentation, and advanced data augmentation are used simultaneously, and one of the advanced data augmentation methods (Mosaic augmentation, CutMix augmentation, and Copy-Paste augmentation) is selected based on the value of aug_choice.

[0075] Those skilled in the art will understand that the numerical ranges used for determination described above are a further detailed explanation of this application in conjunction with specific embodiments, but it should not be construed as the specific implementation of this application being limited to the above numerical ranges. Adjustments can be made reasonably according to actual circumstances, and as long as the range is reasonable, other numerical ranges do not deviate from the spirit and scope of this application. In some other embodiments, 0.4 can be used to determine whether occlusion is applied. If the random value is less than 0.4, occlusion is applied; otherwise, occlusion is not applied. 0.6 can also be used to determine whether advanced data augmentation is applied. If aug_choice < 0.2, Mosaic augmentation is used; if 0.2 ≤ aug_choice < 0.4, CutMix augmentation is used; if 0.4 ≤ aug_choice < 0.6, Copy-Paste augmentation is used; if aug_choice ≥ 0.6, advanced data augmentation is not used.

[0076] This process ensures that every data sample in the labeled dataset undergoes standard data augmentation in each training round, guaranteeing basic data augmentation. Simultaneously, occlusion is applied to approximately 30% of the data samples in the overall labeled dataset, and Mosaic augmentation, CutMix augmentation, and Copy-Paste augmentation are applied to approximately 45% of the data samples in a 1:1:1 ratio.

[0077] In some other implementations, the random.random() function can be used only once to generate a random value to determine whether to apply occlusion, Mosaic enhancement, CutMix enhancement, or Copy-Paste enhancement. That is, the same data sample can be enhanced with a maximum of one of the following enhancements at the same time as standard data enhancement: occlusion, Mosaic enhancement, CutMix enhancement, and Copy-Paste enhancement.

[0078] After enhancement processing, each image generates new enhanced samples, which are then added to the dataset, forming a larger and more diverse set of enhanced samples. This enhanced sample set becomes an important data source for subsequent model training. In this way, the model can be exposed to more training scenarios, thereby improving the accuracy and robustness of recognition, especially in the face of different working environments, occlusion, or background changes, enabling the model to maintain high recognition capabilities.

[0079] It should be noted that occlusion simulations of spare parts during installation, using methods such as CoarseDropout, randomly occludes a portion of the image, allowing the model to learn how to handle partially occluded spare parts and avoid relying solely on complete image features. Mosaic augmentation stitches multiple images into a new image, adding different backgrounds or objects during the augmentation process to simulate the actual installation scenarios of spare parts in complex backgrounds or on different devices, helping to enhance the model's adaptability to various scenarios. CutMix augmentation swaps portions of two images, enabling the model to learn how spare parts change in different environments, simulating spare part installation on different types of devices or in different environments. Copy-Paste augmentation copies a portion of an image from one image to another, simulating the installation scenario of multiple spare parts, helping the model learn when spare parts are installed together on the same device. Standard data augmentation includes image rotation, scaling, cropping, flipping, and color perturbation, which helps the model better adapt to spare part images of different sizes, angles, and lighting conditions.

[0080] In step S105, the augmented sample set is divided into a training set, a validation set, and a test set according to a certain ratio. The training set is used for model training, the validation set is used to adjust the model's hyperparameters, and the test set is used for final performance evaluation of the model. This division ensures the diversity and representativeness of each part of the data and avoids model overfitting. For example, the augmented sample set is divided into the training set, validation set, and test set in a 7:2:1 ratio.

[0081] The SpareNet-Pro model is trained using an augmented sample set. The training process includes calculating the loss through forward propagation and then adjusting the model parameters through backpropagation. Specifically, in each training epoch, training data is fed into the network from the input layer, processed by the backbone network for feature extraction, then fused through a multi-scale feature fusion network, and finally output as a prediction result through a classification head. For example, in this embodiment, the number of training epochs is 50, the batch size is 32, the learning rate (lr) is 0.001, and the target size of the input image is 224×224 pixels.

[0082] Meanwhile, to further optimize the model's recognition capabilities, this application's embodiments design a custom loss function, defining the OcclusionWeightedLoss class. This loss function is based on weighted cross-entropy and weights the sample loss according to the occlusion flag. Specifically, during data augmentation, if a data sample is occluded, such as through CoarseDropout, the loss of that data sample is multiplied by a large loss amplification factor to increase the loss weight of the data sample. The initial value of the loss amplification factor is 1.5, which can be further optimized through parameter tuning on the validation set to improve the model's performance in occluded images.

[0083] Meanwhile, during training, the model calculates gradients based on the loss function, uses the AdamW optimizer to optimize the model parameters, and dynamically adjusts the learning rate based on the performance on the validation set to achieve backpropagation and gradient updates, thereby avoiding overfitting or underfitting during training.

[0084] Furthermore, after each round of training, the model is evaluated using a validation set to examine its performance on unknown data and adjust relevant hyperparameters such as the learning rate, batch size, and weight decay. Simultaneously, after each evaluation, the current model's weights are saved, especially the weights that performed best on the validation set, for use after model training is complete. Finally, the trained model is saved as a file for convenient inference and testing in practical applications. For example, in this embodiment, the trained model achieves a recognition accuracy of 98%.

[0085] In step S106, the trained SpareNet-Pro model is used to identify the spare part image to determine the category and confidence level of the spare part in the image. Exemplarily, in this embodiment, the identification result may further include at least one of the spare part's serial number, URL, and image.

[0086] It should be noted that in actual use, there may be situations where multiple spare parts on the device are very close together, which could result in multiple spare parts appearing in the image to be identified. Therefore, in this embodiment, after capturing the image of the spare part to be identified, it is first cropped to remove the portion that needs to be identified before uploading.

[0087] In summary, this application constructs a SpareNet-Pro model based on a visual AI model and combines it with various data augmentation techniques to achieve efficient spare parts identification in complex environments. By manually annotating and augmenting the collected spare parts images, diverse augmented sample sets are generated, simulating the actual working conditions of spare parts in different installation environments. After training the model using the augmented sample sets, the SpareNet-Pro model can accurately identify the category and confidence level of the spare parts to be identified. This application effectively improves the accuracy and robustness of spare parts identification, especially under partial occlusion, size variations, and complex backgrounds, ensuring the efficient application of the system in practical scenarios such as industrial equipment maintenance and inventory management.

[0088] Those skilled in the art will understand that the embodiments of this application are based on the YOLOv8 model. The above is a further detailed description of this application in conjunction with specific embodiments. However, it should not be assumed that the specific implementation of this application is limited to the YOLOv8 model. Other versions of the YOLO model, such as YOLOv7, YOLOv9 and YOLOv10, do not depart from the spirit and scope of this application.

[0089] It should be noted that, although the embodiments in this application are based on... Figure 1 Steps S101 to S106 are described sequentially, but this does not mean that steps S101 to S106 must be performed in a strict order. The reason this embodiment follows this order is... Figure 1 The order in which steps S101 to S106 are described is provided to facilitate understanding of the technical solutions of the embodiments of this application by those skilled in the art. In other words, in the embodiments of this application, the order of steps S101 to S106 can be appropriately adjusted according to actual needs.

[0090] Based on the same concept, another embodiment of this application also provides a spare parts identification system based on a visual AI model for implementing the above identification method. The identification system includes:

[0091] The front-end system, developed based on the Flask framework, is mainly used for user interface interaction, including the front-end UI interface, which is used to receive user input and display the recognition results;

[0092] The backend system, developed based on the Flask framework, is mainly used for data processing and model invocation, and communicates with the frontend system through a web interface.

[0093] The database uses MySQL to store the data required by the spare parts identification system. It communicates with the backend system through SQL statements and database connection library. The database mainly includes spare parts information table, identification result table, user information table and backend management table. The spare parts information table is used to store at least one of the spare parts' category, number, URL and image. The identification result table is used to record the identification results of each uploaded image during the identification process. The user information table is used to store user login information and usage records. The backend management table is used by the administrator to manage user permissions.

[0094] Mobile applications, developed based on the Android platform, are primarily used for user interface interaction, including the mobile UI, which receives user input and displays the recognition results. At the same time, mobile applications communicate with the backend system through mobile API interfaces.

[0095] In actual operation, users can upload images of spare parts to be identified to the backend system through the frontend UI and mobile application, respectively. The backend system receives the images of the spare parts to be identified and calls the trained SpareNet-Pro model to perform image recognition. The recognition results are uploaded to the database, the frontend system, and the mobile application. At this time, users can query the recognition results through the UI, including the confidence level, and may also include at least one of the spare part's category, number, URL, and image.

[0096] Furthermore, the front-end system preloads images by embedding spare parts images into the front-end UI, thereby speeding up web page loading, reducing query waiting time, and improving user experience.

[0097] Image preloading is a common performance optimization method. Its purpose is to ensure that images or other resources are loaded and displayed quickly when a user accesses a page or application, reducing page loading latency. For example, in this embodiment, based on the Flask framework, spare part image files are stored in the / static directory, and the front-end system can then directly access these static files to achieve image preloading.

[0098] In some other implementations, the following methods can be used: <link rel="preload">The `<image>` tag allows browsers to begin loading image resources early in the page load cycle. Image preloading can also be controlled using JavaScript code. By creating an `Image` object, images can be preloaded into the browser's memory, ensuring they are displayed quickly when the user needs them. Furthermore, image preloading can be implemented using front-end frameworks or libraries like React and Vue, as well as UI frameworks like Bootstrap. This approach allows for automatic or manual control over image loading timing. For example, using React's `React.lazy()` and the `Suspense` component allows images to be loaded only when needed, or state management can be used to control image loading.

[0099] In summary, the spare parts identification system in this application provides an efficient and reliable spare parts identification service through the collaborative work of a front-end system, a back-end system, a database, and a mobile application. The front-end system is responsible for user interaction, the back-end system handles image processing and model invocation, the database stores and manages relevant data, and the mobile application provides a convenient user interface. The entire system ensures rapid response in spare parts image uploading, identification, and result feedback through efficient data communication and coordination, making it suitable for practical applications such as industrial equipment maintenance and inventory management.

[0100] While the present invention has been illustrated and described with reference to certain preferred embodiments, those skilled in the art should understand that the above description is a further detailed explanation of the invention in conjunction with specific embodiments, and should not be construed as limiting the specific implementation of the invention to these descriptions. Various changes in form and detail can be made by those skilled in the art, including several simple deductions or substitutions, without departing from the spirit and scope of the invention.

Claims

1. A spare parts identification method based on a visual AI model, characterized in that, Includes the following steps: Acquire images of spare parts; The spare parts images are annotated using a manual annotation method to generate an annotated dataset; A SpareNet-Pro model for spare parts identification is constructed. The SpareNet-Pro model is an improvement on the YOLO model, which retains the backbone network and classification head of the YOLO model and adds a multi-scale feature fusion network. The labeled dataset is augmented to generate an augmented sample set, simulating the actual working conditions of spare parts when they are installed on the equipment; The SpareNet-Pro model is trained using the enhanced sample set; The trained SpareNet-Pro model is used to identify the image of the spare part to be identified, and the category and confidence level of the spare part are obtained.

2. The spare parts identification method as described in claim 1, characterized in that, The construction of the SpareNet-Pro model for spare parts identification includes: The backbone network serves as the feature extraction network for the SpareNet-Pro model, used to extract low-level, mid-level, and deep features of the image input to the SpareNet-Pro model. The multi-scale feature fusion network performs multi-scale feature fusion of low-level, mid-level and deep-level features through upsampling and feature concatenation. The classification head is used for image classification tasks. It outputs a spare part category prediction by connecting a fully connected layer after global average pooling.

3. The spare parts identification method as described in claim 2, characterized in that, The backbone network uses the C2f module to extract low-level, mid-level, and deep features from the images input to the SpareNet-Pro model.

4. The spare parts identification method as described in claim 3, characterized in that, The YOLO model mentioned is the YOLOv8 model.

5. The spare parts identification method as described in claim 1, characterized in that, The augmentation process involves performing standard data augmentation on each data sample in the labeled dataset during each training round, and determining, based on a random value, to use at most two of the following: advanced data augmentation and occlusion, to generate the augmented sample set.

6. The spare parts identification method as described in claim 5, characterized in that, The advanced data augmentation is Mosaic augmentation, CutMix augmentation, or Copy-Paste augmentation.

7. The spare parts identification method as described in claim 6, characterized in that, The step of training the SpareNet-Pro model using the enhanced sample set includes: The augmented sample set is divided into a training set, a validation set, and a test set; Model training was performed using the AdamW optimizer. Save the optimal model weights.

8. The spare parts identification method as described in claim 7, characterized in that, The model training also includes weighting the loss of the augmented sample set using a loss function. If the occlusion is detected in the data sample during the augmentation process, the loss of the data sample is multiplied by a loss amplification factor. The loss amplification factor is used to increase the loss weight of the data sample to improve the recognition ability of the SpareNet-Pro model when encountering partially occluded objects.

9. The spare parts identification method as described in claim 8, characterized in that, The initial value of the loss amplification factor is 1.5, and the loss amplification factor is adjusted during the model training process based on the error of the enhanced samples with the occlusion on the validation set.

10. A spare parts identification system based on a visual AI model, characterized in that, For implementing the identification method according to any one of claims 1 to 9, the spare parts identification system comprises: A front-end system, developed based on the Flask framework, includes a front-end UI interface, which is used for user interface interaction. The backend system, developed based on the Flask framework, is used for data processing and model invocation. The database uses MySQL to store the data required by the spare parts identification system. A mobile application, developed based on the Android platform, includes a mobile UI interface and is used to upload an image of a spare part to be identified and receive the category and confidence level of the spare part to be identified. The backend system communicates with the frontend system via a web interface, communicates with the mobile application via a mobile API interface, and communicates with the database via SQL statements and a database connection library.

11. The spare parts identification system as described in claim 10, characterized in that, The database includes a spare parts information table, and may also include at least one of an identification result table, a user information table, and a backend management table. The spare parts information table is used to store at least one of the spare parts' category, number, URL, and image. The identification result table is used to record the identification results of each uploaded image during the identification process. The user information table is used to store the user's login information and usage records. The backend management table is used by the administrator to manage the user's permissions.

12. The spare parts identification system as described in claim 11, characterized in that, The front-end system and the mobile application upload the image of the spare part to be identified to the back-end system through the front-end UI and the mobile UI respectively. The back-end system receives the image of the spare part to be identified and calls the trained SpareNet-Pro model to perform image recognition. The recognition result is uploaded to the database, the front-end system and the mobile application. The recognition result includes confidence level, and also includes at least one of the spare part's category, number, URL and image.

13. The spare parts identification system as described in claim 12, characterized in that, The front-end system achieves image preloading by embedding spare parts images into the front-end UI interface.