Cabinet equipment segmentation method and device, electronic equipment and storage medium

The modified DeepLabv3 model with MobileNetV2 and CA mechanism addresses low accuracy in machine cabinet device segmentation, enhancing segmentation efficiency and accuracy for precise device identification.

CN120318505APending Publication Date: 2025-07-15INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510219742.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the accuracy of segmentation detection of internal equipment of the cabinet based on the CNN model is not high, and it is difficult to achieve effective segmentation and identification.

Method used

The improved DeepLabv3 model is adopted, the MobileNetV2 network is used as the backbone network, and the channel attention mechanism (CA module) is added to the model, combined with the dense hollow space pyramid pooling (DenseASPP) module, feature extraction and information fusion are optimized, and segmentation accuracy is improved.

Benefits of technology

In an environment with limited resources, the efficiency and accuracy of equipment segmentation within the cabinet are improved, and the precise identification and management of equipment in complex scenarios is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318505A_ABST
    Figure CN120318505A_ABST
Patent Text Reader

Abstract

The invention provides a cabinet equipment segmentation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring an internal image of a to-be-segmented cabinet; the internal image of the cabinet to be segmented is input into the semantic segmentation model, an equipment recognition result output by the semantic segmentation model is obtained, and the construction process of the semantic segmentation model comprises the steps that a backbone network of a DeepLabv3 model is replaced with a MobileNetV2 network, and the improved DeepLabv3 model is obtained; and adding a channel attention mechanism (CA) module in the improved DeepLabv3 model to obtain a semantic segmentation model. The semantic segmentation model is obtained by improving the DeepLabv3 model, the computing resource demand is reduced through a lightweight structure and efficient data processing capacity, efficient data processing is kept at the same time, and higher efficiency and accuracy are achieved on the semantic segmentation task of the equipment in the cabinet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a method, device, electronic device and storage medium for segmenting cabinet equipment. Background Art

[0002] With the rapid expansion of data centers and communication infrastructures, the types and quantities of equipment inside cabinets are also continuously increasing. The equipment inside cabinets includes servers, switches, routers, as well as various connection devices and cables.

[0003] Existing methods for segmenting and recognizing equipment inside cabinets generally directly rely on a Convolutional Neural Network (CNN) model. They obtain image spatial features through a feature extraction network, use dilated convolution or pyramid pooling for context information fusion, and adopt an encoder-decoder structure to restore image details, thereby realizing the process of segmenting and recognizing equipment. Due to the increasing number of equipment inside cabinets, various equipment often densely arranged together, resulting in low accuracy of directly segmenting and detecting based on the CNN model.

[0004] How to improve the accuracy of detection and achieve effective segmentation and recognition of equipment inside cabinets is an important issue that the industry urgently needs to solve currently. Summary of the Invention

[0005] The present invention provides a method, device, electronic device and storage medium for segmenting cabinet equipment, aiming to solve the defect of low accuracy of directly segmenting and detecting equipment inside cabinets based on the CNN model in the prior art, improve the accuracy of detection, and achieve effective segmentation and recognition of equipment inside cabinets.

[0006] The present invention provides a method for segmenting cabinet equipment, including the following steps: Obtain an image inside the cabinet to be segmented; Input the image inside the cabinet to be segmented into a semantic segmentation model, and obtain the equipment recognition result output by the semantic segmentation model. The semantic segmentation model is trained based on cabinet interior image samples and the equipment category labels corresponding to the cabinet interior image samples; The construction process of the semantic segmentation model includes: Use the MobileNetV2 network to replace the backbone network of the DeepLabv3 model to obtain an improved DeepLabv3 model; Add a Channel Attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0007] A segmentation method for cabinet equipment provided by the present invention. The improved DeepLabv3 model includes a backbone network, an encoder, and a decoder. The CA module is added to the improved DeepLabv3 model to obtain the semantic segmentation model, including: The CA module is added between the backbone network and the spatial pyramid pooling ASPP module in the encoder, and the CA module is added between the backbone network and the decoder to obtain the semantic segmentation model.

[0008] A segmentation method for cabinet equipment provided by the present invention. The internal image of the cabinet to be segmented is input into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model, including: Based on the MobileNetV2 network in the semantic segmentation model, the internal image of the cabinet to be segmented is input for feature extraction to obtain the deep feature map of the internal image of the cabinet to be segmented and the shallow feature map of the internal image of the cabinet to be segmented; Based on the CA module between the backbone network and the ASPP module, information interaction and weight allocation are performed between channels of the deep feature map to obtain a first enhanced feature, and atrous spatial pyramid pooling is performed on the first enhanced feature based on the encoder to obtain high-level semantic features; Based on the CA module between the backbone network and the decoder, information interaction and weight allocation are performed between channels of the shallow feature map to obtain a second enhanced feature; The second enhanced feature and the high-level semantic features are fused, and the fused features are decoded based on the decoder to obtain the device recognition result.

[0009] A segmentation method for cabinet equipment provided by the present invention. Before the internal image of the cabinet to be segmented is input into the semantic segmentation model, it further includes: The spatial pyramid pooling ASPP module in the semantic segmentation model is replaced with a dense atrous spatial pyramid pooling DenseASPP module.

[0010] A segmentation method for cabinet equipment provided by the present invention. The DenseASPP module is used to adjust the atrous convolution parameters of the DeepLabv3 model based on the image feature complexity of the internal image of the cabinet to be segmented.

[0011] A segmentation method for cabinet equipment provided by the present invention. Before replacing the backbone network of the DeepLabv3 model with the MobileNetV2 network, it further includes: Replace the standard convolution in the MobileNetV2 network with dilated convolution to expand the receptive field of the convolution kernel in the MobileNetV2 network.

[0012] According to a segmentation method of a cabinet device provided by the present invention, the training method of the semantic segmentation model includes: Divide the validation set and the test set from the cabinet internal image sample library; Based on the cabinet internal image samples in the validation set and their corresponding device category labels, train the initial semantic segmentation model to obtain a pre-trained semantic segmentation model; Based on the test set, test the pre-trained semantic segmentation model, and determine that the recognition rate of the pre-trained semantic segmentation model is greater than a preset recognition rate threshold to obtain the semantic segmentation model.

[0013] The present invention also provides a segmentation device for cabinet devices, including the following modules: An image acquisition module, configured to acquire an internal image of the cabinet to be segmented; An image segmentation module, configured to input the internal image of the cabinet to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model, and the semantic segmentation model is trained based on the cabinet internal image samples and the device category labels corresponding to the cabinet internal image samples; The construction process of the semantic segmentation model includes: Use the MobileNetV2 network to replace the backbone network of the DeepLabv3 model to obtain an improved DeepLabv3 model; Add a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the segmentation method of the cabinet device as described in any one of the above.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the segmentation method of the cabinet device as described in any one of the above.

[0016] The segmentation method, device, electronic device, and storage medium for cabinet equipment provided by the present invention obtain an improved DeepLabv3 model by using the MobileNetV2 network as the backbone network of the DeepLabv3 model on the basis of the DeepLabv3 model. As the backbone network, MobileNetV2 reduces the computational resource requirements with a lightweight structure and efficient data processing capabilities, while maintaining efficient data processing, and has higher efficiency and accuracy in the semantic segmentation task of internal cabinet equipment. At the same time, a CA module is added to the improved DeepLabv3 model to obtain a semantic segmentation model. The integrated CA attention mechanism processes the information between channels and incorporates directional position information, enhances the feature representation ability of the network through one-dimensional feature encoding operations, improves the ability to capture key information in complex scenarios, and realizes the effective segmentation and recognition of internal cabinet equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic flowchart of the segmentation method for cabinet equipment provided by the present invention.

[0019] Figure 2 is a schematic diagram of the semantic segmentation model provided by the present invention.

[0020] Figure 3 is a schematic diagram of the structure of the channel attention mechanism module provided by the present invention.

[0021] Figure 4 is a schematic diagram of the structure of the segmentation device for cabinet equipment provided by the present invention.

[0022] Figure 5 is a schematic diagram of the structure of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention fall within the protection scope of the present invention.

[0024] Figure 1It is a schematic flow chart of the method for partitioning cabinet equipment provided by the present invention. As Figure 1 shown, the method includes the following: Step 110: Obtain the internal image of the cabinet to be partitioned; Step 120: Input the internal image of the cabinet to be partitioned into the semantic segmentation model, and obtain the device recognition result output by the semantic segmentation model. The semantic segmentation model is trained based on the internal image samples of the cabinet and the device category labels corresponding to the internal image samples of the cabinet; The construction process of the semantic segmentation model includes: Replace the backbone network of the DeepLabv3 model with the MobileNetV2 network to obtain an improved DeepLabv3 model; Add a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0025] The execution subject of the method for partitioning cabinet equipment provided by the present invention can be an electronic device, a component in the electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a Network Attached Storage (NAS), or a personal computer (PC), etc. The present invention does not make specific limitations.

[0026] Next, taking a computer executing the method for partitioning cabinet equipment provided by the present invention as an example, the technical solution of the present invention will be described in detail.

[0027] In step 110, the internal image of the cabinet to be partitioned is obtained.

[0028] Specifically, the internal image of the cabinet to be partitioned refers to the image of the cabinet taken after the cabinet door is opened. The image contains various types of equipment installed in the cabinet, including servers, switches, disk arrays, routers, and various connection devices and cables.

[0029] Since there are many devices and complex connections in the cabinet, accurately identifying and partitioning the cabinet equipment helps to effectively manage the assets of the devices, clearly master information such as the model, quantity, and location of the devices, facilitate asset registration and inventory, avoid asset omission or duplicate purchase, and improve the asset utilization rate and the refined level of management.

[0030] In addition, when a fault occurs, by analyzing the distribution of devices obtained after segmentation and recognition, it is possible to quickly determine which specific device and which component have problems. Maintenance personnel can directly locate the fault point, improving the maintenance efficiency, reducing the downtime, and minimizing the impact on the business.

[0031] In step 120, the internal image of the cabinet to be segmented is input into the semantic segmentation model, and the device recognition result output by the semantic segmentation model is obtained.

[0032] A semantic segmentation model for segmenting internal devices of the cabinet is pre-constructed. Specifically, the semantic segmentation model is optimized and constructed based on the DeepLabv3 model.

[0033] The DeepLabv3 model aims to solve the semantic segmentation task, that is, to assign each pixel in the image to a specific category, such as precisely dividing different objects in the image. Specifically, the DeepLabv3 model realizes the precise segmentation of objects with different scales in the image by introducing techniques such as atrous convolution, multi-scale output, spatial pyramid pooling, and decoder modules.

[0034] The optimization process of the original DeepLabv3 model includes: Taking the MobileNetV2 network as the backbone network of the DeepLabv3 model to obtain an improved DeepLabv3 model, and adding a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0035] The MobileNetV2 network is a lightweight convolutional neural network model. The MobileNetV2 network inherits from the MobileNetV1 network, combines the depthwise separable convolution technology, and introduces the inverted residual and linear bottleneck structures, breaking through the performance bottleneck of the previous version. This improvement not only reduces the model size but also improves the processing speed.

[0036] The inverted residual structure and depthwise separable convolution of MobileNetV2 help to extract multi-scale features, which is particularly important for segmenting targets with different sizes and shapes such as internal devices of the cabinet. Moreover, the linear bottleneck and skip connections of MobileNetV2 help to retain more detailed information, which is crucial for boundary optimization in the semantic segmentation of internal devices of the cabinet.

[0037] Taking the MobileNetV2 network as the backbone network of the DeepLabv3 model, the improved DeepLabv3 model has higher efficiency and accuracy in the semantic segmentation task of the internal equipment of the cabinet compared to the original DeepLabv3 model. This is mainly due to the efficiency, lightweight characteristics, and multi-scale feature extraction ability of the MobileNetV2 network. The improved model can process the segmentation task more quickly while maintaining a high segmentation accuracy, which is of great significance for application scenarios with high real-time requirements and limited resources.

[0038] Optionally, to improve the processing accuracy, the 7th and 8th layers of the MobileNetV2 network are optimized. Dilated Convolution is introduced to expand the receptive field of the convolutional kernel without increasing the number of parameters or computational burden. Dilated Convolution is achieved by inserting "holes" between standard convolutional kernels, enabling the convolutional kernel to cover a larger input area, thereby improving the recognition accuracy of device interfaces and micro-component details.

[0039] Furthermore, by adjusting the stride of the 7th layer of the MobileNetV2 network from the standard 2 to 1, the size of the feature map is ensured to be maintained, enabling the network to capture more detailed information, which is particularly important for the recognition of complex device interfaces and tiny components. The 7th layer configuration uses 48 3x3 dilated convolutional kernels with a dilation rate of 2 and a stride of 1. Such a configuration can carefully process the input feature map and enhance the recognition ability of small-size details.

[0040] The 8th layer maintains the same dilation rate and improves the network's learning ability and feature extraction ability by increasing the number of convolutional kernels to 64, thereby maintaining sensitivity to details at a deeper level. These adjustments ensure a significant improvement in the segmentation accuracy of the model without increasing the computational burden, achieving a more detailed and accurate segmentation operation for the equipment in the communication cabinet. Through these innovative adjustments, the model provides an efficient and practical solution for communication device management in resource-constrained environments.

[0041] After obtaining the improved DeepLabv3 model, a channel attention mechanism CA module is added to the improved DeepLabv3 model to obtain a semantic segmentation model.

[0042] The Channel Attention (CA) module enhances the feature map by learning the importance of each channel. This helps the model pay more attention to the features useful for classification or segmentation tasks, thereby improving the model's performance. In the DeepLabv3 model, adding the CA module can further improve its accuracy in tasks such as semantic segmentation. At the same time, the CA module recalibrates the feature map by assigning different weights to different channels. This recalibration helps enhance the model's ability to express useful features while suppressing unimportant features. This is of great significance for improving the model's generalization ability and robustness.

[0043] Specifically, in the DeepLabv3 model, the CA module can be added after the backbone network. The backbone network extracts features from the input cabinet interior image to be segmented, and after obtaining the feature map, the CA module enhances the feature map by learning the importance of each channel, which can improve the subsequent segmentation and recognition performance.

[0044] After obtaining the semantic segmentation model, based on the pre-determined cabinet interior image samples and the device category labels corresponding to the cabinet interior image samples, model training is carried out to obtain the trained semantic segmentation model.

[0045] Input the cabinet interior image to be segmented into the trained semantic segmentation model to identify the devices inside the cabinet, and obtain the device recognition result output by the semantic segmentation model.

[0046] The segmentation method of cabinet devices provided by the present invention, based on the DeepLabv3 model, uses the MobileNetV2 network as the backbone network of the DeepLabv3 model to obtain an improved DeepLabv3 model. As the backbone network, MobileNetV2 reduces the computational resource requirements with a lightweight structure and efficient data processing ability, while maintaining efficient data processing, and has higher efficiency and accuracy in the semantic segmentation task of cabinet interior devices. At the same time, the CA module is added to the improved DeepLabv3 model to obtain the semantic segmentation model. The integrated CA attention mechanism processes the information between channels and incorporates the directional position information. By performing one-to-one one-dimensional feature encoding operations, it enhances the network's feature representation ability, improves the ability to capture key information in complex scenarios, and realizes the effective segmentation and recognition of cabinet interior devices.

[0047] In one embodiment, the improved DeepLabv3 model includes a backbone network, an encoder, and a decoder. The CA module is added to the improved DeepLabv3 model to obtain the semantic segmentation model, including: adding the CA module between the backbone network and the Atrous Spatial Pyramid Pooling (ASPP) module in the encoder, and adding the CA module between the backbone network and the decoder to obtain the semantic segmentation model.

[0048] The structural schematic diagram of the semantic segmentation model with the CA module added is as Figure 2 shown in the semantic segmentation model schematic diagram provided by the present invention. Specifically, the CA module is added between the backbone network and the Atrous Spatial Pyramid Pooling (ASPP) module, and between the backbone network and the decoder.

[0049] In the semantic segmentation model in related methods, the decoder does not make full use of the shallow feature maps, which limits the ability to capture local details. For this reason, an improved CA channel attention mechanism is introduced, which comprehensively considers the information between channels and the position information related to directions, significantly improving the prediction accuracy and the ability to capture details of the model.

[0050] Among them, the structural schematic diagram of the CA channel attention mechanism module can be as Figure 3 shown in the structural schematic diagram of the channel attention mechanism module provided by the present invention. Among them, C is the number of channels, which represents different feature channels in the feature map, and these channels contain different feature information of the image, such as color, texture, shape, etc.; H is the height, representing the vertical dimension of the feature map, that is, the number of pixels in the vertical direction of the feature map; W is the width, representing the horizontal dimension of the feature map, that is, the number of pixels in the horizontal direction of the feature map.

[0051] The feature encoding process is optimized. By using pooling kernels with sizes of (H, 1) and (1, W), the input feature map can be encoded in the horizontal X and vertical Y directions. This method can retain key spatial information better than two-dimensional global pooling, especially when dealing with the complex equipment layout inside the cabinet. Next, the average pooling results in the two directions are concatenated, and feature transformation is performed through 1×1 convolution, normalization, and ReLU non-linear processing to generate an intermediate feature tensor that synthesizes spatial information, enhancing the recognition ability of the attention mechanism.

[0052] In addition, to balance the complexity and performance of the model, a reduction factor r can be used, which is mainly used to reduce the number of channels in the intermediate layer of the model to alleviate the computational burden and reduce the complexity of the model. This not only speeds up the operation speed but also maintains the performance required for processing semantic segmentation tasks. In the application, it is necessary to find a balance between the lightweight of the model and the segmentation accuracy. Specifically, r = 12 can be used as a compromise choice, and then different r values can be set through experiments at the initial stage of model training to evaluate their impact on performance and computational efficiency, and optimize and adjust through verification loss and accuracy metrics to achieve the best effect.

[0053] In one embodiment, the internal image of the cabinet to be segmented is input into the semantic segmentation model, and the device recognition result output by the semantic segmentation model is obtained, including: based on the MobileNetV2 network in the semantic segmentation model, feature extraction is performed on the input internal image of the cabinet to be segmented to obtain the deep feature map and the shallow feature map of the internal image of the cabinet to be segmented; based on the CA module between the backbone network and the ASPP module, information interaction and weight assignment are performed between channels of the deep feature map to obtain a first enhanced feature, and atrous spatial pyramid pooling is performed on the first enhanced feature based on the encoder to obtain high-level semantic features; based on the CA module between the backbone network and the decoder, information interaction and weight assignment are performed between channels of the shallow feature map to obtain a second enhanced feature; the second enhanced feature is fused with the high-level semantic features, and the fused features are decoded based on the decoder to obtain the device recognition result.

[0054] In the semantic segmentation model, MobileNetV2 is selected as the backbone network. The input internal image of the cabinet to be segmented is first processed by the MobileNetV2 network. This network extracts the deep features and shallow features of the image through techniques such as depthwise separable convolution and inverted residual structure. Among them, the deep features contain the high-level semantic information of the image, which is crucial for understanding the overall content and structure of the image. The shallow features retain more detailed information of the image, such as edges and textures, which are very helpful for accurately segmenting the target objects in the image.

[0055] Between the backbone network and the ASPP module, and between the backbone network and the decoder, a CA module is introduced respectively. By learning the importance of each channel, the CA module conducts inter-channel information interaction and weight assignment for deep features and shallow features. This helps the model to focus more on the features useful for the segmentation task while suppressing unimportant features. After being processed by the CA module, the deep features obtain the first enhanced feature, which is optimized in the channel dimension and focuses more on the information useful for the segmentation task. Similarly, after being processed by the CA module, the shallow features obtain the second enhanced feature, which enhances the channel features related to the segmentation task while retaining the detailed information.

[0056] The first enhanced feature undergoes atrous spatial pyramid pooling processing through the ASPP module to capture multi-scale context information and obtain high-level semantic features. The use of the SPP module helps the model to better understand the complex structure and context information in the image, thereby improving the accuracy of segmentation. The second enhanced feature is fused with the high-level semantic features. This step combines the detailed information and high-level semantic information of the image, making the segmentation result more accurate and complete. The fused features are decoded through the decoder. The decoder restores the feature map to the size of the original image through upsampling and convolution operations. Finally, the result output by the decoder is the device recognition result, which shows the accurate segmentation and recognition of the devices inside the cabinet.

[0057] In one embodiment, before inputting the image inside the cabinet to be segmented into the semantic segmentation model, it further includes: replacing the atrous spatial pyramid pooling ASPP module in the semantic segmentation model with a dense atrous spatial pyramid pooling DenseASPP module.

[0058] In the DeepLabV3+ model, although the ASPP module expands the receptive field through atrous convolution to capture multi-scale features, there are problems of information loss and low efficiency when dealing with the complex scenarios of communication machine cabinets. To overcome these limitations, the DenseASPP structure is introduced, integrating a dynamic dilation rate adjustment mechanism to adapt to different scenario requirements.

[0059] By cascading and densely connecting atrous convolution layers with different dilation rates, DenseASPP can generate features covering a very large range and cover these scale ranges in a very dense manner, having a larger receptive field than the ASPP module.

[0060] Moreover, the atrous convolution layers in DenseASPP are organized in a cascading manner, the dilation rate of each layer increases layer by layer, and the output of each layer will be connected to the input feature map and all the outputs of the lower layers, having richer feature sampling than the ASPP module.

[0061] In one embodiment, the DenseASPP module is used to adjust the dilation convolution parameters of the DeepLabv3 model based on the image feature complexity of the internal image of the cabinet to be segmented.

[0062] Specifically, the DenseASPP module first quickly scans the image through content-aware adjustment technology, identifies the feature density and distribution of key regions, and determines the corresponding dilation rate. Then, based on the complexity of the image features, it optimizes the parameters of the dilation convolution in real time, making the dilation rate and the convolution kernel size change flexibly to improve the ability to capture details and avoid over-computation of simple regions. This mechanism reduces the consumption of computing resources in simple regions while improving the processing accuracy of complex regions, enhancing the adaptability and efficiency of the model, and is particularly suitable for communication machine room cabinet scenarios with high resolution and large complexity differences.

[0063] In one embodiment, before replacing the backbone network of the DeepLabv3 model with the MobileNetV2 network, it further includes: replacing the standard convolution in the MobileNetV2 network with dilation convolution to expand the receptive field of the convolution kernel in the MobileNetV2 network.

[0064] Specifically, in the specific adjustment, the 7th and 8th layers of MobileNetV2 can be optimized. Replacing the standard convolution in the 7th and 8th layers with dilation convolution expands the receptive field of the convolution kernel without increasing the number of parameters or the computational burden. Dilation convolution is achieved by inserting "holes" between standard convolution kernels, enabling the convolution kernel to cover a larger input area, thereby improving the recognition accuracy of device interfaces and micro-component details.

[0065] In one embodiment, a method for training a semantic segmentation model includes: dividing a validation set and a test set from a database of internal images of cabinets; training an initial semantic segmentation model based on the internal image samples of cabinets in the validation set and their corresponding device category labels to obtain a pre-trained semantic segmentation model; testing the pre-trained semantic segmentation model based on the test set, determining that the recognition rate of the pre-trained semantic segmentation model is greater than a preset recognition rate threshold, and obtaining the semantic segmentation model.

[0066] First, prepare a database containing a large number of internal image samples of cabinets. These images should cover various device categories, perspective changes, etc. to ensure the generalization ability of the model.

[0067] Randomly divide the validation set and the test set from the cabinet internal image sample library. The validation set is used to adjust and optimize the model parameters during the training process, while the test set is used to evaluate the final performance of the model after the training is completed. Usually, the ratio of the validation set and the test set can be set according to specific requirements. For example, 70% of the samples can be used for training, and a part of the remaining 30% can be used as the validation set (such as 15%), and the remaining part as the test set (such as 15%). However, it should be noted that this ratio is not fixed and can be adjusted according to the size of the dataset and the actual situation.

[0068] Use the image samples in the validation set and their corresponding device category labels to train the initial semantic segmentation model. During the training process, adjust the model parameters through the backpropagation algorithm to enable the model to accurately classify the input image at the pixel level.

[0069] After the training is completed, use the test set to test the pre-trained semantic segmentation model. Compare the recognition rate on the test set with the preset recognition rate threshold. If the recognition rate is greater than the threshold, it is considered that the model has met the application requirements, and the final semantic segmentation model can be obtained.

[0070] Next, the segmentation device for cabinet equipment provided by the present invention will be described. The segmentation device for cabinet equipment described below can be correspondingly referred to the cabinet equipment segmentation method described above.

[0071] As Figure 4 shown, the device includes: An image acquisition module 410 for acquiring an internal image of the cabinet to be segmented; An image segmentation module 420 for inputting the internal image of the cabinet to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model. The semantic segmentation model is trained based on the cabinet internal image samples and the device category labels corresponding to the cabinet internal image samples; The construction process of the semantic segmentation model includes: Replace the backbone network of the DeepLabv3 model with the MobileNetV2 network to obtain an improved DeepLabv3 model; Add a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0072] The segmentation device for cabinet equipment provided by the present invention, based on the DeepLabv3 model, uses the MobileNetV2 network as the backbone network of the DeepLabv3 model to obtain an improved DeepLabv3 model. As the backbone network, MobileNetV2 reduces the computational resource requirements with a lightweight structure and efficient data processing capabilities, while maintaining efficient data processing, and has higher efficiency and accuracy in the semantic segmentation task of the internal equipment of the cabinet. At the same time, a CA module is added to the improved DeepLabv3 model to obtain a semantic segmentation model. The integrated CA attention mechanism processes the information between channels and incorporates directional position information, enhances the feature representation ability of the network through one-dimensional feature encoding operations, improves the ability to capture key information in complex scenarios, and realizes the effective segmentation and recognition of the internal equipment of the cabinet.

[0073] In one embodiment, the image segmentation module 420 is specifically configured to: The improved DeepLabv3 model includes a backbone network, an encoder, and a decoder. Adding the CA module to the improved DeepLabv3 model to obtain the semantic segmentation model includes: Adding the CA module between the backbone network and the spatial pyramid pooling ASPP module in the encoder, and adding the CA module between the backbone network and the decoder to obtain the semantic segmentation model.

[0074] In one embodiment, the image segmentation module 420 is further specifically configured to: Inputting the image of the cabinet interior to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model, including: Based on the MobileNetV2 network in the semantic segmentation model, extracting features from the input image of the cabinet interior to be segmented to obtain a deep feature map and a shallow feature map of the image of the cabinet interior to be segmented; Based on the CA module between the backbone network and the ASPP module, performing information interaction and weight allocation between channels on the deep feature map to obtain a first enhanced feature, and performing atrous spatial pyramid pooling on the first enhanced feature based on the encoder to obtain high-level semantic features; Based on the CA module between the backbone network and the decoder, performing information interaction and weight allocation between channels on the shallow feature map to obtain a second enhanced feature; Fusing the second enhanced feature with the high-level semantic features, and decoding the fused features based on the decoder to obtain the device recognition result.

[0075] In one embodiment, the image segmentation module 420 is further specifically configured to: Before inputting the internal image of the cabinet to be segmented into the semantic segmentation model, it further includes: Replacing the Atrous Spatial Pyramid Pooling (ASPP) module in the semantic segmentation model with a Dense Atrous Spatial Pyramid Pooling (DenseASPP) module.

[0076] In one embodiment, the image segmentation module 420 is further specifically configured to: The DenseASPP module is used to adjust the atrous convolution parameters of the DeepLabv3 model based on the image feature complexity of the internal image of the cabinet to be segmented.

[0077] In one embodiment, the image segmentation module 420 is further specifically configured to: Before replacing the backbone network of the DeepLabv3 model with the MobileNetV2 network, it further includes: Replacing the standard convolution in the MobileNetV2 network with atrous convolution to expand the receptive field of the convolutional kernel in the MobileNetV2 network.

[0078] In one embodiment, the image segmentation module 420 is further specifically configured to: The training method of the semantic segmentation model includes: Dividing a validation set and a test set from the internal image sample library of the cabinet; Training an initial semantic segmentation model based on the internal image samples in the validation set and their corresponding device category labels to obtain a pre-trained semantic segmentation model; Testing the pre-trained semantic segmentation model based on the test set, and determining that the recognition rate of the pre-trained semantic segmentation model is greater than a preset recognition rate threshold to obtain the semantic segmentation model.

[0079] Figure 5 An example of the physical structure diagram of an electronic device is shown in Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the segmentation method of the cabinet device, and the method includes: obtaining the internal image of the cabinet to be segmented; Input the internal image of the cabinet to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model. The semantic segmentation model is trained based on the internal image samples of the cabinet and the corresponding device category labels of the internal image samples of the cabinet; The construction process of the semantic segmentation model includes: Replace the backbone network of the DeepLabv3 model with the MobileNetV2 network to obtain an improved DeepLabv3 model; Add a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0080] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0081] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the segmentation method of the cabinet device provided by the above-mentioned various methods. The method includes: obtaining an internal image of the cabinet to be segmented; Input the internal image of the cabinet to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model. The semantic segmentation model is trained based on the internal image samples of the cabinet and the corresponding device category labels of the internal image samples of the cabinet; The construction process of the semantic segmentation model includes: Replace the backbone network of the DeepLabv3 model with the MobileNetV2 network to obtain an improved DeepLabv3 model; Add a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0082] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the segmentation method of the cabinet device provided by the above-mentioned various methods. The method includes: obtaining an internal image of the cabinet to be segmented; Inputting the internal image of the cabinet to be segmented into a semantic segmentation model to obtain a device recognition result output by the semantic segmentation model. The semantic segmentation model is trained based on internal cabinet image samples and device category labels corresponding to the internal cabinet image samples; The construction process of the semantic segmentation model includes: Replacing the backbone network of the DeepLabv3 model with the MobileNetV2 network to obtain an improved DeepLabv3 model; Adding a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

[0083] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for partitioning a cabinet device, characterized in that, Including: Obtain the internal image of the cabinet to be segmented; Input the internal image of the cabinet to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model. The semantic segmentation model is trained based on the internal image samples of the cabinet and the device category labels corresponding to the internal image samples of the cabinet; The construction process of the semantic segmentation model includes: Replace the backbone network of the DeepLabv3 model with the MobileNetV2 network to obtain an improved DeepLabv3 model; Add a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

2. The partitioning method of the cabinet device according to claim 1, characterized in that, The improved DeepLabv3 model includes a backbone network, an encoder, and a decoder. Adding the CA module to the improved DeepLabv3 model to obtain the semantic segmentation model includes: Add the CA module between the backbone network and the spatial pyramid pooling ASPP module in the encoder, and add the CA module between the backbone network and the decoder to obtain the semantic segmentation model.

3. The partitioning method of the cabinet device according to claim 2, wherein The step of inputting the internal image of the cabinet to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model includes: Based on the MobileNetV2 network in the semantic segmentation model, extract features from the input internal image of the cabinet to be segmented to obtain the deep feature map and the shallow feature map of the internal image of the cabinet to be segmented; Based on the CA module between the backbone network and the ASPP module, perform information interaction and weight allocation among channels on the deep feature map to obtain a first enhanced feature, and perform dilated spatial pyramid pooling on the first enhanced feature based on the encoder to obtain high-level semantic features; Based on the CA module between the backbone network and the decoder, perform information interaction and weight allocation among channels on the shallow feature map to obtain a second enhanced feature; Fuse the second enhanced feature with the high-level semantic features, and decode the fused features based on the decoder to obtain the device recognition result.

4. The partitioning method of the cabinet device according to claim 1, characterized in that, Before inputting the internal image of the cabinet to be segmented into the semantic segmentation model, it further includes: Replace the spatial pyramid pooling ASPP module in the semantic segmentation model with a dense dilated spatial pyramid pooling DenseASPP module.

5. The partitioning method of the cabinet device according to claim 4, characterized in that, The DenseASPP module is used to adjust the dilated convolution parameters of the DeepLabv3 model based on the image feature complexity of the internal image of the cabinet to be segmented.

6. The partitioning method of the cabinet device according to claim 1, wherein Before replacing the backbone network of the DeepLabv3 model with the MobileNetV2 network, it further includes: Replace the standard convolution in the MobileNetV2 network with dilated convolution to expand the receptive field of the convolution kernel in the MobileNetV2 network.

7. The partitioning method of the cabinet device according to claim 1, characterized in that, The training method of the semantic segmentation model includes: Divide a validation set and a test set from the internal image sample library of the cabinet; Based on the internal cabinet image samples in the validation set and their corresponding device category labels, train the initial semantic segmentation model to obtain the pre-trained semantic segmentation model; Based on the test set, test the pre-trained semantic segmentation model, determine that the recognition rate of the pre-trained semantic segmentation model is greater than the preset recognition rate threshold, and obtain the semantic segmentation model.

8. A splitting device for cabinet equipment, characterized in that, It includes: An image acquisition module for acquiring an internal cabinet image to be segmented; An image segmentation module for inputting the internal cabinet image to be segmented into the semantic segmentation model to obtain the device recognition result output by the semantic segmentation model, and the semantic segmentation model is trained based on the internal cabinet image samples and the device category labels corresponding to the internal cabinet image samples; The construction process of the semantic segmentation model includes: Replace the backbone network of the DeepLabv3 model with the MobileNetV2 network to obtain an improved DeepLabv3 model; Add a channel attention mechanism CA module to the improved DeepLabv3 model to obtain the semantic segmentation model.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the segmentation method of the cabinet device according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the segmentation method of the cabinet device according to any one of claims 1 to 7.