Cabinet U-bit detection method based on two-stage visual identification and perspective correction
Through two-stage visual recognition and perspective correction technology, combined with feature pyramid network and CoordinateAttention module, the problem of low accuracy of cabinet U-position recognition is solved, high-precision and low-cost U-position detection are achieved, and continuous optimization is carried out through active learning mechanisms.
Patent Information
- Application Number
- CN202510480823.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art has problems such as low accuracy, poor environmental adaptability, and high deployment and maintenance costs in cabinet U-position identification and vacancy rate evaluation.
The cabinet U-position detection method based on two-stage visual recognition and perspective correction is adopted to generate standard rectangular images through the first stage irregular cabinet detection and perspective correction. The second stage is to use the feature pyramid network (FPN) and CoordinateAttention (CA) modules to perform U-position state detection, and an active learning mechanism is introduced for model optimization.
It significantly improves the accuracy and robustness of cabinet U-position detection, reduces deployment costs and maintenance difficulties, realizes online learning and continuous optimization capabilities, and adapts to complex perspective environments.
Smart Images

Figure CN120147622A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and in particular to a cabinet U-position detection method based on two-stage visual recognition and perspective correction. Background Art
[0002] With the continuous expansion of the construction scale of data center computer rooms, the utilization efficiency of cabinet U-position resources has become an important factor affecting the operation and maintenance and deployment efficiency of IT assets. Currently, the acquisition of cabinet U-position occupancy information mainly relies on manual inspections or the deployment of sensors such as infrared, pressure, and current to achieve status detection. However, these methods generally have the following deficiencies:
[0003] Dependence on manual labor: Traditional manual inventory methods are inefficient and error-prone, and it is difficult to achieve dynamic and high-frequency updates; High deployment cost: Using physical sensor solutions requires installing devices at each U-position, with cumbersome construction, high costs, and complex maintenance; Poor robustness: Due to actual scenarios such as the possible tilted installation, occlusion, and offset of the shooting angle of the cabinet, existing image recognition methods often cannot stably output standardized images, resulting in insufficient U-position detection accuracy; Weak model generalization: Existing image algorithms mostly adopt single-stage detection or end-to-end solutions, and it is difficult to balance the dual tasks of "overall cabinet recognition" and "fine-grained U-position localization"; Lack of a closed-loop optimization mechanism: Once the traditional vision model is trained, it is fixed, lacking a mechanism for manual participation and self-learning of low-confidence results, and unable to adapt to data drift and complex scenario expansion in long-term operation and maintenance.
[0004] Therefore, how to design a cabinet U-position detection method with high accuracy, low deployment cost, adaptability to complex perspective environments, and the ability of online learning and continuous optimization has become the core problem urgently to be solved in the current technical field. Summary of the Invention
[0005] The present invention aims to overcome the problems in the prior art such as low accuracy, poor environmental adaptability, and high deployment and maintenance costs in cabinet U-position recognition and vacancy rate assessment, and proposes a cabinet U-position detection method based on two-stage visual recognition and perspective correction.
[0006] To achieve the above object, the present invention is realized through the following technical solutions: A cabinet U-position detection method based on two-stage visual recognition and perspective correction, comprising the following steps: I. Obtaining cabinet image data Obtain cabinet image data, which is collected by a camera deployed at the top or middle of the computer room or uploaded by a manual inspection terminal; Preferably, training data preparation, annotation, and training Data is collected from the cabinets in the computer room by manual or the inherent cameras in the computer room, and operations such as data preprocessing are carried out. Two models are trained in stages: an irregular cabinet target detection model (hereinafter referred to as Model 1) and a corrected cabinet vacant U-position target detection model (hereinafter referred to as Model 2).
[0007] Label the original data.
[0008] For Model 1, label the vertex coordinates (x1, y1, x2, y2, x3, y3, x4, y4) of the quadrilateral of the cabinet outer contour. When the cabinet is completely visible, label the actual four corners. When partially occluded, infer the complete contour based on the visible part.
[0009] For Model 2, label the vacant U-position and non-vacant U-position areas. The labels are divided into two types: vacant (vacant) and non-vacant (non-vacant). Among them, vacant is an area not covered by continuous 2U or more (such as U5-U6), and non-vacant is the area covered by equipment, and the boundary is strictly aligned with the grid line.
[0010] Next, preprocess the training data, perform operations including geometric distortion, color perturbation, occlusion simulation, etc., to expand and enhance the dataset.
[0011] For Model 2, perform Mosaic enhancement. When splicing 4 images, force to include at least 1 vacant sample to alleviate the problem of class imbalance.
[0012] Use the preprocessed data for training.
[0013] For Model 1, adopt rotation algorithm. The IoU calculation is based on the minimum circumscribed rotated rectangle. In the inference part, the output inference box is an irregular quadrilateral, and the 4 vertex coordinates (x1, y1, x2, y2, x3, y3, x4, y4) of each box are regressed.
[0014] Loss function. The traditional loss function consists of classification loss, coordinate loss, and confidence loss. The improved version calculates the regression error of each vertex, the rotation angle error, and the scale loss respectively.
[0015] For Model 2, insert CoordinateAttention (CA) in the Feature Pyramid Network (FPN) layer to enhance the positioning ability for slender devices (such as 1U servers) The regression loss adopts EIoULoss to optimize the consistency of the length and width of the bounding box.
[0016] Save the trained model locally for convenient subsequent use. After deployment in the computer room, incremental training will be performed every once in a while (for example, once a week) according to the results of manual review to improve the model accuracy and adapt to the actual situation in the computer room.
[0017] II. The first-stage visual recognition model: Cabinet detection Perform irregular cabinet detection on the image data through the first-stage visual recognition model, and output the four vertex coordinates of the cabinet. Specifically, use an improved object detection algorithm to perform the first detection on the cabinet area and locate the irregularly distributed cabinet targets.
[0018] Obtain the photos taken by the camera or manually, organize the photo data, input it into the queue, and wait for detection.
[0019] Traverse the pictures, predict the 4 vertices of the pictures, retain the results with a confidence level greater than the threshold, and save the prediction results locally.
[0020] III. Image perspective correction Perform image perspective correction based on the four vertex coordinates to generate the cabinet image data of a standard rectangle. Specifically, perform geometric correction on the detected cabinet area to eliminate the perspective distortion caused by the shooting angle.
[0021] Obtain the 4 vertices of the irregular target detection, calculate the maximum width and height of the recognized irregular quadrilateral, and use them as the total width and height of the target picture.
[0022] Use cv2.getPerspectiveTransform and cv2.warpPerspective to perform perspective transformation according to the size of the target picture and perform cropping. The new picture is stored as wraped_img.
[0023] Traverse all the pictures, perform perspective transformation and cropping in sequence, and store them locally for the second object detection.
[0024] IV. The second-stage visual recognition model: U-position status detection Input the corrected image data into the second-stage visual recognition model to identify the occupancy status of the U-positions and construct a U-position status matrix. Specifically, based on the object detection model, perform refined object detection on the corrected cabinet image and count the occupancy status of the U-positions.
[0025] Determine the detection grid size. Since only when 2U or more are vacant can it be considered vacant, it is necessary to define a grid to detect whether all vacant targets are misdetected in a sliding manner.
[0026] Performing object detection inference on each image will result in outcomes labeled with "vacant" and "non-vacant" tags, and at the same time, confidence information for each image will be returned, etc.
[0027] Filter the detection results, retain the prediction boxes with confidence greater than the threshold (e.g., 0.6) to balance the recall rate and false positive rate; perform weighted NMS on overlapping boxes (weight = confidence × IoU) to suppress duplicate detections.
[0028] Construct a U-bit status matrix and initialize a fully vacant matrix , where 1 represents vacant and 0 represents occupied. For each non-vacant detection box, calculate the covered U-bit range and mark it as 0 at the corresponding U-bit position. If the device spans multiple U-bits, round up.
[0029] For each vacant detection box, verify whether the covered U-bit is greater than the sliding grid size and whether all values in matrix S are 1. If it passes, retain the vacant interval.
[0030] V. Calculation of vacancy rate Based on the U-bit status matrix, calculate the vacancy rate of the computer cabinet; VI. Output and visualization of detection results Output the detection results to the computer room management platform and visualize them in the form of a heat map; VII. Model optimization and active learning mechanism Manually review the detection results with confidence lower than the preset threshold and use the results for incremental training and continuous optimization of the model.
[0031] Preferably, the first-stage visual recognition model adopts a structure that supports irregular object detection, and the output is the four vertex coordinates of an irregular quadrilateral. During the training process, it is optimized using the following combined weighted loss function, and the expression is: ; Where: , , , , are the weight coefficients of each loss term, used to balance the magnitudes of different losses; is the classification error; is the vertex regression error; is the rotation angle error; is the scale error; is the confidence error; Among them, The expression is: ; Where: represents the number of samples in a batch, i.e., the number of images input for one training; represents the th sample being calculated; represents the th vertex of the bounding box; represents the coordinates of the th sample predicted by the model for the th vertex; ; represents the true vertex coordinates of the annotation.
[0032] The expression is: ; where: represents the number of samples in a batch, i.e., the number of images input for one training; represents the th sample being calculated; represents the rotation angle of the th target predicted by the model; represents the true rotation angle; The expression is: ; where: represents the dimension of the scale, i.e., the width and height of the target; represents the predicted scale; represents the true scale.
[0033] Specifically, the vertex position loss calculates the regression error of each vertex of the irregular quadrilateral box; the angle loss considers the influence brought by the target rotation angle and adds the calculation of the angle error of each vertex relative to the center point; the scale loss considers the difference between the actual area of the polygon and the area of the circumscribed rectangle of the box.
[0034] For Model 2, CoordinateAttention (CA) is inserted into the Feature Pyramid Network (FPN) layer to enhance the positioning ability for slender devices (such as 1U servers):
[0035] where GAP is global average pooling, GMP is global max pooling, is element-wise multiplication.
[0036] The classification loss adopts Focalloss to alleviate the class imbalance problem:
[0037] Among them, is the class balance coefficient, which adjusts the loss weights of positive and negative samples. is the focusing parameter. The larger it is, the more obvious the loss attenuation of easy samples is, and the model pays more attention to hard samples; is to adjust the direction of the predicted probability, and always represents the predicted probability of the model for the correct class.
[0038] Preferably, the regression part in the combined weighted loss function further adopts the extended intersection over union loss function, and the calculation formula is as follows: ; Among them: is the intersection over union, which measures the overlapping degree of two boxes; represents the center point distance penalty term, that is, the square of the Euclidean distance between the center point of the predicted box and the center point of the true box; represents the square of the length of the diagonal of the minimum closed region; and represent the square of the width and height of the minimum closed region.
[0039] Through manual inspection and photographing, the photos are uploaded to the computer room management platform for analysis; or the cabinet images are obtained in real time through the cameras deployed in the computer room, and preprocessing operations are performed on the images. Further, the technical points of image preprocessing include Gaussian filtering denoising, adaptive histogram equalization, etc.
[0040] In particular, Gaussian filtering denoising is to smooth the original image using a Gaussian kernel function to eliminate the interference caused by uneven illumination or sensor noise. The kernel function formula: ;
[0041] Among them, refers to the weight value of the Gaussian filter at , represents the position coordinates of the current pixel relative to the center of the filter, is the standard deviation of the Gaussian distribution, which controls the smoothness of the filter, is the exponential function, which is used to generate the bell-shaped curve of the Gaussian distribution, is the normalization coefficient, which ensures that the sum of the Gaussian kernel is 1.
[0042] In particular, adaptive histogram equalization enhances the image contrast through the CLAHE algorithm, improving the detail clarity of the cabinet edges and U-position devices. By dividing the image into local sub-blocks, restricting the contrast enhancement amplitude of each sub-block to avoid noise amplification, and smoothing the transition between blocks through bilinear interpolation, thus while significantly improving the local details and contrast of the image, effectively suppressing the generation of over-enhancement and artifacts.
[0043] Preferably, the acquisition device for cabinet image data includes cameras deployed at the top and middle of the computer room to obtain top-down and horizontal dual-view images, and includes a manual inspection terminal equipped with a wide-angle anti-distortion lens; the acquisition device supports image resolution to meet the requirements of whole-cabinet imaging and has the ability to synchronize data for timed acquisition and automatic upload to the platform.
[0044] Preferably, the image perspective correction includes: For the irregular quadrilateral formed by the four vertices obtained from the target detection in the first stage, calculate the width and height mapped to the standard rectangle according to its vertices to determine the target image size; Use the getPerspectiveTransform and warpPerspective functions of OpenCV to perform geometric transformation to eliminate the perspective distortion caused by the shooting angle; Crop the transformation result into a standard rectangular cabinet image warped_img, and the image is stored locally for the target detection processing of the second-stage model.
[0045] Preferably, based on the target detection model, perform refined target detection on the corrected cabinet image and count the occupancy status of U positions.
[0046] Determine the detection grid size. Since only when 2U or more are vacant can it be considered vacant, a grid needs to be defined to slide to detect whether all vacant targets are misdetected.
[0047] According to the actual height (unit: pixel) of the corrected cabinet image and the standard U-bit height (44.45 mm), calculate the pixel height corresponding to each U: ; Where: represents the pixels occupied by each bit; is the standard bit height; is the height of the entire real cabinet, and only when the length of the continuously uncovered area is not less than 2U, this area is counted as the vacant state.
[0048] Furthermore, calculate the prediction confidence level output by the second-stage visual recognition model. If the confidence level is higher than 0.6, it is directly adopted; If the confidence level is between 0.4 and 0.6, submit the corresponding detection result for manual review; The results after review are recorded and added to the training sample set for subsequent incremental training of the model; Perform the determination of the vacant state through sliding window operation. The window covers no less than 2U positions. If the prediction results within the window range are all vacant, the overall area is marked as vacant.
[0049] Preferably, based on the recognition result of the U-bit occupancy status output by the second-stage visual recognition model, the determination of the vacant status by sliding window operation includes: Construct a status vector with a length of , where 0 indicates that the bit is occupied and 1 indicates vacant; When the device detection area spans multiple bits, round up according to the actual number of occupied U-bits, and assign 0 to the corresponding positions in the status vector; Perform a sliding window process on the status vector. If all the status values within the sliding area are 1, it is determined that the area covered by the window is a valid vacant segment; The vacancy rate is calculated according to the following formula: .
[0050] Preferably, before training the second-stage visual recognition model, the following image preprocessing operations are performed on the image data, including: Perform Gaussian filtering denoising and CLAHE enhancement on the original image to improve the local contrast and clarity of the image; Randomly perform geometric perturbation operations, including perspective transformation, rotation, and scaling; Color perturbation is performed based on the HSV color space; Occlusion simulation generates occlusion samples by superimposing irregular black blocks; When training the second-stage model, the Mosaic image stitching enhancement strategy is further adopted. The specific improvement points include: a. Ensure that each group of stitched images contains at least one image sample with a vacant U-bit to balance the class ratio; b. When constructing the candidate sample set, assign a higher selection probability to the samples with vacant U-bits; c. Based on the detection accuracy of each category in the current training cycle, dynamically optimize the scaling ratio and spatial arrangement of the stitched images; Embed a CA module in the Feature Pyramid Network (FPN) of the second-stage visual recognition model to fuse channel and spatial location features, thereby enhancing the model's recognition ability for slender targets.
[0051] Moreover, embed a CoordinateAttention (CA) module in the Feature Pyramid Network (FPN) of the second-stage visual recognition model to fuse channel and spatial location features, thereby enhancing the model's recognition ability for slender targets (such as 1U servers).
[0052] Preferably, the following post-processing operations are performed on the U-bit detection results output by the second-stage visual recognition model: The Weighted Non-Maximum Suppression (WeightedNMS) algorithm is used to merge duplicate detection boxes, and the weights are jointly determined by the IoU and the detection confidence; Generate a heat map of the vacancy rate of the cabinet U-bit. The heat map supports the visual display of multi-dimensional vacancy rate data and provides a historical trend comparison function; If the vacancy rate fluctuation of a certain cabinet exceeds the set threshold within a continuous monitoring period, a space anomaly alarm is triggered; The manual review results will be automatically added to the training set by the system, and the online continuous optimization of the model structure is realized through an active learning mechanism; The active learning mechanism includes: pushing the detection results with the detection confidence between 0.4 and 0.6 to the manual review system. After the review and confirmation, an annotation file is automatically generated and incorporated into the training data set, and an incremental training task and a model parameter update process are triggered every set period (such as weekly).
[0053] By adopting a two-stage visual recognition structure, that is, in the first stage, the irregular target detection is used to accurately locate the cabinet area, the present invention effectively solves the perspective distortion problem caused by the shooting angle, cabinet tilt or partial occlusion, and significantly improves the accuracy and robustness of the cabinet U-bit detection; in the second stage, the refined U-bit state detection is further performed on the corrected image, and the Feature Pyramid Network (FPN) and CoordinateAttention (CA) module are introduced, which effectively improves the recognition accuracy of slender targets (such as 1U servers). At the same time, the present invention adopts a pure visual detection scheme to replace the traditional sensor detection and manual inspection, avoiding the problems of high cost, complex maintenance and low efficiency of manual inspection brought by the sensor scheme. Only a small number of cameras are needed to realize one-to-many real-time dynamic monitoring, greatly reducing the deployment cost and maintenance difficulty and improving the automation level. In addition, the detection method proposed by the present invention has strong adaptability, can be compatible with cabinets of various brands and specifications, and the vacancy rate determination threshold can be flexibly customized, significantly reducing the system expansion cost. The present invention also introduces an active learning and manual review mechanism. Through the manual review and incremental training of low-confidence results, the online continuous optimization of the model is realized, ensuring long-term stability and detection accuracy, and adapting to data drift and complex environmental changes during long-term operation. Finally, the present invention also provides an intuitive U-bit state heat map and historical trend analysis tool, which can quickly discover the imbalance and abnormal fluctuation of the cabinet resource distribution and trigger an automatic alarm mechanism, greatly improving the response speed and overall operation and maintenance efficiency of the computer room management personnel. Therefore, the present invention is significantly superior to the prior art in terms of detection accuracy, economy, automation degree, compatibility, long-term stability and operation and maintenance efficiency, and is particularly suitable for the intelligent management of modern data centers. Description of the Drawings
[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them: Figure 1 is the method flow chart of the present invention; Figure 2 is the overall structure schematic diagram of the embodiment of the present invention; Figure 3 is an example diagram of the original image of the cabinet of the embodiment of the present invention; Figure 4 is an example diagram of the perspective transformation of the cabinet of the embodiment of the present invention; Figure 5 is an example diagram of the detection result of irregular objects of the embodiment of the present invention; Figure 6 is an example diagram of the detection result of vacant objects of the embodiment of the present invention. Detailed Embodiments
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present invention in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the described embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.
[0056] As Figure 1 shown, it is an embodiment of the present invention. This embodiment provides a cabinet U-position detection method based on two-stage visual recognition and perspective correction, including the following steps: Step 1: Obtain cabinet image data, which is collected by a camera deployed on the top or middle of the computer room or uploaded by an artificial inspection terminal; Step 2: Perform irregular cabinet detection on the image data through the first-stage visual recognition model, and output the four vertex coordinates of the cabinet; Step 3: Perform image perspective correction based on the four vertex coordinates to generate standard rectangular cabinet image data; Step 4: Input the corrected image data into the second-stage visual recognition model to identify the occupancy status of the U-position and construct a U-position status matrix; Step 5: Calculate the vacancy rate of the cabinet based on the U-position status matrix; Step 6: Output the detection result to the computer room management platform and visually display it in the form of a heat map; Step 7: Manually review the detection results with confidence levels below the preset threshold, and use the results for incremental training and continuous optimization of the model.
[0057] Example 1: Image data acquisition method Obtain cabinet image data, which is collected by cameras deployed at the top or middle of the computer room, or uploaded by manual inspection terminals. In this example, the data of the front image of the cabinet is obtained by combining manual inspection shooting and fixed camera collection. As Figure 2 shown, the original cabinet image can be obtained by shooting with cameras deployed in the computer room. The cameras are preferably installed at the top or middle of the computer room passage to provide coverage of top-down and front (horizontal) views respectively. For example, a Hikvision camera can be deployed above the front of the cabinet to obtain a top-down view of the cabinet front; the camera parameters are set to high-definition mode to ensure that the resolution meets the recognition requirements. At the same time, manual inspection personnel can use a mobile terminal equipped with an anti-distortion wide-angle lens (such as a Canon brand) to shoot each cabinet one by one. When shooting, keep a horizontal distance of about 1.5 meters from the front of the cabinet, and control the vertical angle of the camera within ±30° to ensure that the obtained image covers the entire cabinet and has less distortion.
[0058] Through the above two methods, the collected image data covers a variety of typical scenarios, including different shooting angles (top-down, horizontal view, etc.) and different lighting conditions (normal lighting, emergency lighting, etc.). To ensure the robustness of the algorithm, it is required that the collected original images are clear and contain the complete cabinet target (the complete visible area of the cabinet reaches more than 90% of the image). In actual deployment, the fixed camera can be set to collect images periodically. For example, it automatically takes a front view photo of the cabinet every 30 minutes and transmits the image to the server through the network (such as the RTSP protocol); the images taken by manual inspection can be uploaded through a mobile application (it is recommended to use the JPEG format and the image quality compression is not less than 90%). If necessary, preprocessing operations can be performed on the obtained original images, such as Gaussian filtering for noise reduction, adaptive histogram equalization, etc., to improve the image quality and the accuracy of subsequent recognition. The acquisition of the above image data provides a reliable data basis for subsequent model training and real-time detection.
[0059] Example 2: Model training and labeled data construction In this example, the above-collected cabinet images are used to train two visual recognition models: Model 1 is used for detecting irregular targets in the cabinet area, and Model 2 is used for detecting the occupancy status of cabinet U positions. To construct the dataset required for training, it is necessary to annotate the images and perform data augmentation processing. The specific steps are as follows: Data Annotation and Label Definition: First, use professional annotation tools (such as LabelImg, etc.) to annotate the original images, generating annotation files (such as JSON or YOLO format) containing the target positions and categories. For Model 1, the outer contour of the cabinet in each image is annotated with an irregular quadrilateral, that is, the pixel coordinates of the four vertices of the cabinet are marked. If the cabinet is completely visible, its actual four corner positions are annotated; if part of the cabinet is occluded, the complete rectangular contour is inferred based on the visible part and the vertex positions are annotated. For Model 2, the occupancy status of each U-position on the front of the cabinet needs to be annotated, and the labels are divided into two categories: "vacant" (unoccupied) and "non-vacant" (occupied). Specifically, if there is no equipment occlusion in a continuous area of 2 U-heights or more on a certain section of the cabinet, the continuous area is annotated as vacant (for example, the vacant U5 - U6 area); any area with equipment installed and covered is annotated as non-vacant, and all marked boundaries are strictly aligned with the horizontal dividing lines of the cabinet U-positions. If it is found that a certain equipment occupies a height exceeding 1.5 U, it is considered to occupy two adjacent U-positions (rounding up to mark the occupied number of U-positions); small objects that only cover less than half of a U-height can be not marked as occupied. Through polygon or rectangle annotation tools, the four-corner coordinate information of the cabinet and the status labels of each U-position area can be obtained simultaneously, providing training samples for the two-stage model.
[0060] Dataset Augmentation and Enhancement: To improve the robustness and generalization ability of the model, various data augmentation operations are performed on the training data after annotation. First, geometric distortion augmentation is introduced: randomly perform perspective transformation, rotation, scaling, etc. on the images to simulate the situations of shooting at different angles and the inclination of the cabinet, so that the model can adapt to view changes. Second, color perturbation is carried out: randomly adjust the hue, saturation, and brightness of the images in the HSV color space to enhance the model's adaptability to different lighting and color temperature conditions. Third, occlusion simulation is adopted: superimpose black rectangular blocks or other occluders at random positions in the images to simulate the situation of temporary occlusion in front of the cabinet (such as people or sundries passing by), so as to improve the detection performance of the model in the presence of partial occlusion. In addition, for the data of Model 2, we especially adopt the Mosaic data augmentation strategy: expand the data samples by splicing four different front images of the cabinet into one image. At least 1 vacant area sample is forced to be included in each composite image spliced by Mosaic to alleviate the problem of class imbalance between vacant and non-vacant categories. After the above processing, ensure that the scale of the final dataset used for training is not less than 2000 diverse image samples.
[0061] Model Training: Use the enhanced data to train Model 1 and Model 2 respectively. Model 1 can adopt an improved YOLO architecture or other object detection networks that can output irregular quadrilateral detection frames. During training, a combined loss function is used to simultaneously optimize indicators such as classification accuracy and vertex position accuracy; Model 2 can adopt existing lightweight object detection networks (such as YOLOv5 or YOLOv8, etc.) to identify the U-bit status. During the training process, the labeled cabinet vertex coordinates and U-bit status are used as supervision signals to continuously adjust the model parameters to make them converge. When the training reaches the expected accuracy, export the model weight files (such as.pt format or.onnx format) respectively for inference deployment. Thus, the offline training and data construction work of the two-stage model are completed.
[0062] Example 3: Implementation of Irregular Object Detection and Confidence Mechanism in the First Stage In this example, for the detection of the entire cabinet, the trained first-stage model is used to perform irregular object detection on the collected images to accurately locate the position and shape of the cabinet. As Figure 4As shown in the figure, Model 1 performs positioning detection on the cabinet in the original image and outputs the four vertex coordinates of the outer contour of the cabinet. The specific implementation process is as follows: First, the input original image is scaled and normalized according to the requirements of the model (for example, scaled to a resolution of 640×640 and the pixel values are normalized to the interval [0,1]), and then sent into the neural network of Model 1 for inference. Model 1 adopts an improved YOLO object detection architecture. The special feature is that its detection head supports predicting irregular quadrilateral boxes instead of the traditional axis-aligned rectangular boxes. That is to say, the model directly regresses the image coordinate positions of the four corner points of the cabinet, so as to accurately depict the inclined rectangular contour of the cabinet in the image. To improve the detection accuracy, a composite loss function is introduced during model training, and weights are given to the classification error, vertex position error, rotation angle error, scale error, and confidence error respectively for joint optimization, so that the model can better learn the inclination angle and shape features of the cabinet target. After the model infers the candidate regions of the cabinet, a confidence mechanism is applied to screen the effective detection results. Each prediction output comes with a confidence score, indicating the credibility of the detected area as a cabinet. In this embodiment, the confidence threshold is set to 0.6, that is, when the confidence of the candidate region of the cabinet output by the model is lower than 0.6, it is regarded as an unreliable detection and discarded; the results higher than this threshold are considered to be reliable cabinet identifications. If there are multiple candidate detection frames (for example, when there are multiple cabinets in the image), the non-maximum suppression (NMS) algorithm can be used to optimize the overlapping detection frames, and only the one with the highest confidence or the frames with a large overlap degree are merged. Thus, it is ensured that only one accurately positioned polygon detection frame is obtained for each cabinet in each image. The four vertex coordinates obtained by the detection are sorted and numbered in clockwise or counterclockwise order for use in subsequent steps. Through the above irregular object detection mechanism, the position and shape of the cabinet in the image can be accurately extracted, providing a reliable basis for perspective correction.
[0063] Specifically, the first-stage visual recognition model adopts a structure that supports irregular object detection, and the output is the four vertex coordinates of an irregular quadrilateral. During the training process, it is optimized using the following combined weighted loss function, and the expression is: ; Where: , , , , are the weight coefficients of each loss term, used to balance the magnitudes of different losses; is the classification error; is the vertex regression error; is the rotation angle error; is the scale error; is the confidence error; Where, The expression is: ; Where: represents the number of samples in a batch, that is, the number of images input for one training; represents the th sample being calculated; represents the th vertex of the bounding box; represents the coordinates of the th sample predicted by the model for the th vertex ; represents the true vertex coordinates of the annotation.
[0064] The expression is: ; Where: represents the number of samples in a batch, that is, the number of images input for one training; represents the th sample being calculated; represents the rotation angle of the th target predicted by the model; represents the true rotation angle; The expression is: ; Where: represents the dimension of the scale, that is, the width and height of the target; represents the predicted scale; represents the true scale.
[0065] Specifically, the vertex position loss calculates the regression error of each vertex of the irregular quadrilateral box; the angle loss considers the influence brought by the target rotation angle and adds the calculation of the angle error of each vertex relative to the center point; the scale loss considers the difference between the actual area of the polygon and the area of the circumscribed rectangle of the box.
[0066] For Model 2, insert CoordinateAttention (CA) in the Feature Pyramid Network (FPN) layer to enhance the positioning ability for slender devices (such as 1U servers):
[0067] Where GAP is global average pooling, GMP is global max pooling, is element-wise multiplication.
[0068] The classification loss uses Focal loss to alleviate the problem of class imbalance:
[0069] Among them, is the class balance coefficient, which adjusts the loss weights of positive and negative samples. is the focusing parameter. The larger it is, the more obvious the loss attenuation of easy samples is, and the model pays more attention to hard samples; is to adjust the direction of the predicted probability, and always represents the predicted probability of the model for the correct class.
[0070] The regression part in the combined weighted loss function further adopts the extended intersection over union loss function, and the calculation formula is as follows: ; Among them: is the intersection over union, which measures the overlapping degree of two boxes; represents the center point distance penalty term, that is, the square of the Euclidean distance between the center point of the predicted box and the center point of the true box; represents the square of the length of the diagonal of the minimum closed region; and represent the square of the width and height of the minimum closed region.
[0071] Example 4: Image perspective correction method and size calculation method Based on the obtained four corner coordinates of the cabinet, this embodiment performs perspective distortion correction on the cabinet image, restores the standard rectangular view of the front of the cabinet, and establishes the corresponding relationship between the pixel size and the actual size. As Figure 3 shown, after performing perspective transformation on the original tilted image of Figure 2 , the front top-down correction diagram of the cabinet can be obtained. The specific implementation steps are as follows: First, based on the four vertex coordinates of the cabinet obtained in Embodiment 3, calculate the width and height of the irregular quadrilateral corresponding to the cabinet in the original image. When calculating, measure the distances between the opposite sides of the quadrilateral respectively: on the one hand, calculate the lengths of the upper edge and the lower edge (i.e., the distances between the two corners at the top of the cabinet and the distances between the two corners at the bottom), and take the maximum value as the target width after the cabinet is corrected; on the other hand, calculate the lengths of the left edge and the right edge, and take the maximum value as the target height. Thus, determine the pixel size (target width × target height) of the target image after perspective transformation. Next, based on the correspondence between the four vertex coordinates of the original image and the four vertex coordinates of the corrected rectangular image, calculate the perspective transformation matrix. The cv2.getPerspectiveTransform function in the OpenCV library can be used to obtain the 3×3 perspective transformation matrix, and then the cv2.warpPerspective function is used to perform a projective transformation on the original image to map the cabinet area into a regular rectangular shape. The transformed image is cropped according to the calculated target width and height to obtain a rectangular image containing only the front of the cabinet (i.e., the perspective correction image, as Figure 3 shown). The upper and lower edges and the left and right edges of the cabinet in this image are parallel to the image border, without the shrinkage deformation caused by the original perspective.
[0072] After completing the perspective correction, it is also necessary to establish the relationship between the pixel size of the image and the actual physical size of the cabinet in order to determine the unit scale of the U position. Specifically, the total pixel height of the cabinet image after correction can be compared and calculated with the actual total U height of the cabinet. Assume that the detected cabinet is a standard 42U cabinet, and the actual height of each 1U is 44.45 mm, then the total height of the cabinet is approximately 42×44.45 = 1866.9 mm. Measure the pixel height (denoted as H_px) from the top to the bottom of the cabinet in the perspective-corrected image, then the pixel ratio corresponding to each millimeter is H_px divided by 1866.9. Further, the pixel height h_px of each 1U in the image is approximately (H_px / 1866.9) × 44.45. In a simple case, the total pixel height of the corrected image can also be directly divided into several equal parts according to the number of Us, and the pixel value corresponding to each equal part is 1U height. After determining the pixel size corresponding to 1U in the above manner, a vertical grid division is established on the corrected image, and the height of each grid is 1U. In this way, the pixel scale of the cabinet U position is obtained, providing a basis for subsequent per-U detection and vacancy rate calculation. It should be noted that if the cabinet is not of standard height or the number of Us is known, the relationship between pixels and actual dimensions can also be converted by pre-measuring the calibration points on the cabinet, and the above calculation method does not limit the specific cabinet model. The key is that through perspective correction and size calculation, the cabinet image can be standardized, enabling subsequent algorithms to accurately analyze by the U position unit.
[0073] Specifically, the image perspective correction includes: For the irregular quadrilateral formed by the four vertices obtained from the first-stage object detection, calculate the width and height mapped to the standard rectangle based on its vertices to determine the size of the target image; Use the getPerspectiveTransform and warpPerspective functions in OpenCV to perform geometric transformation to eliminate the perspective distortion caused by the shooting angle; Crop the transformation result into a standard rectangular cabinet image warped_img, and the image is stored locally for the second-stage model to perform object detection processing.
[0074] Furthermore, in the standard rectangular image obtained after perspective transformation, calculate the pixel height per U-bit through the following formula: ; Where: represents the pixels occupied by each bit; is the height of the standard bit; is the height of the entire real cabinet, and only when the length of the continuous uncovered area is not less than 2U, this area is counted as the vacant state.
[0075] Example 5: Second-stage U-bit occupancy status detection and status matrix construction After completing the perspective correction of the cabinet image, this example uses the second-stage model to detect the U-bit occupancy status of the corrected front image of the cabinet, identify the vacant or occupied situation of each U-bit, and construct the corresponding status matrix. As Figure 5 shown, Model 2 pairs Figure 3After performing detection on the obtained front image of the cabinet, the vacant areas and equipment-occupied areas can be marked. The specific implementation steps are as follows: First, input the corrected front image of the cabinet obtained in Embodiment 4 into Model 2 for object detection. Model 2 has learned the features of different U-height areas inside the cabinet during training and can identify which areas are vacant and which areas have equipment. The model outputs several detection boxes, each with a class label (vacant or non-vacant) and a confidence score attached. Usually, each cabinet device corresponds to a non-vacant detection box (possibly with a height of 1U or several Us), while continuous vacant spaces (at least 2U high) will correspond to a vacant detection box. Next, filter and optimize the detection results output by the model. First, apply a confidence threshold filter to discard detection boxes with a confidence lower than the preset threshold (such as 0.6) to balance the recall rate and precision and reduce the impact of false positives. Then, use the weighted non-maximum suppression (NMS) algorithm to merge overlapping candidate regions for the remaining detection boxes. Specifically, calculate a weighted score according to the overlap degree (IoU) between detection boxes and their respective confidences (for example, confidence × IoU as the weight factor), and fuse the highly overlapping detection boxes with the same class into a single box. Through this strategy, duplicate detections that the model may output can be eliminated, ensuring that only one detection result is retained for the same device or the same vacant area, thereby improving the accuracy of status judgment. After obtaining stable detection results, construct a cabinet U-position status matrix. Assume that the cabinet has a total of N U-positions (for example, N = 42), then establish a one-dimensional array of length N or the corresponding matrix representation, where each element represents the status of the corresponding U-position. Initially, assign all elements of the matrix the value 1, indicating that it is assumed to be vacant. Subsequently, update this matrix according to the detection results: For each detection box with the class of non-vacant, calculate its vertical coverage range in the cabinet and change the corresponding position in the matrix to 0 to indicate occupancy. When specifically implementing, utilize the pixel-to-U-height mapping relationship calculated in Embodiment 4 to map the top and bottom pixel positions of the non-vacant detection box to the specific U-number range. For example, if a detection box ranges from the 10th pixel row to the 30th pixel row in the corrected image and its height covers approximately 1.5U, then mark the starting U-position and the two adjacent U-positions below it in the corresponding matrix as 0 (occupied). And so on, each detected equipment area will set the U-grid positions it covers to 0 in the matrix. If the equipment coverage range has a non-integer number of U-positions, use the ceiling function to ensure that as long as more than half of a U is occupied, it is counted as one U-position. For each detection box of the vacant class, further verify whether the corresponding vacant area in the matrix is indeed empty and meets the condition of being continuous for more than 2U.Specifically, calculate the number of Us corresponding to the height span of the vacant detection box in the corrected image. For example, if a certain vacant box covers U5 to U7, the span is 3Us. Only when its span is not less than 2Us and the matrix within this range was previously all 1 (i.e., not marked as occupied by any device) can it be finally confirmed that the detection box represents a real vacant area. If the height covered by the vacant box is less than 2Us (such as only 1U high), it is regarded as an invalid detection or a transient gap according to the pre-set rules and is not separately counted as a vacant state. Through this sliding window determination mechanism, it is ensured that only vacant segments with two or more consecutive U positions are recorded as valid vacancies, filtering out possible false detections or meaningless zero-star vacancies. After the above process, the cabinet U-position status matrix S is finally obtained, where the elements with a value of 1 correspond to vacant U positions, and the elements with a value of 0 correspond to U positions occupied by devices. This matrix comprehensively describes the current occupancy situation of each U position in the cabinet.
[0076] Specifically, calculate the prediction confidence level output by the visual recognition model in the second stage. If the confidence level is higher than 0.6, it is directly adopted; If the confidence level is between 0.4 and 0.6, the corresponding detection result is submitted for manual review; The results after review are recorded and added to the training sample set for subsequent incremental training of the model; The vacant state is determined through a sliding window operation. The window covers no less than 2U positions. If the prediction results within the window range are all vacant, the overall area is marked as vacant.
[0077] Example 6: Vacancy rate calculation method and sliding window determination mechanism After obtaining the cabinet U-position status matrix S, the vacancy rate of the cabinet is calculated in this example, and the sliding window mechanism is further used to ensure the stability of the vacancy judgment. As Figure 6As shown in the figure, the cabinet vacancy rate is defined as the proportion of vacant U positions in the cabinet and can be calculated as follows: Count the number of U positions with a value of 1 (vacant) in the statistical matrix S, divide it by the total number of U positions N in the cabinet, and multiply by 100% to obtain the vacancy rate percentage. For example, for a 42U cabinet, if 10 positions are vacant, the vacancy rate is approximately 10 / 42 ≈ 23.8%. In actual calculation, the count of vacant U positions can be directly obtained from the matrix constructed in the previous embodiment, so as to quickly obtain the vacancy rate index. It should be noted that in this embodiment, a sliding window judgment mechanism is incorporated into the identification process of the vacant area to ensure the reliability and effectiveness of the calculated vacancy rate. Specifically, in Embodiment 5, only areas with 2U or more consecutive vacancies will be counted as truly vacant. We can understand this rule as applying a sliding window with a height of 2U in the vertical direction of the cabinet to scan the U position sequence of the cabinet: Only when all the consecutive U positions covered by the sliding window are empty (no equipment, occupancy marked as 0), it is determined that there is an effective vacant segment. Whenever there is occupancy within the sliding window, the window is not counted as vacant. By moving the sliding window, all vacant intervals that meet the condition can be detected, and the count of vacant U positions can be corrected accordingly. This mechanism effectively avoids counting single isolated vacant positions into the vacancy rate, thereby improving the accuracy of the statistical result in depicting the actual available space. The finally output vacancy rate takes into account both the total amount of vacant space in the cabinet and the continuity of the vacant space, making the calculation of the vacancy rate more in line with the requirements of the computer room operation and maintenance for the evaluation of available space.
[0078] Specifically, based on the U position occupancy status recognition result output by the second-stage visual recognition model, the determination of the vacant status through the sliding window operation includes: Construct a status vector with a length of , where 0 indicates that the position is occupied and 1 indicates vacant; When the device detection area spans multiple positions, round up according to the actual number of occupied U positions, and assign 0 to the corresponding positions in the status vector; Perform a sliding window process on the status vector. If all the status values within the sliding area are 1, it is determined that the area covered by the window is an effective vacant segment; The vacancy rate is calculated according to the following formula: .
[0079] Embodiment 7: Display method of the detection result (heat map) In this embodiment, the cabinet vacancy rate and U-position status matrix obtained from the above detection are connected to the computer room management platform, and the cabinet vacancy situation is displayed in an intuitive heat map manner, which is convenient for operation and maintenance personnel to consult and analyze. The system associates the vacancy rate of each cabinet with its spatial position and presents it in a color-coded matrix or graph on the monitoring interface. For example, the platform can draw a top-down layout diagram of the computer room or a list of cabinets, and mark the vacancy rates of each cabinet in the form of a heat map: cabinets with a high vacancy rate are displayed in cool colors (such as blue or green), and cabinets with a low vacancy rate and a high equipment occupancy rate are displayed in warm colors (such as red), and the depth of the color reflects the specific vacancy percentage. It can be seen at a glance which cabinets are almost empty (the color is bluish / light) and which cabinets are approaching saturation (the color is reddish / deep). This multi-dimensional data display method is like a heat distribution map, which can help administrators quickly locate the uneven distribution of cabinet resources. In addition, the platform interface can also provide a historical data trend analysis function. By recording and plotting the vacancy rate curve of each cabinet over time to form a time series graph or a dynamic heat map, operation and maintenance personnel can observe the change trend of the cabinet vacancy situation over a period of time. For example, when the vacancy rate of a certain cabinet continues to decline, it may indicate that new equipment is being continuously installed; on the contrary, an increase in the vacancy rate indicates that equipment is being moved out or removed. If it is detected that the vacancy rate of a certain cabinet shows abnormal fluctuations within a short period of time (for example, the fluctuation exceeds 20% within 24 hours, or there are continuous multiple up and down changes), the system can trigger an alarm to indicate that there may be unstable equipment removal / insertion or data anomalies. This combination of visual display and early warning mechanism greatly improves the intuitiveness and real-time nature of cabinet resource monitoring, providing strong support for the operation and maintenance of the data center.
[0080] Specifically, Gaussian filtering denoising and CLAHE enhancement are performed on the original image to improve the local contrast and clarity of the image; Randomly perform geometric perturbation operations, including perspective transformation, rotation, and scaling; Color perturbation is performed based on the HSV color space; Occlusion simulation generates occlusion samples by superimposing irregular black blocks; When training the second-stage model, the Mosaic image splicing enhancement strategy is further adopted to ensure that each group of spliced images contains at least one image sample with a vacant U-position to balance the class ratio; Moreover, a CoordinateAttention (CA) module is embedded in the Feature Pyramid Network (FPN) of the second-stage visual recognition model to fuse channel and spatial position features, thereby enhancing the model's recognition ability for slender targets (such as 1U servers).
[0081] Example 8: Continuous optimization strategy of the model (manual review, active learning, incremental training) This embodiment introduces the continuous optimization mechanism of the detection model after deployment to ensure that the model maintains high accuracy and adapts to the new environment during long-term operation. This optimization strategy is mainly achieved by introducing manual review feedback and active learning, that is, adding a cycle of manual correction and incremental training on the basis of automatic model detection to continuously improve the model performance. Specifically, the system initiates a manual review process for low-confidence detection results. Low confidence usually means that the confidence score given by the model falls within a relatively low intermediate range, such as between 0.4 and 0.6, indicating that the model is not very certain about this judgment. For such results, the system will send the corresponding cabinet image segment and detection marks to the manual annotation platform or the operation and maintenance personnel interface for manual verification and correction. The manual review personnel confirm whether the U-position area is vacant or occupied by equipment according to the actual situation, or mark the equipment missed by the model, and then submit the review result. After receiving the manual annotation feedback, the system adds these corrected samples to the model training set and automatically generates new annotation data that conforms to the training format. When a certain number of manually corrected samples are accumulated or the optimization interval is reached regularly (such as weekly), the incremental training process of the model is triggered. Incremental training uses the weights of the previously trained model as the initial value and continues to train the model on the training set with new samples added. On the one hand, continuing to learn new data features enables the model to adapt to the gradual changes in the computer room environment (such as lighting changes, equipment type replacement, etc.); on the other hand, by fine-tuning with a reduced learning rate (such as set to 1 / 10 of the initial learning rate), it can avoid forgetting the existing knowledge and overfitting to the new data. After one incremental training is completed, the updated model parameters are obtained, and the new model is deployed to replace the old model for subsequent image inference. By repeating this cycle, the model can be continuously iteratively optimized. It should be noted that this continuous optimization strategy of active learning can effectively reduce the workload of comprehensive manual annotation, focus human efforts on a small number of difficult samples where the model is uncertain, and thus efficiently improve the model accuracy. In the method of the present invention, by automatically extracting a certain number (such as 100) of the latest cabinet images from the computer room management platform every week for manual verification and incremental training, the model can continuously adapt to the changes in the cabinet layout or equipment installation, and still maintain a high recognition accuracy and robustness for the cabinet U-position status after long-term operation. This human-machine collaborative optimization method ensures the practicality and reliability of the system after deployment, and further improves the engineering feasibility of the cabinet U-position detection method in the actual data center operation and maintenance.
[0082] Specifically, the weighted non-maximum suppression (Weighted NMS) algorithm is used to merge the duplicate detection boxes, and the weight is jointly determined by the IoU and the detection confidence; Generate a heat map of the vacant rate of the cabinet U-position. The heat map supports the visual display of multi-dimensional vacant rate data and provides a historical trend comparison function; If the vacancy rate of a certain cabinet fluctuates beyond the set threshold within a continuous monitoring period, a space anomaly alarm is triggered; The manual review results will be automatically added to the training set by the system, and the online continuous optimization of the model structure will be achieved through an active learning mechanism;
[0083] The active learning mechanism includes: pushing the detection results with detection confidence between 0.4 and 0.6 to the manual review system, automatically generating an annotation file and incorporating it into the training data set after review and confirmation, and triggering an incremental training task and a model parameter update process every set period (such as weekly).
[0084] In summary, the present invention comprehensively improves the accuracy, automation, and scalability of the U-bit status detection of computer room cabinets by constructing an overall solution of "two-stage visual recognition + perspective correction + active learning optimization". In the first stage, an improved YOLO structure is adopted to achieve accurate detection of irregular cabinet boundaries, and the correction process of standard rectangular images is completed through geometric transformation, solving the problem of U-bit misalignment recognition caused by factors such as angle deviation, tilt, and occlusion in traditional images. In the second stage, the Mosaic enhancement strategy and the CoordinateAttention attention mechanism are integrated to focus on the occupancy determination of slender devices such as 1U, and cooperate with the sliding window mechanism and the state matrix construction method to effectively avoid false detection and missed detection in vacancy determination.
[0085] At the same time, the present invention innovatively introduces a human-machine collaboration mechanism based on the confidence interval: when the confidence of the model output is between 0.4 and 0.6, the system automatically submits it for manual review, and the review result is iteratively used as training data. Combined with a periodic incremental training mechanism (such as automatic tuning every week), it ensures that the model always has continuous adaptability to on-site changes and high-precision prediction capabilities. In addition, the detection results can be displayed in real time on the computer room management platform in the form of heat maps, historical trend charts, etc., significantly improving the resource visualization level and operation and maintenance response efficiency.
[0086] The overall solution takes into account multiple advantages such as "convenient deployment, lightweight model, accurate recognition, data closed-loop, and continuous self-evolution of the algorithm", significantly superior to traditional manual inspection and physical sensor solutions, showing strong practicality in terms of cost control, maintenance efficiency, engineering versatility, and platform integration capabilities, and is particularly suitable for the intelligent management application scenarios of U-bit resources in modern data centers, edge computer rooms, and mixed deployment environments of multi-brand cabinets.
[0087] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0088] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, without being executed in the order shown or discussed.
[0089] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A cabinet U position detection method based on dual-stage visual recognition and perspective correction, characterized in that: The following steps are involved: Obtain cabinet image data, which is collected by a camera deployed on the top or in the middle of the computer room, or photographed and uploaded by a manual inspection terminal; The first-stage visual recognition model detects irregular cabinets in the image data and outputs the coordinates of the four vertices of the cabinet; Perform image perspective correction based on the four vertex coordinates to generate standard rectangular cabinet image data; The corrected image data is input into the second-stage visual recognition model to identify the occupancy state of the U position and construct the U position state matrix; Based on the U-position state matrix, calculating the vacancy rate of the cabinet; The test results are output to the computer room management platform and visualized in the form of a heat map; Detection results with confidence levels below the preset threshold are manually reviewed and used for incremental training and continuous optimization of the model.
2. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 1 is characterized in that: The first-stage visual recognition model adopts a structure that supports irregular target detection, outputs the coordinates of the four vertices of an irregular quadrilateral, and uses the following combined weighted loss function for optimization during training, expressed as: ; in: , , , , is the weight coefficient of each loss term, which is used to balance the magnitude of different losses; is the classification error; is the vertex regression error; is the rotation angle error; is the scale error; is the confidence error; in, The expression is: ; in: Represents the number of samples in a batch, that is, the number of images input for one training; Represents the number being calculated samples; Represents the bounding box vertices; Represents the model prediction The sample The coordinates of the vertices ; Represents the real vertex coordinates of the annotation; The expression is: ; in: Represents the number of samples in a batch, that is, the number of images input for one training; Represents the number being calculated samples; The model predicts the The rotation angle of the target; Represents the true rotation angle; The expression is: ; in: The dimensions representing the scale, i.e. the width and height of the object; represents the scale of the prediction; Indicates the true scale.
3. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 2 is characterized in that: The regression part of the combined weighted loss function further adopts an extended intersection-over-union loss function, and the calculation formula is as follows: ; in: is the intersection-over-union ratio, which measures the degree of overlap between two boxes; represents the center point distance penalty term, that is, the square of the Euclidean distance between the center point of the predicted box and the center point of the true box; Represents the square of the diagonal length of the minimum closed area; and Represents the square of the width and height of the minimum closed area.
4. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 1 is characterized in that: The acquisition device for obtaining cabinet image data includes cameras deployed at the top and middle of the computer room to obtain top-down and horizontal dual-viewing angle images, and includes a manual inspection terminal equipped with a wide-angle anti-distortion lens; The acquisition device supports image resolution that meets the imaging requirements of the entire cabinet and has the ability to synchronize data collected at regular intervals and automatically uploaded to the platform.
5. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 1 is characterized in that: The image perspective correction comprises: For the irregular quadrilateral formed by the four vertices obtained by the first stage of target detection, the width and height of the standard rectangle are calculated based on its vertices to determine the target image size; By calculating the projection transformation matrix from the image platform to the target plane and applying the matrix to perform geometric transformation; the visual change algorithm is implemented through the function solving algorithm of the OpenCV library; The transformation result is cropped into a standard rectangular cabinet image warped_img, which is stored locally for target detection processing by the second-stage model.
6. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 5 is characterized in that: In the standard rectangular image obtained after perspective transformation, the pixel height per U bit is converted by the following formula: ; in: Represent each bits occupied by pixels; It is standard Height of the bit; It is the actual height of the entire cabinet, and the area is considered empty only when the length of the continuous uncovered area is not less than 2U.
7. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 1 is characterized in that: The method further comprises: Calculate the prediction confidence of the output of the second-stage visual recognition model. If the confidence is higher than 0.6, it is directly accepted. If the confidence level is between 0.4 and 0.6, the corresponding test results will be submitted for manual review; The audited results are recorded and added to the training sample set for subsequent incremental training of the model; The vacancy status is determined by a sliding window operation, where the window covers no less than 2U positions. If the prediction results within the window are all vacant, the entire area is marked as vacant.
8. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 7 is characterized in that: Based on the U-position occupancy status recognition result output by the second-stage visual recognition model, the vacancy status determination by sliding window operation includes: The construction length is The state vector , where 0 means The bit is occupied, 1 means it is vacant; When the device detection area spans multiple When the number of bits is U, it is rounded up according to the number of bits actually occupied, and the corresponding position in the state vector is assigned a value of 0; Performing sliding window processing on the state vector, if all state values in the sliding area are 1, determining that the area covered by the window is a valid vacant segment; The vacancy rate is calculated according to the following formula: 。 9. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 1 is characterized in that: The following preprocessing operations are performed on the image data before the second stage visual recognition model training: Perform Gaussian filtering denoising and CLAHE enhancement on the original image; Perform geometric perturbations, including perspective transformations, rotations, and scaling; Perform color perturbation in HSV color space; Occlusion is simulated by superimposing irregular black blocks; In the second stage of model training, the Mosaic image splicing enhancement strategy is adopted, which includes: Make sure that the stitched image contains samples of vacant U positions; Increase the selection probability of vacant U-position samples; Dynamically optimize the scaling and arrangement of stitched images based on category accuracy; A channel attention CA module is embedded in the feature pyramid network FPN of the second stage model to enhance the recognition ability of slender targets.
10. The cabinet U position detection method based on dual-stage visual recognition and perspective correction according to claim 9 is characterized in that: The following post-processing operations are performed on the U-position detection results output by the second-stage visual recognition model: Use the weighted non-maximum suppression algorithm to merge duplicate detection boxes, where the weight is determined by IoU and detection confidence; Generate a heat map of the cabinet U-space vacancy rate, which supports the visualization of multi-dimensional vacancy rate data and provides a historical trend comparison function; If the vacancy rate fluctuation of a cabinet exceeds the set threshold during the continuous monitoring period, a space abnormality alarm is triggered; The manual review results will be automatically added to the training set by the system to achieve online continuous optimization of the model structure through the active learning mechanism; The active learning mechanism includes: pushing the detection results with a detection confidence between 0.4 and 0.6 to the manual review system, automatically generating a labeling file after review and confirmation and incorporating it into the training data set, and triggering incremental training tasks and model parameter update processes every set period.
Citation Information
Patent Citations
Supplementary method and device of shared power bank, and computer readable storage medium
CN110264634A
License plate recognition method and system in open parking space
CN111429727A
Real-time cabinet U position occupation condition detection method based on inspection robot
CN115810158A
Rolling metal surface defect automatic labeling method based on multi-task self-adaptive model
CN119444759A
Method and system for constructing, training and evaluating visual large model of power equipment
CN119692420A
Cited By
Cabinet resource detection and identification method, system, equipment and medium
CN120876998A
Layout detection method, device and equipment for sample disc
CN121391850A