Small target identification method and system based on cascade hierarchy detection and self-comparison
Through the small target recognition method of cascade hierarchical detection and self-comparison, the YOLOv8s network and Transformer encoder are used to solve the problem of small and medium-sized target recognition of drone inspections, and efficient and accurate detection of defects of distribution equipment is achieved.
Patent Information
- Application Number
- CN202510820955.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-19
AI Technical Summary
It is difficult for the prior art to effectively identify small target defects in distribution network equipment during drone inspections, especially in complex environments. Traditional methods have problems such as dilution of small target features and insufficient semantics, resulting in high recognition difficulty and high false detection rate.
A small-objective recognition method based on cascade hierarchical detection and self-comparison is adopted. Through the linkage of the device identification module, the small-objective defect identification module and the feature library, the YOLOv8s network is used for feature enhancement and Euclidean distance judgment, and combined with sliding window cutting and Transformer encoder for accurate identification.
It improves the recognition accuracy of small target defects, reduces the error detection rate, enhances the robustness and stability of the model, adapts to complex background environments, and improves the efficiency and accuracy of drone inspections.
Smart Images

Figure CN120356125A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular, to a small target recognition method and system based on cascade hierarchical detection and self-comparison. Background Art
[0002] With the continuous growth of power demand, the scale of the distribution network is constantly expanding, and its structure has become increasingly complex. Ensuring the safe and stable operation of the distribution network is crucial for social production and people's daily lives. Some small targets such as insurance pins, line clamp insulation covers, and screws are key protection components of distribution network equipment, and their status is directly related to the safety of the equipment and the entire distribution network. In actual operation, due to various factors such as the natural environment (such as wind, rain, vibration), equipment aging, and overload, some small targets may fall off. Once these small targets fall off or are lost, the relevant equipment may not be able to perform the protection function properly, thereby triggering power outages, affecting the reliability of power supply, and even possibly causing equipment damage and economic losses. Therefore, the recognition of small targets is very important. Traditional distribution network inspections mainly rely on manual labor. However, manual inspections are inefficient, labor-intensive, and it is difficult to comprehensively and timely detect problems such as the falling off of small targets such as insurance pins in complex environments. In recent years, unmanned aerial vehicle (UAV) inspection technology has gradually been applied to distribution network inspections, and it can quickly and efficiently obtain image information of distribution network equipment.
[0003] The current mainstream target detection models are mainly of two types. One directly inputs the entire panorama of a picture and then directly uses the detection model to detect it. In this case, there will be a situation where small targets account for a very small proportion in the image, resulting in increased recognition difficulty and also being relatively difficult to observe. The other uses two-stage detection (first roughly detecting components and then finely detecting small targets). For example, the small target defect detection method, system, equipment, and storage medium disclosed in the prior art. This solution detects the distribution network line image by using a region of interest (ROI) detection network and an improved defect detection network to obtain the ROI box image, further obtaining the small target defect detection region image, and finally realizing the detection of small target defects. However, neither of the above two methods pays more attention to small targets when extracting features or performing feature fusion. That is, due to the lack of a feature enhancement mechanism, there will be problems such as the dilution of features and insufficient semantics of small targets in the deep network. As a result, in the process of the model extracting features and fusing features from the shallow layer to the deep layer, some shallow layer feature information will be discarded or lost. The key spatial details of the picture are mainly in the shallow layer, and this part of the information will help locate the region of interest and identify small objects. If this part of the information is missing, it will greatly affect the overall recognition effect. Summary of the Invention
[0004] To address the deficiencies in the existing technology, the present invention provides a small target recognition method and system based on cascaded hierarchical detection and self-comparison. Through the linkage of three parts: the device recognition module, the small target defect recognition module, and the caching and comparison of the feature library, the types of small target defects in the distribution network scenario can be effectively recognized.
[0005] The present invention adopts the following technical solutions.
[0006] In a first aspect, the present invention provides a small target recognition method based on cascaded hierarchical detection and self-comparison, the method comprising: Collect the aerial images captured by the drone for inspection and annotate the aerial images; Input the annotated images into the device recognition module for device area detection and image cropping to obtain area images containing only the regions with small targets; Input the area images into the cascaded small target defect recognition module for defect recognition, and perform feature encoding on the recognized small target defects to obtain the encoded defect features; Calculate the Euclidean distance between the defect features and the historical small target defect features currently updated in the feature library. If the Euclidean distance is greater than the preset threshold, it means that the defect feature is a false detection and will be excluded; if it is not greater than the preset threshold, it means that the defect feature is correct, output the corresponding small target defect recognition result, and cache the defect feature in the feature library.
[0007] Optionally, the content of annotating the eligible images includes: adjustment of the image size of the target box, selection of the position box, and category information of the device and small targets.
[0008] Optionally, the device recognition module includes a target detection module and an image cropping module; The target detection module detects the device areas in the input images where small targets may appear to obtain the region of interest box images; the image cropping module crops the detected region of interest images to obtain area images containing only the regions with small targets.
[0009] Optionally, both the device recognition module and the small target defect recognition module are composed of the YOLOv8s network; Construct a device image sample set based on the annotated historical images to train and optimize the YOLOv8s network to obtain the device recognition module, and output the corresponding area image dataset containing only the regions with small targets; Construct a small target image sample set based on the area image dataset to train and optimize the YOLOv8s network to obtain the small target defect recognition module.
[0010] Optionally, the YOLOv8s network is obtained by replacing the C2f feature fusion module in the neck network of the YOLOv8 network with a C2f-att module based on the attention mechanism; The C2f-att module includes a first convolutional layer Conv1 connected in sequence, a segmentation layer, an EMA attention layer, a Bottleneck deep feature extraction module, a Concat connection layer, and a second convolutional layer Conv2.
[0011] Optionally, the process of the C2f-att module for processing image features is as follows: For the feature T input to the C2f-att module, it is first input into the first convolutional layer Conv1 for channel adjustment to obtain the adjusted feature ; The segmentation layer divides the adjusted feature into two parts in terms of channels to obtain the feature and the feature ; The feature is input into the EMA attention layer to extract multi-scale feature information, and the output feature is obtained; The Bottleneck deep feature extraction module stacks the feature multiple times through multiple Bottlenecks to obtain the deep feature ; The Concat connection layer concatenates the feature with the deep feature to obtain the feature ; The second convolutional layer Conv2 performs a convolution operation on the feature to adjust back the number of channels and output the final feature F.
[0012] Optionally, the EMA attention layer includes a third convolutional layer, a fourth convolutional layer, a Concat connection layer, a GAP pooling layer, a first fully connected layer FC1, a first fully connected layer FC2, and a Multiply weighted network layer; The third convolutional layer and the fourth convolutional layer respectively perform feature transformations of different scales on the feature output by the segmentation layer to obtain the corresponding feature maps and respectively; The Concat connection layer concatenates the feature maps and to obtain the feature Y; The channel attention of the feature Y is calculated through the GAP pooling layer, the first fully connected layer FC1, and the first fully connected layer FC2 to obtain the channel attention weight vector ; Finally, the Multiply weighted network layer is used to perform channel reweighting on the feature Y to obtain the feature with enhanced attention. .
[0013] Optionally, the expression of the channel attention weight vector is as follows:
[0014] In the formula, GAP represents global average pooling; and represent the connection operations of the first fully connected layer FC1 and the first fully connected layer FC2 respectively; is the Sigmoid function; is the RELU function; indicates that the shape of the vector s is one-dimensional and the length of this dimension is C, where C represents the number of channels of the image.
[0015] Optionally, the loss function used for training and optimizing the YOLOv8s network is: DIoU loss function.
[0016] Optionally, a Transformer encoder based on the attention mechanism is used to perform feature encoding on the identified small target defects.
[0017] Optionally, the expression for calculating the Euclidean distance between the defect feature and the historical small target defect feature currently updated in the feature library is as follows:
[0018] In the formula, represents the Euclidean distance between the defect feature and the historical small target defect feature currently updated ; , , represents the dimension of the feature, and are respectively the data of the i-th dimension in the feature and the feature .
[0019] Optionally, the update steps of the historical small target defect feature include: Multiple small target pictures are obtained and encoded respectively to obtain their corresponding original defect features; The mean value of the original defect features corresponding to each small target picture is extracted, and the feature obtained by mean value extraction is used as the initial feature of the historical small target defect feature in the feature library; Each time a new defective feature is cached in the feature library, weighted averaging is performed on the initial feature and the new defective feature to obtain an updated historical small target defective feature.
[0020] In a second aspect, the present invention provides a small target recognition system based on cascaded hierarchical detection and self-comparison, which runs the steps of the method according to any one of the first aspects of the present invention. The system includes: An acquisition and annotation unit for acquiring aerial images of drone inspections and annotating the aerial images; A detection and cropping unit for inputting the annotated images into a device recognition module for device area detection and image cropping, and obtaining region images that only contain areas with small targets; An identification and encoding unit for inputting the region images into a cascaded small target defect recognition module for defect recognition, and performing feature encoding on the recognized small target defects to obtain encoded defect features; A calculation and discrimination unit for calculating the Euclidean distance between the defect feature and the currently updated historical small target defect feature in the feature library. If the Euclidean distance is greater than a preset threshold, it means that the defect feature is a false detection and will be excluded; if it is not greater than the preset threshold, it means that the defect feature is correct, outputs the corresponding small target defect recognition result, and caches the defect feature in the feature library.
[0021] In a third aspect, the present invention provides a terminal, including a processor and a storage medium; The storage medium is used for storing instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first aspects of the present invention.
[0022] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to any one of the first aspects of the present invention.
[0023] The beneficial effects of the present invention are as follows. Compared with the prior art: 1. Through the linkage of the three parts of the device recognition module, the small target defect recognition module, and the caching and comparison of the feature library, the present invention splits the detection process into two stages: rough positioning and precise recognition, avoiding the omission or misjudgment of small targets in one-time detection. By means of region of interest cropping and secondary detection, the proportion of small targets in the image is increased, and the types of small target defects in the distribution network scenario can be effectively recognized, improving the detection accuracy.
[0024] 2. The device recognition module and the small target defect recognition module of the present invention both use the YOLOv8s network for detection and recognition; they are improved based on the original YOLOv8 network structure to enhance the model's perception ability for small targets; mainly in the Neck feature enhancement part, all the original C2f modules are replaced with C2f-att modules with an attention mechanism to strengthen the cross-channel and spatial information fusion ability, making the model more sensitive to the boundaries of small targets, especially being able to more effectively capture fine-grained target information in low-resolution or dense scenarios.
[0025] 3. The C2f-att module provided by the present invention further improves the structure of the EMA attention layer on the basis of the original C2f module to strengthen the discriminative ability of features, and the feature matching mechanism enhances the system's continuous learning ability and generalization ability; by improving the structure of the EMA attention layer, higher attention weights are assigned to small targets, making it pay more attention to small targets during the feature processing and fusion stages compared to before, thus retaining more shallow information, achieving more effective positioning and recognition, and having good recognition scalability.
[0026] 4. The present invention significantly reduces the false detection rate in complex backgrounds and enhances the long-term stability of the system by dynamically updating the historical small target defect features in the feature library and further discriminating the recognition results of the small target defect recognition module using the Euclidean distance.
[0027] 5. The present invention constructs a sample set for model training using the annotated images. Through the cooperation of high-quality data construction and fine annotation, the efficiency and accuracy of model training are optimized. At the same time, using the DIoU loss function is more conducive to the convergence of the model, improving the accuracy of small target detection and enhancing the stability and application effect in actual inspection scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic flow chart of the small target recognition method based on cascaded hierarchical detection and self-comparison in the present invention; Figure 2 It is a flow chart of the device recognition module and a schematic structure diagram of the target detection module in the present invention; Figure 3 It is a schematic structure diagram of the traditional YOLOv8 network; Figure 4 It is a schematic structure diagram of the YOLOv8s network improved in the present invention; Figure 5 It is a schematic structure diagram of the C2f-att network in the present invention; Figure 6 It is a schematic structure diagram of the improved EMA network in the present invention; Figure 7This is a schematic diagram of the structure of the Detect head of YOLOv8s in the present invention; Figure 8 This is a comparative diagram showing the recognition effects of YOLOv8 and YOLOv8s with the example of the insurance pin falling off recognition in the present invention; Figure 9 This is a comparative diagram showing the recognition effects of YOLOv8 and YOLOv8s with the example of the screw breakage and falling off recognition in the present invention; Figure 10 This is a structural principle block diagram of the small target recognition system based on cascade level detection and self-comparison in the present invention. Specific implementation mode
[0029] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The embodiments described in the present invention are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] Embodiment 1: Refer to Figure 1 The embodiment of the present invention provides a small target recognition method based on cascade level detection and self-comparison, which specifically includes the following steps: Step 1: Collect the aerial images of the drone inspection and annotate the aerial images; Specifically, based on the original aerial images of the drone inspection, the images are annotated by means of manual recognition. The annotation content covers the adjustment of the image size of the target box, the selection of the position box, and the category information of the equipment and small targets.
[0031] Step 2: Input the annotated images into the equipment recognition module for equipment area detection and picture cropping to obtain area pictures that only contain areas with small targets; Refer to Figure 2 As shown, the equipment recognition module includes a target detection module and a picture cropping module; Among them, the target detection module detects the equipment areas where small targets may appear in the input pictures to obtain the region of interest box images; the picture cropping module crops the detected region of interest images to obtain area pictures that only contain areas with small targets.
[0032] In this embodiment, the image cropping module uses a sliding window mechanism for regional cropping. In this embodiment, the window size is defined as 896×896, and the original image resolution is not scaled. According to the marking information, a local device image containing small targets is cropped from the whole image, that is, the device area where small targets may appear is first identified in the original image, and then the device area is cropped to retain more small target information.
[0033] Step 3: Input the regional image into the cascaded small target defect recognition module for defect recognition, and perform feature encoding on the recognized small target defects to obtain the encoded defect features. Step 4: Calculate the Euclidean distance between the defect features and the currently updated historical small target defect features in the feature library. If the Euclidean distance is greater than the preset threshold, it means that the defect features are misdetected and will be excluded; if it is not greater than the preset threshold, it means that the defect features are correct, output the corresponding small target defect recognition results, and cache the defect features in the feature library.
[0034] In a preferred but non-limiting embodiment, the device recognition module consists of a YOLOv8s network; a device image sample set is constructed based on the labeled historical images to train and optimize the YOLOv8s network to obtain the device recognition module. Among them, the construction of the sample set includes: intercepting single-frame images that meet the recognition scenario from the UAV aerial video as the data set, dividing the training set and the test set according to requirements, and at the same time correctly annotating the intercepted data. After correctly selecting the data and processing the data format, an image sample set is obtained.
[0035] In a preferred but non-limiting embodiment, the small target defect recognition module also consists of a YOLOv8s network. Specifically, the labeled historical images are input into the trained device recognition module, and a regional image data set containing only the areas with small targets is output; a small target image sample set is constructed based on the regional image data set to train and optimize the YOLOv8s network to obtain the small target defect recognition module.
[0036] It should be noted that the method provided by the present invention adopts a more targeted data processing and modeling process. First, based on the original UAV inspection images, manual recognition and annotation are used to directly annotate the original images, that is, the equipment areas and small target areas where small targets may exist are manually annotated in the original images, so as to construct a high-quality image dataset containing small target labels. Subsequently, in the training stage, according to the information of the annotated equipment areas, the areas of the original images that only contain small targets can be separately cropped. In this way, during training, the original images and the cropped images are input into the model for training together. Such annotation and training strategies not only help the model accurately identify various small targets, but also enable the model to have stronger robustness and adaptability in complex scenarios. At the same time, the present invention uses the annotated small target detection dataset specific to the distribution network equipment environment for model training. Because the general distribution network equipment in the distribution network equipment environment is relatively large in volume (such as power generation towers), the volume ratio of small targets (such as screws on power generation towers) in this environment is even smaller, and the recognition difficulty is greater. Therefore, the data of this scenario is annotated and specifically applied and trained. Finally, the model can be more focused on the feature extraction and recognition tasks of tiny targets, significantly improving the detection accuracy of small targets. Through the cooperation of high-quality data construction and fine annotation, the efficiency and accuracy of model training are optimized, the convergence speed of the model is accelerated, and the stability and application effect in the actual inspection scenario are improved.
[0037] Referring to Figure 3 and Figure 4 , the YOLOv8s network provided by the present invention is obtained by replacing the C2f feature fusion module in the neck network of the YOLOv8 network with a C2f-att module based on the attention mechanism; wherein, the C2f-att module includes a first convolutional layer Conv1 connected in sequence, a segmentation layer, an EMA attention layer, a Bottleneck deep feature extraction module, a Concat connection layer, and a second convolutional layer Conv2.
[0038] Referring to Figure 5 , the process of the C2f-att module when processing image features is as follows: For the feature T input into the C2f-att module, it is first input into the first convolutional layer Conv1 for channel adjustment to obtain the adjusted feature ; The segmentation layer divides the adjusted feature into two parts in terms of channels to obtain the feature and the feature ; the feature is input into the EMA attention layer to extract multi-scale feature information, and the output feature is ; The Bottleneck depth feature extraction module stacks the features through multiple Bottlenecks for multiple times to obtain depth features ; The Concat connection layer concatenates the features with the depth features and then obtains the feature ; The second convolutional layer Conv2 performs a convolution operation on the feature so as to adjust back to the number of channels and output the final feature F.
[0039] Specifically, for the input feature T, after being input into the C2f-att module, it is first input into the convolutional layer for channel adjustment to obtain the adjusted feature :
[0040] Then, the adjusted feature is input into the splitting layer to be divided into two parts in terms of channels, obtaining the feature and the feature :
[0041] Among them, is the main branch input into the subsequent EMA, is the shortcut branch directly participating in the subsequent Concat connection layer.
[0042] The main branch is input into the subsequent EMA attention layer for attention calculation:
[0043] Among them, is the output feature of the EMA attention layer; The output feature of the EMA is input into the Bottleneck depth feature extraction module of the main branch. The features are stacked through multiple Bottlenecks to obtain deep features : ; Among them, represents the number of layers of the Bottleneck; in this embodiment, N takes the value of 3.
[0044] Then, the Concat connection layer is used to concatenate the main branch and the shortcut branch to obtain the concatenated feature :
[0045] Finally, the second convolutional layer of Conv2 performs a convolution operation on the obtained result to adjust the number of channels back:
[0046] where F is the final output of the C2f-att module.
[0047] In a preferred but non-limiting embodiment, EMA (Efficient Multi-scale Attention) is originally a lightweight multi-scale attention mechanism. The main idea is to fuse multi-scale feature information and perform channel attention modeling, enabling the network to automatically focus on more useful feature regions (especially small targets) while maintaining computational efficiency. The present invention has made structural improvements to the EMA structure to adapt to the model of the present invention, such as Figure 6 The overall structure of the improved EMA attention layer is shown as including a third convolutional layer, a fourth convolutional layer, a Concat connection layer, a GAP pooling layer, a first fully connected layer FC1, a first fully connected layer FC2, and a Multiply weighted network layer; Among them, the third convolutional layer and the fourth convolutional layer respectively perform feature transformations of different scales on the features output by the segmentation layer to obtain corresponding feature maps respectively and ; the Concat connection layer concatenates the feature maps and to obtain feature Y; the channel attention of feature Y is calculated through the GAP pooling layer, the first fully connected layer FC1, and the first fully connected layer FC2 to obtain the channel attention weight vector ; finally, the Multiply weighted network layer performs channel re-weighting on feature Y to obtain the feature with enhanced attention .
[0048] Specifically, first the feature map obtained by the segmentation layer after channel division is input into the EMA attention layer, and the input feature map is divided into multiple scales; then, through two convolutional layers perform per-scale feature transformations, allowing each sub-channel to pass through a separate convolution operation to capture receptive fields of different scales:
[0049] Among them, represents the feature maps corresponding to two different scale branches, represents that the convolutional kernel size is The convolution operation indicates that the convolution kernel size is of the convolution operation.
[0050] After that, the output is fed into the Concat connection layer for splicing and channel attention weighting, and the feature maps of each scale and are spliced to obtain the feature Y, and the result is input into the subsequent GAP layer and fully connected layer to calculate the channel attention:
[0051]
[0052] In the formula, GAP represents global average pooling, and respectively represent the connection operations of the first fully connected layer FC1 and the first fully connected layer FC2, is the Sigmoid function, is the RELU function, indicates that the shape of the vector s is one-dimensional and the length of this dimension is C. Similarly, indicates that the shape of the vector Y is three-dimensional, and the lengths of the three dimensions are C, H, and W respectively, corresponding to the number of channels, width, and height of the image.
[0053] Finally, it is input into the Multiply weighted network layer for channel re-weighting to obtain the weighted feature
[0054] The feature with enhanced attention is sent back to the main path of C2f-att for further processing.
[0055] A preferred but non-limiting embodiment uses a Transformer encoder based on the attention mechanism to encode the features of the identified small target defects.
[0056] Next, with reference to Figure 4 the recognition process of the YOLOv8s network of the present invention is further described. The image is first input into the backbone of YOLOv8s to extract image features, and by Figure 4The main process shown is as follows: The original image Image is input into the first convolutional layer of the backbone network for an initial convolution operation. In the second convolution, it is equivalent to performing a downsampling operation. Then, it is input into the C2f layer for the first feature extraction operation. After that, it is input into the third and fourth convolutional layers again for further downsampling, and the result is output to the second C2f layer to obtain higher-level features. Finally, it is input into the fifth convolutional layer and the third C2f layer for downsampling and obtaining higher-level features, and the final result is input into the sppf module for multi-scale context aggregation. Among them, the feature processing method inside the sppf is: Start performing 3 consecutive pooling operations on the input to obtain three feature maps. The input feature map is X:
[0057]
[0058]
[0059] Here, is the feature map obtained after the pooling operation, represents the convolutional kernel size, represents the stride, represents the padding.
[0060] Then, a concatenation operation is performed. The original input and the three pooling results are concatenated along the channel dimension to obtain the feature :
[0061] Finally, convolution fusion is performed to obtain multi-scale context information
[0062]
[0063] The image features processed by the backbone network are subsequently input into the neck network for further processing. Among them, the main process of the neck network is as follows: the high-level semantic features output by sppf are first upsampled through the Upsample layer, and then concatenated with the features output by the C2f layer from a shallower layer through the Concat layer. After that, the merged vector is input into the C2f-att module for fusion processing. Here, since an attention mechanism is added to C2f, the parameter weights can be continuously learned and optimized, and higher weights are added to small target parts to make the model pay more attention to small targets. After fusion, it continues to be upsampled through the Upsample layer, and is merged with the features of the shallowest layer through Concat, and then feature fusion is performed through the C2f-att module again to form the top layer of the feature pyramid. Subsequently, the features here are divided into two paths for output. One path is directly input into the Detect detection head for prediction, and the other path is input into the convolutional layer, which is equivalent to downsampling and feature extraction. The features obtained here are merged with the features output by the previous C2f-att in the Concat module, and the merged and concatenated features are input into the next C2f-att. After being processed by this C2f-att, the output is again divided into two paths. One path is input into the second Detect detection head, and the other path continues to be output to the following convolutional layer to obtain deeper feature information. This feature is continuously input into the following Concat layer and concatenated with the output of SPPF to obtain new features. This part of the features is finally input into the C2f-att layer and ultimately input into the third Detect detection head for prediction.
[0064] The three layers of features from the neck network are respectively sent into the three Detections of the head detection head. The detection head of YOLOv8s is consistent with the original detection head structure of YOLOv8. Its main workflow is still to first perform a convolution operation, process the input feature map through the convolutional layer to obtain the class probability distribution at each position. Here, the sigmoid activation function is mainly used to normalize the classes predicted by each grid. Then, a bounding box regression operation is performed, and the model will regress to obtain the bounding box coordinates of the center of each grid, including the offset of the center point, width, and height. Here, the sigmoid activation function is still used to normalize the position of the box so that the output can be directly mapped to the original image coordinate system. Finally, each predicted box will generate a confidence value, indicating the probability that the box contains the target, that is, whether the box contains an object. Each of these three Detect detection heads completes predictions at different scales, and finally the multi-scale detection results are concatenated and output.
[0065] A preferred but non-limiting embodiment. In the processing of the loss function, this model optimizes the loss function to DIoU. Compared with the traditional IoU Loss which only focuses on the overlapping region of two bounding boxes, for two bounding boxes that have no overlap but are very close in center, IoU is still 0 and cannot provide gradients, which will greatly affect the training effect of small object recognition. In this embodiment, DIoU introduces the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box, encouraging the predicted bounding box to converge to the ground truth bounding box faster, even if they do not overlap initially. The calculation formula of the DIoU loss function is as follows:
[0066] In the formula, the predicted bounding box , the ground truth bounding box ; is the ratio of the intersection to the union of the two bounding boxes, where is the diagonal length of the smallest circumscribed box containing the predicted bounding box and the ground truth bounding box, is the square of the diagonal length of the smallest circumscribed rectangle containing the predicted bounding box and the ground truth bounding box; is the square of the Euclidean distance between the center of the predicted bounding box and the center of the ground truth bounding box, and its calculation formula is:
[0067] The square of the Euclidean distance of the center points in this formula introduces the DIoU loss function. By adding the center distance term, it forces the predicted bounding box to quickly approach the center of the ground truth bounding box and prevents small object deviation; at the same time, since small objects have a small area and are sensitive to position, this method will pay more attention to the position error and improve the detection accuracy.
[0068] For the Detect detection head, its structure is as Figure 7 shown, mainly including two branches. One part is for class prediction, and the other part is for the regression of the bounding box and the calculation of the DIoU loss. When the features output by the neck network are input into the Detect detection head, the features will be divided into two parts. Its ultimate goal is to perform class prediction and bounding box regression respectively. Both parts will first go through two consecutive Conv convolutional layers to further extract features and perform non-linear transformation, so as to enhance the expression ability. Finally, through Conv2D, its function is similar to a linear output layer, which is used to generate the final prediction result, finally obtain the bounding box and calculate the DIoU loss to optimize the parameters, and at the same time obtain the class prediction cls.
[0069] After being processed by the above-mentioned object detection module, the model will obtain the device area information where small targets may appear. Subsequently, it will be input into the image cropping module of the device recognition module. This module adopts a sliding window strategy for area cropping. Here, the window size is defined as 896×896, and the original image resolution is not scaled. That is, in the original image, first identify the device area where small targets may appear, and then crop this device area to retain more small target information. Therefore, after passing through the device recognition module, the picture will first be input into the object detection module, namely the YOLOv8s network, for detection. The YOLOv8s network will detect the area containing small targets, and input the output picture with markings into the image cropping module. This module adopts a sliding window mechanism and crops the local device picture containing small targets from the entire picture according to the marking information.
[0070] Subsequently, the device picture will be input into the small target defect recognition module. This module is mainly composed of an improved YOLOv8s model. The device area picture will pass through the backbone, neck, and head of YOLOv8s in sequence, and finally obtain the picture and annotation information of small target defects. The recognition effect of the obtained picture of small target defects here is better than that of common existing small target detection models and the YOLOv8 model. This is mainly because small targets usually occupy a small area of the image and have weak features, and are easily "submerged" during the downsampling process. Moreover, during small target detection, the background or other large targets are prone to false detection. The present invention considers that the neck of YOLOv8 is the core of multi-scale feature fusion. The C2f-att module adds an attention mechanism in the C2f feature fusion module without increasing too much computational complexity, and the attention mechanism can automatically enhance the response intensity of the small target area according to the context, enabling subsequent layers to pay more attention to small targets. At the same time, it can suppress irrelevant areas and improve the "signal-to-noise ratio" of the small target area, thereby improving the model recognition effect.
[0071] The detected small target defects will be encoded using an image encoder. In this embodiment, a Transfomer encoder is selected because it uses a self-attention mechanism, which can allow each position to interact with all other positions, capture global dependencies. At the same time, since there will be subsequent similarity calculations, which are also similarity calculations between small targets, Transformer can better fuse context information into the encoding and improve the discriminative power. The calculation method of the Transformer encoder is as follows: For a feature map block corresponding to a small target area , after being linearly mapped into a token sequence ; First, perform Patch Embedding and Positional Encoding. After dividing the picture into multiple small intervals (patches), the following calculations are carried out:
[0072]
[0073] Among them, the linear matrix , the position encoding , the features represented by each Patch , the encoded features .
[0074] After that, through the multi-head attention and feed-forward network of the Transformer encoder, the output for the k-th layer is:
[0075] Among them, is the output of the k-th layer, FFN represents the feed-forward neural network, and in addition, MSA is the multi-head attention mechanism, and its calculation formula is:
[0076] Among them, respectively represent the three feature matrices of query, key, and value in the attention calculation process For each head h, there is:
[0077] Among them, The matrix is used to calculate the corresponding matrix.
[0078] Then, after concatenating multiple heads and then mapping:
[0079] Finally, perform Global Pooling and Embedding Projection to generate the final encoding:
[0080]
[0081] The output dimension of the Transformer encoding , the LayNorm layer is LN, and its N is the number of patches.
[0082] Calculate the Euclidean distance between the feature vector encoded by the Transformer encoder and the feature vector in the pre-constructed target feature library. If the distance exceeds the set threshold, it is considered that the target is a new target or a changing target, which means that the detected defect is a false detection, and the target will be excluded; if the recognition is successful and the confidence is high (not lower than the set threshold), add the feature to the feature library and dynamically update the recognition system to enhance the memory and expansion capabilities of the model. Among them, the expression of the Euclidean distance is as follows:
[0083] In the formula, represents the defect feature and the defect feature of the current updated historical small target The Euclidean distance of; , , represents the dimension of the feature, and are respectively the features and feature The data of the i-th dimension in.
[0084] A preferred but non-limiting embodiment, the update step of the historical small target defect feature in step 4 includes: Step 4.1: Obtain multiple small target pictures and encode them respectively to obtain their corresponding original defect features; Step 4.2: Extract the mean value of the original defect features corresponding to each small target picture, and use the feature obtained by the mean value extraction as the initial feature of the historical small target defect feature in the feature library; Step 4.3: Each time a new defect feature is cached in the feature library, perform weighted average processing on the initial feature and the new defect feature to obtain the updated historical small target defect feature.
[0085] Specifically, the construction method of the feature library is that in the initialization stage of the feature library, multiple small target pictures are manually selected and input into the encoder for encoding to obtain the corresponding feature vectors, and then input into the feature library, which is equivalent to initializing the feature library with these multiple feature vectors. Subsequently, as the number of input pictures and vectors increases during the training process, new vectors are continuously added to the feature library according to the Euclidean distance and the threshold.
[0086] It should be noted that in the process of feature comparison and feature library establishment in the embodiments of the present invention, a multi-stage feature aggregation and dynamic weight fusion mechanism and an Euclidean distance dynamic threshold determination mechanism for false detection suppression are introduced. The dynamic weight fusion mechanism is mainly reflected in the feature comparison module designed in the present invention. The feature library not only caches historical features, but also aggregates features by using a staged weighted average mechanism. Among them, the mean values of the basic defect features and the historical cached features are extracted respectively. When fusing, unequal weighting coefficients are introduced (such as 1.0, 0.95, 0.95, that is, the weighting coefficient of the initial feature of the historical small target defect feature is 1.0, and the weighting coefficients of the first and second cached features are both 0.95), and finally, an adaptive correction is carried out through a normalization coefficient (such as 2.9). Compared with the existing simple feature averaging scheme, this structure can suppress the drift of dynamic cached features while maintaining the semantic expression ability of basic defects, effectively improving the comparison stability and recognition robustness. For the Euclidean distance dynamic threshold determination mechanism for false detection suppression, when comparing defect features with the feature library, instead of using a fixed threshold to judge false detection, the comparison threshold range is adaptively adjusted based on the offset trend of the current feature clustering center. If it is found that the trend in the historical feature set changes, the threshold range is dynamically shrunk to avoid false positive phenomena caused by "feature drift". The specific calculation process mainly relies on the feature library which is mainly used to cache the defect features of each small target, and the dimension of the feature is 1 x (K + N) x 1024, where K is the feature of the basic small target and N is the number of subsequently added cached small targets. Before each Euclidean distance calculation, the features of all small targets need to be merged into a cluster feature, which represents the position of this feature cluster of the small target. If the distance between the newly detected feature and this feature cluster in the feature space is less than a preset threshold, it means that the newly detected feature belongs to this feature cluster. When merging, the mean value is taken for the first K features, and the mean value is taken for the last N - 1 features. This adaptive threshold strategy can significantly reduce the false detection rate in complex backgrounds and enhance the long-term stability of the system in actual deployment.
[0087] In summary, taking the insurance pin falling off as an example, this embodiment further illustrates the overall recognition process as follows: After the drone receives an image, such as Figure 1As shown, first, enter the device recognition module to obtain the positions of devices where small target areas may appear, such as cross arms and insulators. Then, crop the detected device areas, and then input the cropped images into the cascaded small target defect recognition module. This solution enables the small target defect recognition model to recognize smaller targets at a larger resolution size, and this targeted design is more efficient than only outputting the region of interest. The region of interest often contains a large number of useless or highly overlapping candidate regions. Next, the detected small target defects will be encoded using a Transformer encoder. Then, the encoded defect features will be compared with the historical small target defect features stored in the feature library using the Euclidean distance. If the Euclidean distance is greater than the preset threshold, it means that the detected defect is a false detection and will be eliminated. If it is less than the preset threshold, it means that the detection is correct, the defect recognition result will be output, and the defect features will be cached in the feature library.
[0088] Figure 8 And Figure 9 The traditional YOLOv8 and the improved YOLOv8s are used to respectively show the overall recognition effect diagrams of two small target defect detection tasks, showing the recognition effect diagrams of the insurance pin and the screw breakage and detachment, the recognition of the region of interest and the recognition of defects for the region of interest. For the recognition effect diagram of the YOLOv8 model, the model will detect many regions of interest where small objects may exist in the picture, but in fact, there are no small targets to be recognized inside these regions. At the same time, when the picture of the region containing the small target is separately input into the YOLOv8 model, the model will not recognize the position of the small target. For the model of the present invention, first, the YOLOv8s model will recognize which regions may contain small targets and will separately recognize the regions of interest where small targets may appear, so as to detect the specific positions of the small targets. By comparing the effect diagrams of the two models, it can be seen that the model architecture and optimization method proposed by the present invention can obtain better results in small target defect detection.
[0089] The beneficial effects of the present invention are as follows. Compared with the prior art: 1. Through the linkage of the three parts of the device recognition module, the small target defect recognition module, and the caching and comparison of the feature library, the present invention splits the detection process into two stages: rough positioning and precise recognition, avoiding the omission or misjudgment of small targets in one-time detection. By means of interested region cropping and secondary detection, the proportion of small targets in the image is increased, and the types of small target defects in the distribution network scenario can be effectively recognized, improving the detection accuracy.
[0090] 2. The device recognition module and the small target defect recognition module of the present invention both use the YOLOv8s network for detection and recognition. They are improved based on the original YOLOv8 network structure to enhance the model's perception ability for small targets. Specifically, in the Neck feature enhancement part, all the original C2f modules are replaced with C2f-att modules with an attention mechanism to strengthen the cross-channel and spatial information fusion ability, making the model more sensitive to the boundaries of small targets, especially in low-resolution or dense scenarios, where it can more effectively capture fine-grained target information.
[0091] 3. The C2f-att module provided by the present invention further improves the structure of the EMA attention layer on the basis of the original C2f module to enhance the discriminative ability of features. The feature matching mechanism enhances the system's continuous learning ability and generalization ability. By improving the structure of the EMA attention layer, higher attention weights are assigned to small targets, enabling it to pay more attention to small targets during the feature processing and fusion stages compared to before, thereby retaining more shallow information, achieving more effective positioning and recognition, and having good recognition scalability.
[0092] 4. The present invention dynamically updates the historical small target defect features in the feature library and uses the Euclidean distance to further discriminate the recognition results of the small target defect recognition module, which can significantly reduce the false detection rate in complex backgrounds and enhance the long-term stability of the system.
[0093] 5. The present invention uses the labeled images to construct a sample set for model training. Through the cooperation of high-quality data construction and fine annotation, the efficiency and accuracy of model training are optimized. At the same time, the DIoU loss function is adopted, which is more conducive to the convergence of the model, improves the accuracy of small target detection, and enhances the stability and application effect in actual inspection scenarios.
[0094] Embodiment 2: As Figure 10 shown, the present invention provides a small target recognition system based on cascaded hierarchical detection and self-comparison. The system is used to implement the steps of the method in Embodiment 1 above. Specifically, the system includes: An acquisition and annotation unit, which is used to acquire the aerial images of drone inspections and annotate the aerial images; A detection and cropping unit, which is used to input the annotated images into the device recognition module for device area detection and image cropping, and obtain the area images that only contain small targets; An identification and encoding unit, which is used to input the area images into the cascaded small target defect recognition module for defect recognition, and perform feature encoding on the identified small target defects to obtain the encoded defect features; A calculation and discrimination unit is configured to calculate the Euclidean distance between the defect feature and the currently updated historical small target defect feature in the feature library. If the Euclidean distance is greater than a preset threshold, it means that the defect feature is a false detection and will be excluded. If it is not greater than the preset threshold, it means that the defect feature is correct, the corresponding small target defect recognition result is output, and the defect feature is cached in the feature library.
[0095] The small target recognition system based on cascade hierarchical detection and self-comparison provided by the embodiments of the present invention and the small target recognition method based on cascade hierarchical detection and self-comparison provided by Embodiment 1 are based on the same technical concept, and can produce the beneficial effects as described in Embodiment 1. The content not described in detail in this embodiment can be referred to in Embodiment 1.
[0096] Embodiment 3: A terminal provided by an embodiment of the present invention includes a processor and a storage medium, and the terminal is an embedded computer system device. The storage medium of the terminal is used to store instructions, and the memory includes a non-volatile storage medium and an internal memory. Among them, the non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium, and the database is used to store instruction data. The processor of the terminal is configured to operate according to the instructions provided by the storage medium to execute the steps of the small target recognition method based on cascade hierarchical detection and self-comparison according to any one of Embodiment 1.
[0097] Embodiment 4: A computer-readable storage medium provided by an embodiment of the present invention stores a computer program thereon, and when the program is executed by a processor, the steps of the method according to any one of Embodiment 1 are implemented.
[0098] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.
[0099] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0100] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0101] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific embodiments of the present invention. Any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A small target recognition method based on cascaded hierarchical detection and self-comparison, characterized in that the method Including: Collecting the aerial images captured by the inspection drone and annotating the aerial images; Inputting the annotated images into the device recognition module for device area detection and picture cropping to obtain area pictures containing only the regions with small targets; Inputting the area pictures into the cascaded small target defect recognition module for defect recognition, and performing feature encoding on the recognized small target defects to obtain the encoded defect features; Calculating the Euclidean distance between the defect features and the historical small target defect features currently updated in the feature library. If the Euclidean distance is greater than the preset threshold, it means that the defect feature is a false detection and will be excluded; If it is not greater than the preset threshold, it means that the defect feature is correct, outputting the corresponding small target defect recognition result, and caching the defect feature into the feature library.
2. The small target recognition method based on cascade hierarchical detection and self-comparison according to claim 1, wherein The content of annotating the qualified images includes: adjusting the image size of the target box, selecting the position box, and the category information of the device and small targets.
3. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 1, wherein The device recognition module includes a target detection module and a picture cropping module; The target detection module detects the device areas where small targets may appear in the input pictures to obtain the region of interest box images; The picture cropping module crops the detected region of interest images to obtain area pictures containing only the regions with small targets.
4. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 3, characterized in that, Both the device recognition module and the small target defect recognition module are composed of the YOLOv8s network; Constructing a device image sample set based on the annotated historical images to train and optimize the YOLOv8s network to obtain the device recognition module, and outputting the corresponding area picture dataset containing only the regions with small targets; Constructing a small target image sample set based on the area picture dataset to train and optimize the YOLOv8s network to obtain the small target defect recognition module.
5. The small target recognition method based on cascade hierarchical detection and self-comparison according to claim 4, characterized in that, The YOLOv8s network is obtained by replacing the C2f feature fusion module in the neck network of the YOLOv8 network with the C2f-att module based on the attention mechanism; The C2f-att module includes a first convolutional layer Conv1, a segmentation layer, an EMA attention layer, a Bottleneck deep feature extraction module, a Concat connection layer, and a second convolutional layer Conv2, which are connected in sequence.
6. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 5, characterized in that The process of the C2f-att module when processing image features is as follows: For the feature T input into the C2f-att module, it is first input into the first convolutional layer Conv1 for channel adjustment to obtain the adjusted feature ; The segmentation layer divides the adjusted feature into two parts in terms of channels to obtain the feature and the feature ; The feature is input into the EMA attention layer to extract multi-scale feature information, and the output feature ; The Bottleneck depth feature extraction module stacks the feature multiple times through multiple Bottlenecks to obtain the depth feature ; The Concat connection layer concatenates the features with the deep features to obtain the feature ; The second convolutional layer Conv2 performs a convolution operation on the feature to adjust back the number of channels and output the final feature F.
7. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 6, wherein The EMA attention layer includes a third convolutional layer, a fourth convolutional layer, a Concat connection layer, a GAP pooling layer, a first fully connected layer FC1, a first fully connected layer FC2, and a Multiply weighted network layer; The third convolutional layer and the fourth convolutional layer respectively features output by the segmentation layer after performing feature transformations of different scales, respectively obtain corresponding feature maps and ; The Concat connection layer concatenates the feature maps and to obtain feature Y; the channel attention of feature Y is calculated through the GAP pooling layer, the first fully connected layer FC1 and the first fully connected layer FC2 to obtain the channel attention weight vector ; finally, the Multiply weighted network layer is used to perform channel reweighting on feature Y to obtain the feature with enhanced attention .
8. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 7, characterized in that The channel attention weight vector has the following expression: Wherein, GAP represents global average pooling; and respectively represent the connection operations of the first fully connected layer FC1 and the first fully connected layer FC2; is the Sigmoid function; is the RELU function; indicates that the shape of the vector s is one-dimensional and the length of this dimension is C, where C represents the number of channels of the image.
9. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 5, characterized in that The loss function used when training and optimizing the YOLOv8s network is: the DIoU loss function.
10. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 1, wherein Using a Transformer encoder based on the attention mechanism to perform feature encoding on the recognized small target defects.
11. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 1, characterized in that, The expression for calculating the Euclidean distance between the defect features and the historical small target defect features currently updated in the feature library is as follows: In the formula, represents the defect feature and the Euclidean distance from the historical small target defect feature updated currently ; , , represents the dimension of the feature, and are respectively the data of the i-th dimension in the feature and the feature .
12. The small target recognition method based on cascaded hierarchical detection and self-comparison according to claim 1 or 11, characterized in that The update steps of the historical small target defect features include: Obtaining multiple small target pictures respectively and performing encoding on them to obtain their corresponding original defect features; Performing mean extraction on the original defect features corresponding to each small target picture, and using the features obtained by mean extraction as the initial features of the historical small target defect features in the feature library. Each time a new defective feature is cached in the feature library, a weighted average process is performed on the initial feature and the new defective feature to obtain an updated historical small target defective feature.
13. A small target recognition system based on cascaded hierarchical detection and self-comparison, which runs the small target recognition method based on cascaded hierarchical detection and self-comparison according to any one of claims 1-12, characterized in that, The system includes: An acquisition and annotation unit, configured to acquire aerial images of drone inspections and annotate the aerial images; A detection and cropping unit, configured to input the annotated images into a device recognition module for device area detection and image cropping, so as to obtain area images containing only small targets; An identification and encoding unit, configured to input the area images into a cascaded small target defect identification module for defect identification, and perform feature encoding on the identified small target defects to obtain encoded defect features; A calculation and discrimination unit, configured to calculate the Euclidean distance between the defect features and the currently updated historical small target defective features in the feature library. If the Euclidean distance is greater than a preset threshold, it means that the defect features are misdetected and will be excluded; if it is not greater than the preset threshold, it means that the defect features are correct, output the corresponding small target defect identification results, and cache the defect features in the feature library.
14. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-12 are implemented.
Citation Information
Patent Citations
Substation equipment defect identification method based on cascade detection model
CN114627360A
Defect detection method for high-power LED module packaging and related equipment
CN119831996A
Vehicle width detection method, device and equipment based on Yolov8 and storage medium
CN120107331A
Cited By
Defect detection model training method, defect detection method and device
CN121392671A