Navigation mark type and health state detection method based on improved YOLOv9

By optimizing YOLOv9 using CGAN, Canny algorithm, affine transformation, and RT-DETR model, the problems of class imbalance and complex scene recognition in navigation mark detection are solved, achieving efficient detection of navigation mark types and multiple health states, and improving detection accuracy and system reliability.

CN121190880APending Publication Date: 2025-12-23长江上海航道处 +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511463974.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing technologies for detecting the health status of navigation aids suffer from class imbalance, making it difficult to effectively identify multiple health statuses. Furthermore, the navigation aid image dataset is too limited to adapt to complex and ever-changing real-world application scenarios.

Method used

Conditional Generative Adversarial Network (CGAN) is used to generate damaged standard samples. The stain dataset is expanded using the Canny algorithm, affine transformation algorithm is used to add skewed samples, and geometric transformation and image stitching methods are introduced. The YOLOv9 model is optimized by combining the AIFI module and Inner-CIoU loss function in the RT-DETR model.

Benefits of technology

It improves the accuracy and efficiency of navigation mark detection, enabling accurate identification of navigation marks under different conditions, reducing reliance on manual operation, enhancing system reliability, and expanding the scope of application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190880A_ABST
    Figure CN121190880A_ABST
Patent Text Reader

Abstract

The invention relates to a navigation mark type and health state detection method based on improved YOLOv9, and belongs to the technical field of computer vision and navigation mark inspection. According to the method, in order to solve the problem of insufficient navigation mark data sets, a shape-damaged navigation mark data set is generated by using a CGAN technology, a navigation mark data set containing stains is generated through expansion by using a Canny algorithm, and a skew navigation mark data set is increased by using an affine transformation algorithm; in addition, traditional data enhancement methods such as geometric transformation and image splicing are introduced to simulate the imaging effects of the navigation mark at various angles and positions, so that the diversity of the image is improved. And the improved YOLOv9 model is utilized to improve the recognition accuracy of the small-size navigation mark. According to the method, the high-efficiency performance is maintained, meanwhile, the detection capability on fine and complex features is greatly improved, and accurate recognition on the navigation mark under different conditions is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision and navigation beacon inspection, and relates to a navigation beacon type and health state detection method based on an improved YOLOv9. BACKGROUND

[0002] Navigation beacon is the abbreviation of navigation aid, which is a key water facility for guiding ships to sail safely, positioning, and indicating dangerous areas and channel boundaries. Its main function is to help ship drivers determine the ship's position and heading, and ensure the ship sails in safe waters. As a core infrastructure to ensure the safety of ship navigation, the normal working state of navigation beacon is crucial. Once the navigation beacon fails or is damaged, such as light out, buoy drift, marker body tilt or damage, it will not be able to provide accurate navigation information for ships, which will easily lead to the ship losing direction and deviating from the planned route, thereby significantly increasing the risk of major navigation accidents such as running aground, grounding, collision, etc. Therefore, timely and accurate monitoring and identification of the health state of navigation beacon is an important prerequisite for ensuring the smoothness of the channel and protecting the safety of life and property.

[0003] However, in the development of an automated navigation beacon health state identification system, several technical challenges are faced. First, in the raw data set used for model training, there is a serious class imbalance problem. The number of damaged, dirty, and tilted beacon samples is much smaller than that of normal beacons. This data imbalance will affect the training effect of the deep learning model, making it difficult to fully learn the abnormal features. Secondly, most existing data sets are composed of large target navigation beacon images taken at close range, while the actual application scenarios are complex and varied, requiring not only to handle close-up images, but also to effectively identify aerial navigation beacon images that are smaller in size and have more varied perspectives, which puts higher requirements on the generalization ability and multi-scale target detection ability of the model.

[0004] Some explorations have been made in the existing technology for the automated detection and identification of navigation beacons. For example: (1) Ge Haipeng proposed an EfficientNet-b0-based navigation beacon image classification algorithm and an improved model based on Faster-Rcnn navigation beacon detection algorithm in the paper "Intelligent Recognition Technology of Navigation Beacon Appearance State Based on Deep Learning", which focuses on intelligent recognition of the appearance state of navigation beacons, and divides the damage degree into four categories: basic shape retention, main structure damage, top beacon damage, and paint damage. However, it mainly focuses on the classification of physical damage of the appearance, and does not cover other health state dimensions.

[0005] (2) In his paper "Night Navigation Light Quality Recognition Technology Based on Deep Learning", Guo Jiaming proposes an improved YOLOv5 model for detecting the color information of navigation lights from complex backgrounds. This method combines the Deepsort algorithm to accurately locate the position of the navigation light, and uses a specially designed light flash classification model (Network F) to identify the different flashing frequencies of the navigation light. By integrating the obtained navigation light color and flashing frequency information, the final accurate detection of the navigation light quality is achieved. However, this method mainly targets night optical signals and cannot evaluate the appearance health status of the navigation mark during the day.

[0006] (3) In his paper "Navigation Mark Body Damage Detection Technology Based on Deep Learning", Zhang Can proposes a lightweight FMB-Yolo model specifically for detecting the damage of navigation mark bodies. This model effectively reduces the number of parameters and overall size, reducing operating costs and resource consumption without sacrificing performance. The research further deploys the FMB-Yolo model to edge devices and significantly speeds up the detection speed through TensorRT technology, enabling it to quickly analyze and process input data, significantly shortening processing time. However, the detection range of this method is limited to physical damage of the mark body.

[0007] (4) In the paper "Navigation Mark Inclination Detection Method Based on Machine Vision", Zhao Jian Sen et al. disclose the construction of an R-YOLO navigation mark inclination detection model and propose a lightweight detection method based on machine vision. Based on the YOLOv5 framework, the CSL (Circular Smooth Label) technology is introduced, which cleverly converts the regression problem of navigation mark inclination angle into a classification problem. This method also focuses on a single posture anomaly problem.

[0008] In summary, the existing technology has made certain progress in the automatic recognition of navigation mark health status, but there are generally limitations, i.e. most researches only detect a certain specific health status, such as appearance damage, night light quality or mark body inclination, lacking a unified solution that can comprehensively and comprehensively evaluate the health of navigation marks in multiple key dimensions such as color, shape, light quality, etc.

[0009] Therefore, there is an urgent need in the field for a more comprehensive navigation mark detection method to meet the demand for intelligent and refined monitoring of modern navigation support systems. SUMMARY

[0010] Therefore, the purpose of the present application is to provide a navigation mark type and health status detection method based on an improved YOLOv9, which effectively identifies and processes navigation mark images to improve the accuracy and efficiency of navigation mark detection. The following two problems are mainly solved: First, to solve the problem of insufficient navigation mark data sets. The Conditional Generative Adversarial Network (CGAN) technology is used to generate damaged sample images. The Canny algorithm is used for edge detection to expand the sample set with stains. Affine transformation algorithm is applied to increase the number of skewed navigation mark samples. In addition, traditional data enhancement methods such as geometric transformation and image stitching are introduced to simulate the imaging effect of navigation marks at various angles and positions, thereby improving the diversity of images.

[0011] Second, to improve the recognition accuracy and precision of aerial images. After optimizing the network for small target navigation mark recognition in aerial images, the improved YOLOv9 model can more effectively capture the detailed features of small targets, significantly improving the accuracy of small size navigation mark recognition. Through these optimizations, the model maintains high performance while significantly improving the detection ability of subtle and complex features, ensuring accurate recognition of navigation marks under different conditions.

[0012] To achieve the above purposes, the present application provides the following technical solutions: A navigation mark type and health state detection method based on improved YOLOv9 only needs to input pictures or videos into the system without manual intervention, and the system can automatically generate detection results. The method specifically includes the following contents: 1) Screening the data set: screening the collected navigation mark original data set to remove blurred, erroneous or incomplete data.

[0013] 2) Enhance the data set: The original data set has obvious deficiencies, with single data form, small quantity, and serious imbalance between normal navigation mark data and abnormal navigation mark data. Therefore, different enhancement strategies are used for different types of abnormal navigation mark data sets, including traditional enhancement methods and algorithm-based enhancement means.

[0014] The Conditional Generative Adversarial Network (CGAN) technology is used to generate damaged navigation mark data set; the Canny algorithm is used to expand the navigation mark data set containing stains; the affine transformation algorithm is used to increase the skewed navigation mark data set; traditional data enhancement methods such as geometric transformation and image stitching are used to simulate the imaging effect of navigation marks at various angles and positions.

[0015] 3) Divide the data set: divide the enhanced navigation mark data set into training set, validation set and test set according to the proportion; 4) Constructing a buoy health state detection model based on improved YOLOv9: In view of the problems that the buoy target in the aerial image is small, the pixel ratio in the image is small, the detailed features are easy to be lost, and the complex aerial background noise is large, the AIFI (Attention-based Intra-scale Feature Interaction) module in the RT-DETR (Real-Time DEtection TRansformer) model is used to replace the SPPFELAN module in the backbone network of YOLOv9, and Inner-CIoU is introduced as a loss function, so that different size targets can be self-adjusted.

[0016] 5) Training the model: input the training set into the buoy health state detection model based on improved YOLOv9 for training, evaluate the model performance through the validation set, and obtain the optimal model.

[0017] 6) Test: use the optimal model to predict the buoy image to be detected, and accurately identify the buoy type and health state in the image.

[0018] Further, the CGAN technology is used to generate the damaged buoy data set, specifically including: first, define the required hyperparameters for training, including the number of training rounds, batch size and learning rate; then, use transforms to preprocess the image, adjust the size, convert to a tensor and perform normalization operation; subsequently, load two data sets of normal buoys and abnormal buoys, and use DataLoader to create corresponding data loaders, including a normal buoy loader and an abnormal buoy loader; then, define the loss function, discriminator and generator; in the training stage, loop through the real damaged image batch in the abnormal buoy loader, generate a conditional label, and set the abnormal buoy label to 1; then train the discriminator, process real damaged images, generate fake damaged images and real normal images, and calculate the corresponding losses d_loss_real_broken, d_loss_fake_broken and d_loss_real_normal, and then obtain the total loss d_loss of the discriminator; after completing the discriminator training, the generator training is performed; after the whole training process, the enhanced damaged buoy data is finally generated.

[0019] Further, the CGAN learns through the adversarial training between the generator and the discriminator. The generator tries to generate samples that look real and meet the given conditions, while the discriminator tries to distinguish the generated samples from the real samples and judge whether they meet the corresponding conditions. During the training process, the generator and the discriminator constantly optimize their own parameters to improve their respective performances, and finally reach a balanced state. The input of the CGAN has two parts, one part is the random noise distribution, and the other part is the conditional label input. It introduces additional conditional information on the basis of GAN, and the application is an expansion on the basis of normal navigation marks. Then the damage is the introduced conditional information.

[0020] The generator is used for generating realistic samples, and the optimization goal of the generator is to find the minimum value so that the generated samples can deceive the discriminator, and the loss function of the generator is represented as:

[0021] wherein, is a random noise vector; is conditional information, used for guiding the generator to generate samples under specific conditions; is a generated sample; is the output probability of the discriminator to the generated sample; represents sampling based on the distribution and represents the expectation of the real data z, and the real data comes from the data distribution .

[0022] The discriminator is used for judging the authenticity of real samples or generated samples; the optimization goal of the discriminator is to minimize so that the discriminator can accurately distinguish real samples from generated samples, and the loss function of the discriminator is represented as:

[0023] wherein, is a real sample; is conditional information; is the output probability of the discriminator to the real sample, is the output probability of the discriminator to the generated sample; represents sampling from the real data joint distribution and . represents sampling based on the distribution and .

[0024] ​The optimization objectives of the generator and the discriminator are combined to obtain a comprehensive optimization objective:

[0025] wherein, is a game objective function; represents the expected loss of the discriminator on the real sample; represents the expected loss of the generator on the generated sample.

[0026] Further, the Canny algorithm is used to expand the buoy data set containing stains, specifically including: using the Canny algorithm to extract and segment the stains of the stained buoy, and then superimposing the segmented stains on the normal buoy to expand the buoy data set containing stains. First, the image is denoised by a Gaussian filter, and the formula is:

[0027] wherein, represents the Gaussian function value at the position of the two-dimensional plane with coordinates σ is the standard deviation; Then, the image gradient amplitude and direction are calculated, the horizontal and vertical direction gradients are calculated using the Sobel operator, the gradient amplitude , and the direction Non-maximum suppression is then performed to retain the local gradient maximum point, wherein, and are the gradients of the image in the horizontal and vertical directions respectively; finally, the final edge is determined through double-threshold detection and edge connection, the high threshold and the low threshold are used to screen edge points.

[0028] Further, an affine transformation algorithm is used to increase the skewed buoy data set, specifically including: the generation of the skewed buoy data image is also based on the processing of normal buoy images, the YOLOv9 original model is used to detect and extract the buoy body or buoy light in the input normal buoy image, and after extraction, the affine transformation is used to rotate a certain angle to obtain the skewed data image of the buoy body or buoy light.

[0029] Affine transformation is a linear mapping, which can be represented as a linear transformation (rotation, scaling, shearing) plus a translation. Affine transformation preserves the collinearity between points and the distance ratio between parallel lines. In two-dimensional space, affine transformation can be represented by the following matrix equation:

[0030] wherein, is the coordinate of the point before transformation; are the transformed point coordinates; are the coefficients of the linear transformation part, corresponding to rotation, scaling and shearing; are the axis and axis translation vectors, respectively representing axis and axis distances.

[0031] Further, a buoy health state detection model based on improved YOLOv9 is constructed: the buoy target in the aerial image is small, the pixel ratio in the image is small, the detailed features are easy to be lost, and it is greatly disturbed by complex aerial background noise. The AIFI module in the RT-DETR model is used to replace the SPPELAN module in the backbone network of YOLOv9, so as to improve the accuracy of the model. At the same time, the Inner-CIoU loss function is introduced to realize self-adjustment of different size targets, so as to better deal with the problem of target overlap and effectively reduce the problem of target false detection.

[0032] The SPPELAN module in the backbone network is replaced by the AIFI module, which specifically includes: placing the AIFI module at the last layer of the backbone network; the AIFI module relies on the multi-head self-attention mechanism to promote the fine interaction between feature maps within the same scale, and it works with the cross-scale feature fusion module (CCFM) based on CNN to form the core architecture of the RT-DETR model encoder; RT-DETR (Real-Time Detection Transformer) is a real-time target detection model based on the Transformer architecture.

[0033] The cross-scale feature fusion module (CCFM) based on CNN specifically includes: first, performing a flattening operation on feature maps of different sizes to flatten their spatial dimensions, thereby obtaining a two-dimensional tensor; then, performing a transpose operation to flexibly adjust the dimension order of the tensor according to the specific needs of subsequent operations, so that the dimension arrangement of the data adapts to the AIFI module.

[0034] Further, the AIFI module includes a multi-head self-attention mechanism, an FFN, a residual connection and layer normalization.

[0035] Further, the loss function Inner-CIoU is introduced to improve the detection ability of small targets; the calculation formula of the Inner-CIoU loss function is:

[0036] In the formula, is the linearized CIoU loss, is the intersection over union between the two groups of boundary boxes, Inner-IoU loss for inner region.

[0037] The present application has the advantages that the present application uses deep learning image processing technology to realize intelligent detection of the health status of navigation marks, aiming at the problem of low efficiency of artificial inspection of the health status of navigation marks. The following effects are obtained: 1) Constructing a navigation mark data set: through data enhancement, a series of data sets of various types of navigation marks in normal, damaged, stained and skewed conditions are constructed, providing a good foundation for subsequent navigation mark type and health status detection.

[0038] 2) Improve recognition efficiency: the improved YOLOv9 has significantly improved recognition rate in navigation mark recognition tasks compared to the original model, and can quickly and accurately identify navigation marks in different scenarios, providing strong technical support for maritime safety.

[0039] 3) Improve system reliability: the present application uses advanced image processing algorithms and machine learning techniques to reduce dependence on manual operation, thereby enhancing the reliability and stability of the entire system and reducing the possibility of human error.

[0040] 4) Expand the application range: although the present application is mainly aimed at navigation mark detection, the techniques and methods used are also applicable to other fields that require accurate image recognition and position calibration, and have wide application prospects.

[0041] Other advantages, objects and features of the present application will be described in the following specification to some extent, and will be apparent to those skilled in the art based on the study of the following, or can be taught from the practice of the present application. The objects and other advantages of the present application can be achieved and obtained by the following specification. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to make the purpose, technical scheme and advantages of the present application clearer, the preferred detailed description of the present application will be given below in combination with the drawings, in which: Figure 1 The improved YOLOv9-based navigation mark type and health status detection method flowchart provided by the present application; Figure 2 The improved YOLOv9 network structure block diagram. DETAILED DESCRIPTION

[0043] The advantages and effects of the present application can be easily understood by those skilled in the art from the description of the present application. The present application can also be implemented or applied by different specific embodiments, and various modifications or changes can be made based on different views and applications without departing from the spirit of the present application. It should be noted that the drawings provided in the following examples only illustrate the basic concept of the present application in a schematic manner, and the following examples and features in the examples can be combined with each other without conflict.

[0044] The drawings are only used for illustrative description, and the representation is only a schematic diagram, not a physical diagram, and cannot be understood as a limitation of the present application; in order to better illustrate the embodiments of the present application, some components in the drawings are omitted, enlarged or reduced, and do not represent the actual size of the product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings can be omitted.

[0045] The same or similar reference numerals in the drawings of the embodiments of the present application correspond to the same or similar components; in the description of the present application, it should be understood that if the terms "upper", "lower", "left", "right", "front", "back" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore the terms describing the positional relationship in the drawings are only used for illustrative description, and cannot be understood as a limitation of the present application, for those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0046] Please refer to Figures 1-2 The embodiment of the present application provides a buoy type and health state detection method based on improved YOLOv9. In the buoy health state detection, the data set used is derived from artificial shooting images, and the number of normal data set is much larger than that of abnormal data set, which has a significant imbalance problem. By expanding the buoy data set, the problem of data set imbalance caused by the absence of abnormal buoy data is effectively solved. At the same time, the improved YOLOv9 network recognition method is innovatively put forward, which realizes the accurate recognition of the type and health state of the buoy, and the specific implementation is as follows: Step 1: Environment preparation. Build a suitable running environment, use Python 3.8 and above versions. YOLOv9 and its dependent libraries can be installed by using git clone and pip commands.

[0047] Step 2: Screening of data set. The collected original data set of buoy is screened to remove the fuzzy, wrong or incomplete data, and the data quality is ensured.

[0048] Step 3: Enhancement of the dataset (1) Use conditional generative adversarial network (CGAN) technology to generate sample images of damaged beacons.

[0049] First, define the hyperparameters required for training, including the number of training rounds, batch size, and learning rate. Next, use transforms to preprocess the images, adjusting the size, converting to tensors, and normalizing. Then, load the normal beacon and abnormal beacon datasets and use DataLoader to create corresponding data loaders, including the normal beacon loader and the abnormal beacon loader. After that, define the loss function, discriminator, and generator. In the training phase, loop through the real broken image batch in the abnormal beacon loader, generate the conditional label, and set the abnormal beacon label to 1. Then train the discriminator, process real broken images, generate fake broken images, and real normal images, and calculate the corresponding losses d_loss_real_broken, d_loss_fake_broken, and d_loss_real_normal, and then get the total loss of the discriminator d_loss. After completing the discriminator training, train the generator. After the entire training process, the enhanced broken beacon data is generated.

[0050] (2) Perform edge detection using the Canny algorithm to extract the stains on the beacon, and then overlay them on clean beacons to expand the sample set containing stains. First, read the image containing stains, then frame the area where the stains are located in the image. Next, convert the framed image to a grayscale image, which facilitates subsequent edge detection operations. Then initialize the Canny edge detector and set the relevant parameters, and use the detector to perform edge detection on the grayscale image to obtain an image containing edge information of the stains. According to the detected edge image, create a binary mask that can clearly distinguish between the stain area and the non-stain area. Use this binary mask to accurately extract the stain area from the original image. Finally, overlay the extracted stain area on the clean image to complete the entire operation process.

[0051] (3) Affine transformation algorithm is used to rotate the recognized beacon body and beacon light at a certain angle, thereby expanding the sample quantity of skewed beacons and beacon lights. First, the original Yolov9 model is used to detect the target in the picture. The model will output the coordinates of the detection box and the category label, so as to clearly define the position of the beacon and the light. Then, according to the coordinates of the detection box, the beacon body or beacon light area is extracted from the original picture. After that, the Canny edge detection algorithm is applied to the extracted area to find the contour, and a mask is created according to the contour to extract the image area in the contour. Finally, the rotation matrix is defined, and the extracted beacon light or beacon body area is rotated through affine transformation, and then the rotated beacon light edge is spliced to the same position on the original picture, so as to obtain the skewed beacon light or skewed beacon body of the beacon picture.

[0052] (4) In addition, traditional data enhancement methods such as geometric transformation and image stitching are combined to effectively improve the diversity of images.

[0053] Step 4: Data set annotation. Use Labelimg software to annotate the data set, and divide it into training set, validation set and test set according to certain proportion; During the annotation process, whole annotation and local annotation are combined, the whole annotation is used to annotate the beacon category information, and the local annotation is used to annotate the local information such as beacon body and beacon light. According to the beacon specification and the requirements of relevant departments, the health status of the beacon is divided into 6 categories: normal beacon body, normal beacon light, damaged beacon body, skewed beacon body, beacon body stain, and skewed beacon light. According to the type of beacon, it is divided into 8 categories: red buoy, black buoy, white buoy, dangerous water area buoy, left and right navigation buoy, cross current buoy, special buoy, and anchor prohibited buoy.

[0054] Step 5: Configure the data set file. A yaml file needs to be created to configure the information of the data set. Create a new hangbiao.yaml file under the directory.

[0055] Step 6: Build a beacon health state detection model based on improved YOLOv9; Replace the SPPELAN module in the Bottleneck network with the AIFI module based on RT-DETR, and introduce Inner-CIoU as the loss function.

[0056] Replace the SPPELAN module in the main network with the AIFI module: AIFI is a key module of the RT-DETR model, which is a new model of real-time detection Transformer. The model introduces a hybrid encoder design to optimize the processing of multi-scale features by separating intra-scale interaction and cross-scale feature fusion. The AIIFI module is placed at the last layer of the backbone network, which can provide more rich and in-depth feature representation than traditional convolution methods, and the performance of the model is improved through the fusion of the two. The AIFI module relies on the attention mechanism to promote the fine interaction between feature maps within the same scale, and works with the cross-scale feature fusion module (CCFM) based on CNN to form the core architecture of the RT-DETR model encoder. The AIFI module improves the understanding and performance of the model for the target detection task, while maintaining high accuracy and reducing computational cost.

[0057] AIFI extracts high-level features in images through self-attention mechanisms, which can process specific data and capture global dependencies. In the process of beacon target detection and classification, the relationship between beacons and the environment around the inland is very important, and AIFI can perceive the global context and reduce the loss of detailed information. The framework is shown in Figure 2 Figure 2 , where s3, s4, s5 are three different sizes of feature maps. First, the flattening operation needs to be performed on these feature maps to flatten the spatial dimensions and obtain two-dimensional tensors. Then, the transpose operation is used to flexibly adjust the dimension order of the tensor according to the specific needs of subsequent operations, so that the dimension arrangement of the data adapts to the subsequent process. AIFI is part of the efficient hybrid encoder in RT-DETR, so after completing the above operations, it enters the relevant operation links of the encoder. The structure shown in the dashed box on the right is similar to the Transformer encoder, which is a highly efficient and complex feature processing unit. This unit performs deep processing on the input feature map through a series of precisely designed operations, aiming to extract more representative and discriminative feature information.

[0058] Improvement of loss function: Inner-CIoU loss function is used as the loss function of YOLOv9, which is an improvement of the traditional CIoU loss function. The introduction of the Inner-CIoU loss function improves the detection ability of small targets; the calculation formula of the Inner-CIoU loss function is:

[0059] In the formula, is the linearized CIoU loss, is the intersection over union between the two groups of bounding boxes, is the Inner-IoU loss of the internal region.​

[0060] (1) Construction of AIFI module The AIFI implementation mainly includes initialization, forward propagation, and two-dimensional sine-cosine position encoding. During initialization, AIFI inherits the TransformerEncoderLayer class in torch.nn module of PyTorch, without adding additional member variables. Its initialization parameters cover input feature channel number, feedforward network internal layer dimension, number of multi-head attention mechanism heads, dropout probability, activation function, and whether to apply LayerNorm before sub-layer, etc. During the forward propagation of AIFI, the input tensor of shape [B, C, H, W] is first flattened to [B, HxW, C] to adapt to the Transformer layer, then two-dimensional sine-cosine position encoding is constructed and added to the input tensor, then the parent class forward method is called to let the data pass through the multi-head attention mechanism and the feedforward neural network, which includes residual connection and normalization operation during the process, and finally the processed tensor is restored to [B, C, H, W] shape. The method of two-dimensional sine-cosine position encoding is to create a grid, calculate the sine and cosine values with different frequencies, and assign a unique encoding to each position to help the model learn the spatial relationship of the input data.

[0061] (2) Construction of Inner-CIoU loss function The bbox_iou function is constructed, which mainly calculates the intersection over union (IoU) between two sets of bounding boxes and supports multiple extended IoU calculation methods. The function first determines whether to use the inner region according to the Inner parameter. If it is used, the bounding box coordinates are first converted to (x, y, width, height) format, and then the bounding box coordinates of the inner region, the intersection and union areas are calculated; if not, the coordinate format is converted as needed before calculating the intersection and union areas. Then calculate the basic IoU value, and then according to whether CIoU, DIoU, GIoU, MDPIoU parameters are True, respectively calculate different types of extended IoU values. In loss tal_dual.py, the loss function is enabled, and Inner=True, CloU=True are set, that is, the Inner-CIoU loss function is used.

[0062] Step 7: Dataset division, model training. The buoy dataset after data augmentation is divided into test set, training set and validation set according to the ratio of 1:2:1. Set the path of the pre-trained model, input image size (imgsz=320), number of training rounds (epoch=300), etc. Run the train.py file in bash to train the model using the command.

[0063] Step 8: Change the parameters to get the optimal model according to the evaluation indexes such as recall rate, average precision mean, etc. The evaluation indexes only need to call the corresponding functions in sklearn and torchmetrics.

[0064] Step 9: After the model training is completed and the expected performance is achieved, the model is used to predict the images to be detected, accurately identifying the types and health status of the beacons in the images.

[0065] This method not only overcomes the problems of data imbalance and multi-scene recognition, but also realizes synchronous and efficient detection of beacon types and their multiple health states (including color, shape, light quality, etc.) within a single model framework.

[0066] Finally, it should be pointed out that the above examples are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and all should be included in the scope of the claims of the present application.

Claims

1. A method for detecting the type and health status of navigational aids based on an improved YOLOv9 system, characterized in that, The method includes: Filter the dataset: Filter the collected raw data of navigation marks and remove blurry, erroneous or incomplete data; Augmenting the dataset: A dataset of damaged navigation marks is generated using CGAN technology, where CGAN stands for Conditional Generative Adversarial Network; the Canny algorithm is used to augment the dataset with smudges; an affine transformation algorithm is used to add a dataset of skewed navigation marks; and geometric transformations and image stitching methods are used to simulate the imaging effects of navigation marks at various angles and positions. Dataset partitioning: The enhanced beacon dataset is divided into training, validation, and test sets according to a set ratio; A beacon health status detection model based on an improved YOLOv9 is constructed: the AIFI module in the RT-DETR model is used to replace the SPPFELAN module in the backbone network of YOLOv9, and Inner-CIoU is introduced as the loss function to achieve self-adjustment for targets of different sizes; where AIFI represents attention-based intra-scale feature interaction. Training the model: The training set is input into the improved YOLOv9-based beacon health status detection model for training. The model performance is evaluated through the validation set to obtain the optimal model. Test: Use the optimal model to predict the navigation mark images to be detected, and accurately identify the types and health status of the navigation marks in the images.

2. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 1, characterized in that, The CGAN technique is used to generate a dataset of damaged beacons. Specifically, the process includes: First, defining the hyperparameters required for training; then, using transforms to preprocess the images, resizing, converting them to tensors, and normalizing them; next, loading two datasets, one for normal beacons and one for abnormal beacons, and creating corresponding data loaders, including a normal beacon loader and an abnormal beacon loader; then, defining the loss function, discriminator, and generator; during the training phase, iterating through batches of real damaged images in the abnormal beacon loader, generating conditional labels, and setting the abnormal beacon label to 1; then training the discriminator, processing real damaged images, generating fake damaged images, and real normal images respectively, and calculating the corresponding losses d_loss_real_broken, d_loss_fake_broken, and d_loss_real_normal, thus obtaining the total loss d_loss of the discriminator; after completing the discriminator training, training the generator; after the entire training process, the enhanced damaged beacon data is finally generated.

3. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 1 or 2, characterized in that, CGAN learns through adversarial training between the generator and the discriminator; The generator is used to generate realistic samples, and the optimization goal of the generator is to find the minimum value. This allows the generated samples to fool the discriminator, and the generator's loss function... Represented as: in, It is a random noise vector; It is conditional information used to guide the generator in producing samples under specific conditions; These are samples generated by the generator; It is the output probability of the discriminator for the generated samples; Represents distribution pairs and sampling, This represents the expected value of the real data z, where the real data comes from the data distribution. ; The discriminator is used to determine the authenticity of a sample or generated sample; the optimization objective of the discriminator is to minimize... This enables the discriminator to accurately distinguish between real and generated samples; the discriminator's loss function... Represented as: in, These are real samples; It is conditional information; It is the output probability of the discriminator for the real sample. It is the output probability of the discriminator for the generated samples; Indicates the distribution of the real data from the joint distribution. and Perform sampling; Represents distribution pairs and sampling; Combining the optimization objectives of the generator and discriminator yields a comprehensive optimization objective: in, It is the objective function of the game; This represents the expected loss of the discriminator on the real samples; This represents the generator's expected loss on the generated samples.

4. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 1, characterized in that, The Canny algorithm is used to expand and generate a navigation mark dataset containing stains. Specifically, the Canny algorithm is used to extract and segment the stains from the stained navigation marks, and then the segmented stains are merged and superimposed onto normal navigation marks, thereby expanding the navigation mark dataset containing stains. First, a Gaussian filter is used to reduce noise in the image. The formula is: in, Represents the coordinates on a two-dimensional plane as The Gaussian function value at the location, σ Standard deviation; Next, the magnitude and direction of the image gradient are calculated. The Sobel operator is used to calculate the horizontal and vertical gradients, and the gradient magnitude is... ,direction Then, non-maximum suppression is performed, preserving the points where the local gradient is maximized. and These represent the gradients of the image in the horizontal and vertical directions, respectively; finally, the final edges are determined through dual-threshold detection and edge connection, with the high threshold... and low threshold Used to filter edge points.

5. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 1, characterized in that, An affine transformation algorithm is used to augment the skewed navigation beacon dataset. Specifically, the YOLOv9 original model is used to detect and extract the skewed ...

6. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 1, characterized in that, The SPPELAN module in the backbone network is replaced with the AIFI module. Specifically, the AIFI module is placed in the last layer of the backbone network. The AIFI module relies on a multi-head self-attention mechanism to promote fine interaction between feature maps at the same scale. It works in conjunction with the CNN-based cross-scale feature fusion module to form the core architecture of the RT-DETR model encoder. RT-DETR is a real-time object detection model based on the Transformer architecture.

7. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 6, characterized in that, The CNN-based cross-scale feature fusion module specifically includes: first, flattening feature maps of different sizes to flatten their spatial dimensions, thereby obtaining a two-dimensional tensor; then, performing a transpose operation, and flexibly adjusting the dimensional order of the tensor according to the specific needs of subsequent operations, so that the dimensional arrangement of the data can be adapted to the AIFI module.

8. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 6, characterized in that, The AIFI module includes a multi-head self-attention mechanism, FFN, residual connections, and layer normalization.

9. The method for detecting navigational aid types and health status based on improved YOLOv9 according to claim 1, characterized in that, The Inner-CIoU loss function is introduced to improve the detection capability of small targets; the formula for calculating the Inner-CIoU loss function is: In the formula, For linearized CIoU loss, The intersection-union ratio between the two sets of bounding boxes. This represents the Inner-IoU loss for the inner region.

Citation Information

Cited By

  • Pipe inner surface defect detection method based on improved YOLOv11

    CN122115433A