Method and system for estimating pose of operation and maintenance ship based on deep learning
By setting feature points on the offshore wind turbine tower and using visual sensors and an improved YOLOv8 network, fast and accurate attitude estimation of maintenance vessels was achieved, solving the problems of computational latency and high cost in existing technologies and meeting the needs of real-time motion control.
Patent Information
- Application Number
- CN202511075447.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-12-19
AI Technical Summary
Existing technologies struggle to acquire the pose information of maintenance vessels in real time when ships are swaying at sea. This results in high computational latency for maritime visual SLAM methods, easy failure during point cloud segmentation, poor robustness, high cost of UAV-UAV collaborative positioning, severe communication latency, and inability to meet real-time motion control requirements.
Multiple feature points are set on the tower of the offshore wind turbine. The tower image is obtained by the visual sensor on the maintenance vessel. The YOLOv8 network is improved and trained to obtain the pose estimation model of the maintenance vessel. The visual sensor is used to replace the high-precision gyroscope to achieve fast and accurate pose estimation.
It reduces operation and maintenance costs, improves measurement accuracy and detection speed, meets the requirements of real-time motion control, and the visual sensor can intuitively observe the hull status and quickly detect problems.
Smart Images

Figure CN121170002A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of ship pose estimation technology, and in particular to a deep learning-based method and system for estimating the pose of maintenance ships. Background Technology
[0002] The offshore wind power industry is developing rapidly, and offshore wind power generation has become an important direction for renewable energy development. With the continuous commissioning of offshore wind turbines, offshore maintenance vessels need to regularly maintain the wind turbine generators. During offshore maintenance operations, the vessels experience varying degrees of swaying due to wave action, posing significant safety hazards for the safe transfer of personnel and equipment. Existing technologies have developed various offshore berthing systems to address this issue, among which berthing systems based on a six-degree-of-freedom parallel platform are gradually becoming a research hotspot. This berthing system uses the platform's real-time pose as feedback information to control the movement of six electric cylinders in real time, maintaining the stability of the platform and creating a stable platform for the maintenance vessel's personnel. This ensures the platform remains relatively stationary relative to the offshore wind turbine, thereby guaranteeing the safety of personnel and equipment. Previously, the hull's pose information was measured using gyroscopes; however, gyroscopes are expensive, and due to ship vibrations and sea temperature differences, gyroscopes are prone to measurement errors during long-term continuous operation. Summary of the Invention
[0003] The main objective of this application is to propose a low-cost and accurate deep learning-based method and system for estimating the pose of maintenance vessels.
[0004] To achieve the above objectives, one aspect of this application proposes a deep learning-based method for estimating the pose of a maintenance vessel, comprising the following steps:
[0005] Multiple feature points are set on the tower of the offshore wind turbine;
[0006] The tower image is acquired using a visual sensor on the maintenance vessel, and the tower image includes the feature points.
[0007] The preset YOLOv8 network was improved and trained to obtain the pose estimation model of the maintenance vessel.
[0008] The tower image is input into the maintenance vessel pose estimation model to obtain the maintenance vessel pose information.
[0009] In some embodiments, the setting of multiple feature points on the tower of the offshore wind turbine specifically involves:
[0010] The feature points are respectively set on the tower support columns on the left wing of the tower bottom near the ship, the right wing of the tower bottom near the ship, the left side of the tower middle section near the ship, the right side of the tower middle section near the ship, the left side of the tower support column of the maintenance vessel docking point, and the right side of the tower support column of the maintenance vessel docking point.
[0011] In some embodiments, the improvement and training of the preset YOLOv8 network to obtain the pose estimation model of the maintenance vessel specifically includes:
[0012] Construct a feature point extraction layer and a pose estimation layer;
[0013] By adding the feature point extraction layer and the pose estimation layer after the output layer of the YOLOv8 network, the improved YOLOv8 network is obtained.
[0014] A training dataset is collected, and the improved YOLOv8 network is trained using the training dataset to obtain the pose estimation model of the maintenance vessel.
[0015] In some embodiments, the step of collecting a training dataset and training the improved YOLOv8 network using the training dataset to obtain the ship pose estimation model specifically includes:
[0016] Collect several tower image samples including the aforementioned feature points, and obtain the pose sample of the maintenance vessel corresponding to each tower image sample;
[0017] The tower image samples and the maintenance vessel pose samples are divided to obtain a training set and a test set;
[0018] The improved YOLOv8 network is trained using the training set, and its performance is verified using the test set, resulting in the trained pose estimation model for the maintenance vessel.
[0019] In some embodiments, the maintenance vessel pose estimation model includes a feature point extraction layer and a pose estimation layer. The step of inputting the tower image into the maintenance vessel pose estimation model to obtain the maintenance vessel pose information specifically includes:
[0020] The tower image is input into the pose estimation model of the maintenance vessel to locate feature points, and the bounding box corresponding to the feature point is output. The bounding box includes the bounding box category, the pixel coordinates of the feature point, and the confidence score.
[0021] The bounding box category, the feature point pixel coordinates, and the confidence score are input into the feature point extraction layer, and the sorted feature point pixel coordinates are output.
[0022] The sorted feature point pixel coordinates are input into the pose estimation layer, and the pose information of the maintenance vessel is output.
[0023] In some embodiments, the step of inputting the bounding box category, the feature point pixel coordinates, and the confidence score into the feature point extraction layer and outputting the sorted feature point pixel coordinates specifically includes:
[0024] Set a confidence threshold;
[0025] Determine whether the confidence score corresponding to the bounding box is greater than the confidence threshold;
[0026] When the confidence level corresponding to the bounding box is greater than the confidence threshold, the bounding box is output.
[0027] When the confidence level corresponding to the bounding box is less than or equal to the confidence level threshold, the bounding box of the previous frame is output.
[0028] The feature point pixel coordinates corresponding to the output bounding box are sorted according to the category, and the sorted feature point pixel coordinates are output.
[0029] In some embodiments, the pose estimation layer includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. The activation function of the first fully connected layer is f1(x) = x, and the activation function of the second fully connected layer is... The activation function of the third fully connected layer is f3(x) = x.
[0030] To achieve the above objectives, another aspect of this application proposes a deep learning-based pose estimation system for maintenance vessels, comprising:
[0031] The feature point setting module is used to set multiple feature points on the tower of an offshore wind turbine.
[0032] The tower image acquisition module is used to acquire tower images of the tower through a vision sensor on the maintenance vessel, the tower image including the feature points;
[0033] The network improvement module is used to improve and train the preset YOLOv8 network to obtain the pose estimation model of the maintenance ship.
[0034] The pose estimation module is used to input the tower image into the maintenance vessel pose estimation model to obtain the maintenance vessel pose information.
[0035] To achieve the above objectives, another aspect of this application provides a workstation device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0036] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above.
[0037] The embodiments of this application include at least the following beneficial effects: The deep learning-based method and system for estimating the pose of maintenance vessels first sets multiple feature points on the tower of an offshore wind turbine. Then, a visual sensor on the maintenance vessel acquires an image of the tower, which includes feature points. A pre-set YOLOv8 network is then improved and trained to obtain a pose estimation model for the maintenance vessel. Finally, the tower image is input into the pose estimation model to obtain the pose information of the maintenance vessel. This application sets feature points on the tower and uses a visual sensor instead of a high-precision gyroscope to acquire tower images including feature points, which can significantly reduce maintenance costs. Furthermore, based on the improved YOLOv8 network, a pose estimation model for the maintenance vessel that can estimate the pose of the maintenance vessel based on the tower image is trained, resulting in accurate measurement and fast detection speed. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments of this application are described below. It should be understood that the drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions in this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of the structure of an offshore wind power generation system provided in one embodiment of this application;
[0040] Figure 2 A flowchart illustrating the steps of a deep learning-based pose estimation method for maintenance vessels provided in one embodiment of this application;
[0041] Figure 3 This is a schematic diagram of the structure of a maintenance vessel pose estimation model provided in one embodiment of this application;
[0042] Figure 4 This is a logical schematic diagram illustrating the acquisition of the position and orientation information of a maintenance vessel according to one embodiment of this application.
[0043] Figure 5 A schematic diagram of the structure of a deep learning-based pose estimation system for maintenance vessels provided in one embodiment of this application;
[0044] Figure 6 This is a schematic diagram of the hardware structure of a workstation device provided in one embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0047] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0048] Mooring pier: In offshore wind power operation and maintenance, the mooring pier is a key piece of equipment used to connect the maintenance vessel and the wind turbine for the safe transfer of personnel and equipment.
[0049] TensorRT: A high-performance deep learning inference optimizer and runtime library launched by NVIDIA, designed specifically for deploying deep learning models on NVIDIA GPUs. Its core goal is to transform trained models (such as those trained by PyTorch / TensorFlow) into a highly optimized inference engine, achieving low-latency, high-throughput real-time inference, and is widely used in fields with stringent real-time requirements such as autonomous driving, industrial quality inspection, and medical imaging.
[0050] Pose: Describes the complete information about the position and orientation (orientation) of an object in three-dimensional space.
[0051] The offshore wind power industry is developing rapidly, and offshore wind power generation has become an important direction for renewable energy development. With the continuous commissioning of offshore wind turbines, offshore maintenance vessels need to regularly maintain the wind turbine generators. During offshore maintenance operations, the vessels experience varying degrees of swaying due to wave action, posing significant safety hazards for the safe transfer of personnel and equipment. Existing technologies have developed various offshore berthing systems to address this issue, among which berthing systems based on a six-degree-of-freedom parallel platform are gradually becoming a research hotspot. This berthing system uses the platform's real-time pose as feedback information to control the movement of six electric cylinders in real time, maintaining the stability of the platform and creating a stable platform for the maintenance vessel's personnel. This ensures the platform remains relatively stationary relative to the offshore wind turbine, thereby guaranteeing the safety of personnel and equipment. Previously, the hull's pose information was measured using gyroscopes; however, gyroscopes are expensive, and due to ship vibrations and sea temperature differences, gyroscopes are prone to measurement errors during long-term continuous operation.
[0052] Related technologies propose a method, device, computer equipment, and storage medium for visual SLAM pose estimation at sea, which uses the temporal relationship between dynamic keyframes to estimate the ship's pose. However, keyframe selection needs to simultaneously meet fixed thresholds for time (≤1 second) and the number of matching points (≥L1 / L2). Under high sea states, increased ship sway may cause multiple consecutive frames to fail to meet the number of matching points threshold (feature points disappear rapidly), triggering frequent relocalization. The joint optimization of three keyframes (K1-K2-K3) requires the calculation of the essential matrix, triangulation, and reprojection error, resulting in high single-processing latency, which cannot meet the requirements of real-time control of the berthing platform.
[0053] Existing technologies propose a 6D pose estimation method for ship targets based on point cloud data, including an end-to-end method using point cloud data and an improved deep learning network. However, the point cloud segmentation stage requires foreground points within the 3D bounding box. If critical parts of the ship are occluded by waves (such as the lower part of the hull), the lack of foreground points will cause the segmentation mask to fail. In the two-stage processing flow of RPN+RCNN, the point cloud region pooling and multi-level sampling layer operations result in high computational complexity, which cannot meet the requirements of real-time motion control.
[0054] Existing technologies propose a robust UAV-UAV cooperative relative visual positioning method and system. This method combines the UAV and UAV, using the UAV's pose ground truth value to obtain the ship's relative pose. However, the model requires pre-collecting the UAV's pose ground truth value, but obtaining high-precision ground truth values in real-world maritime environments is extremely costly. Communication between the UAV and the ship relies on long-range wireless communication methods; the latency of long-range communication leads to asynchronous ground truth transmission, resulting in significant communication delays. Furthermore, the process involves four stages of sequential computation: adaptive corner detection, graph neural network feature enhancement, cost matrix adjustment, Hungarian matching, and KF-GRU optimization, introducing computational delays that cannot meet the requirements of real-time motion control.
[0055] In view of this, this application proposes a deep learning-based method for estimating the pose of an operation and maintenance vessel. First, multiple feature points are set on the tower of an offshore wind turbine. Then, a visual sensor on the operation and maintenance vessel acquires an image of the tower, which includes the feature points. A pre-set YOLOv8 network is then improved and trained to obtain a pose estimation model for the operation and maintenance vessel. Finally, the tower image is input into the pose estimation model to obtain the vessel's pose information. This application sets feature points on the tower and uses a visual sensor instead of a high-precision gyroscope to acquire tower images including feature points, significantly reducing operation and maintenance costs. Furthermore, based on the improved YOLOv8 network, a pose estimation model capable of estimating the vessel's pose based on tower images is trained, resulting in accurate measurement and fast detection speed.
[0056] like Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of an offshore wind power generation system for performing a deep learning-based pose estimation method for maintenance vessels, provided as an embodiment of this application.
[0057] The offshore wind power generation system of this application embodiment includes an offshore wind turbine generator set, a docking system, an industrial camera 6, and a workstation 6. The industrial camera fixed on the maintenance vessel obtains image information of the tower 1 of the offshore wind turbine generator set and sends the image information to the workstation 5. The vision algorithm in the workstation 5 calculates the pose of the maintenance vessel based on the feature points in the image information and sends the pose of the maintenance vessel to the servo controller 3 of the docking system. After obtaining the pose of the maintenance vessel, the servo controller 3 calculates the corresponding control command according to the built-in control algorithm and sends it to the servo driver 4. The servo controller 4 controls the six-degree-of-freedom parallel platform 2 to realize a stable and reliable docking system and ensure the safe transfer of personnel and equipment.
[0058] Reference Figure 2 , Figure 2This is a flowchart illustrating the steps of a deep learning-based pose estimation method for maintenance vessels according to one embodiment of this application. This application proposes a deep learning-based pose estimation method for maintenance vessels, which may include, but is not limited to, the following steps S101 to S104:
[0059] Step S101: Set multiple feature points on the tower of the offshore wind turbine;
[0060] It should be noted that in this embodiment of the application, multiple obvious feature points are installed on the tower of the offshore wind turbine. Then, using a fixed camera, multiple sets of tower images with feature points are taken in different poses of the maintenance vessel. The pose of the maintenance vessel is estimated in subsequent steps using these tower images with feature points. This eliminates the need for expensive high-precision gyroscopes, significantly reducing maintenance costs, and the overall method is simpler and faster.
[0061] As an optional implementation, step S101 can be further divided into the following steps S1011:
[0062] Step S1011: Set feature points on the tower support columns on the left wing of the tower bottom near the ship, the right wing of the tower bottom near the ship, the left side of the tower middle section near the ship, the right side of the tower middle section near the ship, the left side of the tower support column of the maintenance vessel docking point, and the right side of the tower support column of the maintenance vessel docking point.
[0063] In some optional embodiments, this application embodiment installs six feature points on the tower. These six feature points need to ensure that the maintenance vessel can clearly capture the positions of the six feature points at the docking point. The positions and structures of the feature points are as follows:
[0064] (1) Location:
[0065] Feature 1: The bottom of the tower is located on the port side of the ship, about 3 meters above the maximum wave splash zone;
[0066] Feature point 2: The bottom of the tower is symmetrical to feature point 1, located on the right wing of the ship.
[0067] Feature 3: The middle section of the tower is located on the left side of the ship, about 6-8 meters above the berthing point of the maintenance vessel;
[0068] Feature point 4: The middle section of the tower is located on the right side of the ship, symmetrical to feature point 3;
[0069] Feature 5: On the tower support column to the left of the maintenance vessel's berthing point;
[0070] Feature point 6: On the right tower support column of the maintenance vessel's berthing point.
[0071] (2) Feature point design:
[0072] Diameter: 0.5m;
[0073] Style: Bright red circle;
[0074] Material: Anodized aluminum plate (3mm thick) covered with high-intensity reflective film;
[0075] Fixing method: welding.
[0076] Step S102: Obtain tower images of the tower using visual sensors on the maintenance vessel. The tower images include feature points.
[0077] It should be noted that industrial cameras fixed on the maintenance vessel capture images of the offshore wind turbine towers, including the aforementioned installation feature points. This embodiment uses a visual sensor instead of a high-precision gyroscope, significantly reducing maintenance costs. While fiber optic gyroscopes can only transmit the vessel's position and orientation to the docking system, visual sensors not only obtain the maintenance vessel's position and orientation but also allow personnel to visually observe the vessel's status through images. This enables faster and more intuitive detection of problems such as malfunctions or deviations from the tower.
[0078] Step S103: Improve and train the preset YOLOv8 network to obtain the pose estimation model of the maintenance ship.
[0079] Specifically, this application embodiment improves the YOLOv8 network by adding a feature point extraction layer and a pose estimation layer after the output layer of the YOLOv8 network. This allows for the training of a new neural network that can directly obtain the pose of the maintenance vessel from a tower image with six feature points, i.e., the maintenance vessel pose estimation model.
[0080] As an optional implementation, step S103 can be further divided into the following steps S1031 to S1033:
[0081] Step S1031: Construct a feature point extraction layer and a pose estimation layer;
[0082] Step S1032: Add a feature point extraction layer and a pose estimation layer after the output layer of the YOLOv8 network to obtain the improved YOLOv8 network;
[0083] As a further optional implementation, the pose estimation layer includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. The activation function of the first fully connected layer is f1(x) = x, and the activation function of the second fully connected layer is... The activation function of the third fully connected layer is f3(x) = x.
[0084] Specifically, such as Figure 3The diagram shows the structure of the pose estimation model for a maintenance vessel. The feature point extraction layer outputs the sorted pixel coordinates of the feature points. The pose estimation layer consists of three fully connected layers, using the ordered feature points obtained from the feature point extraction layer as input. The activation functions of the first and third fully connected layers are f1(x) = x and f3(x) = x, respectively, while the activation function of the second fully connected layer is...
[0085] Step S1033: Collect the training dataset and train the improved YOLOv8 network using the training dataset to obtain the pose estimation model of the maintenance ship.
[0086] As an optional implementation, step S1033 can be further divided into the following steps S10331 to S10333:
[0087] Step S10331: Collect several tower image samples including feature points, and obtain the pose sample of the maintenance vessel corresponding to each tower image sample.
[0088] Step S10332: Divide the tower image samples and maintenance vessel pose samples to obtain training set and test set;
[0089] Step S10333: Train the improved YOLOv8 network using the training set, and verify the performance of the improved YOLOv8 network using the test set to obtain the trained pose estimation model of the maintenance ship.
[0090] Specifically, firstly, an industrial camera, stationary relative to the maintenance vessel, is used to capture images of the fixed tower with feature points under different weather conditions, times, and shooting angles (i.e., tower image samples). The acquired images undergo image processing to make the feature points more prominent. Simultaneously, the pose information of the maintenance vessel corresponding to each image is acquired (i.e., maintenance vessel pose samples). This pose information is normalized, and the data is divided into training and testing sets in a 7:3 ratio. Next, the YOLOv8 network is trained using the existing training and testing sets. Based on the output predicted pose and the maintenance vessel pose samples, the training parameters are adjusted to obtain the optimal training model, i.e., the maintenance vessel pose estimation model. Finally, the trained maintenance vessel pose estimation model is exported to the format required by TensorRT and deployed using the TensorRT platform. Finally, the deployed TensorRT model is embedded into the maintenance ship's workstation. The workstation receives video streams from industrial cameras in real time, performs real-time inference on the video streams, and transmits the inference results to the servo controller, thereby achieving stable control of the docking platform.
[0091] Step S104: Input the tower image into the maintenance vessel pose estimation model to obtain the maintenance vessel pose information.
[0092] As an optional implementation, the ship pose estimation model includes a feature point extraction layer and a pose estimation layer. Step S104 can be further divided into the following steps S1041 to S1043:
[0093] Step S1041: Input the tower image into the pose estimation model of the maintenance vessel to locate feature points, and output the bounding boxes corresponding to the feature points. The bounding boxes include the bounding box category, the pixel coordinates of the feature points, and the confidence score.
[0094] Step S1042: Input the bounding box category, feature point pixel coordinates and confidence score into the feature point extraction layer, and output the sorted feature point pixel coordinates;
[0095] Specifically, such as Figure 4 The diagram shows the logic for acquiring the pose information of the maintenance vessel. First, the tower image with six feature points is passed to the YOLOv8 input layer, from which the bounding box (BOX) is obtained at the YOLOv8 output layer.
[0096]
[0097] The bounding boxes (BOXs) of the six feature points output by YOLOv8 at time t are used as input to the feature point extraction layer. Each bounding box (BOX) contains three elements: i represents the category of the bounding box (BOX) used to distinguish the differences between feature points, [u... i (t),v i [(t)] represents the center pixel coordinates (i.e., the pixel coordinates of the feature point) at time t, CONF i (t) represents the time t of BOX i The confidence level of (t).
[0098] As an optional implementation, step S1042 can be further divided into the following steps S10421 to S10425:
[0099] Step S10421: Set the confidence threshold;
[0100] Step S10422: Determine whether the confidence level corresponding to the bounding box is greater than the confidence level threshold;
[0101] Step S10423: When the confidence level of the bounding box is greater than the confidence level threshold, output the bounding box.
[0102] Step S10424: When the confidence level corresponding to the bounding box is less than or equal to the confidence threshold, output the bounding box of the previous frame.
[0103] Step S10425: Sort the feature point pixel coordinates corresponding to the output bounding box according to the category, and output the sorted feature point pixel coordinates.
[0104] Specifically, confidence level is a key indicator used to determine whether the bounding box (BOX) can be trusted. When the confidence level exceeds the confidence level threshold, it means that the bounding box (BOX) is trusted. If the confidence level threshold is too low, it may lead to incorrect feature point recognition. If the confidence level threshold is too high, it may lead to the loss of feature point recognition. Therefore, an appropriate confidence level is also a key performance indicator in feature recognition.
[0105] In some alternative embodiments, such as Figure 4 As shown, to ensure accuracy and recall, this embodiment sets the confidence threshold to 0.9, and applies the BOX based on the confidence threshold. i (t) Make a judgment in the feature point extraction layer:
[0106]
[0107] At time step t, if the currently detected confidence level CONF i If (t) is higher than the confidence threshold (e.g., 0.9), then the currently detected bounding box (BOX) is used directly. i (t) is the output; if the current confidence level CONF i If (t) is lower than or equal to the confidence threshold (e.g., 0.9), the current detection result is discarded, and the target bounding box (BOX) from the previous frame is used. i (t-1) is the output.
[0108] Then, in the feature point extraction layer, the center pixel coordinates [u] are determined based on the category information i in each bounding box (BOX). i (t),v i The feature points are sorted (t) and the sorted pixel coordinates are used as the output of the feature extraction layer.
[0109] Step S1043: Input the sorted feature point pixel coordinates into the pose estimation layer and output the pose information of the maintenance ship.
[0110] Specifically, the pose estimation layer includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. The input feature point pixel coordinates undergo a linear transformation through the first fully connected layer, mapping the feature point coordinates to a high-dimensional space. The activation output of the first fully connected layer is then input to the second fully connected layer to further extract high-level features. Finally, the output of the second fully connected layer is input to the third fully connected layer to complete the final pose information mapping, outputting six-DOF pose information.
[0111] The above describes the deep learning-based pose estimation method for maintenance vessels according to embodiments of this application. It can be recognized that, compared with traditional methods for measuring hull pose information, embodiments of this application have the following advantages:
[0112] 1. Install multiple obvious feature points on the tower of offshore wind turbines, and use visual sensors to replace high-precision gyroscopes to obtain tower images with feature points for maintenance vessel attitude estimation. This can significantly reduce maintenance costs, and the overall method is simpler and faster.
[0113] Second, while obtaining the position and attitude of the maintenance vessel, the use of visual sensors also allows staff to intuitively observe the status of the vessel through images, enabling them to more quickly and intuitively identify problems when the vessel malfunctions or deviates from the tower.
[0114] Third, the YOLOv8 network is improved by adding a feature point extraction layer and a pose estimation layer after the output layer of the YOLOv8 network. This enables the direct acquisition of the maintenance vessel pose from a tower image with six feature points, with a stable detection speed of over 100fps, which is faster than other vision algorithms.
[0115] Reference Figure 5 This application also provides a deep learning-based system for estimating the pose of maintenance vessels, including:
[0116] The feature point setting module is used to set multiple feature points on the tower of an offshore wind turbine.
[0117] The tower image acquisition module is used to acquire tower images of the tower through vision sensors on the maintenance vessel. The tower images include feature points.
[0118] The network improvement module is used to improve and train the preset YOLOv8 network to obtain the pose estimation model of the maintenance ship.
[0119] The pose estimation module is used to input tower images into the pose estimation model of the maintenance vessel to obtain the pose information of the maintenance vessel.
[0120] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0121] This application also provides a workstation device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This workstation device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0122] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0123] Please see Figure 6 , Figure 6 The hardware structure of a workstation device according to another embodiment is illustrated. The workstation device is equipped with a processor 7, a memory 11, and a graphics card 10. The memory 11 is used to store the program of the depth vision algorithm, the processor 7 is used to execute the algorithm program, and the graphics card 10 is used to accelerate the algorithm. The workstation device is also equipped with two network interfaces for communication. Network interface 1 is used to receive video signals from industrial cameras, and network interface 2 is used to transmit the pose of the maintenance ship calculated by the neural network to the servo controller.
[0124] The processor 7 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0125] The memory 11 can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 11 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 11 and is called and executed by the processor 7 to execute the methods described in the embodiments of this application.
[0126] Network interface 1 and network interface 2 are used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0127] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0128] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0129] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0130] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0131] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0132] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0133] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0135] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0136] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0137] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0138] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0139] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0141] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0142] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A deep learning-based operation ship pose estimation method, characterized in that, The method comprises the following steps: a plurality of feature points are arranged on a tower of an offshore wind turbine; a tower image of the tower is acquired by a visual sensor on a service ship, the tower image comprising the feature points; a preset YOLOv8 network is improved and trained to obtain a service ship pose estimation model; the tower image is input into the service ship pose estimation model to obtain service ship pose information.
2. The method of claim 1, wherein, The plurality of feature points arranged on the tower of the offshore wind turbine are specifically: the feature points are arranged on the left wing of the tower cylinder bottom near the ship side, the right wing of the tower cylinder bottom near the ship side, the left side of the tower cylinder middle near the ship, the right side of the tower cylinder middle near the ship, the tower cylinder support column on the left side of the service ship parking point, and the tower cylinder support column on the right side of the service ship parking point.
3. The method of claim 1, wherein, The preset YOLOv8 network is improved and trained to obtain the service ship pose estimation model, specifically comprising: a feature point extraction layer and a pose estimation layer are constructed; the feature point extraction layer and the pose estimation layer are added after the output layer of the YOLOv8 network to obtain the improved YOLOv8 network; a training data set is collected, and the improved YOLOv8 network is trained by the training data set to obtain the service ship pose estimation model.
4. The method of claim 3, wherein, The training data set is collected, and the improved YOLOv8 network is trained by the training data set to obtain the service ship pose estimation model, specifically comprising: a plurality of tower image samples comprising the feature points are collected to obtain service ship pose samples corresponding to each tower image sample; the tower image samples and the service ship pose samples are divided to obtain a training set and a test set; the improved YOLOv8 network is trained by the training set, and the performance of the improved YOLOv8 network is verified by the test set to obtain the trained service ship pose estimation model.
5. The method of claim 1, wherein, The service ship pose estimation model comprises a feature point extraction layer and a pose estimation layer, and the tower image is input into the service ship pose estimation model to obtain service ship pose information, specifically comprising: the tower image is input into the service ship pose estimation model for feature point positioning, and a bounding box corresponding to the feature points is output, the bounding box comprising a bounding box category, feature point pixel coordinates, and a confidence; the bounding box category, the feature point pixel coordinates, and the confidence are input into the feature point extraction layer to output sorted feature point pixel coordinates; the sorted feature point pixel coordinates are input into the pose estimation layer to output the service ship pose information.
6. The method of claim 5, wherein, The bounding box category, the feature point pixel coordinates, and the confidence are input into the feature point extraction layer to output sorted feature point pixel coordinates, specifically comprising: a confidence threshold is set; it is determined whether the confidence corresponding to the bounding box is greater than the confidence threshold; when the confidence corresponding to the bounding box is greater than the confidence threshold, the bounding box is output; when the confidence corresponding to the bounding box is less than or equal to the confidence threshold, the bounding box of the previous frame is output; The feature point pixel coordinates corresponding to the output bounding box according to the category are sorted, and the sorted feature point pixel coordinates are output.
7. The method of claim 3, wherein, The pose estimation layer comprises a first full connection layer, a second full connection layer and a third full connection layer, an activation function of the first full connection layer is f1(x)=x, an activation function of the second full connection layer is An activation function of the third full connection layer is f3(x)=x.
8. A deep learning-based operation ship pose estimation system, characterized by, Comprise: A feature point setting module is configured to set a plurality of feature points on a tower of an offshore wind turbine generator; A tower image acquisition module is configured to acquire a tower image of the tower through a visual sensor on a service ship, the tower image comprising the feature points; A network improvement module is configured to improve and train a preset YOLOv8 network to obtain a service ship pose estimation model; A pose estimation module is configured to input the tower image into the service ship pose estimation model to obtain service ship pose information.
9. A workstation apparatus, characterized by The workstation device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method in any one of claims 1 to 7.
Citation Information
Patent Citations
Pose estimation system and method for visual carrier landing navigation on mobile platform
CN105021184A
Camera pose estimation method and system, electronic equipment and readable medium
CN117455994A
Ship meeting situation judgment method based on ship posture recognition
CN118675120A
Unmanned ship autonomous berthing and unberthing method and system based on visual laser fusion
CN118691999A
Method and system for visual inspection of offshore wind turbine generators
EP3974647A1