Marine unmanned vehicle fusion positioning method based on fixed end and vehicle end vision

By installing fixed-end cameras inside the ship's cabin and depth cameras in front of the vehicle, and combining YOLOv7 and the federated Kalman filter algorithm, the fixed-end and vehicle-end vision fusion positioning was achieved, which improved the positioning accuracy and stability of marine unmanned vehicles in complex warehouse environments and solved the problems of inflexible positioning inside the cabin and feature matching failure during convoy driving.

CN120991849APending Publication Date: 2025-11-21713TH RES INST OF CHINA STATE SHIPBUILDING CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510947658.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

The limited texture features and repetitive structural features inside ship cabins make existing positioning methods inflexible in dynamic storage environments. Identification positioning depends on deployment density and accuracy, and SLAM systems are prone to vehicle positioning drift. When a group of vehicles travels in formation, feature matching fails due to occlusion, increasing installation and maintenance costs.

Method used

A fixed-end and vehicle-end vision fusion localization method is adopted. By installing a fixed-end camera on the inner wall of the cabin and a depth camera in front of the vehicle, combined with the YOLOv7 target detection model and the federated Kalman filter algorithm, multi-view vehicle recognition and localization are achieved, and the localization information of the fixed-end and vehicle-end vision SLAM systems is fused.

Benefits of technology

It improves the positioning accuracy and stability of marine unmanned vehicles in complex warehouse environments, solves the problem of positioning loss when vehicles are traveling in convoy, reduces installation and maintenance costs, and provides stable positioning information support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120991849A_ABST
    Figure CN120991849A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a marine unmanned vehicle fusion positioning method based on fixed end and vehicle end vision. The method comprises the following steps: S1, arranging a camera; s2, calibrating a camera; s3, image acquisition and model training; step S4, vehicle image recognition; s5, calculating a three-dimensional coordinate; step S6, vehicle end vision SLAM is carried out; and S7, fusion positioning: carrying out information fusion on the output of the two distributed positioning subsystems of the fixed end and the vehicle end. According to the embodiment of the invention, multi-fixed-end visual identification positioning and vehicle-end visual matching positioning are fused together, so that the positioning precision of the vehicle can be improved under the conditions of fewer cabin texture features and repeated structural features, the stability and accuracy of positioning of the ship material guarantee unmanned vehicle in a complex storage environment are improved, and the positioning accuracy of the ship material guarantee unmanned vehicle is improved. And the problems of feature matching failure and positioning loss caused by mutual shielding of vehicles during group vehicle formation driving can be solved, and guarantee is provided for single unmanned vehicle intelligent, cluster control and formation driving.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of unmanned vehicle fusion positioning for ships, and particularly relates to an unmanned vehicle fusion positioning method for ships based on fixed end and vehicle end vision. BACKGROUND

[0002] Unmanned vehicle group intelligence technology for ship material support can realize self-positioning of vehicle group and cooperative operation between vehicles, and is the only way for the development of material support. Among them, vehicle positioning is a key technology for realizing autonomous navigation of unmanned support vehicles, and is also an important technical foundation for realizing cooperative transfer of unmanned vehicles and unmanned intelligent warehousing. At present, indoor material support vehicle positioning usually adopts identification positioning, simultaneous localization and mapping (SLAM), or a combination of the above two methods and IMU to realize positioning. Among them, the identification positioning method through ground guide lines, two-dimensional codes, laser reflection stickers and the like has simple principle and low implementation cost, but the positioning accuracy is affected by the installation accuracy and layout density of the identification, the fixed installation position reduces the flexibility of the vehicle automatic operation driving path, and the wear and obstruction of the identification affect the accuracy and reliability of the positioning, increasing the installation and maintenance cost. The simultaneous localization and mapping technology is suitable for scenes with obvious structural features, rich texture features and large environmental changes. This is because the vehicle needs to detect loops when passing through the same position repeatedly to correct the accumulated positioning error, and the repeated structural features and similar texture features in the cabin may cause the SLAM system to fail to effectively distinguish different similar environment areas, resulting in misjudgment of loop, so that the cumulative error is continuously enlarged, eventually leading to vehicle positioning drift and affecting the accuracy of vehicle autonomous navigation. The IMU estimates the pose by measuring the acceleration and angular velocity of the vehicle, and the positioning method combined with IMU can effectively solve the problem of temporary loss of vehicle positioning, but the ship is usually in a rocking working condition, and the rolling and pitching of the ship body cause the attitude information output by the IMU to be unstable, which brings great interference to the IMU observation value, so the IMU sensor is usually not used for cabin positioning of ship material support vehicles.

[0003] The cabin of the ship has less texture features, a large number of goods and strong structural repeatability, the type and placement position of the goods are often changed, and the cabin passage is usually narrow in order to improve the utilization efficiency of the storage space, so accurate positioning information needs to be provided for the unmanned vehicle to safely pass through in the process of material transportation guarantee. There are two major shortcomings in using fixed markers for positioning: first, the layout position of the markers is fixed, which is difficult to adapt to the complex and changeable cabin environment; second, the positioning accuracy is seriously dependent on the layout density and accuracy of the marker code, and if the damage or obstruction of the marker is not found in time, it is easy to cause the vehicle positioning to be lost or drift, which brings high installation and maintenance costs to unmanned warehousing logistics. The SLAM system does not need to lay fixed markers, but using this method alone for positioning of the unmanned vehicle in the cabin storage environment also has defects: first, the storage condition of the cabin changes dynamically, so the positioning map needs to be updated frequently, which increases the operation and maintenance cost; second, the texture features of the cabin inner wall are less, and the structural features of the goods are repeated, which easily causes the vehicle positioning to drift, and when the group vehicles are driving in formation, there is mutual obstruction between individuals, which affects the effective data collection of the vehicle-mounted sensor. SUMMARY

[0004] In view of the above problems existing in the prior art, the embodiment of the present application provides a ship unmanned vehicle fusion positioning method based on fixed end and vehicle end vision to solve the technical problems of less texture features, repeated structural features, inflexible marker positioning in dynamic storage environment, unstable vehicle end SLAM positioning, feature matching failure and positioning loss due to mutual obstruction between vehicles when the group vehicles are driving in formation in the prior art, and to realize real-time positioning of the group vehicles in complex warehouse environment and provide protection for single vehicle intelligence, cluster control and formation driving of the unmanned vehicle.

[0005] The embodiment of the present application provides a ship unmanned vehicle fusion positioning method based on fixed end and vehicle end vision, which comprises the following steps:

[0006] Step S1, camera arrangement, a plurality of fixed end cameras are arranged at a high position of the inner wall of the cabin to ensure that the field of view of the fixed end cameras covers the cabin space, and at least one vehicle end depth camera is installed in front of the vehicle to realize visual odometry;

[0007] Step S2, camera calibration, calibrate a plurality of fixed end cameras to obtain the internal and external parameters of each fixed end camera;

[0008] Step S3, image acquisition and model training, vehicle images under multiple viewing angles of the cabin application scene are acquired based on the fixed end cameras, a vehicle image annotation dataset is constructed, and vehicle recognition model training is completed based on a YOLOv7 target detection model;

[0009] Step S4, vehicle image recognition, the color image collected by the fixed-end camera in real time is transmitted to the vehicle recognition model as input, the vehicle recognition model recognizes and locates the vehicle in the image to obtain the two-dimensional pixel coordinates (x c ,y c ) of the center point of the vehicle bounding rectangle;

[0010] Step S5, three-dimensional coordinate calculation, the conversion relationship between the coordinate systems is determined through the internal and external parameters of the fixed-end camera, the two-dimensional pixel coordinates (x c ,y c ) are converted to the cabin coordinate system to obtain the three-dimensional space coordinates of the vehicle;

[0011] Step S6, vehicle-end visual SLAM, a point-plane-based visual odometry is constructed, and the vehicle-end pose estimation is completed based on the visual SLAM of the vehicle-end depth camera;

[0012] Step S7, fusion positioning, the outputs of the two distributed positioning subsystems of the fixed end and the vehicle end are fused based on the federated Kalman filter algorithm.

[0013] In an embodiment, the fixed-end camera is a depth camera or a binocular camera.

[0014] In an embodiment, the fixed-end camera is a monitoring camera or a monocular camera.

[0015] In an embodiment, in step S2, the Zhang calibration method is used to obtain the internal parameter matrix K, the rotation matrix R and the translation vector T of each depth camera.

[0016] In an embodiment, in step S3, based on the vehicle image collected by the fixed-end camera, the vehicle is photographed at multiple angles and different distances under multiple backgrounds and multiple positions, and annotation is completed, and the original sample image is processed including random scaling, flipping, cropping and adding noise to expand the target sample data set.

[0017] In an embodiment, in step S5, when multiple fixed-end cameras detect the same target vehicle at the same time, the output coordinates of the multiple fixed-end cameras are fused based on the weighted least squares method to obtain the vehicle positioning result based on the fixed-end camera.

[0018] In an embodiment, the fixed-end camera closer to the target vehicle has a higher weight in the weighted least squares method.

[0019] In an embodiment, step S5 includes,

[0020] the pixel point in the depth image is converted to the coordinate system of the color camera through the conversion relationship between the coordinate systems to obtain the three-dimensional space coordinates of the point in the camera coordinate system;

[0021] Based on the installation extrinsic parameters of the fixed-end camera in the cabin coordinate system, the coordinate points in the camera coordinate system are transformed to the cabin coordinate system to obtain the three-dimensional coordinates of the vehicle output by the fixed-end camera.

[0022] In one embodiment, step S5 includes,

[0023] By calibrating the intrinsic and extrinsic parameters of the fixed-end camera, the plane equation of the cabin floor in the camera coordinate system is obtained, and then the position of any point on the plane in the camera coordinate system is obtained.

[0024] The three-dimensional spatial coordinates of the target vehicle are obtained by utilizing the coordinate transformation relationship between the camera coordinate system and the cabin coordinate system.

[0025] In one embodiment, step S6 includes,

[0026] Extract the feature point set P from the scene, and classify the feature points based on the distance invariant principle;

[0027] The cyclic feature point set is used to find three feature points with constant relative distance between frame k-1 and frame k, and put them into the same point set. The pose transformation matrix is ​​calculated for each motion state point set, and point sets with similar results are put into the same set.

[0028] Using the walls and floor of the cabin as high-dimensional features to provide constraints for visual odometry, dynamic points in the scene are removed, and camera pose calculation is optimized.

[0029] Compared with existing technologies, the beneficial effects of the fusion positioning method for marine unmanned vehicles based on fixed-end and vehicle-end vision provided by the embodiments of the present invention are as follows: As a fusion positioning method, the embodiments of the present invention realize the positioning of unmanned vehicles for material support in the dynamic storage environment of ship cabins by constructing two positioning subsystems: fixed-end vehicle recognition and vehicle-end visual SLAM. By fusing multi-fixed-end visual recognition positioning and vehicle-end visual matching positioning, the positioning accuracy of vehicles can be improved under the condition that there are few texture features and repetitive structural features in the cabin. This improves the stability and accuracy of the positioning of unmanned vehicles for marine material support in complex storage environments, and can solve the problem of feature matching failure and positioning loss caused by mutual occlusion between vehicles when a group of vehicles is driving in formation. This provides a guarantee for the intelligence of unmanned vehicles, cluster control and formation driving. Attached Figure Description

[0030] Figure 1 A schematic diagram illustrating the workflow of a fusion positioning method for marine unmanned vehicles based on fixed-end and vehicle-end vision provided in an embodiment of the present invention;

[0031] Figure 2 A schematic diagram of the data flow involved in a fusion positioning method for marine unmanned vehicles based on fixed-end and vehicle-end vision provided in an embodiment of the present invention;

[0032] Figure 3 A schematic diagram of the vehicle-end vision odometry positioning principle based on point and surface features is provided for a fusion positioning method for marine unmanned vehicles based on fixed end and vehicle end vision provided in an embodiment of the present invention.

[0033] Figure 4 The diagram illustrates the principle of federated Kalman filter fusion positioning for a marine unmanned vehicle based on fixed-end and vehicle-end vision, as provided in this embodiment of the invention. Detailed Implementation

[0034] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Various embodiments and features of this application are described herein with reference to the accompanying drawings.

[0036] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.

[0037] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application, which have the features described in the claims and are therefore all within the scope of protection defined herein.

[0038] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.

[0039] Specific embodiments of this application are described below with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this application, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to ascertain the true intent based on the user's historical operations, and to avoid unnecessary or redundant details that would obscure this application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in various ways with substantially any suitable detailed structure.

[0040] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.

[0041] The principles and features of the present invention are described below with reference to the accompanying drawings. The embodiments described are for illustrative purposes only and are not intended to limit the scope of the invention. The following description, in conjunction with... Figures 1-4The preferred embodiments of the present invention will be described in further detail below:

[0042] like Figures 1-4 As shown, this embodiment of the invention provides a fusion positioning method for marine unmanned vehicles based on fixed-end and vehicle-end vision, including the following steps:

[0043] Step S1, camera arrangement: Several fixed-end cameras are set up high on the inner wall of the cabin to ensure that the field of view of the fixed-end cameras covers the cabin space. At least one vehicle-end depth camera is installed in front of the vehicle to realize visual odometry.

[0044] Step S2, camera calibration, calibrating multiple fixed-end cameras and obtaining the intrinsic and extrinsic parameters of each fixed-end camera. In one embodiment, the intrinsic parameter matrix K, rotation matrix R, and translation vector T of each depth camera are obtained by Zhang's calibration method.

[0045] Step S3, Image Acquisition and Model Training: Vehicle images from multiple perspectives in the cabin application scenario are acquired using a fixed-end camera, and a vehicle image annotation dataset is constructed. The vehicle recognition model is trained based on the YOLOv7 object detection model. In one embodiment, vehicle images inside the cabin are acquired using a fixed-end camera. The vehicle is photographed from multiple angles and at different distances under multiple backgrounds and locations and annotated. The original sample images are processed including random scaling, flipping, cropping, and adding noise to expand the target sample dataset.

[0046] Step S4, vehicle image recognition: The color image captured in real time by the fixed-end camera is input to the vehicle recognition model. The vehicle recognition model identifies and locates the vehicle in the image, obtaining the two-dimensional pixel coordinates (x, y, z) of the center point of the vehicle's circumscribed rectangle. c ,y c );

[0047] Step S5, 3D coordinate calculation: Determine the transformation relationship between coordinate systems using the intrinsic and extrinsic parameters of the fixed-end camera, and convert the 2D pixel coordinates (x, y, z) to 2D coordinates. c ,y c The vehicle's three-dimensional spatial coordinates are obtained by converting to the cabin coordinate system. When multiple fixed-end cameras detect the same target vehicle at the same time, the output coordinates of the multiple fixed-end cameras are fused based on the weighted least squares method to obtain the vehicle positioning result based on the fixed-end cameras. Among them, the fixed-end camera that is closer to the target vehicle has a higher weight in the weighted least squares method.

[0048] Step S6, Vehicle-side Visual SLAM: Constructing a point-and-surface-based visual odometry system. Visual SLAM based on the vehicle-side depth camera completes vehicle pose estimation. In one embodiment, firstly, a feature point set P is extracted from the scene, and the feature points are classified based on the distance-invariant principle. Then, the feature point set is iterated, and three feature points with invariant relative distances are found between frame k-1 and frame k, and these are placed into the same point set. A pose transformation matrix is ​​calculated for each motion state point set, and point sets with similar results are placed into the same set. Secondly, the walls and ground in the cabin are used as high-dimensional features to provide constraints for the visual odometry system, removing dynamic points in the scene and optimizing camera pose calculation.

[0049] Step S7, Fusion Positioning: Based on the federated Kalman filter algorithm, the outputs of the two distributed positioning subsystems, the fixed end and the vehicle end, are fused together.

[0050] This invention provides a fusion localization method for marine unmanned vehicles based on fixed-end and vehicle-end vision. It integrates fixed-end image recognition and localization with onboard visual SLAM using federated Kalman filtering. By combining two information sources with vehicle wheel speedometer data, it obtains real-time location information of the vehicle cabin. The workflow diagram is shown below. Figure 1 As shown in the diagram, the data flow is as follows: Figure 2 As shown. In this embodiment of the invention, several fixed-end cameras are installed high on the interior wall of the cabin, covering the cabin space; one vehicle-end depth camera is installed at the front of the vehicle to achieve visual odometer readings.

[0051] First, the target vehicle is identified and located using multiple fixed-end cameras, and the positioning results acquired by the multiple cameras are fused using the weighted least squares method. Second, vehicle-side localization is achieved through visual SLAM, and dynamic points are removed based on visual odometry using point and area features to reduce the impact of moving objects on localization. Finally, the fixed-end and vehicle-side localization data are fused using the federated Kalman filter algorithm to obtain the vehicle's positioning information in the cabin coordinate system. This method mainly includes functional units such as camera parameter calibration, vehicle image dataset acquisition, recognition model training, vehicle image recognition, 3D coordinate fusion calculation, vehicle-side visual SLAM localization, and Kalman filter fusion localization. Among these, the fixed-end camera parameter calibration, vehicle image dataset acquisition, and recognition model training units are preparatory steps for implementing this invention; if the system state does not change, they only need to be completed once. Vehicle image recognition, 3D coordinate fusion calculation, vehicle-side visual SLAM localization, and Kalman filter fusion localization are the key steps for implementing this invention.

[0052] This invention constructs two visual positioning subsystems: a fixed-end system and a vehicle-mounted system. The fixed-end system integrates visual positioning results from multiple cameras, improving the accuracy of fixed-end visual positioning. The vehicle-mounted system employs point-area feature visual odometry to reduce the impact of moving objects on visual positioning. Finally, the two independent positioning subsystems are fused based on a federated Kalman filter system, improving the system's fault tolerance. This provides stable and reliable positioning information for unmanned vehicles used for marine material support in indoor warehousing environments with weak textures, easy occlusion, and dynamically changing storage types and methods, ensuring vehicle-to-vehicle intelligence, collaborative transportation, and platooning operations.

[0053] 1) Integrating multiple fixed-end cameras for vehicle identification and positioning

[0054] Vehicle identification within a cabin is achieved using the YOLOv7 target detection model. Global vehicle localization within the cabin coordinate system is realized based on the detection box coordinates and camera characteristics. When multiple cameras simultaneously observe the same target vehicle, a fused global vehicle localization is obtained using weighted least squares, improving the accuracy of fixed-end visual localization. Fixed-end cameras can utilize depth cameras or binocular cameras to obtain accurate localization data. Alternatively, vehicle identification and localization can be achieved using surveillance cameras or monocular cameras, sacrificing some accuracy to reduce costs.

[0055] 2) Visual odometry is constructed using point and area features based on vehicle-mounted depth cameras.

[0056] Considering that weakly textured planes such as indoor floors and bulkheads have a large probability of appearing in the camera's field of view and account for a large proportion of the field of view, we combine the point features and spatial surface features of the images acquired by the depth camera, remove dynamic points in the scene through image feature point classification and surface feature constraints, construct a visual odometry, and reduce the interference of moving objects (such as people, other vehicles, etc.) on the system during visual SLAM operation.

[0057] 3) Positioning adaptable to complex dynamic indoor scenarios or scenarios where satellite signals are denied.

[0058] Two different positioning subsystems are designed. Based on multi-source information, a loosely coupled, decentralized filtering approach is used to achieve fusion positioning. Even if one positioning method fails, the system can still output positioning results, demonstrating strong fault tolerance. Furthermore, the fusion framework is highly scalable, facilitating the addition of other positioning subsystems to adapt to more complex application scenarios. This invention is applicable to various complex dynamic scenarios, such as cabin scenes with easy occlusion, repetitive structures, weak textures, and mutual occlusion between vehicles, as well as deck scenes where satellite signals are denied.

[0059] Example 1

[0060] This invention implements fusion positioning of marine unmanned vehicles based on a fixed-end depth camera and a vehicle-mounted depth camera. The implementation process is as follows:

[0061] 1) Camera calibration: Calibrate multiple fixed-end depth cameras and obtain the intrinsic and extrinsic parameters of each camera.

[0062] 2) Image acquisition and model training: Vehicle images from multiple perspectives in the cabin application scenario are acquired using a depth camera installed at a fixed end. A vehicle image annotation dataset is constructed, and a vehicle recognition model is trained based on the YOLOv7 object detection model.

[0063] 3) Vehicle Image Recognition: Color images captured in real-time by a fixed camera are input to the vehicle recognition model. The model identifies and locates vehicles in the images. Based on the normalized center coordinates (x, y) of the target detection box and the length and width w, h of the detection box, the center point of the vehicle's bounding rectangle is calculated, and its two-dimensional pixel coordinates (x, y) are output. c ,y c ).

[0064] 4) 3D coordinate fusion calculation: Calculate the coordinates (x, y) in the depth camera color images of the same vehicle. c ,y c The system retrieves the corresponding spatial coordinates of the vehicle and fuses the recognition and positioning results from multiple fixed-end cameras to output the final three-dimensional coordinates of the vehicle. First, the pixels in the depth image are transformed to the coordinate system of the color camera using coordinate system transformation relationships, obtaining the three-dimensional spatial coordinates of that point in the camera coordinate system. Then, based on the camera's installation extrinsic parameters in the cabin coordinate system, it is transformed to the cabin coordinate system to obtain the vehicle's three-dimensional coordinates output by that camera. If multiple cameras simultaneously detect the same target vehicle, the output coordinates of the multiple cameras are fused using a weighted least squares method to obtain the vehicle positioning result based on the fixed-end camera.

[0065] 5) Vehicle-side visual SLAM: Construct a point-and-surface-based visual odometry system and use a depth camera-based visual SLAM system to complete vehicle pose estimation.

[0066] 6) Fusion positioning: The outputs of the two distributed positioning subsystems are fused based on the federated Kalman filter algorithm.

[0067] In step 1), the intrinsic parameter matrix K, rotation matrix R, and translation vector T of each depth camera are obtained by Zhang's calibration method.

[0068] In step 2), vehicle images inside the cabin are acquired using a fixed-end depth camera. The vehicle is photographed from multiple angles and distances under various backgrounds and locations, and annotated. The original sample images are then processed by random scaling, flipping, cropping, and noise addition to expand the target sample dataset. The error function is shown in the following formula:

[0069] L total =λ1Lcls +λ2L box +λ3L obj

[0070] In the formula, L cls L box and L obj λ1, λ2, and λ3 are the classification loss, localization loss, and confidence loss of the network, respectively; λ1, λ2, and λ3 are the weights corresponding to each part of the loss.

[0071] The network training results are evaluated using the mean accuracy as the metric, as shown in the following formula:

[0072]

[0073] In the formula, TP, FP, and FN represent the number of samples that were detected and correctly identified, the number of samples that were detected but incorrectly identified, and the number of samples that were not identified, respectively.

[0074] During training, the error function between the output of the neural network and the manually labeled true values ​​is calculated. The parameters of the neural network are continuously adjusted using a large amount of data to reduce the error function until the error function converges, thus completing the training.

[0075] In step 3), the model trained in step 2) is used for inference to detect targets in the real-time acquired images at the fixed end. YOLOv7 includes a backbone network, a neck network, and a detection output. The backbone network continuously extracts features from the input image and obtains these features after multiple convolutions. The neck network fuses the features output by the backbone network based on feature pyramids and path aggregation structures. The detection output generates the final detection result using non-maximum suppression, outputting the category, confidence score, and normalized coordinates of the detection box. The coordinates are then transformed to the image pixel coordinate system using the following formula:

[0076]

[0077] In the formula, x and y are the normalized coordinates of the detection box center, respectively; x c y c These are the center coordinates of the detection box in the pixel coordinate system; w img h img The width and height of the input image.

[0078] In step 4), firstly, the vehicle coordinates identified by a single depth camera are calculated. Solving for the vehicle's 3D coordinates requires aligning the depth map to the color image to obtain the spatial coordinate values ​​corresponding to all pixels in the color image. This involves the transformation relationships between the pixel coordinate system, image coordinate system, camera coordinate system, and cabin coordinate system. The final pixel coordinates (x...) c ,y c ) and cabin coordinates (X w ,Yw Z w The conversion formula can be expressed as:

[0079]

[0080] In the formula, For the camera intrinsic parameter matrix, Let f be the extrinsic parameter matrix from the camera coordinate system to the cabin coordinate system, and d be the camera focal length. x d y These represent the actual physical lengths of each pixel on the x and y axes, respectively. Substituting the vehicle's center pixel coordinates into the above formula yields the three-dimensional coordinates of the vehicle within the cabin as identified by a single depth camera.

[0081] If N depth cameras at different locations identify the same target vehicle, then the N localization results are fused using weighted least squares. Assume that each camera i provides a localization result P of the vehicle inside the cabin. i (x i ,y i The objective function is:

[0082]

[0083] In the formula, x and y are the fused vehicle positioning, and ω i Let be the normalized weight for the i-th camera observation. Since the depth camera measurement accuracy is inversely proportional to the measurement distance, the distance from the camera to the vehicle is used to assign weights to each camera, so that cameras closer to the vehicle have higher weights. Then the weight of the i-th camera observation can be expressed as:

[0084]

[0085] In the formula, d i Let be the Euclidean distance from the target vehicle to the i-th camera. Therefore, the vehicle localization output by the fixed-end depth camera is:

[0086] In step 5), a visual SLAM system is used based on the vehicle's depth camera to complete the global localization estimation of the vehicle. The visual SLAM system includes depth camera data acquisition, visual odometry, backend optimization, and loop closure detection. The following mainly introduces visual odometry based on point and surface features. Since the vehicle camera inevitably captures a large amount of weakly textured walls and floors inside the cabin, it is difficult to extract rich features to generate the feature descriptors required by the SLAM system. However, by adding surface features, more stable and richer ensemble information can be provided for visual odometry.

[0087] First, a feature point set P is extracted from the scene, and the feature points are classified based on the distance invariance principle. The feature point set is iterated, and three feature points with invariant relative distance (same motion state) are found between frame k-1 and frame k, and these are placed into the same point set. The pose transformation matrix is ​​calculated for each motion state point set, and point sets with similar results are placed in the same set to reduce subsequent pose estimation errors. Second, the walls and floor in the cabin can be used as high-dimensional features to provide constraints for traditional visual odometry, removing dynamic points in the scene and optimizing camera pose calculation. Color and depth images of the scene are acquired using an onboard depth camera. The color image and depth image are aligned using camera intrinsic parameters to obtain the position of each pixel in the cabin space, creating a 3D point cloud map. After downsampling the point cloud, the nearest neighbor points are used to calculate the position P of each point. i (x pi ,y pi ,z pi Local normal vector of ) The equation of the plane can be obtained as follows:

[0088]

[0089] In the formula, the plane parameter D pi It can be determined from the coordinates of the point and the local normal vector. Clustering is performed on the points using the DBSCAN algorithm. Finally, the virtual plane parameters of each cluster are obtained using SVD decomposition. Let the previous frame, current frame, and next frame be k and k+1 frames respectively. Since the camera can be approximated as moving at a constant speed for a short period, the rotation and translation matrices between frames k and k+1 are the same as those between frames k-1 and k. Therefore, the pose transformation between frames k and k+1 is as follows:

[0090]

[0091] In the formula, p k p k+1 Let P be the spatial location of frame point P in frames k and k+1, respectively. Let P be the rotation matrix and translation vector between the two frames, respectively. The position of point P in frame k+1 can be obtained using the above formula, and then the parameters of the plane equation can be estimated using the coordinates of the three points. The planar parameters in the scene are updated every 10 frames using the local vector method. A point (x, y, z) in the plane of frame k-1 is transformed into a point (x′, y′, z′) in frame k using the pose transformation matrix, as shown in the following equation:

[0092]

[0093] The plane equation for the k-th frame is:

[0094] [n k D k[x′ y′ z′ 1] T =0

[0095] In the formula, n k Let be the plane normal vector of the k-th frame.

[0096] From the above two equations, we can obtain:

[0097] [Rn k n k t+D k [xyz 1] T =0

[0098] Comparing the above equation with the plane equation of the (k-1)th frame, we can obtain:

[0099] n k-1 =Rn k

[0100] n k t = D k-1 -D k

[0101] The rotation and translation matrix has three degrees of freedom. The camera pose transformation matrix [R,t] can be directly calculated from the plane parameters using the above formula based on the three mutually perpendicular normal vectors. If there are no three pairs of mutually perpendicular planes, then the dynamic and static feature point sets are distinguished based on the plane parameters. The feature points are clustered based on the principle that the relative distance between two points remains unchanged. The pose transformation is calculated for each feature point. The pose transformation of the static feature point set should minimize the difference error between the plane parameters of the current frame and the previous frame. The error function is as follows:

[0102]

[0103] In the formula, Err is the sum of the differences between the plane parameters; n is the number of planes matched in the scene; [n k D k ] represents the plane parameters of the k-th frame.

[0104] If the planar parameter errors of the dynamic point set and the static point set in the current frame are similar, meaning that not enough planes have been extracted to distinguish between the static and dynamic feature point sets, then the camera pose transformation matrix is ​​solved using ORB feature points based on PnP and the random sample consensus algorithm. Four 3D-2D point pairs are randomly selected, and the 3D points are denoted as p. i 2D point is denoted as p i Calculate the reprojection error for each point and minimize the total error, as shown in the following formula:

[0105]

[0106] In the formula, K is the camera intrinsic parameter matrix. The camera pose transformation matrix can be obtained by solving the above formula using the nonlinear least squares method.

[0107] After the front-end odometer design is completed, back-end optimization is performed using graph optimization methods to eliminate the impact of noise during vehicle operation. Loop closure detection is completed using a bag-of-words model. The vehicle is repositioned based on the correlation between current and historical data to eliminate accumulated errors. The camera pose is transformed to the center of the vehicle body and further transformed to the cabin coordinate system to output the vehicle positioning information.

[0108] In step 6), positioning observation models for both the fixed end and the vehicle end are established to obtain state estimates and covariance matrices. Based on federated Kalman filtering, the vehicle-side visual SLAM and the fixed end recognition and positioning results are weighted and fused. The fused global state estimate provides more accurate vehicle positioning information.

[0109] In the design of the federated Kalman filter, since the vehicle wheel speedometer is completely autonomous, it is set as a common reference system to provide local motion information such as vehicle speed and angular velocity for state prediction. Local filter 1 is designed to perform Kalman filtering on the motion state quantities output by the wheel speedometer and the positioning data output by the fixed terminal system. Local filter 2 is designed to perform Kalman filtering on the motion state quantities output by the wheel speedometer and the positioning data output by the vehicle-side visual SLAM.

[0110] The vehicle state vector of the positioning system is as follows:

[0111] x k =[x,y,θ,v x ,v y ]

[0112] In the formula, x k θ is the state vector; x and y are the global coordinates of the vehicle, respectively; θ is the heading angle.

[0113] In local filter 1, the state estimate and covariance matrix of the fixed-end identification and positioning are obtained through Kalman filtering prediction and update steps. The state prediction model of the vehicle state is as follows:

[0114]

[0115] In the formula, The state estimate of local filter 1 at time k; x 1,k-1 Let F be the state variable of filter 1 at time k-1, F be the state transition matrix, B be the control input matrix, and u be the state input matrix. k The control input is fed back from the wheel speed gauge; w k The noise is a process noise and follows a Gaussian distribution.

[0116] The covariance prediction model is as follows:

[0117]

[0118] In the formula, P 1,k-1 Q1 is the covariance matrix of filter 1 at time k-1, representing the uncertainty of vehicle state estimation; Q1 is the process noise covariance matrix, representing the uncertainty in the state update process.

[0119] The observation model for filter 1 based on the fixed-end depth camera for identification and localization is as follows:

[0120]

[0121] In the formula, Y 1,w Y 1,w H1 represents the vehicle coordinates obtained from observations by a fixed-end camera; H1 is the observation matrix; v k The noise is observed to follow a Gaussian distribution.

[0122] The Kalman gain of local filter 1 is calculated using the following formula:

[0123]

[0124] In the formula, K1 is the Kalman gain of local filter 1, which indicates the weighted relationship between the predicted information and the observed information; R1 is the observation noise covariance matrix, which is used to describe the measurement error of fixed-end identification and positioning.

[0125] The positioning information output by the fixed-end visual recognition is used as observation data, and the predicted state estimate is corrected using Kalman gain:

[0126]

[0127] In the formula, For the predicted vehicle status; These are observations based on the predicted state.

[0128] Finally, the covariance matrix is ​​updated to describe the uncertainty of the state estimate:

[0129]

[0130] In the formula, I is the identity matrix.

[0131] Similar to filter 1, the state prediction and update steps of vehicle-side visual SLAM are as follows:

[0132] The vehicle state prediction model and covariance prediction model are as follows:

[0133]

[0134]

[0135] In the formula, x is the state prediction value of filter 2 at time k; 2,k-1 State estimation for vehicle-side visual SLAM at time k-1; u k P is the control quantity from the wheel speed gauge. 2,k-1 Q1 is the covariance matrix of filter 2 at time k-1; Q2 is the process noise covariance matrix.

[0136] The update steps correct the predicted state estimate based on the vehicle-mounted visual positioning system. The current observation value is:

[0137] z 2,k =H2x 2,k +v k

[0138] In the formula, z 2,k H1 represents the positioning observations obtained from visual SLAM, and H2 represents the vehicle's positioning in the cabin coordinate system; H3 represents the observation matrix; v k To observe noise.

[0139] Kalman gain is expressed as follows:

[0140]

[0141] In the formula, H2 is the observation matrix of the vehicle-side visual SLAM; R2 is the covariance matrix of the observation noise.

[0142] The updated state estimate and covariance matrix are as follows:

[0143]

[0144] This completes the design of the local filters for the two positioning subsystems. A federated Kalman filter is then used to estimate the state output of the two local filters. and Perform weighted fusion. Estimate the state after fusion. The covariance matrix P f as follows:

[0145]

[0146] By outputting the fused state estimates, the global positioning of the vehicle in the cabin coordinate system can be obtained.

[0147] Example 2

[0148] This invention provides a fusion positioning system for unmanned marine vehicles based on a fixed-end monocular camera and a vehicle-mounted depth camera. This implementation eliminates the need for a dedicated depth camera; existing monitoring cameras in the warehouse environment can perform image acquisition and positioning calculations, sacrificing some positioning accuracy but reducing implementation costs. The implementation process is as follows:

[0149] Implementation method two involves fusion positioning of marine unmanned vehicles based on a fixed-end monocular camera and a vehicle-mounted depth camera. This method eliminates the need for a depth camera; existing monitoring cameras in the warehouse environment can be used to complete image acquisition and positioning calculations from the fixed-end camera, sacrificing some positioning accuracy but reducing implementation costs. The implementation process is as follows:

[0150] 1) Camera calibration: Calibrate a fixed monocular or surveillance camera to obtain the camera's intrinsic and extrinsic parameters;

[0151] 2) Image acquisition and model training: Based on a fixed monocular or surveillance camera, vehicle images are acquired from multiple perspectives, a vehicle image annotation dataset is constructed, and a vehicle recognition model is trained based on the YOLOv7 object detection model.

[0152] 3) Vehicle Image Recognition: Images captured in real-time by a fixed camera are input to the vehicle recognition model. The model identifies and locates vehicles in the images. Based on the normalized center point coordinates (x, y) of the target detection box and the length and width w, h of the detection box, the center point of the vehicle's bounding rectangle is calculated and its two-dimensional pixel coordinates (x, y) are output. c ,y c );

[0153] 4) Three-dimensional coordinate calculation: that is, calculating the coordinates (x, y) in the image. c ,y c The corresponding spatial coordinate information is obtained by using the calibrated camera intrinsic and extrinsic parameters to obtain the plane equation of the cabin floor in the camera coordinate system. The position of any point on this plane in the camera coordinate system can be obtained. By using the coordinate transformation relationship between the camera coordinate system and the cabin coordinate system, the three-dimensional spatial coordinates of the vehicle can be obtained.

[0154] 5) Vehicle-side visual SLAM: Visual SLAM based on depth cameras is used to estimate the vehicle's pose;

[0155] 6) Fusion positioning: The outputs of the two distributed positioning subsystems are fused based on the federated Kalman filter algorithm.

[0156] Example 2 is basically the same as Example 1 in terms of steps. The difference is that in step 4), a single camera calculates the position of the identified vehicle. Since a monocular or monitoring camera cannot directly obtain the depth value of a spatial point, the fixed-end positioning subsystem uses a ranging method based on calibration parameters to calculate the vehicle position.

[0157] The key to this measurement method lies in the camera parameter calibration, and the resulting intrinsic parameter matrix is ​​shown in the following formula:

[0158]

[0159] Based on the camera mounting extrinsic parameters rotation matrix R, translation vector T, and the ground equation in the camera coordinate system, the coordinates of any point on the ground in the camera coordinate system can be obtained. Then, using the coordinate transformation relationship between the camera coordinate system and the cabin coordinate system, the coordinates of that point in the cabin coordinate system can be calculated. The calculation steps can be divided into three steps: ground equation calculation, two-dimensional coordinate point extraction and back projection, and three-dimensional coordinate solution.

[0160] A point (X) in the cabin coordinate system W ,Y W Z W The rotation matrix R and translation vector T obtained through extrinsic parameter calibration are transformed into the camera coordinate system (X). C ,Y C Z C Then, based on the camera's own parameters, (X) C ,Y C Z C Converting to pixel coordinates (u,v), the final conversion formula between pixel coordinates and cabin coordinates is as follows:

[0161]

[0162] Since the installation position and angle of the fixed-end camera in the cabin coordinate system are known, the plane equation of the cabin floor in the camera coordinate system can be obtained through calibration. Let the coordinates of a point on the cabin floor in the cabin coordinate system be O. w =(0,0,0) T The normal vector is By utilizing the coordinate transformation relationship between camera coordinates and cabin coordinates, the coordinates of the center point of the ground in the camera coordinate system can be obtained as O. C =(x o ,y o ,z o ) T The normal vector is Therefore, the ground equation in the camera coordinate system can be written as:

[0163] n1(x-x0)+n2(y-y0)+n3(z-z0)=0

[0164] Two-dimensional center point extraction and backprojection are used to obtain the center point of the target vehicle in the image, and thus obtain the three-dimensional coordinates of the vehicle center in the camera coordinate system represented by pixel depth z. The center point of the vehicle detection box is selected and denoted as (x...). c ,y c Let the pixel depth of this point be z, then the corresponding 3D coordinates of this pixel are... c x f x c y f yAll of these can be obtained from the camera intrinsic parameter matrix. The goal of solving for the 3D center point is to calculate the coordinates of the center point of the bottom of the target vehicle. Substituting the 3D coordinate values ​​in the camera coordinate system into the ground equation, the depth can be expressed as:

[0165]

[0166] Based on the transformation relationship between the cabin coordinate system and the camera coordinate system, the normal vector is obtained. and Relationship:

[0167]

[0168] Combining the above two equations, the pixel depth z of the vehicle detection frame center point can be obtained. Using the conversion relationship between pixel coordinates and cabin coordinates, the vehicle pixel coordinates (x, y, z) can be calculated. c ,y c The three-dimensional coordinates in the cabin coordinate system.

[0169] If N cameras at different locations identify the same target vehicle, the N positioning results are fused using the weighted least squares method to finally output the vehicle's positioning in the cabin coordinate system. The calculation steps are the same as in Example 1.

[0170] The above embodiments are merely exemplary embodiments of the present invention and are not intended to limit the present invention. The scope of protection of the present invention is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to the present invention within its spirit and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of the present invention.

Claims

1. A fusion positioning method for marine unmanned vehicles based on fixed-end and vehicle-end vision, characterized in that, Includes the following steps: Step S1, camera arrangement: Several fixed-end cameras are set up high on the inner wall of the cabin to ensure that the field of view of the fixed-end cameras covers the cabin space. At least one vehicle-end depth camera is installed in front of the vehicle to realize visual odometry. Step S2, camera calibration, calibrate multiple fixed-end cameras, and obtain the intrinsic and extrinsic parameters of each fixed-end camera; Step S3, Image Acquisition and Model Training: Vehicle images from multiple perspectives in the cabin application scenario are acquired using a fixed-end camera, a vehicle image annotation dataset is constructed, and the vehicle recognition model is trained based on the YOLOv7 object detection model. Step S4, vehicle image recognition: The color image captured in real time by the fixed-end camera is input to the vehicle recognition model. The vehicle recognition model identifies and locates the vehicle in the image, obtaining the two-dimensional pixel coordinates (x, y, z) of the center point of the vehicle's circumscribed rectangle. c ,y c ); Step S5, 3D coordinate calculation: Determine the transformation relationship between coordinate systems using the intrinsic and extrinsic parameters of the fixed-end camera, and convert the 2D pixel coordinates (x, y, z) to 2D coordinates. c ,y c The vehicle's three-dimensional spatial coordinates are obtained by converting to the cabin coordinate system; Step S6, vehicle-side visual SLAM, construct a point-and-surface-based visual odometry, and complete vehicle pose estimation based on the vehicle-side depth camera's visual SLAM. Step S7, Fusion Positioning: Based on the federated Kalman filter algorithm, the outputs of the two distributed positioning subsystems, the fixed end and the vehicle end, are fused together.

2. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 1, characterized in that: The fixed-end camera is a depth camera or a binocular camera.

3. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 1, characterized in that: The fixed-end camera is a surveillance camera or a monocular camera.

4. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 1, characterized in that: In step S2, the intrinsic parameter matrix K, rotation matrix R, and translation vector T of each depth camera are obtained by Zhang's calibration method.

5. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 1, characterized in that: In step S3, vehicle images inside the cabin are acquired using a fixed-end camera. The vehicle is photographed from multiple angles and at different distances under multiple backgrounds and locations, and the images are labeled. The original sample images are processed, including random scaling, flipping, cropping, and adding noise, to expand the target sample dataset.

6. The method for fusion positioning of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 1, characterized in that: In step S5, when multiple fixed-end cameras detect the same target vehicle simultaneously, the output coordinates of the multiple fixed-end cameras are fused based on the weighted least squares method to obtain the vehicle positioning result based on the fixed-end cameras.

7. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 6, characterized in that: Fixed-end cameras that are closer to the target vehicle have higher weights in weighted least squares.

8. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 2, characterized in that: Step S5 includes, By transforming the pixels in the depth image to the coordinate system of the color camera through the transformation relationship between coordinate systems, the three-dimensional spatial coordinates of the point in the camera coordinate system are obtained. Based on the installation extrinsic parameters of the fixed-end camera in the cabin coordinate system, the coordinate points in the camera coordinate system are transformed to the cabin coordinate system to obtain the three-dimensional coordinates of the vehicle output by the fixed-end camera.

9. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 3, characterized in that: Step S5 includes, By calibrating the intrinsic and extrinsic parameters of the fixed-end camera, the plane equation of the cabin floor in the camera coordinate system is obtained, and then the position of any point on the plane in the camera coordinate system is obtained. The three-dimensional spatial coordinates of the target vehicle are obtained by utilizing the coordinate transformation relationship between the camera coordinate system and the cabin coordinate system.

10. The method for fusing localization of marine unmanned vehicles based on fixed-end and vehicle-end vision according to claim 1, characterized in that: Step S6 includes, Extract the feature point set P from the scene and classify the feature points based on the distance invariant principle; The cyclic feature point set is used to find three feature points with constant relative distance between frame k-1 and frame k, and put them into the same point set. The pose transformation matrix is ​​calculated for each motion state point set, and point sets with similar results are put into the same set. Using the walls and floor of the cabin as high-dimensional features to provide constraints for visual odometry, dynamic points in the scene are removed, and camera pose calculation is optimized.