A method and device for ship video positioning integrating artificial intelligence and BeiDou satellite navigation

By integrating artificial intelligence with BeiDou satellite navigation, and utilizing camera equipment and digital elevation maps, the distortion of video data is corrected and combined with BeiDou satellite navigation, thus solving the problem of low accuracy in ship video positioning and achieving high-precision ship positioning and visualization.

CN117132649BActive Publication Date: 2026-07-17HENAN PROVINCIAL COMM PLANNING & DESIGN INST CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HENAN PROVINCIAL COMM PLANNING & DESIGN INST CO LTD
Filing Date
2023-08-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Current technologies rely solely on simple video target tracking, resulting in low accuracy in ship video positioning and an inability to determine specific ship information.

Method used

By employing an artificial intelligence-integrated approach with BeiDou satellite navigation, video data is collected through camera equipment, and distortion correction and calibration are performed. Combined with digital elevation maps, the pixel positions of ships are identified and mapped to the world coordinate system. By comparing with BeiDou satellite navigation positioning information, latitude and longitude are corrected and elevation data is obtained, thereby improving positioning accuracy.

Benefits of technology

It improves the accuracy of ship video positioning, enables distance and depth perception of the environment in the video data field of view, and can visualize the actual geographical location of the ship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117132649B_ABST
    Figure CN117132649B_ABST
Patent Text Reader

Abstract

This application provides a method and apparatus for ship video positioning based on artificial intelligence and BeiDou satellite navigation. The method includes: acquiring video data of ship navigation through shore-based camera equipment; performing Zhang's calibration on the shore-based camera equipment; correcting distortion in the acquired video data; intelligently identifying the ship in the video data's field of view; identifying the ship's visual pixel coordinate system position based on the camera equipment's location; fusing a geographic information model to map the ship's positioning information from the pixel coordinate system to the geographic coordinate system; correcting the position information on a digital elevation map (DEM); correcting latitude and longitude and acquiring elevation data; and combining the ship's geographic location information with BeiDou satellite positioning information to visually display the ship's position in the real-time field of view. This application improves the accuracy of ship video positioning and provides a visual representation through this method of ship video positioning based on artificial intelligence and BeiDou satellite navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of maritime positioning technology, specifically a method and device for ship video positioning that integrates artificial intelligence with BeiDou satellite navigation. Background Technology

[0002] Artificial intelligence (AI) is a new technological science that studies and develops methods to simulate and extend human-like behavior. It comprises several parts, including AI theory, methods, technologies, and application systems. The development of the internet and the continuous improvement of computer performance have enabled AI to make significant progress in areas such as reinforcement learning, deep learning, and machine learning, leading to numerous research directions such as intelligent robots, speech recognition, pattern recognition, image recognition, expert systems, and natural language processing, resulting in a diversified development trend for AI. With the maturity of technologies such as data, algorithms, and computing power, AI, currently in its nascent stage, is beginning to truly solve problems and generate tangible economic benefits.

[0003] The BeiDou Navigation Satellite System consists of three segments: space segment, ground segment, and user segment. It can provide high-precision, high-reliability positioning, navigation, and timing services to various users around the world, 24 / 7. It also has short message communication capabilities and has initially achieved regional navigation, positioning, and timing capabilities. The positioning accuracy is at the decimeter and centimeter level, the velocity measurement accuracy is 0.2 meters per second, and the timing accuracy is 10 nanoseconds.

[0004] Currently, research combining artificial intelligence with the BeiDou Navigation Satellite System to solve ship video positioning problems is relatively limited. Relying solely on simple video target tracking results in low positioning accuracy for ships. Therefore, overcoming these technical problems and shortcomings is a key issue that needs to be addressed. Summary of the Invention

[0005] To overcome the problems of low accuracy and inability to determine specific ship information in existing technologies that rely solely on video target tracking, this application provides a ship video positioning method and apparatus that integrates artificial intelligence with BeiDou satellite navigation, employing the following technical solution:

[0006] Firstly, this application provides a method for ship video positioning based on artificial intelligence and BeiDou satellite navigation, including:

[0007] Step 1: Collect video data of the ship's navigation using pre-set camera equipment;

[0008] Step 2: The camera device is calibrated using Zhang's calibration algorithm to obtain the internal and external parameters of the camera device, which are then used to correct the distortion of the video data acquired by the camera device.

[0009] Step 3: Use intelligent ship analysis algorithms to identify the pixel positions of the ship in the field of view, and use the pixel coordinate system to map the pixel position information to the position information in the world coordinate system;

[0010] Step 4: Since the camera model can convert coordinates in three-dimensional geographic space into image point coordinates in two-dimensional image space, elevation information is lost. At the same time, the position of the ship in geographic space is also deviated. Here, we assume that the target is in contact with the ground. The monitoring camera emits a spectral ray, and the target centroid is used as the image space positioning point. The spectral ray is intersected with the digital elevation map (DEM) representing the terrain through the spatial positioning point to estimate the three-dimensional coordinates of the target in geographic space. Then, the position of the target positioning point in the image space is estimated in the three-dimensional geographic space.

[0011] Step 5: Correct latitude and longitude and obtain elevation data;

[0012] Step 6: Preprocess the ship navigation video data, identify the ships in the video data, and compare them in real time with the Beidou satellite navigation and positioning information, i.e., visualize the display.

[0013] Step 7: Compare the video positioning information with the BeiDou satellite navigation positioning information. If a match is found, mark the ship information in the video. If video positioning information exists but BeiDou satellite navigation positioning is lost, mark the ship as not having enabled BeiDou satellite navigation positioning settings. Other situations are unlikely and will not be considered.

[0014] Furthermore, the analysis of video data of ship navigation in steps 3 and 4 includes: a ship identification and analysis algorithm for video data from shore-based video surveillance equipment; and mapping pixel position information to position information in the world coordinate system using a pixel coordinate system.

[0015] Furthermore, step 4, obtaining the position information of the feature points of the different keyframes in three-dimensional space, further includes: converting the pixel coordinates of the feature points into coordinates in three-dimensional space, and using triangulation to correct the position information of the feature points on the digital elevation map.

[0016] Furthermore, step 5, which involves correcting latitude and longitude and obtaining elevation data, includes:

[0017] (1) Collect reference points at known locations in the camera equipment, including latitude and longitude information and elevation data of the reference points.

[0018] (2) Select some key frames and feature points extracted from the key frames, and match the feature points with reference points to establish the correspondence between the feature points and the reference points.

[0019] (3) Based on the correspondence between the feature point and the reference point, the latitude and longitude of the camera device are corrected, and the position of the feature point is mapped to the accurate latitude and longitude coordinates; based on the correspondence between the feature point and the reference point, the elevation data corresponding to the feature point is obtained.

[0020] Further, step 6, which involves preprocessing the video data of the ship's navigation environment to identify the ship in the video data, includes: obtaining the latitude and longitude of the camera device to determine its position, and extracting keyframes from the video data's field of view based on the camera device's position. A multi-target tracking method fusing YOLO and Deepsort frameworks is employed to read the keyframes in the video data's field of view, obtain the ship's position within the keyframes, track the ship's motion trajectory in the pixel coordinate system using a target tracking algorithm based on consecutive keyframes, and identify the detected ship in each frame using a convolutional neural network.

[0021] Secondly, this application also provides a ship video positioning device that integrates artificial intelligence with BeiDou satellite navigation, including:

[0022] The ship acquisition module is used to collect video data of ship navigation through preset camera equipment;

[0023] The camera equipment distortion correction module is used to calibrate the camera equipment using Zhang's calibration algorithm, obtain the internal and external parameters of the camera equipment, and perform distortion correction on the video data acquired by the camera equipment.

[0024] The world coordinate system position information acquisition module for ships is used to identify the pixel position of the ship in the field of view using intelligent ship analysis algorithms, and to map the pixel position information into position information in the world coordinate system using the pixel coordinate system.

[0025] The ship's 3D geospatial position estimation module is used because the camera model can convert coordinates in 3D geospatial space into image point coordinates in 2D image space, resulting in the loss of elevation information. At the same time, the position of the ship in the geospatial space also has deviations. Here, it is assumed that the target is in contact with the ground, the monitoring camera emits a spectral ray, and the target's centroid is used as the image space positioning point. The spectral ray is intersected with the digital elevation map (DEM) representing the terrain through the spatial positioning point to estimate the target's 3D coordinates in geospatial space. Then, the position of the target positioning point in the image space is estimated in 3D geospatial space.

[0026] The elevation data acquisition module is used to correct latitude and longitude and acquire elevation data;

[0027] The ship visualization module is used to preprocess the ship navigation video data, identify the ships in the video data, and compare them in real time with Beidou satellite navigation and positioning information, that is, to visualize the display.

[0028] The ship positioning module is used to compare video positioning information with BeiDou satellite navigation positioning information. If a match is found, the ship information is marked in the video. If video positioning information exists but BeiDou satellite navigation positioning is lost, the ship is marked as not having BeiDou satellite navigation positioning settings enabled. Other situations are impossible and are not considered.

[0029] Thirdly, this application provides an electronic device, comprising:

[0030] One or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, cause the device to perform the method as described in the first aspect.

[0031] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0032] Fifthly, this application provides a computer program that, when executed by a computer, performs the method described in the first aspect.

[0033] In one possible design, the program in the fifth aspect can be stored wholly or partially on a storage medium packaged with the processor, or it can be stored wholly or partially on a memory not packaged with the processor.

[0034] This application has the following beneficial effects:

[0035] This application acquires video data of ship navigation using camera equipment, performs Zhang's calibration on the camera equipment, corrects distortion in the acquired video data, identifies ships within the video data's field of view, determines the ship's visual position based on the camera equipment's location, maps the ship's pixel coordinate system position to its position in the world coordinate system, and combines this with location information from a digital elevation map (DEM) for correction, correcting latitude and longitude and acquiring elevation data; the ship's visual position is then mapped onto the DEM to obtain the ship's latitude, longitude, and elevation information. This application employs an artificial intelligence-integrated BeiDou satellite navigation ship video positioning method, using 3D point clouds to perform 3D reconstruction of the environment within the acquired video data's field of view, mapping the dynamic ship's positioning information from image space to the ship's actual geographical location in geospatial space. This improves the accuracy of distance and depth perception of the environment within the video data's field of view, enhances the accuracy of ship video positioning, and provides a visual representation. Attached Figure Description

[0036] Figure 1 This is an exemplary system architecture diagram to which embodiments of this application can be applied;

[0037] Figure 2 This is a flowchart illustrating the method of an embodiment of this application;

[0038] Figure 3 This is a schematic diagram of the imaging model of the camera device according to an embodiment of this application;

[0039] Figure 4 This is a schematic diagram of the imaging process of the camera device according to an embodiment of this application;

[0040] Figure 5 This is a schematic diagram illustrating the transformation process from the world coordinate system to the camera coordinate system in an embodiment of this application;

[0041] Figure 6 This is a flowchart illustrating step S3 in an embodiment of this application;

[0042] Figure 7 This is a flowchart of the feature point acquisition process for keyframes in an embodiment of this application;

[0043] Figure 8 This is a flowchart illustrating the process of obtaining elevation data according to an embodiment of this application;

[0044] Figure 9 This is a schematic diagram illustrating the actual location information of the vessel in the actual geographical location according to an embodiment of this application;

[0045] Figure 10 This is a schematic diagram of an apparatus according to an embodiment of this application;

[0046] Figure 11 This is a schematic diagram of a computer device according to an embodiment of this application. Detailed Implementation

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0048] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0049] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0050] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0051] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0052] Terminal devices 101, 102, and 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc.

[0053] Server 105 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal devices 101, 102, and 103.

[0054] It should be noted that the ship video positioning method integrating artificial intelligence and BeiDou satellite navigation provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the ship video positioning device integrating artificial intelligence and BeiDou satellite navigation is generally set in the server / terminal device.

[0055] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0056] Continue to refer to Figure 2 The figure shows a flowchart of the ship video positioning method based on artificial intelligence fusion with BeiDou satellite navigation according to this application. The method includes the following steps:

[0057] Step S1: Collect video data of the ship's navigation environment using preset camera equipment.

[0058] The process of collecting video data of the ship's navigation environment through preset camera equipment includes: collecting video data of the ship's course environment using shore-based video surveillance equipment and in-ship video surveillance equipment.

[0059] Step S2: The camera device is calibrated using Zhang's calibration algorithm to obtain the internal and external parameters of the camera device, which are then used to correct the distortion of the video data acquired by the camera device.

[0060] In one possible implementation, the camera device works by using projection to transform real three-dimensional world coordinates into two-dimensional camera coordinates. A schematic diagram of the camera device is shown below. Figure 3 As shown, assume the coordinate system of the camera device is O. c -x c -y c -zc Where z is the front of the camera device, x represents the direction to the right, y represents the direction to the down, O is the optical center, the real-world spatial point P is projected onto the physical imaging plane O′-xyz through the optical center O, and the imaging point is P′. The distance from the pixel plane center O′ to the optical center O is the focal length f of the camera device.

[0061] In one possible implementation, please refer to the imaging process of the camera device. Figure 4 During the imaging process of a camera device, four coordinate systems are involved: world coordinate system, camera coordinate system, image coordinate system, and pixel coordinate system. The transformation from the world coordinate system to the camera coordinate system is a rigid transformation process. The world coordinate system is obtained by rotation and translation to obtain the camera coordinate system. Please refer to [further details]. Figure 5 , Figure 5 The transformation from the world coordinate system to the camera coordinate system is shown. The top right point is a point in 3D space. Assuming that the origin of this point in 3D space is the world coordinate system, the coordinates of this point in 3D space are X. w If the origin of this point in 3D space is the camera coordinate system, then the coordinates of this point in 3D space are X. c During the transformation from the world coordinate system to the camera coordinate system, the world coordinate system can be transformed to coincide with the camera coordinate system by rotating R and translating t.

[0062] The transformation from camera coordinates to image coordinates involves perspective projection, converting from 3D to 2D; the transformation from image coordinates to pixel coordinates involves affine transformation; the transformation formula from world coordinates to pixel coordinates is:

[0063] Suppose that the three-dimensional spatial coordinates of a point P on the calibration plate are P = [X...]. w ,Y w Z w ] T The projected coordinates in the image space are p = [u, v]. T Let P be represented in homogeneous coordinate form. H =[X w ,Y w Z w ,1] T p H =[u,v,1] T The three-dimensional spatial coordinates of point P can be obtained from the imaging model. H With image space coordinates p H The correspondence is as follows:

[0064] z c p H =K[R t]P H (Formula 1)

[0065] The above equation z c The camera's intrinsic parameter matrix is ​​K, and its extrinsic parameter matrix consists of rotation matrix R and translation matrix t.

[0066]

[0067] f x f y Let $\frac{u0}{v0}$ be the camera focal length, $u0$ and $v0$ be the principal point coordinates of the image pixels, and $γ$ be the radial distortion parameter. Since a checkerboard is used as the calibration board here, all points are on the same plane. Therefore, assuming the plane z = 0, the formula can be simplified to:

[0068]

[0069] u and v represent coordinates in the pixel coordinate system, t is the translation vector, and let H = [h1 h2 h3] = λK[r1 r2 t], where λ is a constant factor. Then the correspondence between the three-dimensional spatial coordinates and the image spatial coordinates can be expressed as follows:

[0070] z c p H =HP H (Formula 4)

[0071] The H matrix, also known as the homography matrix, is the key to solving this problem. The H matrix is ​​a homogeneous matrix with 8 unknowns, requiring at least 8 equations to solve. Each pair of corresponding points provides two equations, so at least four pairs of points are needed to find the H matrix.

[0072] From the orthogonality property of rotation matrices, we can obtain: r1 T r2 = 0, r1 T r1 = r2 T Substituting r2 into formula 5, we get:

[0073]

[0074] Where h1 and h2 are the specific parameters of the homography matrix H, let B = K -T K -1 Then the symmetric matrix B can be written as:

[0075]

[0076] The intrinsic parameter matrix K has 5 unknowns, requiring at least 5 equations. A matrix H provides two equations; therefore, to derive matrix K from matrix B, at least 3 homography matrices H are needed, i.e., at least three images. The unknowns in B can be represented as a 6-dimensional vector b:

[0077] Let the i-th column in H be h. iFrom the derivation, we can obtain:

[0078]

[0079] Where h i1 h j1 h i1 h j2 h i2 h j1 h i2 h j2 h i3 h j1 h i1 h j3 h i3 h j2 h i2 h j3 h i3 h j3 Given the specific parameters of the homography matrix H, we can obtain the following from the constraints of Equation 7:

[0080]

[0081] Among them, v 11 v 12 v 22 These are the parameters of the intrinsic parameter matrix.

[0082] Step S3: Use a multi-target tracking method that integrates YOLO and DeepSort frameworks to identify the pixel position of the ship in the field of view, and use the pixel coordinate system to map the pixel position information to the position information in the world coordinate system.

[0083] DeepSort is a target tracking method based on the SORT framework, which combines Kalman filtering and Hungarian algorithms to avoid excessive occlusion by using high-precision detection results.

[0084] X(k|k-1)=A×X(k|k-1) (Formula 9)

[0085] X(k|k-1) is the result predicted using the previous state, X(k-1|k-1) is the optimal result of the previous state, and U(k) is the control variable of the current state. If there is no control variable, it can be 0.

[0086] P(k|k-1)=A×P(k|k-1)×A T +Q (Formula 10)

[0087] P(k|k-1) is the covariance of X(k|k-1), P(k-1|k-1) is the covariance of X(k-1|k-1), A TLet A be the transpose matrix, and Q be the covariance of the system process. Equations 1 and 2 are the first two of the five formulas for the Kalman filter, which are the predictions for the system.

[0088]

[0089] X(k|k)=P(k|k-1)+Kg(k)×[Z(k)-H×X(k|k-1)] (Formula 12)

[0090] Kg represents the Kalman gain, which yields the optimal estimate X(k|k) at state k. However, to ensure the Kalman filter continues operating until the system terminates, we must update the covariance of X(k|k) at state k.

[0091] P(k|k)=[I-Kg(k)×H]×P(k|k-1) (Formula 13)

[0092] Where I is a matrix of 1s, for a single model and single measurement, I = 1. When the system enters state k+1, P(k|k) is P(k-1|k-1) in Equation 13. In this way, the algorithm can continue its autoregressive operation.

[0093] During the assignment process, two indicators are used to integrate motion information and target appearance features. A bipartite graph matching method based on the Hungarian algorithm is employed to correlate the state predicted by the Kalman filter with the target detection measurement values ​​from the target detection algorithm. State target detection value 1:

[0094] d 1 (i,j)=(d j -y i ) T S i -1 (d j -y i )(Formula 14)

[0095] Where, d j y represents the position of the j-th detection box. i S represents the predicted bounding box position of the i-th tracker. i This represents the covariance matrix between the detected position and the average tracked position. The Mahalanobis distance accounts for the uncertainty of state measurement by calculating the standard deviation between the detected position and the average predicted position, and further accounts for it through the inverse χ² value. 2 The 95% confidence intervals calculated from the distribution are used to threshold the Mahalanobis distance. If the Mahalanobis distance of a given association is less than a specified threshold t... (1) If the motion state association is successful, the threshold was set to 9.4877 in the experiment. State target detection value 2:

[0096]

[0097] Where, r k (i) For feature vectors, The range of values.

[0098] For each detection box d j Find an eigenvector r j (The corresponding 128-dimensional feature vector is calculated using Reid's CNN network), with the constraint that ||r j || = 1. For each tracked target, a storage space is built to store the feature vectors of the most recent 100 successfully associated frames for that target. The second metric is to calculate the minimum cosine distance between the feature set of the most recent 100 successfully associated frames for the i-th tracker and the feature vector of the j-th detection result in the current frame.

[0099] The correlation method uses a linear weighted average of the two metrics as the final metric. The overall object detection measurement is as follows:

[0100] c i,j =λd 1 (i,j)+(1-λ)d 2 (i,j) (Formula 16)

[0101] Where, d 1 (i,j) is the Mahalanobis distance, d 2 (i,j) represents the cosine distance, and λ represents the weighting coefficient.

[0102] Note: Fusion is only performed when both metrics meet their respective threshold conditions. Distance metrics work well for short-term predictions and matching, but for long-term occlusion, appearance feature metrics are more effective. For cases with camera motion, λ = 0 can be set. However, the Mahalanobis distance threshold still applies; if the first metric's criteria are not met, the match cannot proceed to step c. (i,j) The integration phase.

[0103] Therefore, the MOG2 algorithm is used to select video frames containing foreground targets based on a differential detection strategy. Cross-frame detection significantly improves detection efficiency. The extracted contours are then used as map symbols for visualization in geospatial representation. Next, the YOLOv3 algorithm based on deep learning is used for target detection, and the DeepSort algorithm is used for multi-target tracking. The video stream data of the selected keyframes is fed into the YOLOv3 detector, which outputs bounding boxes, categories, and confidence scores. This output is then fed into the DeepSort multi-target tracker, where improved recursive Kalman filtering is used to predict and track positions. The Mahalanobis distance and the cosine distance of the depth descriptor are used as the fused metric, and the Hungarian algorithm is used for cascaded matching to output dynamic tracking and localization information. See the detailed process below. Figure 6 The algorithm steps are as follows:

[0104] Step 301: Input the video stream and perform moving target differential detection.

[0105] Step 302: Filter out frames with foreground targets and calculate the Euclidean distance between the bounding rectangles of the preceding and following frames. Frames with a distance greater than the threshold are marked as non-detection frames.

[0106] Step 303: Input the detection frame into the YOLO target detector and output the four-dimensional vector of the bounding box, along with the category and confidence score.

[0107] Step 304: Use the YOLO high-precision detection results as input to the DeepSort multi-target detector, and after matching them with the Kalman prediction information, determine the tracking result.

[0108] Step S4: Since the camera model can convert coordinates in three-dimensional geographic space into image point coordinates in two-dimensional image space, elevation information is lost. At the same time, the position of the ship in the geographic space is also deviated. Here, it is assumed that the target is in contact with the ground. The monitoring camera emits a spectral ray, and the target centroid is used as the image space positioning point. The spectral ray is intersected with the digital elevation map (DEM) representing the terrain through the spatial positioning point to estimate the three-dimensional coordinates of the target in geographic space. Then, the position of the target positioning point in the image space is estimated in the three-dimensional geographic space.

[0109] Please refer to the flowchart for keyframe feature point acquisition. Figure 7 The step of using machine vision algorithms to obtain feature points of the keyframes includes:

[0110] Step 401: Detect the corner points of the keyframes and obtain the key points in the keyframes.

[0111] Step 402: Annotate the key points and obtain the feature points of the key points.

[0112] Step 403: Obtain the same feature points based on the adjacent keyframes to form feature point pairs of the adjacent keyframes.

[0113] Step 404 is used to match feature points of adjacent keyframes.

[0114] In one possible implementation, the digital elevation map (DEM) is acquired as follows: key frames are extracted from the video data field of view based on the position of the camera device; feature points of the key frames are acquired using machine vision algorithms; feature points of adjacent key frames are matched to establish continuous spatiotemporal information between different key frames; based on the continuous spatiotemporal information between different key frames, the positional information of the feature points of different key frames in three-dimensional space is acquired to generate an initial sparse three-dimensional point cloud; the sparse three-dimensional point cloud is interpolated, filled, and fitted to generate a dense three-dimensional point cloud; and the environment in the acquired video data field of view is reconstructed in three dimensions based on the dense three-dimensional point cloud to obtain a digital elevation map (DEM) with elevation data.

[0115] Calculate two-dimensional geographic coordinates using the inverse operation of the homography matrix:

[0116] At this point, the camera projection matrix P is transformed into the homography matrix H, mapping points in the world coordinate system to the image coordinate system. Assuming a point m in the image has x and y coordinates, its corresponding world coordinate point is M. w ,Y w Let M be the coordinates of the world coordinate point.

[0117] m = [x, y, 1] T (Formula 17)

[0118] M = [X] w ,Y w [,0,1] T (Formula 18)

[0119] m = HM (Formula 19)

[0120] Right now

[0121]

[0122] in

[0123]

[0124] The H matrix obtained above is a mapping matrix that transforms object space points on the plane into image space through perspective transformation. To solve for the projection of image space points into object space, it is necessary to invert the H matrix, i.e.

[0125]

[0126] H-1 =(K[r1,r2,t]) -1 (Formula 23)

[0127] When the world coordinate elevation is assumed to be 0, i.e., it is regarded as a plane, the H matrix is ​​obtained by calculating the camera intrinsic parameter matrix K and the extrinsic parameter matrix [r1,r2|t].

[0128] Given a calibrated surveillance camera with intrinsic and extrinsic parameters, and the corresponding target location in image space, construct the view ray: (X0,Y0,Z0)+k(U,V,W), where (X0,Y0,Z0) are the true coordinates of the surveillance camera in 3D geographic space, (U,V,W) are the unit vectors of the observation ray direction emanating from the camera's principal optical axis, and k≥0 represents any distance. Determining the intersection point of the view ray with the 3D scene is typically very complex. However, when the scene information is stored as a DEM, a simple geometric traversal algorithm based on the Bresenham algorithm for drawing digital line segments is used. Consider projecting the view ray perpendicularly onto the DEM grid. Starting from the DEM grid (X0,Y0) where the surveillance camera is located, sequentially check each grid (X,Y) it passes outwards until the elevation value stored in the DEM grid exceeds the Z-direction component of the 3D view ray at that location. The Z-value at (X,Y) can be calculated using the formula:

[0129]

[0130] The previous section derived a rigorous mapping model from 2D image space to 3D geographic space, using DEM to provide third-dimensional information constraints to solve the problem of locating targets in images in 3D geographic space. However, this requires high accuracy of DEM data, and obtaining decimeter-level DEM data is quite difficult, increasing the complexity of the problem. Furthermore, the field of view of surveillance cameras is usually largely planar, and researchers are more concerned with planar areas such as roads and squares. We can assume that the ground is a plane for our study, thus freeing the solution from the constraints of DEM and simplifying the mapping model.

[0131] Let the world coordinates of a point P in space be (X... w ,Y w Z w The coordinates (X, Y) in the camera coordinate system can be converted using the rotation matrix R and the translation vector t. c ,Y c Z c ), coordinates (X) c ,Y c Z c The image coordinates (u,v) and its corresponding image coordinates have the following perspective projection relationship:

[0132]

[0133] In the formula, f x f y d is the camera focal length. x d y Let be the physical pixel dimensions of the camera sensor in the horizontal and vertical directions, u0 and v0 be the principal point coordinates of the image pixels, and K be the intrinsic parameter matrix determined only by parameters related to the internal structure of the camera. [R|t] is the extrinsic parameter matrix determined by the rotation matrix R and translation vector t of the camera relative to the world coordinate system. P is the camera projection matrix. Once the camera projection matrix P is determined, a spatial point in the world coordinate system can be uniquely determined by matrix P to its corresponding image spatial point.

[0134]

[0135] From the above equation, we can obtain that when the image points are projected by the camera projection matrix P, the pseudo-inverse matrix P... -1 When inversely mapped to 3D space, this system of equations has no unique solution; that is, the coordinates of image points cannot uniquely determine the 3D coordinates of spatial points in the world coordinate system. More spatial point information is needed to assist in obtaining the third-dimensional information. This is achieved using the pseudo-inverse matrix P. -1 Solving for the 3D coordinates of a point in space requires solving for a large number of parameters. The matrix has 11 degrees of freedom. Each set of corresponding points in 3D space and image coordinates can provide two equations. To solve this matrix, at least 6 sets of corresponding points need to be selected to complete the task, which is quite cumbersome.

[0136] Step S5: Correct latitude and longitude and obtain elevation data. Please continue to refer to [the relevant documentation / reference]. Figure 8 Step 5, which involves correcting latitude and longitude and obtaining elevation data, includes:

[0137] Step 501: Collect reference points at known locations in the camera device, including latitude and longitude information and elevation data of the reference points.

[0138] Step 502: Select a portion of keyframes and feature points extracted from the keyframes, and match the feature points with reference points to establish the correspondence between the feature points and the reference points.

[0139] Step 503: Based on the correspondence between the feature points and the reference points, the latitude and longitude of the camera device are corrected to map the position of the feature points to accurate latitude and longitude coordinates; based on the correspondence between the feature points and the reference points, the elevation data corresponding to the feature points are obtained.

[0140] Step S6: Preprocess the ship navigation video data, identify the ships in the video data, and compare them in real time with Beidou satellite navigation and positioning information, i.e., visualize the display.

[0141] The preprocessing of the video data of the ship navigation environment to identify the ship in the video data includes: obtaining the latitude and longitude of the camera device to determine the position of the camera device; extracting keyframes in the field of view of the video data based on the position of the camera device; reading the keyframes in the field of view of the video data using a multi-target tracking method that combines YOLO and DeepSort; applying a target detection algorithm to the images in the keyframes to obtain the position of the ship in the keyframes; using a target tracking algorithm to track the motion trajectory of the ship in the pixel coordinate system based on continuous keyframes; and using a convolutional neural network to identify the ship detected in each frame.

[0142] Step S7: Compare the video positioning information with the BeiDou satellite navigation positioning information. If a match is found, mark the ship information in the video. If video positioning information exists but BeiDou satellite navigation positioning is lost, mark the ship as not having enabled BeiDou satellite navigation positioning settings. Other situations are impossible and will not be considered.

[0143] In one possible implementation, the ship's position information acquired by the camera device is only the image coordinates, lacking the latitude, longitude, and elevation information in the actual geographical environment. Please refer to [further details]. Figure 9 This application maps the acquired visual position information of the ship onto a digital elevation map (DEM). Based on the intersection of the visual position of the ship in the camera device and the digital elevation map (DEM), the latitude, longitude and elevation information of the ship in the digital elevation map (DEM) can be obtained, thereby obtaining the real position information of the ship in the actual geographical location.

[0144] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0145] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0146] Continue to refer to Figure 10 The ship video positioning device integrating artificial intelligence and Beidou satellite navigation described in this embodiment includes: a ship acquisition module 1001, a camera equipment distortion correction module 1002, a ship world coordinate system position information acquisition module 1003, a ship three-dimensional geospatial position estimation module 1004, an elevation data acquisition module 1005, a ship visualization display module 1006, and a ship positioning module 1007.

[0147] The ship acquisition module 1001 is used to collect video data of ship navigation through preset camera equipment;

[0148] The camera equipment distortion correction module 1002 is used to calibrate the camera equipment using Zhang's calibration algorithm, obtain the internal and external parameters of the camera equipment, and perform distortion correction on the video data acquired by the camera equipment.

[0149] The world coordinate system position information acquisition module 1003 for ships is used to identify the pixel position of the ship in the field of view using intelligent ship analysis algorithms, and to map the pixel position information into position information in the world coordinate system using the pixel coordinate system.

[0150] The three-dimensional geospatial position estimation module 1004 for ships is used because the camera model can convert coordinates in three-dimensional geospatial space into image point coordinates in two-dimensional image space, resulting in the loss of elevation information. At the same time, the position of the ship mapped in geospatial space is also biased. Here, it is assumed that the target is in contact with the ground, the monitoring camera emits a spectral ray, the centroid of the target is used as the image space positioning point, the spectral ray is intersected with the digital elevation map (DEM) representing the terrain through the spatial positioning point to estimate the three-dimensional coordinates of the target in geospatial space, and then the position of the target positioning point in the image space is estimated in the three-dimensional geospatial space.

[0151] The elevation data acquisition module 1005 is used to correct latitude and longitude and acquire elevation data.

[0152] The ship visualization display module 1006 is used to preprocess the ship navigation video data, identify the ships in the video data, and compare them in real time with Beidou satellite navigation and positioning information, that is, to display them in a visual way.

[0153] The ship positioning module 1007 is used to compare video positioning information with BeiDou satellite navigation positioning information. If a match is found, the ship information is marked in the video. If video positioning information exists but BeiDou satellite navigation positioning is lost, the ship is marked as not having BeiDou satellite navigation positioning settings enabled. Other situations are impossible and are not considered.

[0154] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 11 , Figure 11 This is a basic structural block diagram of the computer device in this embodiment.

[0155] The computer device 11 includes a memory 11a, a processor 11b, and a network interface 11c that are interconnected via a system bus. It should be noted that only the computer device 11 with components 11a-11c is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0156] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0157] The memory 11a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11a may be an internal storage unit of the computer device 11, such as the hard disk or memory of the computer device 11. In other embodiments, the memory 11a may also be an external storage device of the computer device 11, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 11. Of course, the memory 11a may also include both the internal storage unit and the external storage device of the computer device 11. In this embodiment, the memory 11a is typically used to store the operating system and various application software installed on the computer device 11, such as the program code for a ship video positioning method and device that integrates artificial intelligence and BeiDou satellite navigation. Furthermore, the memory 11a can also be used to temporarily store various types of data that have been output or will be output.

[0158] In some embodiments, the processor 11b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 11b is typically used to control the overall operation of the computer device 11. In this embodiment, the processor 11b is used to run program code stored in the memory 11a or process data, for example, to run the program code of the artificial intelligence-integrated BeiDou satellite navigation ship video positioning method and device.

[0159] The network interface 11c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 11 and other electronic devices.

[0160] This application also provides another embodiment, namely, providing a non-volatile computer-readable storage medium storing a program for a ship video positioning method and apparatus based on artificial intelligence fusion with BeiDou satellite navigation. The ship video positioning method and apparatus based on artificial intelligence fusion with BeiDou satellite navigation can be executed by at least one processor to perform the steps of the ship video positioning method and apparatus based on artificial intelligence fusion with BeiDou satellite navigation as described above.

[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0162] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A ship video positioning method integrating artificial intelligence and BeiDou satellite navigation, characterized in that, include: Step 1: Collect video data of ship navigation using pre-set shore-based camera equipment; Step 2: The camera device is calibrated using Zhang's calibration algorithm to obtain the internal and external parameters of the camera device, which are then used to correct the distortion of the video data acquired by the camera device. Step 3: Use intelligent ship analysis algorithms to identify the pixel positions of the ship in the field of view, and use the pixel coordinate system to map the pixel position information to the position information in the world coordinate system; Step 4: Since the camera model can convert coordinates in three-dimensional geographic space into image point coordinates in two-dimensional image space, elevation information is lost. At the same time, the position of the ship in geographic space is also deviated. Here, we assume that the target is in contact with the ground. The monitoring camera emits a spectral ray, and the target centroid is used as the image space positioning point. The spectral ray is intersected with the digital elevation map (DEM) representing the terrain through the spatial positioning point to estimate the three-dimensional coordinates of the target in geographic space. Then, the position of the target positioning point in the image space is estimated in the three-dimensional geographic space. Step 5, based on the digital elevation map (DEM), correct the latitude and longitude and obtain elevation data, including: (1) Collect reference points at known locations in shore-based camera equipment, including latitude and longitude information and elevation data of the reference points; (2) Select some key frames and feature points extracted from the key frames, and match the feature points with reference points to establish the correspondence between the feature points and the reference points; (3) Based on the correspondence between the feature point and the reference point, the latitude and longitude of the camera device are corrected, and the position of the feature point is mapped to the accurate latitude and longitude coordinates; based on the correspondence between the feature point and the reference point, the elevation data corresponding to the feature point is obtained; The digital elevation map is acquired as follows: key frames are extracted from the video data field of view based on the position of the camera device; feature points of the key frames are obtained using machine vision algorithms; feature points of adjacent key frames are matched to establish continuous spatiotemporal information between different key frames; based on the continuous spatiotemporal information between different key frames, the positional information of the feature points of different key frames in three-dimensional space is obtained to generate an initial sparse three-dimensional point cloud; the sparse three-dimensional point cloud is interpolated, filled, and fitted to generate a dense three-dimensional point cloud; based on the dense three-dimensional point cloud, the environment in the acquired video data field of view is reconstructed in three dimensions to obtain a digital elevation map with elevation data. Step 6: Preprocess the ship navigation video data, identify the ships in the video data, and compare them in real time with the Beidou satellite navigation and positioning information, i.e., visualize the display. Step 7: Compare the video positioning information with the BeiDou satellite navigation positioning information. If a match is found, mark the ship information in the video. If video positioning information exists but BeiDou satellite navigation positioning is lost, mark the ship as not having enabled BeiDou satellite navigation positioning settings.

2. The ship video positioning method based on artificial intelligence and BeiDou satellite navigation according to claim 1, characterized in that, The analysis of preset ship navigation video data in steps 3 and 4 includes: A ship identification and analysis algorithm based on video data from shore-based video surveillance equipment; The pixel coordinate system maps pixel position information to position information in the world coordinate system.

3. The ship video positioning method based on artificial intelligence and BeiDou satellite navigation according to claim 1, characterized in that... Step 6, which involves preprocessing the video data of the vessel's navigation to identify the vessel within the video data, includes: The latitude and longitude of the camera device are obtained to determine its location. Keyframes in the video data field of view are extracted based on the location of the camera device. YOLO and DeepSort are used to read the keyframes in the video data field of view. An object detection algorithm is applied to the images in the keyframes to obtain the ship's position in the keyframes. Based on continuous keyframes, an object tracking algorithm is used to track the ship's motion trajectory in the pixel coordinate system. A convolutional neural network is used to identify the detected ship in each frame.

4. A ship video positioning device integrating artificial intelligence and BeiDou satellite navigation, characterized in that, include: The ship acquisition module is used to collect video data of ship navigation through preset shore-based camera equipment; The camera equipment distortion correction module is used to calibrate the camera equipment using Zhang's calibration algorithm, obtain the internal and external parameters of the camera equipment, and perform distortion correction on the video data acquired by the camera equipment. The world coordinate system position module for ships is used to identify the pixel position of the ship in the field of view using intelligent ship analysis algorithms, and to map the pixel position information into position information in the world coordinate system using the pixel coordinate system. The ship's 3D geospatial position estimation module is used because the camera model can convert coordinates in 3D geospatial space into image point coordinates in 2D image space, resulting in the loss of elevation information. At the same time, the position of the ship in the geospatial space also has deviations. Here, it is assumed that the target is in contact with the ground, the monitoring camera emits a spectral ray, and the target's centroid is used as the image space positioning point. The spectral ray is intersected with the digital elevation map (DEM) representing the terrain through the spatial positioning point to estimate the target's 3D coordinates in geospatial space. Then, the position of the target positioning point in the image space is estimated in 3D geospatial space. The elevation data acquisition module, used to correct latitude and longitude and acquire elevation data based on the digital elevation map (DEM), includes: (1) Collect reference points at known locations in shore-based camera equipment, including latitude and longitude information and elevation data of the reference points; (2) Select some key frames and feature points extracted from the key frames, and match the feature points with reference points to establish the correspondence between the feature points and the reference points; (3) Based on the correspondence between the feature point and the reference point, the latitude and longitude of the camera device are corrected, and the position of the feature point is mapped to the accurate latitude and longitude coordinates; based on the correspondence between the feature point and the reference point, the elevation data corresponding to the feature point is obtained; The digital elevation map is acquired as follows: key frames are extracted from the video data field of view based on the position of the camera device; feature points of the key frames are obtained using machine vision algorithms; feature points of adjacent key frames are matched to establish continuous spatiotemporal information between different key frames; based on the continuous spatiotemporal information between different key frames, the positional information of the feature points of different key frames in three-dimensional space is obtained to generate an initial sparse three-dimensional point cloud; the sparse three-dimensional point cloud is interpolated, filled, and fitted to generate a dense three-dimensional point cloud; based on the dense three-dimensional point cloud, the environment in the acquired video data field of view is reconstructed in three dimensions to obtain a digital elevation map with elevation data. The ship visualization module is used to preprocess the ship navigation video data, identify the ships in the video data, and compare them in real time with Beidou satellite navigation and positioning information, that is, to visualize the display. The ship positioning module is used to compare video positioning information with BeiDou satellite navigation positioning information. If a match is found, the ship information is marked in the video; if video positioning information exists but BeiDou satellite navigation positioning is lost, the ship is marked as not having BeiDou satellite navigation positioning settings enabled.

5. An electronic device, characterized in that, include: One or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, the one or more computer programs including instructions that, when executed by the device, cause the device to perform the method as described in any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 3.