underwater vehicle positioning methods, devices, systems, and underwater vehicles
By using an event camera and a multi-medium refraction correction model, the problem of low visual positioning accuracy of traditional frame cameras in underwater environments was solved, enabling AUVs to achieve high-precision pose calculation and safe docking in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CORP LTD
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional frame cameras suffer from low visual positioning accuracy in underwater environments due to motion blur, overexposure, and optical refraction distortion, making it difficult to meet the reliability and safety requirements of AUVs in high-precision docking and recovery missions.
The target event stream is acquired using an event camera. Refraction distortion correction is performed using sub-pixel level center coordinates, and combined with a multi-medium refraction correction model to achieve high-frequency, high-precision pose calculation. The pose is then updated using an extended Kalman filter algorithm.
It provides a highly reliable asynchronous visual positioning solution with extremely low latency and resistance to ambient light interference, ensuring the safety and positioning accuracy of closed-loop control of AUVs under high-speed relative motion and extreme lighting conditions.
Smart Images

Figure CN122176058B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of underwater navigation and positioning technology, and more specifically, to an underwater vehicle positioning method, device, system, and underwater vehicle. Background Technology
[0002] When performing underwater docking, target recovery, and close-range cooperative operations, AUVs (Autonomous Underwater Vehicles) generally rely on optical guidance systems to provide high-precision relative pose information. Traditional solutions in related technologies employ frame rate-based CMOS or CCD cameras to acquire images, extract feature points on the target surface (such as LED markings and structured light patterns), and combine this with the PnP algorithm to calculate the six-degree-of-freedom pose between the AUV and the target platform. To improve recognition robustness, continuously emitting or low-frequency flashing LED arrays are often deployed on the docking target as active guidance sources, and feature regions are extracted using image processing techniques (such as thresholding and morphological filtering).
[0003] However, the underwater environment exhibits strong light absorption and scattering characteristics. Backscattering caused by suspended particles severely reduces image contrast, making feature extraction difficult. Simultaneously, dynamic caustic spots formed by wave refraction create instantaneous, non-uniform bright areas on the target surface, leading to localized overexposure and loss of target details in traditional cameras. Furthermore, frame cameras sample entire frames at a fixed frequency. When AUVs are affected by swells or rapid maneuvers, the motion during exposure causes severe image blurring, inaccurate feature point localization, and a significantly increased PnP calculation failure rate. Their limited dynamic range also makes it difficult to handle scenes with both strong guide lights and dark backgrounds, further exacerbating center shift. More critically, traditional image processing relies on frame synchronization, which has inherent latency (typically tens of milliseconds). In high-speed closed-loop control, this leads to command lag and increases the risk of collisions. Additionally, AUV cameras are typically encapsulated within pressure vessels, requiring light to pass sequentially through water, a glass window, and an air cavity. The difference in refractive index causes nonlinear refractive distortion. If the projection calculation is still performed using a pinhole camera model in the air, it will severely compromise positioning accuracy.
[0004] In summary, among the relevant technologies, traditional frame cameras, due to their inherent frame rate sampling mechanism, limited dynamic range, and fixed delay, are difficult to adapt to the complex underwater environment of high dynamics, strong interference, and multi-medium refraction. This leads to the visual positioning system being unstable under conditions of high-speed movement, drastic changes in lighting, and optical distortion, resulting in a sharp drop in positioning accuracy. This severely restricts the reliability and safety of AUVs in high-precision docking and recovery missions.
[0005] There is currently no effective solution to the above problems. Summary of the Invention
[0006] This application provides an underwater vehicle positioning method, apparatus, system, and underwater vehicle to at least solve the technical problem in the related art of low visual positioning accuracy caused by motion blur, overexposure, and optical refraction distortion in traditional frame cameras under underwater dynamic interference environments.
[0007] According to one aspect of the embodiments of this application, an underwater vehicle positioning method is provided, comprising: acquiring a target event stream collected by an event camera, wherein the event camera is disposed on the underwater vehicle, and the target event stream is used to characterize information on brightness polarity change events at different pixel positions; determining an event point cloud whose event trigger frequency conforms to a target frequency based on the target event stream, and determining sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud, wherein the target frequency is the flashing frequency of a guide light source disposed on a docking target corresponding to the underwater vehicle, and the sub-pixel-level center coordinates correspond one-to-one with the guide light source; mapping the sub-pixel-level center coordinates into a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the sub-pixel-level center coordinates, wherein the refraction distortion correction processing is used to compensate for geometric distortion caused by differences in refractive index when light propagates through different media, and the three-dimensional line-of-sight vector is used to characterize the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source; and determining the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
[0008] Optionally, acquiring the target event stream captured by the event camera includes: acquiring the original event stream captured by the event camera, wherein the original event stream contains multiple brightness polarity change events, as well as the timestamps, pixel positions, and brightness polarity states corresponding to the brightness polarity change events; for each brightness polarity change event in the original event stream, determining the first pixel position and the first timestamp corresponding to the brightness polarity change event, and counting the number of brightness polarity change events within a preset spatial neighborhood of the first pixel position and within a preset time window length of the first timestamp in the original event stream; if the number of events is less than a preset event number threshold, the current brightness polarity change event is determined to be a noise event, and the noise event is filtered out from the original event stream, thereby obtaining the target event stream.
[0009] Optionally, determining the event point cloud whose event triggering frequency matches the target frequency based on the target event stream includes: establishing an event time series corresponding to each pixel position in the target event stream, wherein a series of brightness polarity change events generated by the pixel position are arranged in timestamp order in the event time series; determining the time interval between adjacent events in the event time series, and using a frequency bandpass filter to determine whether the time interval meets the frequency constraint corresponding to the target frequency, wherein the frequency bandpass filter is used to filter brightness polarity change events consistent with the flashing frequency of the guide light source with the target frequency as the center frequency and a preset bandwidth as the passband range, and the frequency constraint is used to characterize that the difference between the instantaneous frequency corresponding to the time interval of adjacent events and the target frequency is within a preset tolerance range; statistically analyzing the proportion of events whose time interval meets the frequency constraint at each pixel position, and determining the signal saliency parameter corresponding to the pixel position based on the proportion, wherein the signal saliency parameter is used to characterize the degree of conformity between the event triggering frequency of the pixel position and the flashing frequency of the guide light source; generating an event point cloud based on the events at pixel positions whose signal saliency parameter is greater than the target saliency threshold, wherein the target saliency threshold is determined based on the turbidity of the underwater environment.
[0010] Optionally, determining the sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud includes: spatially clustering the event point cloud to obtain multiple event clusters, where each event cluster corresponds to a guiding light source; validating each event cluster, and deleting an event cluster if the number of events contained within the event cluster is less than a minimum threshold; for each valid event cluster, calculating the average pixel position corresponding to each brightness polarity change event contained within the event cluster to obtain the sub-pixel-level center coordinates corresponding to the event cluster.
[0011] Optionally, mapping the subpixel-level center coordinates to a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the subpixel-level center coordinates includes: using the camera intrinsic parameter matrix of the event camera to convert the subpixel-level center coordinates into a unit line-of-sight vector in the camera coordinate system; determining the first refractive index ratio between the air medium and the glass medium, and the second refractive index ratio between the glass medium and the water medium in the current underwater environment, and establishing a target mapping function based on the first and second refractive index ratios, wherein the target mapping function is used to characterize the transformation relationship of the direction vector when light propagates through the water-glass-air multilayer medium; and converting the unit line-of-sight vector according to the target mapping function to obtain the three-dimensional line-of-sight vector corresponding to the subpixel-level center coordinates.
[0012] Optionally, determining the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target includes: determining the preset spatial coordinates of the guide light source set on the docking target, and establishing a reprojection error equation based on the preset spatial coordinates and the three-dimensional line-of-sight vector. The reprojection error equation is used to characterize the deviation relationship between the observed three-dimensional line-of-sight vector and the line-of-sight vector reprojected based on the current pose estimate. By minimizing the angular residual between the three-dimensional line-of-sight vector and the reprojection vector, the reprojection error equation is solved to obtain the target pose of the underwater vehicle. The target pose includes the following information: the rotation matrix and translation vector of the underwater vehicle relative to the docking target.
[0013] Optionally, the method further includes: determining the state vector and inertial measurement data of the underwater vehicle at the current moment, wherein the state vector is used to characterize the navigation state parameters and sensor error compensation parameters of the underwater vehicle, and the inertial measurement data is the motion state differential information collected by the inertial measurement unit on the underwater vehicle; using the extended Kalman filter algorithm, predicting the pose prediction value of the underwater vehicle at the next moment based on the state vector and inertial measurement data at the current moment; when the sub-pixel-level center coordinates observed by the event camera are updated, determining the residual between the newly observed sub-pixel-level center coordinates and the predicted pixel coordinates, and correcting the pose prediction value based on the residual to obtain the updated pose estimate, wherein the predicted pixel coordinates are the coordinates of the pose prediction value mapped onto the image plane of the event camera after refraction distortion correction.
[0014] According to another aspect of the embodiments of this application, an underwater vehicle positioning device is also provided, comprising: an event stream acquisition module, configured to acquire a target event stream collected by an event camera, wherein the event camera is mounted on the underwater vehicle, and the target event stream is used to characterize information on brightness polarity change events at different pixel positions; and a guidance feature extraction module, configured to determine, based on the target event stream, an event point cloud whose event triggering frequency matches a target frequency, and determine sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud, wherein the target frequency is the flashing frequency of a guidance light source mounted on the docking target corresponding to the underwater vehicle. The subpixel-level center coordinates correspond one-to-one with the guide light source; the refraction distortion correction module is used to map the subpixel-level center coordinates into a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the subpixel-level center coordinates. The refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media. The three-dimensional line-of-sight vector is used to represent the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source; the target pose determination module is used to determine the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
[0015] According to another aspect of the embodiments of this application, an underwater vehicle positioning system is also provided, including: an underwater vehicle equipped with an event camera, and a docking target corresponding to the underwater vehicle, wherein the docking target is equipped with multiple guide light sources, and the guide light sources flash light according to the target frequency; the underwater vehicle is used to collect target event streams through the event camera, wherein the target event stream is used to characterize information on brightness polarity change events at different pixel positions; based on the target event stream, an event point cloud with event triggering frequencies conforming to the target frequency is determined, and based on the pixel coordinates corresponding to the event point cloud, sub-pixel-level center coordinates are determined, wherein the sub-pixel-level center coordinates correspond one-to-one with the guide light sources; by performing refraction distortion correction processing on the sub-pixel-level center coordinates, the sub-pixel-level center coordinates are mapped into a three-dimensional line-of-sight vector, wherein the refraction distortion correction processing is used to compensate for geometric distortion caused by differences in refractive index when light propagates through different media, and the three-dimensional line-of-sight vector is used to characterize the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source; based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light sources on the docking target, the target pose of the underwater vehicle is determined.
[0016] According to another aspect of the embodiments of this application, an underwater vehicle is also provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes an underwater vehicle positioning method during runtime.
[0017] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored computer program, wherein the device containing the non-volatile storage medium executes an underwater vehicle positioning method by running the computer program.
[0018] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of an underwater vehicle positioning method.
[0019] In this embodiment, a target event stream acquired by an event camera is used. The event camera is mounted on an underwater vehicle, and the target event stream is used to characterize information about brightness polarity change events at different pixel positions. Based on the target event stream, an event point cloud with event trigger frequencies matching the target frequency is determined. Sub-pixel-level center coordinates are then determined based on the pixel coordinates corresponding to the event point cloud. The target frequency is the flashing frequency of the guide light source installed on the docking target corresponding to the underwater vehicle, and the sub-pixel-level center coordinates correspond one-to-one with the guide light source. By performing refraction distortion correction processing on the sub-pixel-level center coordinates, they are mapped to a three-dimensional line-of-sight vector. The refraction distortion correction processing is used to compensate for geometric distortion caused by differences in refractive index when light propagates through different media. The three-dimensional line-of-sight vector is used to characterize the direction from the optical center of the event camera to the underwater guide light source. The direction vector after refraction correction; based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guiding light source on the docking target, the target pose of the underwater vehicle is determined. By introducing an event camera with high temporal resolution and ultra-high dynamic range, combined with a spatiotemporal filtering algorithm for underwater planktonic noise and a multi-medium refraction correction model, high-frequency and high-precision feature extraction of a specific frequency scintillation guiding source is achieved. This provides an AUV with a highly reliable asynchronous visual positioning scheme with extremely low latency, resistance to ambient light interference, and effective compensation for optical distortion in close-range docking or recovery missions. This ensures the closed-loop control safety and positioning accuracy of the AUV under high-speed relative motion and extreme lighting conditions, thereby solving the technical problem of low visual positioning accuracy caused by motion blur, overexposure, and optical refraction distortion of traditional frame cameras in underwater dynamic interference environments. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0021] Figure 1 This is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a method for locating an underwater vehicle, according to an embodiment of this application.
[0022] Figure 2 This is a schematic diagram of a method for locating an underwater vehicle according to an embodiment of this application;
[0023] Figure 3 This is a schematic diagram of a method for asynchronous visual positioning of an autonomous underwater vehicle against dynamic interference, provided in an embodiment of this application.
[0024] Figure 4 This is a schematic diagram of an optically guided retrieval method according to an embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the module architecture of an underwater vehicle positioning device according to an embodiment of this application. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] To facilitate a better understanding of the embodiments of this application by those skilled in the art, some technical terms or nouns involved in the embodiments of this application are explained as follows:
[0029] AUV (Autonomous Underwater Vehicle): An intelligent underwater vehicle that is unmanned, untethered, self-powered, and completes underwater tasks through autonomous decision-making.
[0030] PnP (Perspective-n-Point): Given n 3D world coordinate points and their 2D projections onto the image, calculate the six-degree-of-freedom relative pose between the camera and the object (rotation matrix R + translation vector t).
[0031] Autonomous underwater vehicles (AUVs) rely heavily on optical guidance systems to provide high-precision relative pose information when performing underwater docking, target recovery, and close-range collaborative operations. However, the strong absorption and scattering of light by water makes it difficult for traditional visual guidance methods in related technologies to achieve accurate positioning.
[0032] First, light is subject to severe scattering and absorption when propagating underwater. In particular, backscattering caused by suspended particles generates significant "noise" interference, leading to a substantial decrease in image contrast and making feature extraction difficult. Simultaneously, dynamic caustics caused by refraction from water waves create intensely bright, flickering spots on the object's surface. This instantaneous and non-uniform illumination change easily causes localized overexposure in traditional cameras, resulting in the loss of details on the guided target.
[0033] Secondly, traditional frame-based cameras have inherent flaws in their operating mechanism, making them ill-suited for highly dynamic underwater environments. Frame-based cameras acquire complete image frames at a fixed frequency. When the AUV experiences severe shaking due to swells or rapid movement, the sensor suffers significant motion blur within the exposure time, leading to a precipitous drop in positioning accuracy. Furthermore, traditional cameras have limited dynamic range, often struggling to balance exposure in scenarios with bright underwater guide lights and dark backgrounds, resulting in target centering errors. More critically, traditional cameras output massive amounts of data with a fixed sampling delay. In the docking closed-loop control phase, requiring extremely high response frequencies, tens of milliseconds of image processing latency often cause control command lag, increasing the risk of AUV collisions.
[0034] Finally, underwater visual positioning also faces complex optical refraction problems. AUV cameras are typically mounted inside waterproof pressure vessels. Light passes through a glass window from the water into the air medium, a process that follows Snell's law and produces nonlinear refraction distortion. Traditional pinhole camera models are no longer applicable underwater; without accurate refraction compensation, positioning algorithms will suffer from severe depth estimation errors. Although current acoustic navigation systems can provide long-range positioning, their performance is far inferior to optical navigation in close-range operations due to low update frequency and significant multipath effects.
[0035] To address the aforementioned issues, this application provides relevant solutions to overcome the shortcomings of traditional frame cameras, such as motion blur, insufficient dynamic range, and high data latency, which are prone to occur in complex underwater dynamic environments. By utilizing the asynchronous pixel triggering mechanism of an event camera and combining it with a depth-optimized underwater event processing algorithm, high-frequency pose calculation is achieved. The following is a detailed description.
[0036] According to an embodiment of this application, a method for locating an underwater vehicle is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0037] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal (or electronic device) for implementing an underwater vehicle positioning method is shown. Figure 1 As shown, the computer terminal 10 (or electronic device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0038] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or electronic device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0039] The memory 104 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the underwater vehicle positioning method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned underwater vehicle positioning method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0040] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0041] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or electronic device).
[0042] Under the aforementioned operating environment, this application provides a method for locating an underwater vehicle. Figure 2 This is a schematic diagram of a method for locating an underwater vehicle according to an embodiment of this application, as shown below. Figure 2 As shown, the method includes the following steps:
[0043] Step S202: Obtain the target event stream collected by the event camera, wherein the event camera is set on the underwater vehicle, and the target event stream is used to characterize the information of brightness polarity change events at different pixel positions;
[0044] In this embodiment, the event stream consists of asynchronously triggered pixel-level brightness polarity change events. Each event records the pixel coordinates of the occurrence location, the trigger time, and the polarity information of the increase or decrease in brightness, which is used to characterize the instantaneous response sequence of light intensity changes in the environment, rather than a complete image frame. This step directly relies on the inherent working characteristics of the event camera, extracting only the discrete event signals generated by brightness changes at the pixel level, thereby realizing the original perception and recording of dynamic light intensity disturbances and providing a basic data source for subsequent processing.
[0045] Step S204: Based on the target event stream, determine the event point cloud whose event trigger frequency matches the target frequency, and determine the sub-pixel level center coordinates based on the pixel coordinates corresponding to the event point cloud. The target frequency is the flashing frequency of the guide light source set on the docking target corresponding to the underwater vehicle, and the sub-pixel level center coordinates correspond one-to-one with the guide light source.
[0046] In this embodiment, based on the target event stream, an event point cloud with event triggering frequencies matching the target frequency is determined. Specifically, by analyzing the event triggering time sequence of each pixel in the asynchronous event stream, a set of events whose event occurrence intervals match the fixed flashing frequency of the guide light source on the docking target is identified. This allows for the selection of event points that are caused solely by the guide light source and exhibit periodic brightness variation characteristics. Subsequently, based on the pixel coordinates corresponding to these selected event point clouds, the centroid of their spatial distribution is calculated, yielding sub-pixel-level center coordinates corresponding to each guide light source. This enables high-precision positioning of the guide light source without relying on the integer pixel coordinates of the entire frame image.
[0047] Step S206: By performing refraction distortion correction processing on the sub-pixel level center coordinates, the sub-pixel level center coordinates are mapped into a three-dimensional line-of-sight vector. The refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media. The three-dimensional line-of-sight vector is used to characterize the direction vector after refraction correction from the optical center of the event camera to the underwater guide light source.
[0048] This step takes the sub-pixel-level center coordinates detected by the event camera as input. Based on the path deflection characteristics caused by the inconsistent refractive index when light propagates between different media interfaces such as water, glass and air, the original pixel coordinates are geometrically corrected to eliminate the view position offset caused by the refraction of the medium interface. The final output three-dimensional line-of-sight vector accurately reflects the corrected direction of the light path after propagation in the real physical medium, starting from the optical center of the camera and pointing to the underwater guide light source, ensuring that the line-of-sight direction on which the subsequent pose estimation is based has physical consistency.
[0049] Step S208: Determine the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
[0050] The aforementioned preset spatial coordinates represent the fixed three-dimensional positions of each guide light source on the docking target within its own coordinate system. These values were obtained through calibration before the operation and remain unchanged. By establishing a spatial geometric relationship between each three-dimensional line-of-sight vector and its corresponding preset spatial coordinates, the system can construct the constraint relationship between the observation direction and known spatial points, thereby directly deriving the rotation and translation relationship of the underwater vehicle relative to the docking target, i.e., its pose in six-degree-of-freedom space.
[0051] Through the above steps, by introducing an event camera with high temporal resolution and ultra-high dynamic range, and combining it with a spatiotemporal filtering algorithm for underwater planktonic noise and a multi-medium refraction correction model, high-frequency and high-precision feature extraction of scintillation guidance sources at specific frequencies is achieved. This provides AUVs with a highly reliable asynchronous visual positioning scheme that features extremely low latency, resistance to ambient light interference, and effective compensation for optical distortion in close-range docking or recovery missions. This ensures the safety and positioning accuracy of closed-loop control of AUVs under high-speed relative motion and extreme lighting conditions, thereby solving the technical problem of low visual positioning accuracy caused by motion blur, overexposure, and optical refraction distortion in traditional frame cameras under underwater dynamic interference environments.
[0052] The underwater vehicle positioning method in steps S202 to S208 of the embodiments of this application will be further described below.
[0053] Figure 3 This is a schematic diagram of a method for asynchronous visual positioning of an autonomous underwater vehicle (AUV) against dynamic interference, provided in an embodiment of this application. Figure 3 As shown in the embodiment of this application, preprocessing of subordinate event streams based on spatiotemporal correlation can be performed first to eliminate noise. Then, active guidance features are extracted based on frequency response to lock the LED light source. Next, a multi-medium refraction event projection model is constructed to compensate for optical distortion. Then, initial pose estimation is performed based on asynchronous feature clustering. Finally, high-frequency fusion update of pose is achieved through tightly coupled extended Kalman filtering, forming a complete technical closed loop from raw event data to high-precision pose output. The above process is described in detail below.
[0054] First, to address the pseudo-event noise generated by backscattering of plankton and suspended particles in the underwater environment, the embodiments of this application can perform noise reduction processing on the original asynchronous event stream to obtain the target event stream. The specific steps are as follows.
[0055] In some embodiments of this application, acquiring a target event stream captured by an event camera includes the following steps: acquiring an original event stream captured by the event camera, wherein the original event stream contains multiple brightness polarity change events, as well as timestamps, pixel positions, and brightness polarity states corresponding to the brightness polarity change events; for each brightness polarity change event in the original event stream, determining the first pixel position and the first timestamp corresponding to the brightness polarity change event, and counting the number of brightness polarity change events within a preset spatial neighborhood of the first pixel position and within a preset time window length of the first timestamp in the original event stream; if the number of events is less than a preset event number threshold, determining the current brightness polarity change event as a noise event, and filtering the noise event from the original event stream, thereby obtaining the target event stream.
[0056] Specifically, acquiring the raw event stream collected by the event camera refers to capturing pixel-level brightness polarity change events in the underwater environment in real time through the event camera. Each event contains a timestamp of occurrence, pixel coordinates, and polarity information of brightness increase or decrease, which is used to record the details of light intensity changes in the scene with high dynamic response. Determining the first pixel position and first timestamp of each event is to provide accurate positioning and time reference for subsequent spatiotemporal correlation analysis. The specific formula is expressed as follows.
[0057] For any brightness polarity change event If the following formula is satisfied, the event is retained; otherwise, it is determined to be a noise event:
[0058]
[0059] in, Indicates in Time (i.e., the first timestamp mentioned above) pixel position The event is triggered at the location (i.e., the first pixel position mentioned above) and its polarity is If the brightness increases, If the brightness decreases, , The radius of the spatial neighborhood (i.e., the preset spatial neighborhood range). This is the time threshold (i.e., the preset time window length). The minimum number of events threshold within the activation neighborhood (i.e., the preset event number threshold). This is an indicator function.
[0060] Because underwater pseudo-events (such as plankton scattering) are usually isolated and randomly distributed, while events triggered by real guide light sources exhibit spatiotemporal clustering due to frequency consistency (for example, a continuously flashing LED guide light will trigger multiple events at the same location, while a single suspended particle reflection only produces an isolated event), for each newly generated event signal, the number of its associated events can be searched within a preset spatial neighborhood and a very short time window. For example, the number of neighboring events of the event within a preset spatial neighborhood (such as a 3×3 pixel neighborhood) and a preset time window length (such as within 5 milliseconds) can be counted. When the statistical result is lower than the preset event number threshold (such as 3 events), the event is determined to be noise and is removed, thereby effectively filtering out isolated pseudo-events caused by backscattering, bubble disturbance, or optical noise, and retaining real target feature events with spatiotemporal continuity. Finally, a pure event stream (target event stream) reflecting the guide features and environmental contours is obtained, which significantly improves the accuracy of subsequent event point cloud extraction and sub-pixel center localization, laying a reliable data foundation for refraction distortion correction and 3D pose calculation.
[0061] After obtaining the target event stream, further active guidance feature extraction based on frequency response can be performed. First, based on the target event stream, the event point cloud with event triggering frequencies matching the target frequency can be determined. The specific steps are as follows.
[0062] In some embodiments of this application, determining an event point cloud whose event triggering frequency matches the target frequency based on the target event stream includes the following steps: For each pixel location in the target event stream, an event time series corresponding to the pixel location is established, wherein a series of brightness polarity change events generated by the pixel location are arranged in timestamp order in the event time series; the time interval between adjacent events in the event time series is determined, and a frequency bandpass filter is used to determine whether the time interval meets the frequency constraint corresponding to the target frequency, wherein the frequency bandpass filter is used to filter brightness polarity change events consistent with the flashing frequency of the guiding light source, with the target frequency as the center frequency and a preset bandwidth as the passband range, and the frequency constraint is used to characterize that the difference between the instantaneous frequency corresponding to the time interval of adjacent events and the target frequency is within a preset tolerance range; the proportion of events whose time interval meets the frequency constraint at each pixel location is statistically analyzed, and a signal saliency parameter corresponding to the pixel location is determined based on the proportion, wherein the signal saliency parameter is used to characterize the degree of conformity between the event triggering frequency of the pixel location and the flashing frequency of the guiding light source; an event point cloud is generated based on the events at pixel locations whose signal saliency parameter is greater than the target saliency threshold, wherein the target saliency threshold is determined based on the turbidity of the underwater environment.
[0063] In this embodiment of the application, AUVs can be deployed at a fixed frequency on the docking target (such as a recovery bin). An LED array flashing at the target frequency (as described above) serves as an active guiding light source (this LED array contains at least four guiding light sources), such as... Figure 4 As shown, A, B, C, and D are four LED lights serving as guide lights, and E on the underwater autonomous unmanned vehicle is the event camera. Indicates the coordinate system of the recovery capsule. Indicates the event camera coordinate system.
[0064] This application embodiment utilizes the characteristic that an event camera is only sensitive to changes in brightness, and records the event trigger frequency at each pixel location using a time window statistical method. The system has a built-in frequency bandpass filter to accurately lock and extract frequencies that conform to... The event point cloud fundamentally eliminates static backgrounds, slowly changing underwater natural light, and non-periodic caustic interference from surface waves. For example, precise feature stripping can be achieved by performing time-frequency analysis on the event sequence of each pixel. The specific implementation process is as follows.
[0065] First, event time interval modeling is performed. For the target event stream captured by the event camera, an event time series sorted by timestamp is constructed for each pixel location. For example, for a pixel... The series of asynchronous event streams it generates can be represented as The time interval between adjacent events is defined as:
[0066]
[0067] Because active light sources have a frequency The theoretical event triggering period for blinking should satisfy... (That is, the event is triggered every time the brightness increases or decreases).
[0068] Simultaneously, a frequency bandpass filter can be constructed, creating a filter with a target frequency. Centered on, with a bandwidth of A digital bandpass filter. Utilizing a statistical time window. The proportion of events that satisfy the frequency constraint is used to calculate the signal saliency parameter for each pixel. :
[0069]
[0070] in, As an indicator function, when the difference between the instantaneous frequency and the target frequency is within tolerance... The value is 1 if the signal is within the specified range, and 0 otherwise. This signal saliency parameter quantifies the consistency between the event sequence of the pixel and the flicker frequency of the target. The higher the value, the greater the probability that the pixel belongs to the guiding light source. For example, if 80% of the event intervals of a pixel fall within the passband, its saliency is 0.8, which is significantly higher than the background noise point with only 20% matching.
[0071] At the same time, the target salience threshold can be dynamically set based on the current underwater turbidity. When the water is turbid, the threshold is increased to suppress spurious event noise; when the water is clear, the threshold is decreased to retain the effective signal. For example, when the water is turbid (e.g., the turbidimeter reading exceeds 5), there is more spurious event noise, and in this case, the threshold needs to be increased. (e.g., set to 0.7) 0.8); When the water quality is relatively clear, there is less disturbance, which can reduce [the risk of contamination]. (e.g., set to 0.5) 0.6).
[0072] Then, the target saliency threshold can be used. A mask is applied to all pixels to filter noise and lock onto the target guide light. Only pixels that meet the requirements are retained. The event to which the pixel belongs is used to generate a feature point cloud. (i.e., the event point cloud mentioned above), as shown in the following formula:
[0073]
[0074] Next, using a spatial neighborhood search-based clustering algorithm (such as DBSCAN or neighborhood growth), with feature frequency consistency as a constraint, the extracted feature event point cloud is divided into several independent event clusters, each corresponding to an LED guidance source (guiding light source). For each clustered event cluster, the centroid method is used to calculate the center point; for example, for N events contained within a cluster, its sub-pixel center coordinates are calculated. This can be expressed as the weighted average of the pixel coordinates of all events within the cluster. The specific steps are as follows.
[0075] In some embodiments of this application, determining the sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud includes the following steps: spatially clustering the event point cloud to obtain multiple (at least four) event clusters, wherein each event cluster corresponds to a guiding light source; validating each event cluster, and determining the event cluster as invalid if the number of events contained in the event cluster is less than a minimum number threshold, and deleting the invalid cluster; for each valid event cluster, calculating the average value of the pixel positions corresponding to each brightness polarity change event contained in the event cluster to obtain the sub-pixel-level center coordinates corresponding to the event cluster.
[0076] Specifically, a spatial neighborhood clustering method based on Euclidean distance can be used to cluster event point clouds. This clustering method may include the following steps: first, starting from the frequency-filtered point cloud... Select an unmarked event point as the seed point; then, using this seed point as the center, within a radius... The internal search identifies event points that also satisfy the frequency characteristics and assigns them to the current cluster. The above process is repeated iteratively until there are no more points in the region that meet the conditions. At the same time, the validity of each event cluster can be checked to remove noise. If the number of events in an event cluster is less than the preset threshold (i.e., the minimum number of events threshold mentioned above), the noise will be removed. If there are fewer than 10 events, it is considered residual noise (invalid cluster) and the invalid event cluster is deleted; then, for each valid event cluster... The sub-pixel level center coordinates can be calculated using the centroid method:
[0077]
[0078] For a cluster containing N events, its sub-pixel center coordinates are the weighted average of the pixel coordinates of all events within that cluster. That is, by calculating the arithmetic mean of the pixel positions of each event within the cluster, a floating-point precision center position coordinate is obtained. Essentially, sub-pixel-level positioning is achieved through statistical methods. For example, if a cluster contains 15 events with pixel coordinates distributed between (123.2, 89.7) and (124.8, 90.3), the calculated average coordinates (124.0, 90.0) represent a more precise center position than the original pixel coordinates, significantly improving positioning resolution.
[0079] Through the synergistic effect of the above steps, the problem of false feature merging caused by non-target light source interference, dynamic turbidity disturbance and sensor false triggering in complex underwater environments can be effectively suppressed. This ensures that each sub-pixel-level center coordinate corresponds to a real, stable and high-confidence guiding light source, providing a reliable and unique visual anchor point for subsequent three-dimensional line-of-sight vector correction and accurate pose calculation based on the refraction model. This significantly improves the robustness and accuracy of the visual positioning system under underwater dynamic interference environments.
[0080] Furthermore, addressing the refraction phenomenon of light passing through the "water-pressure vessel glass-camera air cavity" in underwater imaging, this application embodiment establishes a multi-medium refraction correction model to correct refraction distortion. This model no longer uses traditional single-center pinhole projection but introduces a refraction distortion correction factor. By calculating the refraction angle of light at each medium interface, the two-dimensional pixel coordinates of each event cluster are determined. (i.e., sub-pixel-level center coordinates) are mapped to the corrected 3D view vector. This compensates for the geometric deformation caused by refraction, ensuring the physical authenticity of the positioning model.
[0081] Let the pixel coordinates be Its incident vector in the air is According to Snell's Law, light travels through water (refractive index 100%). ) enters the air (refractive index) The refraction relationship of () is:
[0082]
[0083] By refining through iteration or analytical methods, a direction vector from pixel coordinates to underwater object points is established. mapping function :
[0084]
[0085] in, This is the normal vector of the glass plane.
[0086] In this embodiment, the specific steps for refraction distortion correction processing of sub-pixel level center coordinates are as follows.
[0087] In some embodiments of this application, mapping subpixel-level center coordinates to a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the subpixel-level center coordinates includes the following steps: using the camera intrinsic parameter matrix of the event camera, converting the subpixel-level center coordinates into a unit line-of-sight vector in the camera coordinate system; determining the first refractive index ratio between the air medium and the glass medium, and the second refractive index ratio between the glass medium and the water medium in the current underwater environment, and establishing a target mapping function based on the first and second refractive index ratios, wherein the target mapping function is used to characterize the transformation relationship of the direction vector when light propagates through the water-glass-air multilayer medium; and converting the unit line-of-sight vector according to the target mapping function to obtain the three-dimensional line-of-sight vector corresponding to the subpixel-level center coordinates.
[0088] Specifically, the sub-pixel center coordinates can be obtained first using the camera intrinsic matrix of the event camera. Convert to unit line-of-sight vector in camera coordinate system This achieves the transformation from a pixel to a line-of-sight vector within the air. The specific formula is as follows:
[0089]
[0090] in, This is the intrinsic parameter matrix. This indicates the direction of light propagation within the air cavity in front of the camera.
[0091] Next, a first layer of refraction correction (between air and glass) is performed. Specifically, this correction can be performed using the vector form of Snell's law, calculating the angle of incidence of the light at the air-glass interface, based on the air's refractive index. ) and glass (refractive index) Using the refractive index ratio of the two materials and the normal vector of the glass plane, the direction vector of light propagation within the glass medium can be determined. As shown in the following formula:
[0092]
[0093] in, , This is the normal vector of the glass plane.
[0094] Simultaneously, a second layer of refraction correction (between glass and water) is required. Similarly, this refraction correction is performed according to Snell's law. The angle of incidence of the light at the glass-water interface is calculated, based on the glass's refractive index... ) and water (refractive index) The refractive index ratio of the two elements is used to solve for the final ray propagation direction vector in the water medium, combined with the glass plane normal vector.
[0095]
[0096] in, .
[0097] Because the two refractive interfaces of a watertight pressure vessel are parallel to each other, the refractive index of the intermediate glass medium... They cancel each other out when calculating direction vectors (according to...) , In calculating the direction vector (This will disappear over time). Therefore, the above two-layer refraction process can be simplified by combining them to establish a mapping function directly from pixel coordinates to the underwater line-of-sight vector. The combined direction mapping can be simplified as follows:
[0098]
[0099] The mapping function The system takes the air line-of-sight vector, the refractive index ratio of air to water, and the glass plane normal vector as inputs, and outputs the underwater unit direction vector after refraction correction (i.e., the corrected three-dimensional line-of-sight vector).
[0100] Through the above process, the unit line-of-sight vector can be transformed into a vector-level vector, outputting a physically modeled 3D line-of-sight vector. This vector accurately represents the actual optical path direction from the camera's optical center through multiple layers of media to the underwater guide light source, thus completely eliminating the line-of-sight vector inaccuracy problem caused by neglecting the effect of glass window refraction in traditional methods. This significantly improves the accuracy and reliability of vehicle attitude estimation based on visual feedback in underwater dynamic environments.
[0101] After obtaining the corrected 3D direction vectors corresponding to each guiding light source Then, the preset 3D spatial coordinates of each guiding light source on the docking target platform can be used as a reference. (Since it is a cooperative target, it can be obtained through measurement), and the initial attitude estimation of the underwater vehicle is carried out as follows.
[0102] In some embodiments of this application, determining the target pose of an underwater vehicle based on a three-dimensional line-of-sight vector and preset spatial coordinates corresponding to a guide light source on the docking target includes: determining the preset spatial coordinates of a guide light source set on the docking target, and establishing a reprojection error equation based on the preset spatial coordinates and the three-dimensional line-of-sight vector. The reprojection error equation is used to characterize the deviation relationship between the observed three-dimensional line-of-sight vector and the line-of-sight vector reprojected based on the current pose estimate. By minimizing the angular residual between the three-dimensional line-of-sight vector and the reprojection vector, the reprojection error equation is solved to obtain the target pose of the underwater vehicle. The target pose includes the following information: the rotation matrix and translation vector of the underwater vehicle relative to the docking target.
[0103] Specifically, by determining the preset spatial coordinates of the guiding light source set on the docking target, its precise three-dimensional position information in the global reference system can be established as a geometric benchmark for visual positioning. For example, four flashing LED arrays with known spacing can be fixed at the edge of the underwater docking hatch. Their coordinates can be pre-calibrated and stored through high-precision mapping to provide an absolute reference for subsequent pose calculation. Then, combined with the three-dimensional line-of-sight vector obtained by the event camera after refraction distortion correction, a reprojection error equation based on the three-dimensional spherical projection model can be further established. Its essence is to construct the geometric deviation relationship between the observed three-dimensional line-of-sight vector and the theoretical line-of-sight vector obtained by back-projecting the preset spatial coordinates to the camera coordinate system based on the current pose estimate. This equation does not depend on the Euclidean distance of the pixel coordinate error, but takes the residual angle between the two vectors as the optimization target, that is, to calculate the cosine difference between the observation direction and the reprojection direction, so as to better fit the characteristics of underwater vision where direction information is dominant and distance information is ambiguous.
[0104] In this embodiment, the observation vector can be minimized using a PnP algorithm (such as EPnP). The target pose is solved by the angular residual between the reprojection vector and the target pose vector, thereby calculating the rotation matrix of the AUV relative to the target platform. Translation vector The objective function is shown in the following equation:
[0105]
[0106] By iteratively adjusting the rotation matrix and translation vector, the observation lines of all guiding light sources can be made highly consistent with the theoretical reprojection lines in direction. Even if there is event noise caused by water flow disturbance, the system can still converge stably to the optimal pose, avoiding the divergence or multiple solutions caused by perspective distortion or feature sparsity in traditional pixel-based reprojection errors. Ultimately, high-precision and robust vehicle pose estimation is achieved in underwater dynamic interference environments.
[0107] Furthermore, asynchronous pose observations acquired by the event camera can be deeply fused with data from the inertial measurement unit (IMU) on the AUV to establish a high-frequency pose fusion update mechanism based on a tightly coupled extended Kalman filter (EKF) state estimation framework: when the feature point pose increment generated by the event camera arrives, an update step is triggered to correct the residual. Since the event camera is not dependent on the frame rate, the system can output continuous pose information at frequencies in the kilohertz (kHz) range, providing near real-time navigation feedback for the AUV control system. Details are as follows.
[0108] In some embodiments of this application, the method further includes: determining the state vector and inertial measurement data of the underwater vehicle at the current moment, wherein the state vector is used to characterize the navigation state parameters (such as position, velocity, attitude quaternions, etc.) and sensor error compensation parameters (such as the zero bias information of the inertial measurement unit (accelerometer zero bias, gyroscope zero bias)) of the underwater vehicle, and the inertial measurement data is the motion state differential information (such as linear motion acceleration information and angular motion rate information) collected by the inertial measurement unit on the underwater vehicle; using the extended Kalman filter algorithm, based on the state vector and inertial measurement data at the current moment, predicting the pose prediction value of the underwater vehicle at the next moment; when the sub-pixel-level center coordinates observed by the event camera are updated, determining the residual between the newly observed sub-pixel-level center coordinates and the predicted pixel coordinates, and correcting the pose prediction value based on the residual to obtain the updated pose estimate, wherein the predicted pixel coordinates are the coordinates of the pose prediction value mapped onto the image plane of the event camera after refraction distortion correction.
[0109] Specifically, the state vector of the underwater vehicle (AUV) can be established by acquiring its current state vector and the differential motion state information collected by the inertial measurement unit. , including location ,speed Posture Quaternions and IMU bias (accelerometer bias) and gyroscope zero bias ).
[0110] Based on this, the extended Kalman filter algorithm is used to fuse inertial data to continuously predict the pose at the next moment. The prediction equation (based on IMU sampling) is as follows:
[0111]
[0112] in, This represents process noise, used to illustrate the measurement noise of the IMU sensor and the impact of external random disturbances such as ocean currents on the prediction model; The state transition function is represented by the state vector at the current time step. and IMU input (e.g., three-axis acceleration, angular velocity) are used as inputs to calculate the predicted pose at the next moment. (i.e., the above pose prediction value); This indicates the event time interval. Since the timestamp accuracy of the event camera is at the microsecond level, The error is extremely small, and the linearization error is significantly reduced.
[0113] When the event camera recaptures the event point cloud that matches the flicker frequency of the guiding light source and calculates the sub-pixel-level center coordinates, the pose prediction value can be mapped into a 3D line-of-sight vector after refraction distortion correction, and then back-projected onto the image plane to obtain the predicted pixel coordinates. By calculating the spatial residual between this predicted pixel coordinate and the newly observed sub-pixel-level center coordinate, the consistency deviation between event observation and inertial prediction is quantified, thereby driving the Kalman gain to perform optimal correction of the pose prediction value, achieving real-time compensation for positioning errors. The update equation (based on event triggering) is shown below: When a guide feature update is detected (i.e., when a sub-pixel-level center coordinate update occurs), the measurement residual is constructed. Then through Kalman gain The corrected state is shown in the following formula:
[0114]
[0115] in, The measurement model is used to calculate the predicted pixel position (coordinates) of the guide light on the event camera image plane based on the current predicted AUV pose and through a multi-media refraction projection model. This represents the center coordinates of the guide light on the image plane, calculated from the guide source event cluster after preprocessing and frequency extraction. Indicates Kalman gain, , Let denote the covariance matrix, representing the uncertainty in the current state estimate. The Jacobian matrix represents the observed position of the guide light on the event camera plane when the AUV pose changes slightly (i.e., the observed value). How much change will occur? This represents the measurement noise covariance, which indicates the noise level during the event camera observation process.
[0116] In the embodiments of this application, the measurement model By directly establishing a mapping between state variables and pixel coordinates, the process of solving PnP in each iteration is skipped, significantly reducing computational latency and improving robustness when feature points are missing (less than 4 points). Furthermore, during system operation, when the residual of the extended Kalman filter is detected... When the value is too large (e.g., 6 pixels), the PnP algorithm will reposition the system and forcibly correct the state vector to ensure the system's positioning accuracy.
[0117] To enable those skilled in the art to better understand the above process steps in the embodiments of this application, the following examples illustrate the above process steps in specific case scenarios.
[0118] Suppose an AUV is attempting to enter a deep-sea recovery capsule. At the target end, i.e., the entrance to the recovery capsule, is an array of four LED lights, their flashing frequency set to a fixed 10Hz (i.e., a period of 100ms). The AUV is equipped with an event camera (1280 resolution). The system consisted of a 720-axis IMU and a six-axis IMU (sampling frequency 500Hz). The environmental conditions were: water depth 50 meters, with weak ocean currents and significant suspended particles (generating backscatter noise).
[0119] In the above scenario, event camera data can be acquired first. Record as an asynchronous stream ,in, Microsecond-level timestamps For pixel coordinates, The polarity of the brightness change. Simultaneously, IMU data is acquired ( ): Records are synchronized sequences Record the triaxial acceleration and angular velocity. Then, the AUV can be positioned and recovered by performing the following steps, as detailed below.
[0120] Step 1: Underwater event flow denoising (eliminating background clutter)
[0121] As the AUV moves forward, suspended particles in the ocean current reflect light, generating numerous randomly distributed pseudo-events. The algorithm sets the spatial radius. Pixels, time threshold By using a spatiotemporal correlation filter, the system identifies and removes isolated points that have no spatial correlation in a short period of time, thereby improving the background signal-to-noise ratio of the event stream by approximately 60%.
[0122] Step 2: 10Hz Feature Extraction (Locking Guide Lights)
[0123] The system maintains a 200 ms sliding time window. For each pixel coordinate... The algorithm calculates the time difference between two brightness changes. Since the guide light flashes at 10Hz, the time interval between a large number of events detected by the algorithm in a specific area is approximately 50 ms (polarity reversal interval).
[0124] Saliency calculation: If more than 85% of the events within a certain pixel area conform to the 10Hz frequency characteristic, then its saliency is set. (Here, This allows the guide light to be precisely separated from the complex caustic light spots and static background.
[0125] Clustering and centroid calculation are performed on event clusters to obtain sub-pixel coordinates. .
[0126] Step 3: Refraction Compensation (Correcting Visual Distortion)
[0127] Because the light passes through a "water-glass sealed enclosure-air cavity," the target will appear closer than its actual location according to traditional air models. (System call mapping function) With an input water refractive index of 1.33, the detected 2D pixel center points are converted into corrected 3D direction vectors. This eliminated approximately 5% of the depth measurement bias.
[0128] by Figure 4 Taking the left-side guide light B shown in the diagram as an example, its pixel coordinates are... Through multi-medium refraction mapping function After processing, the output underwater unit direction vector is Similarly, the correction values for the other three lights can be obtained. , and .
[0129] Step 4: Asynchronous Pose Initial Estimation (Fast Localization)
[0130] Based on the known geometric dimensions of the four lights on the recovery capsule, their three-dimensional coordinates in the recovery capsule coordinate system are obtained. Taking the B guide light as an example... Similarly, the three-dimensional coordinates of the other three lights can be obtained. , and Finally, the alliance and The rotation matrix is solved using the PnP algorithm. Translation vector This step completes the initial 6-DOF pose calculation of the AUV relative to the recovery capsule within 1 ms (with errors controlled within the centimeter level).
[0131] Step 5: Tightly Coupled Extended Kalman Filter Fusion (Kilometer-Level Control Response)
[0132] Within the last 2 meters of the AUV approaching the recovery capsule, the system enters a high-frequency update mode:
[0133] Prediction: The IMU predicts the inertial motion of the AUV with a period of 2 ms.
[0134] Update: Whenever the event camera captures a new guide light event, the extended Kalman filter immediately calculates the residual between that pixel and the predicted position.
[0135] Results: The system outputs pose feedback to the control motor at a frequency of 1200Hz. Even under ocean current disturbances, the AUV maintained subpixel-level tracking accuracy and successfully entered the recovery capsule with a center deviation of less than 3 cm.
[0136] By applying the technical solution of this application embodiment, an event camera is deployed on the underwater vehicle to collect target event streams. Utilizing the microsecond-level temporal resolution of the event camera, which only responds to changes in brightness polarity, motion blur caused by exposure accumulation under high-speed motion can be effectively avoided by traditional frame cameras. Furthermore, event point clouds matching the target frequency are selected based on the specific flicker frequency of the guide light source, actively suppressing underwater dynamic interference such as non-periodic noise like caustic light and backscattering, thus achieving accurate extraction of target signals. On this basis, the positioning resolution is improved by calculating sub-pixel-level center coordinates, and the geometric distortion caused by light passing through multiple media interfaces such as water, glass, and air is systematically corrected using a physical refraction model. The pixel coordinates are accurately mapped to a three-dimensional line-of-sight vector pointing from the camera's optical center to the actual guide light source, fundamentally eliminating depth estimation errors caused by refraction deviations. Finally, combined with the spatial coordinates of the guide light source preset on the docking target, the high-precision pose of the underwater vehicle is calculated through geometric inverse kinematics. The aforementioned technologies work together to overcome the bottleneck of visual positioning accuracy in complex underwater dynamic environments, which is constrained by motion blur, overexposure, and optical refraction distortion. They enable stable and accurate pose estimation under conditions of strong interference, low signal-to-noise ratio, and multi-medium refraction, significantly improving the safety and reliability of underwater docking operations.
[0137] According to an embodiment of this application, an embodiment of an underwater vehicle positioning device is also provided. Figure 5 This is a schematic diagram of the modular architecture of an underwater vehicle positioning device according to an embodiment of this application. Figure 5 As shown, the device includes:
[0138] The event stream acquisition module 50 is used to acquire the target event stream collected by the event camera, wherein the event camera is set on the underwater vehicle, and the target event stream is used to characterize the information of brightness polarity change events at different pixel positions;
[0139] The guidance feature extraction module 52 is used to determine the event point cloud whose event trigger frequency matches the target frequency based on the target event flow, and to determine the sub-pixel level center coordinates based on the pixel coordinates corresponding to the event point cloud. The target frequency is the flashing frequency of the guidance light source set on the docking target corresponding to the underwater vehicle, and the sub-pixel level center coordinates correspond one-to-one with the guidance light source.
[0140] The refraction distortion correction module 54 is used to map the sub-pixel-level center coordinates into a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the sub-pixel-level center coordinates. The refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media. The three-dimensional line-of-sight vector is used to characterize the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source.
[0141] The target pose determination module 56 is used to determine the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
[0142] Optionally, acquiring the target event stream captured by the event camera includes: acquiring the original event stream captured by the event camera, wherein the original event stream contains multiple brightness polarity change events, as well as the timestamps, pixel positions, and brightness polarity states corresponding to the brightness polarity change events; for each brightness polarity change event in the original event stream, determining the first pixel position and the first timestamp corresponding to the brightness polarity change event, and counting the number of brightness polarity change events within a preset spatial neighborhood of the first pixel position and within a preset time window length of the first timestamp in the original event stream; if the number of events is less than a preset event number threshold, the current brightness polarity change event is determined to be a noise event, and the noise event is filtered out from the original event stream, thereby obtaining the target event stream.
[0143] Optionally, determining the event point cloud whose event triggering frequency matches the target frequency based on the target event stream includes: establishing an event time series corresponding to each pixel position in the target event stream, wherein a series of brightness polarity change events generated by the pixel position are arranged in timestamp order in the event time series; determining the time interval between adjacent events in the event time series, and using a frequency bandpass filter to determine whether the time interval meets the frequency constraint corresponding to the target frequency, wherein the frequency bandpass filter is used to filter brightness polarity change events consistent with the flashing frequency of the guide light source with the target frequency as the center frequency and a preset bandwidth as the passband range, and the frequency constraint is used to characterize that the difference between the instantaneous frequency corresponding to the time interval of adjacent events and the target frequency is within a preset tolerance range; statistically analyzing the proportion of events whose time interval meets the frequency constraint at each pixel position, and determining the signal saliency parameter corresponding to the pixel position based on the proportion, wherein the signal saliency parameter is used to characterize the degree of conformity between the event triggering frequency of the pixel position and the flashing frequency of the guide light source; generating an event point cloud based on the events at pixel positions whose signal saliency parameter is greater than the target saliency threshold, wherein the target saliency threshold is determined based on the turbidity of the underwater environment.
[0144] Optionally, determining the sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud includes: spatially clustering the event point cloud to obtain multiple event clusters, where each event cluster corresponds to a guiding light source; validating each event cluster, and deleting an event cluster if the number of events contained within the event cluster is less than a minimum threshold; for each valid event cluster, calculating the average pixel position corresponding to each brightness polarity change event contained within the event cluster to obtain the sub-pixel-level center coordinates corresponding to the event cluster.
[0145] Optionally, mapping the subpixel-level center coordinates to a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the subpixel-level center coordinates includes: using the camera intrinsic parameter matrix of the event camera to convert the subpixel-level center coordinates into a unit line-of-sight vector in the camera coordinate system; determining the first refractive index ratio between the air medium and the glass medium, and the second refractive index ratio between the glass medium and the water medium in the current underwater environment, and establishing a target mapping function based on the first and second refractive index ratios, wherein the target mapping function is used to characterize the transformation relationship of the direction vector when light propagates through the water-glass-air multilayer medium; and converting the unit line-of-sight vector according to the target mapping function to obtain the three-dimensional line-of-sight vector corresponding to the subpixel-level center coordinates.
[0146] Optionally, determining the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target includes: determining the preset spatial coordinates of the guide light source set on the docking target, and establishing a reprojection error equation based on the preset spatial coordinates and the three-dimensional line-of-sight vector. The reprojection error equation is used to characterize the deviation relationship between the observed three-dimensional line-of-sight vector and the line-of-sight vector reprojected based on the current pose estimate. By minimizing the angular residual between the three-dimensional line-of-sight vector and the reprojection vector, the reprojection error equation is solved to obtain the target pose of the underwater vehicle. The target pose includes the following information: the rotation matrix and translation vector of the underwater vehicle relative to the docking target.
[0147] Optionally, the target pose determination module 56 is further configured to: determine the state vector and inertial measurement data of the underwater vehicle at the current moment, wherein the state vector is used to characterize the navigation state parameters and sensor error compensation parameters of the underwater vehicle, and the inertial measurement data is the motion state differential information collected by the inertial measurement unit on the underwater vehicle; using the extended Kalman filter algorithm, based on the state vector and inertial measurement data at the current moment, predict the pose prediction value of the underwater vehicle at the next moment; when the sub-pixel-level center coordinates observed by the event camera are updated, determine the residual between the newly observed sub-pixel-level center coordinates and the predicted pixel coordinates, and correct the pose prediction value based on the residual to obtain the updated pose estimate, wherein the predicted pixel coordinates are the coordinates of the pose prediction value mapped onto the image plane of the event camera after refraction distortion correction.
[0148] It should be noted that the modules in the above-mentioned underwater vehicle positioning device can be program modules (such as a set of program instructions to implement a certain function) or hardware modules. For the latter, they can be in the following forms, but are not limited to these: each of the above modules is in the form of a processor, or the functions of each of the above modules are implemented by a processor.
[0149] It should be noted that the underwater vehicle positioning device provided in this embodiment can be used to perform... Figure 2 The underwater vehicle positioning method shown above is also applicable to the embodiments of this application, and will not be repeated here.
[0150] According to an embodiment of this application, an embodiment of an underwater vehicle positioning system is also provided. The system includes: an event camera (e.g., [missing information]). Figure 4 The underwater vehicle shown in E) (e.g.) Figure 4 The underwater autonomous unmanned vehicle shown in the figure), and the docking target corresponding to the underwater vehicle (e.g., Figure 4 (The recovery capsule shown).
[0151] The docking target is equipped with multiple (at least four) guiding light sources (such as... Figure 4 The LED light array shown consists of four guiding light sources (A, B, C, and D). The guiding light sources flash and emit light according to the target frequency.
[0152] An underwater vehicle is used to acquire target event streams via an event camera. These event streams characterize information about brightness polarity changes at different pixel locations. Based on the event streams, event point clouds with trigger frequencies matching the target frequency are determined. Sub-pixel-level center coordinates are then determined based on the pixel coordinates of the event point clouds, with each sub-pixel-level center coordinate corresponding to a guiding light source. Refraction distortion correction is applied to the sub-pixel-level center coordinates, mapping them to a three-dimensional line-of-sight vector. This correction compensates for geometric distortion caused by differences in refractive index when light propagates through different media. The three-dimensional line-of-sight vector characterizes the refracted-corrected direction vector from the optical center of the event camera to the underwater guiding light source. Based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guiding light source on the docking target, the target pose of the underwater vehicle is determined.
[0153] It should be noted that the underwater vehicle positioning system provided in this embodiment is capable of executing... Figure 2 The underwater vehicle positioning method shown above is also applicable to the embodiments of this application, and will not be repeated here.
[0154] This application embodiment also provides an underwater vehicle, including: a memory and a processor. The processor is used to run a program stored in the memory, wherein the program executes the following underwater vehicle positioning method: acquiring a target event stream collected by an event camera, wherein the event camera is set on the underwater vehicle, and the target event stream is used to characterize information on brightness polarity change events at different pixel positions; determining an event point cloud whose event trigger frequency matches the target frequency based on the target event stream, and determining sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud, wherein the target frequency is the flashing frequency of the guide light source set on the docking target corresponding to the underwater vehicle, and the sub-pixel-level center coordinates correspond one-to-one with the guide light source; mapping the sub-pixel-level center coordinates to a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the sub-pixel-level center coordinates, wherein the refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media, and the three-dimensional line-of-sight vector is used to characterize the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source; determining the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
[0155] This application embodiment also provides a non-volatile storage medium, which includes a stored computer program. The device containing the non-volatile storage medium executes the following underwater vehicle positioning method by running the computer program: acquiring a target event stream collected by an event camera, wherein the event camera is mounted on the underwater vehicle, and the target event stream is used to characterize information about brightness polarity change events at different pixel locations; determining an event point cloud whose event triggering frequency matches the target frequency based on the target event stream, and determining sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud, wherein the target frequency is the underwater vehicle's... The flashing frequency of the guide light source set on the docking target corresponding to the underwater vehicle is determined, and the sub-pixel-level center coordinates correspond one-to-one with the guide light source. By performing refraction distortion correction processing on the sub-pixel-level center coordinates, the sub-pixel-level center coordinates are mapped into a three-dimensional line-of-sight vector. The refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media. The three-dimensional line-of-sight vector is used to represent the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source. Based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target, the target pose of the underwater vehicle is determined.
[0156] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the underwater vehicle positioning method described in various embodiments of this application: acquiring a target event stream collected by an event camera, wherein the event camera is set on the underwater vehicle, and the target event stream is used to characterize information on brightness polarity change events at different pixel positions; determining an event point cloud whose event trigger frequency matches the target frequency based on the target event stream, and determining sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud, wherein the target frequency is the flashing frequency of the guide light source set on the docking target corresponding to the underwater vehicle, and the sub-pixel-level center coordinates correspond one-to-one with the guide light source; mapping the sub-pixel-level center coordinates to a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the sub-pixel-level center coordinates, wherein the refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media, and the three-dimensional line-of-sight vector is used to characterize the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source; determining the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
[0157] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0158] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0159] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0161] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0162] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0163] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method of positioning an underwater vehicle, characterized by, include: Acquire a target event stream collected by an event camera, wherein the event camera is mounted on an underwater vehicle, and the target event stream is used to characterize information on brightness polarity change events at different pixel locations; Based on the target event stream, determining the event point cloud whose event triggering frequency matches the target frequency includes: for each pixel position in the target event stream, establishing an event time sequence corresponding to the pixel position, wherein in the event time sequence, a series of brightness polarity change events generated by the pixel position are arranged in timestamp order; determining the time interval between adjacent events in the event time sequence, and using a frequency bandpass filter to determine whether the time interval meets the frequency constraint corresponding to the target frequency, wherein the frequency bandpass filter is used to filter and guide the light source flicker frequency with the target frequency as the center frequency and a preset bandwidth as the passband range. Consistent brightness polarity change events, the frequency constraint is used to characterize that the difference between the instantaneous frequency corresponding to the time interval of adjacent events and the target frequency is within a preset tolerance range; the proportion of events at each pixel position whose time interval satisfies the frequency constraint is statistically analyzed, and based on the proportion, the signal saliency parameter corresponding to the pixel position is determined, wherein the signal saliency parameter is used to characterize the degree of conformity between the event trigger frequency of the pixel position and the flashing frequency of the guiding light source; the event point cloud is generated based on the events at the pixel positions whose signal saliency parameter is greater than the target saliency threshold, wherein the target saliency threshold is determined based on the turbidity of the underwater environment; Based on the pixel coordinates corresponding to the event point cloud, sub-pixel level center coordinates are determined, wherein the target frequency is the flashing frequency of the guide light source set on the docking target corresponding to the underwater vehicle, and the sub-pixel level center coordinates correspond one-to-one with the guide light source; By performing refraction distortion correction processing on the sub-pixel-level center coordinates, the sub-pixel-level center coordinates are mapped into a three-dimensional line-of-sight vector. The refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media. The three-dimensional line-of-sight vector is used to characterize the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source. The target pose of the underwater vehicle is determined based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
2. The underwater vehicle positioning method of claim 1, wherein, The target event stream acquired by the event camera includes: Obtain the raw event stream captured by the event camera, wherein the raw event stream contains multiple brightness polarity change events, as well as the timestamp, pixel position and brightness polarity state corresponding to the brightness polarity change events; For each brightness polarity change event in the original event stream, determine the first pixel position and first timestamp corresponding to the brightness polarity change event, and count the number of brightness polarity change events within a preset spatial neighborhood of the first pixel position and within a preset time window length of the first timestamp in the original event stream. If the number of events is less than a preset event number threshold, the current brightness polarity change event is identified as a noise event and the noise event is filtered out from the original event stream to obtain the target event stream.
3. The underwater vehicle positioning method of claim 1, wherein, Determining the sub-pixel-level center coordinates based on the pixel coordinates corresponding to the event point cloud includes: Spatial clustering is performed on the event point cloud to obtain multiple event clusters, wherein each event cluster corresponds to one of the guiding light sources; For each event cluster, a validity check is performed. If the number of events contained in the event cluster is less than a minimum threshold, the event cluster is determined to be an invalid cluster and is deleted. For each valid event cluster, calculate the average value of the pixel positions corresponding to each brightness polarity change event contained in the event cluster to obtain the sub-pixel-level center coordinates corresponding to the event cluster.
4. The underwater vehicle positioning method according to claim 1, characterized in that, Mapping the sub-pixel-level center coordinates into a three-dimensional viewing vector by performing refraction distortion correction processing on the sub-pixel-level center coordinates includes: Using the camera intrinsic parameter matrix of the event camera, the sub-pixel-level center coordinates are converted into a unit line-of-sight vector in the camera coordinate system; A first refractive index ratio between the air medium and the glass medium, and a second refractive index ratio between the glass medium and the water medium in the current underwater environment are determined. Based on the first refractive index ratio and the second refractive index ratio, a target mapping function is established, wherein the target mapping function is used to characterize the transformation relationship of the direction vector when light propagates through the water-glass-air multilayer medium. Based on the target mapping function, the unit gaze vector is transformed to obtain the three-dimensional gaze vector corresponding to the sub-pixel level center coordinates.
5. The underwater vehicle positioning method according to claim 1, characterized in that, Determining the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target includes: Determine the preset spatial coordinates of the guiding light source set on the docking target, and establish a reprojection error equation based on the preset spatial coordinates and the three-dimensional line-of-sight vector. The reprojection error equation is used to characterize the deviation relationship between the observed three-dimensional line-of-sight vector and the line-of-sight vector reprojected based on the current pose estimation value. By minimizing the angular residual between the three-dimensional line-of-sight vector and the reprojection vector, the reprojection error equation is solved to obtain the target pose of the underwater vehicle, wherein the target pose includes: the rotation matrix and translation vector of the underwater vehicle relative to the docking target.
6. The underwater vehicle positioning method according to claim 1, characterized in that, The method further includes: The state vector and inertial measurement data of the underwater vehicle at the current moment are determined, wherein the state vector is used to characterize the navigation state parameters and sensor error compensation parameters of the underwater vehicle, and the inertial measurement data is the motion state differential information collected by the inertial measurement unit on the underwater vehicle. Using the extended Kalman filter algorithm, based on the state vector and inertial measurement data at the current moment, the pose prediction value of the underwater vehicle at the next moment is predicted; When the sub-pixel-level center coordinates observed by the event camera are updated, the residual between the newly observed sub-pixel-level center coordinates and the predicted pixel coordinates is determined, and the pose prediction value is corrected based on the residual to obtain the updated pose estimate, wherein the predicted pixel coordinates are the coordinates of the pose prediction value mapped onto the event camera image plane after the refraction distortion correction processing.
7. A positioning device for an underwater vehicle, characterized in that, include: An event stream acquisition module is used to acquire target event streams collected by an event camera, wherein the event camera is mounted on an underwater vehicle, and the target event stream is used to characterize information about brightness polarity change events at different pixel positions. The guidance feature extraction module is used to determine, based on the target event stream, event point clouds whose event trigger frequencies conform to the target frequency. This includes: establishing an event time sequence corresponding to each pixel position in the target event stream, wherein a series of brightness polarity change events generated at the pixel position are arranged in timestamp order within the event time sequence; determining the time interval between adjacent events in the event time sequence, and using a frequency bandpass filter to determine whether the time interval satisfies the frequency constraint corresponding to the target frequency, wherein the frequency bandpass filter is used to filter brightness polarity change events consistent with the flicker frequency of the guidance light source, with the target frequency as the center frequency and a preset bandwidth as the passband range; the frequency constraint is used to characterize the instantaneous frequency corresponding to the time interval between adjacent events and the target frequency. The difference in target frequencies is within a preset tolerance range; the proportion of events at each pixel location whose time interval satisfies the frequency constraint is statistically analyzed, and a signal saliency parameter corresponding to the pixel location is determined based on the proportion, wherein the signal saliency parameter is used to characterize the degree of conformity between the event triggering frequency of the pixel location and the flashing frequency of the guide light source; an event point cloud is generated based on the events at the pixel locations whose signal saliency parameter is greater than the target saliency threshold, wherein the target saliency threshold is determined based on the turbidity of the underwater environment; and sub-pixel-level center coordinates are determined based on the pixel coordinates corresponding to the event point cloud, wherein the target frequency is the flashing frequency of the guide light source set on the docking target corresponding to the underwater vehicle, and the sub-pixel-level center coordinates correspond one-to-one with the guide light source; The refraction distortion correction module is used to map the sub-pixel-level center coordinates into a three-dimensional line-of-sight vector by performing refraction distortion correction processing on the sub-pixel-level center coordinates. The refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media. The three-dimensional line-of-sight vector is used to characterize the refraction-corrected direction vector from the optical center of the event camera to the underwater guide light source. The target pose determination module is used to determine the target pose of the underwater vehicle based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guide light source on the docking target.
8. A positioning system for an underwater vehicle, characterized in that, include: An underwater vehicle equipped with an event camera, and a docking target corresponding to the underwater vehicle, wherein... The docking target is equipped with multiple guiding light sources, which flash light according to the target frequency. The underwater vehicle is used to acquire a target event stream via the event camera, wherein the target event stream is used to characterize information on brightness polarity change events at different pixel locations; based on the target event stream, it determines an event point cloud whose event triggering frequency conforms to the target frequency, including: for each pixel location in the target event stream, establishing an event time series corresponding to the pixel location, wherein in the event time series, a series of brightness polarity change events generated at the pixel location are arranged in timestamp order; determining the time interval between adjacent events in the event time series, and using a frequency bandpass filter to determine whether the time interval meets the frequency constraint corresponding to the target frequency, wherein the frequency bandpass filter is used to filter brightness polarity change events consistent with the flashing frequency of the guide light source with the target frequency as the center frequency and a preset bandwidth as the passband range, and the frequency constraint is used to characterize that the difference between the instantaneous frequency corresponding to the time interval of adjacent events and the target frequency is within a preset tolerance range; and statistically analyzing the proportion of events at each pixel location whose time interval meets the frequency constraint. Based on the stated ratio, a signal saliency parameter corresponding to the pixel position is determined, wherein the signal saliency parameter characterizes the degree of agreement between the event triggering frequency of the pixel position and the flashing frequency of the guiding light source; an event point cloud is generated based on events at pixel positions for which the signal saliency parameter is greater than a target saliency threshold, wherein the target saliency threshold is determined based on the turbidity of the underwater environment; and sub-pixel-level center coordinates are determined based on the pixel coordinates corresponding to the event point cloud, wherein the sub-pixel-level center coordinates correspond one-to-one with the guiding light source; by performing refraction distortion correction processing on the sub-pixel-level center coordinates, the sub-pixel-level center coordinates are mapped into a three-dimensional line-of-sight vector, wherein the refraction distortion correction processing is used to compensate for the geometric distortion caused by the difference in refractive index when light propagates through different media, and the three-dimensional line-of-sight vector characterizes the refraction-corrected direction vector from the optical center of the event camera to the underwater guiding light source; based on the three-dimensional line-of-sight vector and the preset spatial coordinates corresponding to the guiding light source on the docking target, the target pose of the underwater vehicle is determined.
9. An underwater vehicle, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the underwater vehicle positioning method according to any one of claims 1 to 6.
10. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored computer program, wherein the device containing the non-volatile storage medium executes the underwater vehicle positioning method according to any one of claims 1 to 6 by running the computer program.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the underwater vehicle positioning method according to any one of claims 1 to 6.