A highly dynamic pose estimation method for intelligent agents based on motion-encoded event plane representation

By constructing motion coding plane representation and 3D-2D alignment technology using event cameras, the problem of low pose estimation accuracy when satellite signals are interfered with or in complex environments is solved, high-precision pose tracking is achieved, and the application capabilities of unmanned systems are improved.

CN118570298BActive Publication Date: 2025-09-19SOUTHEAST UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410613621.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-09-19
Estimated Expiration
2044-05-17

AI Technical Summary

Technical Problem

Existing technologies have low pose estimation accuracy and easily accumulated errors when satellite signals are interfered with or in complex and highly dynamic environments, which limits the popularization and promotion of unmanned systems.

Method used

A motion-encoded event plane representation method is adopted to capture dynamic changes through an event camera, construct a consistent plane representation of the scene, combine it with a semi-dense three-dimensional map of the environment, and use 3D-2D alignment technology to achieve high-precision pose estimation.

Benefits of technology

High-precision and robust posture tracking is achieved in high-speed and high-dynamic scenarios, improving the positioning accuracy and efficiency of unmanned systems in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118570298B_ABST
    Figure CN118570298B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for high-dynamic pose estimation of an intelligent agent based on motion-encoded event plane representation, comprising the following steps: 1. reading the event stream output by an event camera, where each event is represented by a four-dimensional vector; 2. constructing a bidirectional linked list to store the triggering time and sequence of pixels in a local neighborhood, and performing stack updates through an asynchronous event-driven thread; 3. constructing a motion-encoded event plane representation to achieve consistent representation of the environment in high-speed, high-dynamic scenes; 4. constructing a semi-dense scene local map in the form of a 3D point cloud; 5. encoding the spatiotemporal constraints of camera motion based on the event plane representation, and using 3D-2D alignment technology to achieve real-time estimation of the six-degree-of-freedom pose. This method, by leveraging the low latency advantage of the event camera and its natural response to scene edges, combined with semi-dense scene information, can achieve high-precision and robust pose tracking in high-speed, high-dynamic scenes, thereby fully unlocking the potential of event cameras in high-speed unmanned system applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of visual navigation technology, and in particular is a method for estimating the high-dynamic pose of an intelligent body based on motion coding event plane representation. Background Art

[0002] Six-degree-of-freedom pose estimation technology for unmanned systems plays a vital role in fields such as mobile robotics, autonomous driving, and virtual reality (VR). Existing pose estimation technologies in practical applications either rely heavily on satellite navigation systems, which are unusable when satellite signals are interfered with and have low accuracy; rely on inertial navigation technology, which is prone to drift due to accumulated errors over time; or rely on traditional vision, which is significantly affected by environmental dynamics such as motion speed and external lighting, greatly limiting the popularization and widespread use of unmanned systems. Therefore, ensuring the accuracy and efficiency of pose estimation in the absence of satellite signals and in complex, highly dynamic environments has become a key bottleneck that needs to be addressed in the field of robotics and unmanned systems.

[0003] As an emerging AI perception paradigm, event cameras break through traditional visual imaging mechanisms by using event-triggered methods to rapidly capture dynamic changes in scenes. These cameras offer ultra-low latency (3µs) and a high dynamic range (120dB), overcoming the limitations of conventional vision and offering significant advantages in high-speed, high-dynamic scenarios. Using event cameras to achieve robust, reliable, and efficient real-time positioning and attitude determination has become a hot topic in industry and science. Current pose estimation algorithms based on event cameras are still in their infancy. Recent algorithms have failed to fully exploit their advantages, lacking in system convergence speed, real-time computational timeliness, and disturbance tolerance. Furthermore, environmental adaptability and system robustness in high-speed, high-dynamic scenarios require further improvement. Furthermore, due to its relatively late start in China, research on pose estimation for unmanned systems remains primarily based on traditional vision, with relatively little research on event cameras. Therefore, conducting research on pose estimation methods based on event cameras in complex and high-dynamic scenes provides technical support for the ubiquitous application of robots and unmanned systems. In the tide of vigorously developing unmanned and intelligent equipment, it has important scientific significance and practical application value for promoting the development of my country's robotics industry and enhancing the strength of unmanned equipment.

[0004] The existing technology comparison is as follows;

[0005] Comparative Patent 1: Visual Inertial Odometry Using an Event Camera, authorization announcement date, 2020 / 4 / 21; authorization announcement date number: CN111052183A, the invention relates to a method for generating a motion-corrected image for a visual inertial odometry including an event camera, wherein the event camera is rigidly connected to an inertial measurement unit (IMU), wherein the event camera includes pixels arranged in an image plane, and the pixels are configured to output an event (e) when a brightness change occurs when there is a brightness change in the scene, wherein each event (e) includes the time when the event is recorded and the position of the corresponding pixel where the brightness change is detected, and the method comprises the following steps: obtaining at least one set (S) of events, wherein, to At least one set (S) includes a plurality of subsequent events (e); IMU data (D) is acquired over the duration of at least one set (S); a motion-corrected image is generated from at least one set (S) of events (e), wherein the motion-corrected image is obtained by assigning a position of each event recorded at an estimated camera pose at a corresponding event time of each event to an adjusted event position, wherein the adjusted event position is obtained by determining a position of the event (ej) with respect to an estimated reference camera pose at a reference time, wherein the estimated camera pose at the event time (tj) and the estimated reference camera pose at the reference time are estimated by means of the IMU data.

[0006] Although it also involves the use of event cameras for pose estimation, the comparison technology also combines the motion data of the inertial measurement unit (IMU) to correct the event image, and the overall algorithm architecture is a combined navigation with tightly coupled events and IMUs; while this application does not use IMU data, but achieves camera pose estimation based on pure events through the motion coding plane representation generated from the events, combined with a completely different algorithm architecture.

[0007] Comparative Patent 2: A Method for Determining the Main Human Posture in Examination Videos Using Motion Coding, Authorization Announcement Date: May 2, 2022, Authorization Announcement Number: CN 110738151 B. This invention discloses a method for determining the main human posture in examination videos using motion coding, comprising the following steps: S1: Motion coding the examinee: Extracting the examinee's posture from each frame of the examinee's video to form a posture sequence, and then motion coding the examinee based on this sequence; S2: Still segment segmentation: Detecting still segments in the examination video using motion coding; S3: Posture classification: Calculating the posture mean of each still segment detected in S2 and creating a posture category array based on the posture mean; S4: Main posture classification: Traversing the posture category array and selecting the category corresponding to the posture with the longest total duration as the examinee's main posture; S5: Determining the reasonable range of motion of the examinee's elbow joint. This method for determining the main human posture in examination videos using motion coding can accurately and quickly determine the examinee's main posture during the examination and the reasonable range of motion of the examinee's elbow joint when the examinee is in the main posture.

[0008] Although also called "motion coding," in the comparative patent, "motion coding" refers to converting the examinee's behavior into a series of codes or labels according to specific rules to facilitate analysis and identification; in this application, "motion coding" refers to estimating the local dynamics of the visual scene from a high-frequency asynchronous event stream, encoding historical events in combination with an exponential decay kernel, and achieving a consistent representation of the scene. Therefore, the meanings and applications of the two are clearly different.

[0009] Comparative Patent 3: A SLAM method and system based on a multimodal semantic framework for dynamic environments, authorization announcement date: 2023 / 10 / 31, authorization announcement number: CN116977628A, this invention discloses a SLAM method and system based on a multimodal semantic framework for dynamic environments. The specific process includes camera image acquisition, lidar data processing, IMU pre-integration, instance segmentation, multimodal fusion, feature map update and global semantic map construction. Through the multimodal fusion of vision, lidar and IMU, the accuracy and computational complexity are balanced; through dynamic semantic understanding, multimodal sensors are assisted in building 3D dynamic maps, ensuring the accuracy and real-time performance of the system. The present invention solves the problem of dynamic and unstructured underground and warehouse navigation and positioning where GNSS cannot function, and realizes unmanned driving precision perception and positioning technology in unstructured environments.

[0010] First, the comparison technology is a SLAM algorithm based on a combination of traditional cameras, lidar, and IMUs, while the present application is a pose estimation algorithm based on an event camera, with a different input modality. Second, the comparison technology uses a semantic framework for perception and positioning, while the present application mainly performs pose calculations through 3D-2D alignment, using a different algorithm. Furthermore, the "dynamic" in the comparison technology refers to dynamic objects in the environment, while the "high dynamic" in the present application refers to challenging scenes caused by high-speed camera motion and complex lighting conditions. Therefore, the comparison patent and the present application method are completely different technologies. Summary of the Invention

[0011] To solve the above technical problems, the present invention proposes a high-dynamic pose estimation method for intelligent agents based on motion-coded event plane representation. It utilizes the dynamic information continuously captured by the event camera to construct a consistent plane representation of the scene. Combined with the semi-dense three-dimensional map of the environment, it can achieve accurate estimation of the six-degree-of-freedom pose of the camera in high-speed and high-dynamic scenes.

[0012] To achieve the above object, the technical solution adopted by the present invention is:

[0013] A method for high-dynamic pose estimation of an intelligent agent based on motion-encoded event plane representation has the following specific steps, which are characterized by:

[0014] Step 1: Read the event stream output by the event camera. Each event is represented by a four-dimensional vector e(x, y, t, p), which contains the pixel coordinates (x, y), trigger time t, and polarity p. The polarity p∈{1,-1} represents the increase or decrease of the brightness of the pixel.

[0015] Step 2: Use a doubly linked list to store the pixel triggering time and response order in the local neighborhood, and construct an asynchronous event-driven thread to perform stack updates;

[0016] Step 3: Based on the event camera and the semi-dense and low-latency response to the scene outline, the pixel-level planar motion of the event is estimated by constructing threads accessing the doubly linked list in real time, and a motion-encoded event plane representation is constructed to achieve a consistent representation of the environment in high-speed and high-dynamic scenes.

[0017] Step 4: Construct a semi-dense local map of the scene at the reference moment in the form of a 3D point cloud using an event-based or image-based composition algorithm or with the help of perception methods such as RGB-D depth cameras:

[0018]

[0019] Step 5: Based on the MER's encoding representation of the spatiotemporal constraints of camera motion, a 3D-2D alignment technique is used to synchronize the local map to the current MER and minimize the global geometric alignment error to achieve high-precision robust estimation of the camera's six-degree-of-freedom pose, thereby fully unleashing the potential of event cameras in high-speed pose tracking applications. The 3D-2D alignment pose estimation algorithm fully utilizes the event camera's low-latency natural response to scene edges without the need for preprocessing operations such as feature extraction or data association.

[0020] As a further improvement of the present invention, in step 2, a bidirectional linked list is constructed to store the triggering time and sequence of pixels in the local neighborhood, and the update of the thread is driven by asynchronous events, including the following process:

[0021] (2-1) The pixel plane of the event camera is divided into multiple subdomains of size (2R+1)×(2R+1), which are defined as local pixel neighborhoods

[0022]

[0023] Among them, (x i ,y i ) is the center pixel of the subdomain, and the neighborhood range parameter R is set to 4. In order to enhance the spatial correlation of pixels near the boundary, a semi-overlapping layout is adopted, where adjacent subdomains are spaced R pixels apart and overlap each other by half their size.

[0024] (2-2) Based on the stacking strategy to reduce the memory and time complexity of the algorithm, for each local pixel neighborhood Using two doubly linked lists As a memory unit, it stores the trigger time of positive (+) and negative (-) polarity events in response order:

[0025]

[0026] As the basic unit of the stack structure, each node corresponds to a pixel Contains a numerical unit to record the latest trigger time t for the corresponding polarity s∈{+1,-1} event last (x, y, p), and two pointer links are connected to the previous and next nodes respectively. The doubly linked list occupies a fixed-size memory unit and is connected in sequence according to the pixel response, which is convenient for forward or backward sequential access starting from any given node;

[0027] (2-3) The stack linked list is updated through an event-driven thread. For each event in the data stream, the k bidirectional linked lists (k≤4) to which the triggering pixel belongs are updated according to its polarity, including the update of the numerical unit and the sequential update of the linked list. When a pixel node is activated, a pointer link is established between the two nodes connected to it. After that, the node is removed and relinked to the end of the linked list, realizing efficient update of the triggering time and response order of pixels in the neighborhood.

[0028] As a further improvement of the present invention, constructing the event plane representation of motion coding in step 3 includes the following process:

[0029] (3-1) For each pixel on the plane, estimate the instantaneous velocity of the scene edge on the plane when the last projection arrives. Based on the semi-dense contour representation of the scene by the event camera, the pixel (x i ,y i ) belongs to the local neighborhood The edge length within is approximately (2R+1), based on its response event e i (x i ,y i ,t i ,p i ), from the linked list to which the node belongs Sequentially search n(2R+1) t i Before and after nodes:

[0030]

[0031] According to the ultra-low latency and continuous response characteristics of the event camera, no matter how fast the scene edge moves on the pixel plane, it will continuously trigger pixels and output events in chronological order and along its movement direction. Therefore, the time span of the node sequence is regarded as the time for the edge to pass through n pixels. If n is set to 3, its instantaneous speed is

[0032]

[0033] The calculation defines the local instantaneous planar motion of the visual scene on its pixels, reflecting the local spatiotemporal scene dynamics perceived by the event camera;

[0034] After (3-2), an exponential decay kernel is introduced to compress the three-dimensional spatiotemporal characteristics of the event representation, highlighting the recent event motion. Combined with the real-time calculated event plane motion, the pixel-level dynamic decay rate is encoded, and the historical event is propagated to the current moment. The motion-encoded event plane representation MER is constructed. The formula is as follows:

[0035]

[0036] Therefore, the intensity of the event plane is a function of the motion history of that location, describing the recent activity of the scene edge at the pixel plane. cur is the current moment, t last is the most recent response time on the pixel (x, y), v(x, y) is the edge speed when the corresponding pixel responds, exp(·) is an exponential function. After each pixel is triggered by the edge and initialized to 1, each unit pixel movement of the edge exponentially decays the MER at a constant rate. The movement distance d of the edge is determined by its speed at the corresponding pixel and the time elapsed. Once the edge moves a certain number of pixels d th , MER will be set to 0 to improve the signal-to-noise ratio and depict the scene outline with uniform thickness. th It is set to 8. In addition, this method also introduces a maximum attenuation rate threshold δ to suppress the influence of background noise. δ is set to 30ms.

[0037] As a further improvement of the present invention, the estimation of the camera's six-degree-of-freedom pose based on the 3D-2D alignment algorithm in step 5 includes the following process:

[0038] (5-1) Convert MER into inverse form and construct the constraint potential field of the observation scene contour in the current posture state:

[0039]

[0040] Based on the estimated camera pose, the map points are in the constrained potential field The projected sampling values ​​in ,represent the spatiotemporal inconsistency of the observations;

[0041] (5-2) Define the projection transformation function W to project the map point to the current view:

[0042]

[0043] in, is the current pose parameter relative to the reference time, parameterized by the attitude and translation vector t represented by the quaternion q. The function T(·) converts the pose parameter into a rigid body transformation matrix, and π(·) projects the map point converted to the current pose onto the pixel plane.

[0044] The goal of the 3D-2D alignment algorithm is to determine the optimal pose parameter θ to better align the map point cloud with the minimum value of the current constraint potential field F, thereby minimizing the spatiotemporal inconsistency of the observations. The objective function is:

[0045]

[0046] The nonlinear optimization problem is solved by the Levenberg-Marquardt algorithm.

[0047] Beneficial effects:

[0048] This invention fully utilizes the low latency advantage of event cameras and their natural response to scene edges to estimate event plane motion, construct a motion-encoded event plane representation, and achieve a prototype representation with dynamic consistency for the environment. The invention also accelerates calculations through a double-linked list and multi-threading to significantly improve computational efficiency and save memory power consumption. Based on the encoding of spatiotemporal constraints of camera motion by event plane representation, semi-dense 3D-2D alignment technology is used to estimate six-degree-of-freedom pose, thereby achieving high-precision and robust pose tracking in high-speed and high-dynamic scenes, thereby fully unlocking the potential of event cameras in high-speed unmanned system applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a flow chart of the high dynamic pose estimation method provided by the present invention.

[0050] Figure 2 This is a schematic diagram of a bidirectional chain representation of the storage pixel trigger time and response sequence provided by the present invention.

[0051] Figure 3 This is a comparison between event plane representation and image frames in high-speed and high-dynamic scenes provided by the present invention. DETAILED DESCRIPTION

[0052] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0053] The present invention discloses a method for estimating high dynamic pose of an intelligent body based on motion coding event plane representation. The method process of the present invention is as follows: Figure 1 As shown, the specific steps include:

[0054] Step 1: Read the event stream output by the event camera. Each event is represented by a four-dimensional vector e(x, y, t, p), which contains the pixel coordinates (x, y), trigger time t, and polarity p. The polarity p∈{1,-1} indicates the increase or decrease of the brightness of the pixel.

[0055] Step 2: Use a doubly-linked list table to store the pixel trigger time and response order within the local neighborhood, and construct an asynchronous event-driven pipeline to perform stack updates. The specific steps include:

[0056] (2-1) The pixel plane of the event camera is divided into multiple subareas of size (2R+1)×(2R+1), which are defined as local pixel neighborhoods.

[0057]

[0058] Among them, (x i ,y i ) is the central pixel of the subdomain, and the neighborhood range parameter R is set to 4. In order to enhance the spatial correlation of pixels near the boundary, a half-overlap layout is adopted, such as Figure 2 As shown, adjacent sub-fields are spaced apart by R pixels and overlap each other by half their size.

[0059] (2-2) Based on the stack strategy, the memory and time complexity of the algorithm are reduced, and for each local pixel neighborhood Using two doubly linked lists As a memory unit, it stores the trigger time of positive (+) and negative (-) polarity events in response order:

[0060]

[0061] As the basic unit of the stack structure, each node corresponds to a pixel Contains a numerical unit (Node data) to record the latest trigger time t for the corresponding polarity s∈{+1,-1} event last (x, y, p), and two pointer links (Link) connected to the previous and next nodes respectively. The doubly linked list occupies a fixed-size memory unit and is connected sequentially according to the pixel response, facilitating forward or backward sequential access from any given node.

[0062] (2-3) The stack linked list is updated through the event-driven thread. For each event in the data stream, the k bidirectional linked lists (k≤4) to which the triggering pixel belongs are updated according to its polarity, including the update of the numerical unit and the order of the linked list. Figure 2 As shown in the figure, when a pixel node is activated, a pointer link is established between the two nodes connected to it, and then the node is removed and relinked to the end of the linked list to achieve efficient update of the triggering time and response order of pixels in the neighborhood.

[0063] Step 3: Based on the event camera and the semi-dense and low-latency response to the scene outline, the synchronous on-demand generation pipeline accesses the doubly linked list to estimate the pixel-level planar motion of the event and construct a motion-encoded event-surface representation (MER), achieving a consistent representation of the environment in high-speed and high-dynamic scenes. The specific steps include:

[0064] (3-1) For each pixel on the plane, estimate the instantaneous velocity of the scene edge on the plane when the last projection arrives. Based on the semi-dense contour representation of the scene by the event camera, the pixel (x i ,y i ) belongs to the local neighborhood The edge length within is approximately (2R+1), based on its response event e i (x i ,y i ,t i ,p i ), from the linked list to which the node belongs Sequentially search n(2R+1) t i Before and after nodes:

[0065]

[0066] Due to the ultra-low latency and continuous response characteristics of the event camera, no matter how fast the scene edge moves on the pixel plane, it will continuously trigger pixels and output events in a time sequence along its direction of movement. Therefore, the time span of this node sequence can be regarded as the time it takes for the edge to pass through n pixels. If n is set to 3, its instantaneous speed can be calculated by

[0067]

[0068] The calculation defines the local instantaneous planar motion of the visual scene at its pixels, reflecting the local spatiotemporal scene dynamics perceived by the event camera. This method enhances the recognition of contours by leveraging event polarity, which reveals the direction of event and luminosity changes and can instantly respond to differences in relative edge motion.

[0069] After (3-2), an exponential decay kernel is introduced to compress and represent the three-dimensional spatiotemporal features of the event, highlighting the recent event motion. Combined with the real-time resolved event plane motion, the pixel-level dynamic decay rate is encoded, and the historical events are propagated to the current moment to construct the motion-encoded event plane representation (MER). The formula is as follows:

[0070]

[0071] Therefore, the intensity of the event plane is a function of the motion history of that location, describing the recent activity of the scene edge at the pixel plane. cur is the current moment, t last is the most recent response time on the pixel (x, y), v(x, y) is the edge speed when the corresponding pixel responds, and exp(·) is an exponential function. After each pixel is triggered by the edge and initialized to 1, each unit pixel movement of the edge will exponentially decay the MER at a constant rate. The movement distance d of the edge is determined by its speed at the corresponding pixel and the time it has passed. Once the edge moves a certain number of pixels d th , MER will be set to 0 to improve the signal-to-noise ratio and depict the scene outline with uniform thickness. th Set to 8. In addition, this method also introduces a maximum attenuation rate threshold δ to suppress the influence of background noise, and δ is set to 30ms.

[0072] Comparison of MER and image frames in high-speed and high-dynamic scenes Figure 3 As shown, it can be seen that MER effectively avoids motion blur in the image and achieves a consistent representation of the environment. Moreover, since MER uses local operations, it has the ability to adapt to scene textures.

[0073] Step 4: Construct a semi-dense local map of the scene at the reference moment in the form of a 3D point cloud using an event-based or image-based composition algorithm or a perception method such as an RGB-D depth camera:

[0074]

[0075] Step 5: Based on the MER's encoding representation of the spatiotemporal constraints of camera motion, a 3D-2D alignment algorithm is used to synchronize the local map to the current MER and minimize the global geometric alignment error to achieve high-precision robust estimation of the camera's six-degree-of-freedom (6-DOF) pose, thereby fully unleashing the potential of event cameras in high-speed pose tracking applications. The 3D-2D alignment pose estimation algorithm fully utilizes the event camera's low-latency natural response to scene edges without the need for preprocessing operations such as feature extraction or data association. The specific steps include:

[0076] (5-1) Convert MER into inverse form and construct the constraint potential field of the observation scene contour in the current posture state:

[0077]

[0078] Based on the estimated camera pose, the map points are in the constrained potential field The projected sampling values ​​in , represent the spatiotemporal inconsistency of the observations.

[0079] (5-2) Define the projection transformation function W to project the map point to the current view:

[0080]

[0081] in, is the current pose parameter relative to the reference time, parameterized by the attitude represented by the quaternion q and the translation vector t. The function T(·) converts the pose parameters into a rigid body transformation matrix, and π(·) projects the map point converted to the current pose onto the pixel plane.

[0082] The goal of the 3D-2D alignment algorithm is to determine the optimal pose parameter θ to better align the map point cloud with the current constraint potential field. The minimum value of is aligned to minimize the spatiotemporal inconsistency of observations. The objective function is:

[0083]

[0084] The nonlinear optimization problem is solved by using the Levenberg-Marquardt algorithm.

[0085] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent variation based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A high-dynamic pose estimation method based on motion coding event plane representation, comprising the following steps, characterized in that: Step 1: Read the event stream output by the event camera. Each event is represented by a four-dimensional vector. , contains the pixel coordinates of the event , trigger time and polarity , where polarity Indicates the increase or decrease in brightness of a pixel; Step 2: Use a doubly linked list to store the pixel triggering time and response order in the local neighborhood, and construct an asynchronous event-driven thread to perform stack updates; In step 2, a bidirectional linked list is constructed to store the triggering time and order of pixels in the local neighborhood. The update of the thread driven by asynchronous events includes the following process: (2-1) Divide the pixel plane of the event camera into multiple subdomain, defined as a local pixel neighborhood : ; in, is the center pixel of the subdomain, and the neighborhood range parameter R is set to 4. In order to enhance the spatial correlation of pixels near the boundary, a semi-overlapping layout is adopted, with each adjacent subdomain spaced R pixels apart and overlapping each other by half the size. (2-2) Based on the stacking strategy to reduce the memory and time complexity of the algorithm, for each local pixel neighborhood Using two doubly linked lists As a memory unit, it stores the triggering time of positive and negative polarity events in response order: ; As the basic unit of the stack structure, each node corresponds to a pixel , contains a numerical unit to record its corresponding polarity The latest trigger time of the event , and two pointer links are connected to the previous and next nodes respectively. The bidirectional linked list occupies a storage unit of fixed size memory and is connected in sequence according to the pixel response, which is convenient for forward or backward sequential access starting from any given node; (2-3) Update the stack linked list through an event-driven thread. For each event in the data stream, update the k bidirectional linked lists (k ≤ 4) to which the triggering pixel belongs according to its polarity. This includes updating the numeric unit and the order of the linked lists. When a pixel node is activated, a pointer link is established between the two nodes connected to it. After that, the node is removed and relinked to the end of the linked list, achieving efficient update of the trigger time and response order of pixels in the neighborhood. Step 3: Based on the event camera and the semi-dense and low-latency response to the scene outline, the pixel-level planar motion of the event is estimated by constructing threads accessing the doubly linked list in real time, and a motion-encoded event plane representation is constructed to achieve a consistent representation of the environment in high-speed and high-dynamic scenes. Step 4: Construct a semi-dense local map of the scene at the reference moment in the form of a 3D point cloud using an event-based or image-based composition algorithm or with the help of an RGB-D depth camera perception method: ; Step 5: Based on MER’s encoding representation of the spatiotemporal constraints of camera motion, 3D-2D alignment technology is used to synchronize the local map to the MER at the current moment and minimize the global geometric alignment error to achieve high-precision robust estimation of the camera’s six-degree-of-freedom pose, thereby fully unleashing the potential of event cameras in high-speed pose tracking applications. The 3D-2D alignment pose estimation algorithm fully utilizes the event camera’s low-latency natural response to scene edges without the need for feature extraction or data association preprocessing operations.

2. The high-dynamic pose estimation method based on motion-coded event plane representation according to claim 1, characterized in that: Constructing the event plane representation of motion coding in step 3 includes the following process: (3-1) For each pixel on the plane, estimate the instantaneous velocity of the scene edge on the plane when the last projection arrives, based on the semi-dense contour representation of the scene by the event camera, pixel Belong to the local neighborhood The edge length inside is approximately , based on its response events , from the linked list to which the node belongs Sequential search indivual Before and after nodes: ; According to the ultra-low latency and continuous response characteristics of the event camera, no matter how fast the scene edge moves on the pixel plane, it will continuously trigger pixels and output events in chronological order and along its movement direction. Therefore, the time span of the node sequence is regarded as the time for the edge to pass through n pixels. If n is set to 3, its instantaneous speed is ; The calculation defines the local instantaneous planar motion of the visual scene on its pixels, reflecting the local spatiotemporal scene dynamics perceived by the event camera; After (3-2), an exponential decay kernel is introduced to compress the three-dimensional spatiotemporal features of the event representation, highlighting the recent event motion. Combined with the real-time resolved event plane motion, the pixel-level dynamic decay rate is encoded, and the historical events are propagated to the current moment. The motion-encoded event plane representation MER is constructed. The formula is as follows: ; The intensity of the event plane is therefore a function of the motion history of that location, describing the recent activity of the scene edge at the pixel plane, where For the current moment, It's a pixel The most recent response time on is the edge speed corresponding to the pixel response, It is an exponential function. After each pixel is triggered by the edge and initialized to 1, each time the edge moves a unit pixel, the MER is exponentially decayed at a constant rate. The moving distance of the edge Determined by its speed at the corresponding pixel and the time it takes, once the edge moves a certain number of pixels MER will be set to 0 to improve the signal-to-noise ratio and depict the scene outline with uniform thickness. Set to 8. In addition, this method also introduces a maximum decay rate threshold , suppress the influence of background noise, Set to 30ms.

3. The high-dynamic pose estimation method based on motion-coded event plane representation according to claim 1, characterized in that: In step 5, estimating the camera's six-degree-of-freedom pose based on the 3D-2D alignment algorithm includes the following steps: (5-1) Convert MER into inverse form and construct the constraint potential field of the observation scene contour in the current posture state: ; Based on the estimated camera pose, the map points are in the constrained potential field The projected sampling values ​​in ,represent the spatiotemporal inconsistency of the observations; (5-2) Define the projection transformation function W to project the map point to the current view: ; in, is the current pose parameter relative to the reference time, which is represented by the quaternion Represents the pose and translation vector Parameterization, function Convert the pose parameters into a rigid body transformation matrix, Project the map points transformed to the current pose onto the pixel plane; The goal of the 3D-2D alignment algorithm is to determine the optimal pose parameters , to better integrate the map point cloud with the current constraint potential field The minimum value of is aligned to minimize the spatiotemporal inconsistency of observations. The objective function is: ; The nonlinear optimization problem is solved by the Levenberg-Marquardt algorithm.

Citation Information

Patent Citations

  • A method for determining the main pose of the human body in examination room videos using motion coding

    CN110738151B

  • Visual-inertial odometry with event camera

    CN111052183A

  • Multi-modal semantic framework-based SLAM (simultaneous localization and mapping) method and system applied to dynamic environment

    CN116977628A

  • Pose estimation method and related device

    CN115997234A

  • Six-degree-of-freedom real-time pose estimation method and device

    CN116797860A