A method for determining pose and related devices

By removing the corner points of the dynamic target in the indoor positioning system and combining IMU and wheel speedometer data, the problem of large visual positioning error in dynamic environments is solved, and more accurate position determination is achieved.

CN114445490BActive Publication Date: 2025-07-18YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011199017.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-31
Publication Date
2025-07-18
Estimated Expiration
2040-10-31

AI Technical Summary

Technical Problem

In indoor environments, vision-based positioning technology cannot provide reliable observations due to dynamic moving objects, resulting in large errors in position calculation, and it is difficult for the prior art to achieve accurate position determination in dynamic environments.

Method used

By detecting corner points on the image frames captured by the camera device, the corner points of the area where the dynamic target is located are eliminated, the eliminated corner points and inertial measurement unit (IMU) and wheel speedometer data are used, and the positioning change is determined in combination with a nonlinear optimization algorithm to reduce the error introduced by dynamic features.

Benefits of technology

It improves the accuracy of position determination in dynamic environments, reduces visual reprojection errors, and enhances the reliability of the positioning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114445490B_ABST
    Figure CN114445490B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a pose determination method, and the method includes: obtaining a first image frame and a second image frame captured by an imaging device, where the first image frame and the second image frame respectively include a dynamic target; respectively performing corner detection on the first image frame and the second image frame to obtain a plurality of first corners of the first image frame and a plurality of second corners of the second image frame; removing the first corners in the area of the dynamic target of the first image frame among the plurality of first corners to obtain the plurality of first corners after removal; removing the second corners in the area of the dynamic target of the second image frame among the plurality of second corners to obtain the plurality of second corners after removal; determining the pose change of the imaging device according to the plurality of first corners after removal and the plurality of second corners after removal. In this embodiment, the corners in the area of the dynamic target among the plurality of corners are removed to ensure that the visual reprojection error is more reliable, thereby overcoming the error in determining the pose change introduced by the dynamic features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more particularly, to a pose determination method and related devices thereof. Background Art

[0002] With the development of mobile robot technology, indoor real-time positioning technology has received extensive attention. Knowing the position of the robot itself can provide real-time information for modules such as planning and control to complete desired tasks. However, in an indoor environment, the global positioning system (GPS) cannot be used for positioning due to unstable signals. Some indoor positioning methods based on signal generating devices, such as ultra-wide band (UWB) and wireless fidelity (WIFI), require the installation of signal generating devices in the usage scenario or are prone to problems such as restricted movement areas and cost issues. Secondly, due to the high cost of laser sensors, indoor positioning technology based on lasers will bring about the problem of excessively high costs.

[0003] Due to the characteristics of camera devices, such as low cost and the ability to capture rich information, vision-based active real-time positioning technology has been proposed. Common positioning schemes include vision-based positioning technology. The main idea of this scheme is to solve the pose of the moving body based on visual feature point matching and global optimization. First, corner points are extracted from the image, and the pose change of the shooting device between the two frames is calculated using the matching corner points between the two frames. The pose change can include position change and rotation angle change. However, in some scenarios, there are dynamically moving objects, and at this time, the visual constraint cannot provide reliable observation values, resulting in large errors and thus affecting the calculation accuracy of the relative pose. Summary of the Invention

[0004] In a first aspect, this application provides a pose determination method, and the method includes:

[0005] Obtain a first image frame and a second image frame captured by a camera device. The first image frame and the second image frame are adjacent image frames captured by the camera device, and the first image frame and the second image frame each include a dynamic target, where a dynamic target is a target that has a displacement relative to the ground when the camera device captures an image frame. The so-called displacement relative to the ground when the camera device captures an image frame means that the dynamic target has moved relative to the ground in space, for example, moving from position A to position B (position A and position B are two different positions). In one implementation, the dynamic target can be a vehicle, such as a sedan, a truck, a bus, a trailer, an incomplete vehicle, and a motorcycle, etc. It should be understood that the first image frame and the second image frame may include the same dynamic target or different dynamic targets; respectively perform corner detection on the first image frame and the second image frame to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame; in some cases, it can be understood that the human eye's recognition of an image is usually completed in a local small area or small window. If the gray level of the image in the window changes significantly when the small window is moved a small range in each direction, then it can be considered that there are corner points in the window. If the gray level of the image in the window does not change when the small window is moved a small range in each direction, then it can be considered that there are no corner points in the window. In some embodiments, the plurality of first corner points and the plurality of second corner points may be Features from accelerated segment test (FAST) corner points, Harris corner points, or Binary Robust Invariant Scalable Keypoints (BRISK) corner points, etc. In this regard, the embodiments of the present invention do not make limitations. Remove the first corner points in the area of the dynamic target included in the first image frame among the plurality of first corner points to obtain the plurality of removed first corner points; remove the second corner points in the area of the dynamic target included in the second image frame among the plurality of second corner points to obtain the plurality of removed second corner points; when there are dynamic corner points (that is, corner points in the area of the dynamic target), the observations of the same corner point in different camera states include not only the parallax introduced by the movement of the camera itself, but also the parallax brought by the movement of the feature point itself. In the embodiments of the present application, by removing the corner points in the area of the dynamic target among the plurality of corner points, the reliability of the visual reprojection error term can be ensured. Determine the pose change of the camera device when capturing the second image frame relative to when capturing the first image frame according to the plurality of removed first corner points and the plurality of removed second corner points.

[0006] The pose change may refer to the change in the position of the imaging device and the change in the rotation angle. The change in position may represent the distance between the position of the imaging device when capturing the second image frame and the position when capturing the first image frame. The change in rotation angle may represent the angular difference between the rotation angle of the imaging device when capturing the second image frame and the rotation angle when capturing the first image frame. When the imaging device is fixed on a vehicle, the pose change of the imaging device can be considered as the pose change of the vehicle.

[0007] Taking the first image frame as the image frame before the second image frame as an example, multiple corner points extracted from the first image frame can be used for optical flow, so as to find the matching corner points between the first image frame and the second image frame, and obtain the optical flow information of the matching corner points. This optical flow information is used to represent the motion information of the matching corner points in these two adjacent images (which can be called parallax from the perspective of image frames). Further, the pose change of the imaging device when capturing the second image frame relative to when capturing the first image frame can be determined based on the optical flow information. The algorithm used for optical flow can be the Lucas-Kanade optical flow algorithm or other algorithms. In addition to optical flow, descriptors or direct methods can also be used to match the corner points.

[0008] When there are dynamic corner points, the observations of the same corner point in different camera states include not only the parallax introduced by the movement of the camera itself, but also the parallax caused by the movement of the corner point itself. In this embodiment, through the above method, for the problem of large positioning errors in a dynamic environment, the corner points in the area where the dynamic target is located are removed from the corner points to ensure that the visual reprojection error is more reliable, thereby overcoming the errors when the pose changes introduced by dynamic features.

[0009] In a possible implementation, the dynamic target is a vehicle.

[0010] In a possible implementation, the method is applied to a target vehicle, and the target vehicle is fixedly provided with the camera device, an inertial measurement unit (IMU), and a wheel speed meter; the method further includes: obtaining the motion state data of the target vehicle measured by the IMU during the period from shooting the first image frame to shooting the second image frame by the camera device; obtaining the wheel speed data of the target vehicle measured by the wheel speed meter during the period from shooting the first image frame to shooting the second image frame by the camera device; correspondingly, determining the pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the remaining multiple first corner points and the remaining multiple second corner points includes: determining the first pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the parallax of the remaining multiple first corner points and the remaining multiple second corner points in their respective image frames, where the first pose change does not include scale information; determining the second pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the motion state data, the wheel speed data, and the first pose change, and the second pose change includes scale information; performing non-linear optimization on the second pose change to obtain the pose change. Among them, the motion state data may be the vehicle acceleration data and angular velocity data measured by an inertial sensor (IMU), and the wheel speed data may be the vehicle wheel rotation speed, steering wheel angle data, etc.

[0011] In one implementation, the remaining multiple first corner points and the remaining multiple second corner points can be aligned with the motion state data and the wheel speed data in terms of time stamps. Among them, due to different data frequencies, the single-frame image data (the remaining corner points) correspond to multiple motion state data and wheel speed data. Then, pre-integration processing can be performed on the motion state data and the wheel speed data of each frame of image to provide an initial pose value for the image. Then, a sliding window method is adopted to perform initialization processing using pure vision data to obtain the poses of the camera device when shooting all image frames within the sliding window and the positions without scale information (which can also be called depth). Combining the motion state data and the wheel speed data to restore the above-mentioned scale information without scale, and recalculating the positions of all corner points, and then performing non-linear optimization on the restored scale information and the inverse depth of the corner points to obtain the pose change.

[0012] In a possible implementation, performing non-linear optimization on the second pose change includes: performing non-linear optimization on the second pose change according to a preset optimization function, where the preset optimization function includes a wheel speed meter residual term.

[0013] For smooth or real - value - approaching optimization of pose changes, non - linear optimization is usually performed. Non - linear optimization refers to finding an optimal set of numerical mappings in a given objective function, that is, x - min f(x). According to the derivative theory, the valid values of x can be obtained by solving the derivative equation Δf(x)=0. When f(x) is a non - linear function, this optimization is non - linear optimization. In the embodiments of the present application, the optimization function can be composed of five terms, and can be specifically expressed as: Optimization function = Prior residual + IMU residual term + Wheel speedometer residual term + Visual reprojection error term + Loop closure detection reprojection error term.

[0014] In the embodiments of the present application, to address the problem of large positioning errors in a dynamic environment, the corner points in the region where the dynamic target is located among multiple corner points are removed to ensure that the visual reprojection error is more reliable, thereby overcoming the positioning errors introduced by dynamic features. And for the disadvantage of large initialization errors in the two - dimensional motion of the visual - inertial system, a wheel speedometer residual term is added. The pose is initialized and calculated through the joint information of vision, IMU, and wheel speedometer. When the IMU convergence is not good, the wheel speedometer can be used for compensation, thereby improving the effect.

[0015] In a possible implementation, the method further includes:

[0016] Detecting the dynamic target in the first image frame and the dynamic target in the second image frame through a pre - trained neural network to obtain the region where the dynamic target included in the first image frame is located, and the region where the dynamic target included in the second image frame is located.

[0017] In one implementation, the current key frame can also be subjected to loop closure detection with the map. The detected frame is the loop closure frame, and the co - visible feature points between the loop closure frame and the frames in the sliding window are found; a reprojection error term is established and added to the non - linear optimization; the key frame that has been optimized and slid out of the window is optimized with its associated frame in four degrees of freedom; the optimized key frame is inserted into the map. In actual use, the map, that is, loop closure detection, can be selected to be turned on or off. When it is turned on, mapping is performed while positioning, and the final obtained positions are all relative to the position of the first fixed camera. Of course, this can be fixed to a specific reference system through subsequent coordinate transformation.

[0018] In a second aspect, the present application provides a pose determination device, and the device includes:

[0019] An acquisition module, configured to acquire a first image frame and a second image frame captured by a camera device, where the first image frame and the second image frame are adjacent image frames captured by the camera device, and the first image frame and the second image frame respectively include dynamic targets;

[0020] A corner extraction module for obtaining a plurality of first corners of the first image frame and a plurality of second corners of the second image frame by performing corner detection on the first image frame and the second image frame respectively;

[0021] A rejection module for rejecting the first corners in the area of the dynamic target included in the first image frame among the plurality of first corners to obtain the plurality of first corners after rejection;

[0022] Rejecting the second corners in the area of the dynamic target included in the second image frame among the plurality of second corners to obtain the plurality of second corners after rejection;

[0023] A positioning module for determining the pose change of the imaging device when shooting the second image frame relative to when shooting the first image frame according to the plurality of first corners after rejection and the plurality of second corners after rejection.

[0024] In a possible implementation, the plurality of first corners and the plurality of second corners include one of the following: FAST corners of accelerated segment test features, Harris corners, and BRISK corners of binary robust invariant scalable key points.

[0025] In a possible implementation, the device is applied to a target vehicle, and the target vehicle is fixedly provided with the imaging device, an inertial measurement unit (IMU), and a wheel speed meter; the acquisition module is used to acquire the motion state data of the target vehicle during the period from shooting the first image frame to shooting the second image frame measured by the IMU;

[0026] Acquiring the wheel speed data of the target vehicle during the period from shooting the first image frame to shooting the second image frame measured by the wheel speed meter;

[0027] Correspondingly, the positioning module is used to determine the first pose change of the imaging device when shooting the second image frame relative to when shooting the first image frame according to the parallax of the plurality of first corners after rejection and the plurality of second corners after rejection in their respective image frames, where the first pose change does not include scale information;

[0028] Determining the second pose change of the imaging device when shooting the second image frame relative to when shooting the first image frame according to the motion state data, the wheel speed data, and the first pose change, where the second pose change includes scale information;

[0029] Performing non-linear optimization on the second pose change to obtain the pose change.

[0030] In a possible implementation, the positioning module is configured to perform non-linear optimization on the second pose change according to a preset optimization function, where the preset optimization function includes a wheel speedometer residual term.

[0031] In a possible implementation, the device further includes:

[0032] A dynamic target detection module, configured to detect dynamic targets in the first image frame and detect dynamic targets in the second image frame through a pre-trained neural network, so as to obtain the regions where the dynamic targets included in the first image frame are located, and the regions where the dynamic targets included in the second image frame are located.

[0033] In a possible implementation, the dynamic target is a vehicle.

[0034] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores program code, where the program code includes instructions for performing some or all of the operations in the method described in the first aspect above.

[0035] In a fourth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a communication device, the communication device is caused to perform some or all of the operations in the method described in the first aspect above.

[0036] In a fifth aspect, a chip is provided. The chip includes a processor, and the processor is configured to perform some or all of the operations in the method described in the first aspect above.

[0037] An embodiment of the present application provides a pose determination method, the method comprising: acquiring a first image frame and a second image frame captured by an imaging device, the first image frame and the second image frame being adjacent image frames captured by the imaging device, the first image frame and the second image frame each including a dynamic target; respectively performing corner detection on the first image frame and the second image frame to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame; removing the first corner points in the plurality of first corner points that are in the region of the dynamic target included in the first image frame to obtain a plurality of first corner points after removal; removing the second corner points in the plurality of second corner points that are in the region of the dynamic target included in the second image frame to obtain a plurality of second corner points after removal; determining a pose change of the imaging device when capturing the second image frame relative to when capturing the first image frame according to the plurality of first corner points after removal and the plurality of second corner points after removal. When there are dynamic corner points, the observations of the same corner point in different camera states include not only the parallax introduced by the movement of the camera itself, but also the parallax caused by the movement of the corner point itself. Through the above method, this embodiment eliminates the corner points in the region of the dynamic target among the plurality of corner points to ensure that the visual reprojection error is more reliable, thereby overcoming the pose change determination error introduced by dynamic features for the problem of large positioning errors in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a functional block diagram of a vehicle provided by an embodiment of the present invention;

[0039] Figure 2 is a functional block diagram of an autonomous driving system provided by an embodiment of the present invention;

[0040] Figure 3 is a flowchart illustration of a pose determination method provided by an embodiment of the present application;

[0041] Figure 4 is a flowchart illustration of a pose determination method provided by an embodiment of the present application;

[0042] Figure 5 is a structural illustration of a pose determination device provided by an embodiment of the present application;

[0043] Figure 6 is a structural illustration of a pose determination device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following will describe embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention.

[0045] In the description, claims and drawings of this application, terms such as "first", "second", "third" and "fourth" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.

[0046] Reference to "an embodiment" in this context means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase may occur in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0047] The terms "component", "module", "system", etc. used in this specification are used to denote computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside in a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media storing various data structures. A component can communicate, for example, through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component in a local system, a distributed system, and / or a network, such as data interacting with other systems through signals over the Internet).

[0048] Figure 1 It is a functional block diagram of vehicle 100 provided by an embodiment of the present invention. In one embodiment, vehicle 100 is configured to be in a fully or partially autonomous driving mode. For example, vehicle 100 can control itself while in the autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment through manual operation, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the possibility of the other vehicle performing the possible behavior, and control vehicle 100 based on the determined information. When vehicle 100 is in the autonomous driving mode, vehicle 100 can be set to operate without interacting with a person.

[0049] Vehicle 100 may include various subsystems, such as a propulsion system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, as well as a power source 110, a computer system 112, and a user interface 116. Optionally, vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. Additionally, each subsystem and component of vehicle 100 may be interconnected by wire or wirelessly.

[0050] Propulsion system 102 may include components that provide powered movement for vehicle 100. In one embodiment, propulsion system 102 may include an engine 118, an energy source 119, a transmission 120, and wheels / tires 121. Engine 118 may be an internal combustion engine, an electric motor, an air compression engine, or other types of engine combinations, such as a hybrid engine composed of a gasoline engine and an electric motor, or a hybrid engine composed of an internal combustion engine and an air compression engine. Engine 118 converts the energy source 119 into mechanical energy.

[0051] Examples of energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other power sources. Energy source 119 may also provide energy for other systems of vehicle 100.

[0052] Transmission 120 may transmit the mechanical power from engine 118 to wheels / tires 121. Transmission 120 may include a gearbox, a differential, and a drive shaft. In one embodiment, transmission 120 may also include other devices, such as a clutch. The drive shaft may include one or more shafts that can be coupled to one or more wheels / tires 121.

[0053] Sensor system 104 may include several sensors that sense information about the environment around vehicle 100. For example, sensor system 104 may include a global positioning system 122 (the positioning system can be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU) 124, a radar system 126, a lidar 128, and a camera 130. Sensor system 104 may also include sensors that monitor the internal systems of vehicle 100 (such as an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). Sensor data from one or more of these sensors may be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). Such detection and identification are key functions for the safe operation of autonomous vehicle 100.

[0054] Among them, with the development of advanced driver assistance systems (ADAS) and autonomous driving technologies, higher requirements are imposed on the performance of the radar system 126, such as distance and angle resolution. The improvement of the distance and angle resolution of the in-vehicle radar system 126 enables the in-vehicle radar system 126 to detect multiple measurement points for a target object when imaging the target, forming high-resolution point cloud data. The radar system 126 in this application can also be referred to as a point cloud imaging radar.

[0055] The global positioning system 122 can be used to estimate the geographical location of the vehicle 100. The IMU 124 is used to sense the position and orientation changes of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope.

[0056] The radar system 126 can use radio signals to sense objects within the surrounding environment of the vehicle 100. In some embodiments, in addition to sensing objects, the radar system 126 can also be used to sense the speed and / or forward direction of the objects.

[0057] The lidar 128 can use lasers to sense objects in the environment where the vehicle 100 is located. In some embodiments, the lidar 128 can include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components.

[0058] The camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a static camera or a video camera. In the embodiments of this application, the camera 130 can also be referred to as an imaging device.

[0059] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 can include various elements, including a steering system 132, a throttle 134, a braking unit 136, a computer vision system 140, a route control system 142, and an obstacle avoidance system 144.

[0060] The steering system 132 can be operated to adjust the forward direction of the vehicle 100. For example, in one embodiment, it can be a steering wheel system.

[0061] The throttle 134 is used to control the operating speed of the engine 118 and thus control the speed of the vehicle 100.

[0062] The braking unit 136 is used to control the deceleration of the vehicle 100. The braking unit 136 can use friction to slow down the wheels / tires 121. In other embodiments, the braking unit 136 can convert the kinetic energy of the wheels / tires 121 into electric current. The braking unit 136 can also take other forms to slow down the rotational speed of the wheels / tires 121 so as to control the speed of the vehicle 100.

[0063] The computer vision system 140 can process and analyze the images captured by the camera 130 to identify objects and / or features in the surrounding environment of the vehicle 100. The objects and / or features may include traffic signals, road boundaries, and obstacles. The computer vision system 140 can use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 can be used to map the environment, track objects, estimate the speed of objects, and so on.

[0064] The route control system 142 is used to determine the driving route of the vehicle 100. In some embodiments, the route control system 142 can combine data from the sensors 138, the global positioning system 122, and one or more predefined maps to determine the driving route for the vehicle 100.

[0065] The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise cross potential obstacles in the environment of the vehicle 100.

[0066] Of course, in one example, the control system 106 can additionally or alternatively include components other than those shown and described. Or some of the above - shown components can also be reduced.

[0067] The vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users through the peripheral device 108. The peripheral device 108 may include a wireless communication system 146, an on - vehicle computer 148, a microphone 150, and / or a speaker 152.

[0068] In some embodiments, the peripheral device 108 provides a means for the user of the vehicle 100 to interact with the user interface 116. For example, the on - vehicle computer 148 can provide information to the user of the vehicle 100. The user interface 116 can also operate the on - vehicle computer 148 to receive user input. The on - vehicle computer 148 can be operated through a touch screen. In other cases, the peripheral device 108 can provide a means for the vehicle 100 to communicate with other devices located inside the vehicle. For example, the microphone 150 can receive audio (e.g., voice commands or other audio inputs) from the user of the vehicle 100. Similarly, the speaker 152 can output audio to the user of the vehicle 100.

[0069] The wireless communication system 146 can communicate wirelessly with one or more devices directly or via a communication network. For example, the wireless communication system 146 can use 3G cellular communication, such as Code Division Multiple Access (CDMA), Enhanced Versatile Disk (EVD), Global System for Mobile Communications (GSM) / General Packet Radio Service (GPRS), or 4G cellular communication, such as Long-Term Evolution (LTE). Or 5G cellular communication. The wireless communication system 146 can communicate with a Wireless Local Area Network (WLAN) using WiFi. In some embodiments, the wireless communication system 146 can communicate directly with a device using an infrared link, Bluetooth, or ZigBee. For other wireless protocols, such as various vehicle communication systems, for example, the wireless communication system 146 can include one or more Dedicated Short Range Communications (DSRC) devices, which can include public and / or private data communication between vehicles and / or roadside stations.

[0070] The power source 110 can supply power to various components of the vehicle 100. In one embodiment, the power source 110 can be a rechargeable lithium-ion or lead-acid battery. One or more battery packs of such a battery can be configured to supply power to various components of the vehicle 100. In some embodiments, the power source 110 and the energy source 119 can be implemented together, as in some all-electric vehicles.

[0071] Some or all of the functions of the vehicle 100 are controlled by the computer system 112. The computer system 112 can include at least one processor 113 that executes instructions 115 stored in a non-transitory computer-readable medium such as a data storage device 114. The computer system 112 can also be multiple computing devices that control individual components or subsystems of the vehicle 100 in a distributed manner.

[0072] The processor 113 can be any conventional processor, such as a commercially available Central Processing Unit (CPU). Alternatively, the processor can be a special-purpose device such as an Application Specific Integrated Circuit (ASIC) or other hardware-based processor. Although Figure 1The functional diagram shows the processor, memory, and other components of computer 110 in the same block, but those of ordinary skill in the art should understand that the processor, computer, or memory may actually include multiple processors, computers, or memories that may or may not be stored within the same physical enclosure. For example, the memory may be a hard disk drive or other storage medium located within an enclosure different from computer 110. Thus, a reference to a processor or computer will be understood to include a reference to a collection of processors or computers or memories that may or may not operate in parallel. Instead of using a single processor to perform the steps described herein, some components, such as the steering component and the deceleration component, may each have their own processor that only performs calculations related to component-specific functions.

[0073] In an embodiment of the present application, the processor 113 may obtain data from the camera 130 and other sensor devices and perform vehicle positioning based on the obtained data.

[0074] In various aspects described herein, the processor may be located remote from the vehicle and communicate wirelessly with the vehicle. In other aspects, some of the processes described herein are executed on a processor disposed within the vehicle while others are executed by a remote processor, including taking the necessary steps to perform a single maneuver.

[0075] In some embodiments, the data storage device 114 may contain instructions 115 (e.g., program logic) that may be executed by the processor 113 to perform various functions of the vehicle 100, including those functions described above. The data storage device 114 may also contain additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the propulsion system 102, the sensor system 104, the control system 106, and the peripheral devices 108.

[0076] In addition to the instructions 115, the data storage device 114 may also store data, such as road maps, route information, the position, direction, speed of the vehicle, and other vehicle data, as well as other information. Such information may be used by the vehicle 100 and the computer system 112 when the vehicle 100 is in autonomous, semi-autonomous, and / or manual modes.

[0077] The user interface 116 is used to provide information to or receive information from the user of the vehicle 100. Optionally, the user interface 116 may include one or more input / output devices within the collection of peripheral devices 108, such as the wireless communication system 146, the vehicle-to-vehicle computer 148, the microphone 150, and the speaker 152.

[0078] The computer system 112 can control the functions of the vehicle 100 based on inputs received from various subsystems (e.g., the propulsion system 102, the sensor system 104, and the control system 106) and from the user interface 116. For example, the computer system 112 can utilize the input from the control system 106 to control the steering system 132 to avoid obstacles detected by the sensor system 104 and the obstacle avoidance system 144. In some embodiments, the computer system 112 can operate to provide control over many aspects of the vehicle 100 and its subsystems.

[0079] Optionally, one or more of the above components can be installed or associated separately from the vehicle 100. For example, the data storage device 114 can exist partially or completely separately from the vehicle 100. The above components can be communicatively coupled together in a wired and / or wireless manner.

[0080] Optionally, the above components are just an example. In actual applications, the components in each of the above modules may be added or deleted according to actual needs. Figure 1 It should not be construed as a limitation to the embodiments of the present invention.

[0081] An autonomous vehicle traveling on a road, such as the vehicle 100 above, can identify objects within its surrounding environment to determine adjustments to its current speed. The objects can be other vehicles, traffic control devices, or other types of objects. In some examples, each identified object can be considered independently, and based on the respective characteristics of the object, such as its current speed, acceleration, distance from the vehicle, etc., can be used to determine the speed adjustment required for the autonomous vehicle.

[0082] Optionally, the autonomous vehicle 100 or a computing device associated with the autonomous vehicle 100 (such as Figure 1 the computer system 112, the computer vision system 140, the data storage device 114) can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, the behavior of each identified object depends on the behavior of the others, so all the identified objects can also be considered together to predict the behavior of a single identified object. The vehicle 100 can adjust its speed based on the predicted behavior of the identified objects. In other words, the autonomous vehicle can determine what steady state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the objects. In this process, other factors can also be considered to determine the speed of the vehicle 100, such as the lateral position of the vehicle 100 on the road being traveled, the curvature of the road, the proximity of static and dynamic objects, etc.

[0083] In addition to providing instructions to adjust the speed of an autonomous vehicle, the computing device can also provide instructions to modify the steering angle of vehicle 100 so that the autonomous vehicle follows a given trajectory and / or maintains a safe lateral and longitudinal distance from an object near the autonomous vehicle (e.g., a sedan in an adjacent lane on the road).

[0084] The above-mentioned vehicle 100 can be a sedan, a truck, a motorcycle, a bus, a ship, an airplane, a helicopter, a lawn mower, a recreational vehicle, a playground vehicle, construction equipment, a tram, a golf cart, a train, a trolley, etc., and the embodiments of the present invention are not particularly limited thereto.

[0085] Scenario Example 1: Autonomous Driving System

[0086] According to Figure 2 , computer system 101 includes a processor 103, and processor 103 is coupled to a system bus 105. Processor 103 can be one or more processors, and each processor can include one or more processor cores. A display adapter 107, which can drive a display 109, and display 109 is coupled to system bus 105. System bus 105 is coupled to an input / output (I / O) bus through a bus bridge 111. An I / O interface 1115 is coupled to the I / O bus. I / O interface 1115 communicates with a variety of I / O devices, such as an input device 117 (e.g., a keyboard, a mouse, a touch screen, etc.), a media tray 1121 (e.g., a compact disc read-only memory (CD-ROM), a multimedia interface, etc.). A transceiver 123 (which can send and / or receive radio communication signals), a camera 155 (which can capture static and dynamic digital video images), and an external USB port 125. Optionally, the interface connected to I / O interface 1115 can be a USB interface.

[0087] Among them, processor 103 can be any conventional processor, including a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, or a combination of the above. Optionally, the processor can be a dedicated device such as an ASIC. Optionally, processor 103 can be a neural network processor or a combination of a neural network processor and the above conventional processors.

[0088] Optionally, in various embodiments herein, the computer system 101 may be located remotely from the autonomous vehicle and may communicate wirelessly with the autonomous vehicle. In other aspects, some of the processes herein are executed on a processor disposed within the autonomous vehicle, and others are executed by a remote processor, including taking actions required to perform a single maneuver.

[0089] The computer 101 may communicate with a software deployment server 149 via a network interface 129. Exemplarily, the network interface 129 is a hardware network interface, such as a network card. The network 127 may be an external network, such as the Internet, or an internal network, such as Ethernet or a virtual private network (VPN). Optionally, the network 127 may also be a wireless network, such as a WiFi network, a cellular network, etc.

[0090] The hard disk drive interface is coupled to the system bus 105. The hardware drive interface is connected to the hard disk drive. The system memory 135 is coupled to the system bus 105. Data running in the system memory 135 may include the operating system 137 and application programs 143 of the computer 101.

[0091] The operating system includes a Shell 139 and a kernel 141. The Shell 139 is an interface between the user and the kernel of the operating system. The shell is the outermost layer of the operating system. The shell manages the interaction between the user and the operating system: waits for the user's input, interprets the user's input to the operating system, and processes various output results of the operating system.

[0092] The kernel 141 consists of those parts of the operating system that manage memory, files, peripherals, and system resources. The kernel 141 directly interacts with the hardware. The operating system kernel typically runs processes and provides inter-process communication, provides CPU time slice management, interrupts, memory management, and IO management, etc.

[0093] The application program 143 includes programs related to controlling the autonomous driving of the vehicle, such as programs for managing the interaction between the autonomous vehicle and road obstacles, programs for controlling the route or speed of the autonomous vehicle, and programs for controlling the interaction between the autonomous vehicle and other autonomous vehicles on the road. The application program 143 also exists on the system of the deploying server 149. In one embodiment, when the application program 143 needs to be executed, the computer system 101 may download the application program 143 from the software deployment server 149.

[0094] Sensor 153 is associated with computer system 101. Sensor 153 is used to detect the environment around computer system 101. For example, sensor 153 can detect animals, cars, obstacles and crosswalks, etc. Further, sensor 153 can also detect the environment around the above-mentioned objects such as animals, cars, obstacles and crosswalks, such as: the environment around the animals, for example, other animals around the animals, weather conditions, brightness of the surrounding environment, etc. Optionally, if computer 101 is located in an autonomous vehicle, the sensor can be a radar system, etc.

[0095] Common positioning solutions include pure vision-based positioning technology, namely visual SLAM. The main idea of this solution is to solve the posture of the mobile body based on visual feature point matching and global optimization. First, features are extracted from the image, and the relative posture transformation between the two frames is calculated using the matching feature points between the two frames. Finally, this information is used to calculate the odometer information.

[0096] In addition to visual SLAM, the positioning method based on the fusion of visual and inertial measurement unit (IMU) data is also widely used. The main idea of this solution is to fuse IMU with visual data to achieve mobile positioning. However, in a dynamic environment, the positioning method based on the fusion of visual and inertial measurement unit (IMU) data cannot provide reliable observation values due to visual constraints, resulting in large errors, thus affecting the positioning accuracy.

[0097] To solve the above problems, refer to Figure 3 , Figure 3 A flow chart of a method for determining a posture provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the posture determination method provided in the embodiment of the present application includes:

[0098] 301. Acquire a first image frame and a second image frame taken by a camera device, where the first image frame and the second image frame are adjacent image frames taken by the camera device, and the first image frame and the second image frame respectively include a dynamic target.

[0099] In an embodiment of the present application, when the vehicle is in an environment without GPS or with unstable GPS, such as a tunnel, an underground parking lot, a high-rise building, or a place with many obstructions, the first image frame and the second image frame taken by the camera device can be obtained, wherein the first image frame and the second image frame can be two consecutive image frames taken by the camera device.

[0100] In the embodiments of the present application, the first image frame and the second image frame each include a dynamic target, where a dynamic target is a target that has a displacement relative to the ground when the imaging device captures the image frame. The so-called displacement relative to the ground when the imaging device captures the image frame means that the dynamic target has moved relative to the ground in space. For example, it has moved from position A to position B (position A and position B are two different positions). In one implementation, the dynamic target can be a vehicle, such as a sedan, a truck, a bus, a trailer, an incomplete vehicle, a motorcycle, and so on.

[0101] It should be understood that the first image frame and the second image frame may include the same dynamic target, and this dynamic target is located in different regions in the first image frame and the second image frame. The first image frame and the second image frame may also include different dynamic targets. For example, the first image frame may include vehicle 1 and vehicle 2, and the second image frame may include vehicle 1 and vehicle 2. The position of vehicle 1 in the first image frame is different from the position of vehicle 1 in the second image frame, and the position of vehicle 2 in the first image frame is different from the position of vehicle 2 in the second image frame. Another example is that the first image frame may include vehicle 1 and vehicle 2, and the second image frame may include vehicle 1 and vehicle 3. Vehicle 3 is not included in the second image frame, vehicle 2 is not included in the first image frame, and the position of vehicle 1 in the first image frame is different from the position of vehicle 1 in the second image frame.

[0102] 302. By respectively performing corner detection on the first image frame and the second image frame, a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame are obtained.

[0103] In the embodiments of the present application, after the first image frame and the second image frame captured by the imaging device are obtained, corner detection may be further performed on the first image frame and the second image frame to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame.

[0104] In one implementation, step 302 may be executed by the processor of the vehicle itself. That is to say, the processor of the vehicle itself may perform corner detection on the first image frame and the second image frame to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame. In another implementation, step 302 may be executed by a server on the cloud side. That is to say, the vehicle may send the first image frame and the second image frame captured by the imaging device to the server on the cloud side, and the server may perform corner detection on the first image frame and the second image frame to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame.

[0105] In some embodiments, the multiple first corner points and the multiple second corner points may be corner points such as Features from accelerated segment test (FAST) corner points, Harris corner points, or Binary Robust Invariant Scalable Keypoints (BRISK) corner points. However, the embodiments of the present invention do not limit this.

[0106] In some cases, it can be understood that the human eye's recognition of an image is usually completed in a local small area or small window. If the grayscale of the image in the window changes significantly when the small window is moved a small range in all directions, then it can be considered that there are corner points in the window. If the grayscale of the image in the window does not change when the small window is moved a small range in all directions, then it can be considered that there are no corner points in the window.

[0107] Regarding how to obtain the multiple first corner points of the first image frame and the multiple second corner points of the second image frame by performing corner detection on the first image frame and the second image frame respectively, it can be based on existing implementations and will not be elaborated here.

[0108] 303. Remove the first corner points in the areas of the dynamic objects included in the first image frame among the multiple first corner points to obtain the multiple first corner points after removal; remove the second corner points in the areas of the dynamic objects included in the second image frame among the multiple second corner points to obtain the multiple second corner points after removal.

[0109] It should be understood that the first corner points in the areas of the dynamic objects included in the first image frame among the multiple first corner points can be removed first, and then the second corner points in the areas of the dynamic objects included in the second image frame among the multiple second corner points can be removed; or, the second corner points in the areas of the dynamic objects included in the second image frame among the multiple second corner points can be removed first, and then the first corner points in the areas of the dynamic objects included in the first image frame among the multiple first corner points can be removed; or the operations of removing the first corner points in the areas of the dynamic objects included in the first image frame among the multiple first corner points and removing the second corner points in the areas of the dynamic objects included in the second image frame among the multiple second corner points can be performed simultaneously.

[0110] In one implementation, the dynamic objects in the first image frame can be detected through a pre-trained neural network, and the dynamic objects in the second image frame can be detected to obtain the regions where the dynamic objects included in the first image frame are located, and the regions where the dynamic objects included in the second image frame are located. Then, the first corner points among the multiple first corner points that are located in the regions of the dynamic objects detected in the first image frame are removed to obtain the removed first corner points, and the second corner points among the multiple second corner points that are located in the regions of the dynamic objects detected in the second image frame are removed to obtain the removed second corner points.

[0111] In one implementation, the optical flow method can also be used to track the corner points and supplement new corner points. Then, a unique identification ID is given to each corner point, and the normalized coordinates, pixel coordinates, pixel velocities, etc. of the corner points in the camera coordinate system are calculated.

[0112] When there are dynamic corner points (i.e., the corner points in the regions of dynamic objects), the observed values of the same corner point in different camera states include not only the parallax introduced by the camera's own movement but also the parallax brought by the movement of the feature point itself. In the embodiments of the present application, by removing the corner points in the regions of the dynamic objects among the multiple corner points, the reliability of the visual reprojection error term can be ensured. Specifically, the visual reprojection optimization term can refer to the following formula:

[0113]

[0114] Wherein, is the coordinate of the l-th landmark point observed in the normalized camera coordinate system of the j-th camera, and π c is the camera internal parameter. In order to obtain the movement information of the camera itself, when there are dynamic corner points, the observed values of the same corner point in different camera states include not only the parallax introduced by the camera's own movement but also the parallax brought by the movement of the corner point itself. To avoid this situation, in this embodiment, the dynamic corner points are removed during image processing to improve the accuracy of the visual residual term.

[0115] Exemplarily, it can be referred to Figure 4 , when the imaging device is a front-view monocular camera, feature extraction, tracking, and dynamic object removal can be performed on the image frames captured by the camera. Then, the visual residual affected by the dynamic feature points (which can also be called dynamic corner points) is removed by using the monocular cross-matrix structure SFM.

[0116] 304. Determine the pose change of the imaging device when shooting the second image frame relative to when shooting the first image frame according to the removed multiple first corner points and the removed multiple second corner points.

[0117] In an embodiment of the present application, after removing the first corner points in the area of the moving target included in the first image frame from the multiple first corner points to obtain the multiple first corner points after removal, and removing the second corner points in the area of the moving target included in the second image frame from the multiple second corner points to obtain the multiple second corner points after removal, the pose change of the imaging device when shooting the second image frame relative to when shooting the first image frame can be determined according to the multiple first corner points after removal and the multiple second corner points after removal.

[0118] The embodiment of the present application can be applied to a target vehicle. The target vehicle is fixedly provided with the imaging device, an inertial measurement unit (IMU), and a wheel speed meter. It can also obtain the motion state data of the target vehicle during the period from shooting the first image frame to shooting the second image frame measured by the IMU, and obtain the wheel speed data of the target vehicle during the period from shooting the first image frame to shooting the second image frame measured by the wheel speed meter. Furthermore, the pose change of the imaging device when shooting the second image relative to when shooting the first image can be determined according to the motion state data, the wheel speed data, the multiple first corner points after removal, and the multiple second corner points after removal. Among them, the motion state data can be the vehicle acceleration data and angular velocity data measured by an inertial sensor (IMU), and the wheel speed data can be the vehicle wheel rotation speed, steering wheel angle data, etc.

[0119] Specifically, in one implementation, the first pose change of the imaging device when shooting the second image frame relative to when shooting the first image frame can be determined according to the parallax of the multiple first corner points after removal and the multiple second corner points after removal in their respective image frames, where the first pose change does not include scale information; according to the motion state data, the wheel speed data, and the first pose change, the second pose change of the imaging device when shooting the second image frame relative to when shooting the first image frame is determined, and the second pose change includes scale information; the second pose change is non-linearly optimized to obtain the pose change. Among them, during the non-linear optimization process, the second pose change can be non-linearly optimized according to a preset optimization function, where the preset optimization function includes a wheel speed meter residual term.

[0120] Nonlinear optimization refers to finding an optimal set of numerical mappings in a given objective function, i.e., x - min f(x). According to the derivative theory, the effective values of x can be obtained by solving the derivative equation Δf(x) = 0. When f(x) is a nonlinear function, this optimization is nonlinear optimization. In the embodiments of the present application, the optimization function can be composed of five terms, and can be specifically expressed as: Optimization function = prior residual + IMU residual term + wheel speedometer residual term + visual reprojection error term + loop closure detection reprojection error term.

[0121] In the embodiments of the present application, the multiple first corner points after removal and the multiple second corner points after removal can be aligned with the motion state data and wheel speed data in terms of time stamps. Among them, due to different data frequencies, a single-frame image data (corner points after removal) corresponds to multiple motion state data and wheel speed data. Then, the motion state data and wheel speed data of each frame of image can be pre-integrated to provide an initial pose value for the image. After that, the sliding window method is adopted to perform initialization processing using pure visual data, and the pose and the position without scale information (which can also be called depth) when the camera device captures all the image frames within the sliding window are obtained. Combining the motion state data and wheel speed data to restore the above-mentioned scale information without scale, and recalculating the positions of all corner points, and then performing nonlinear optimization on the restored scale information and the inverse depth of the corner points to obtain the pose change.

[0122] Specifically, the following formula can be referred to:

[0123]

[0124] Among them, is the position of the k-th frame in the world coordinate system, is the angle of the k-th frame in the world coordinate system, is the velocity of the k-th frame in the world coordinate system, is the pre-integrated position between two frames, is the pre-integrated angle between two frames, is the pre-integrated velocity between two frames, is the acceleration bias of the k-th frame, is the bias of the gyroscope of the k-th frame, is the quaternion representation of the attitude of the k-th frame in the world coordinate system.

[0125] As can be seen from the above formula, the estimated values can include the accelerometer bias b a and the gyroscope bias b ω . Since in two-dimensional motion, the excitation is insufficient and the acceleration bias is difficult to estimate, which will cause the bias value to converge too slowly and lead to inaccurate pose estimation. Therefore, in the embodiments of the present application, a wheel speedometer residual term is introduced, and it can be specifically shown as the following formula:

[0126]

[0127] Among them, is the angle of the k-th frame in the world coordinate system, is the position of the k-th frame in the world coordinate system, is the pre-integrated position between two frames, is the pre-integrated angle between two frames, is the quaternion representation of the pose of the k-th frame in the world coordinate system.

[0128] Exemplarily, referring to Figure 4 , the wheel speed data measured by the wheel speedometer (including vehicle rotation speed data, steering wheel data, etc.) can be pre-integrated to obtain a pre-integration result, and then a pre-integration residual is established based on the pre-integration result.

[0129] In the embodiments of the present application, in view of the disadvantage of large initialization error of the two-dimensional motion of the visual inertial navigation system, a wheel speedometer residual term is added. The pose is initialized and calculated by jointly using the information of vision, IMU, and wheel speedometer. When the IMU convergence is not good, the wheel speedometer can be used for compensation, thereby improving the effect.

[0130] In one implementation, the current key frame can also be subjected to loop closure detection with the map. The detected frame is the loop closure frame, and the co-visible feature points between the loop closure frame and the frames in the sliding window are found; a reprojection error term is established and added to the non-linear optimization; the key frame that has been optimized and slid out of the window and its associated frames are optimized with four degrees of freedom; the optimized key frame is inserted into the map; in actual use, the map, that is, the loop closure detection, can be selected to be turned on or off. When it is turned on, mapping is performed while localizing, and the finally obtained positions are all relative to the position of the first fixed camera. Of course, this can be fixed to a specific reference system by performing a coordinate transformation later.

[0131] An embodiment of the present application provides a pose determination method, and the method includes: obtaining a first image frame and a second image frame captured by an imaging device, where the first image frame and the second image frame are adjacent image frames captured by the imaging device, and the first image frame and the second image frame respectively include dynamic targets; performing corner detection on the first image frame and the second image frame respectively to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame; removing the first corner points in the plurality of first corner points that are in the region of the dynamic target included in the first image frame to obtain the plurality of first corner points after removal; removing the second corner points in the plurality of second corner points that are in the region of the dynamic target included in the second image frame to obtain the plurality of second corner points after removal; determining the pose change of the imaging device when capturing the second image frame relative to when capturing the first image frame according to the plurality of first corner points after removal and the plurality of second corner points after removal. When there are dynamic corner points, the observations of the same corner point in different camera states include not only the parallax introduced by the movement of the camera itself, but also the parallax brought by the movement of the corner point itself. Through the above method, in view of the problem of large positioning error in a dynamic environment, the corner points in the region of the dynamic target are removed from the plurality of corner points to ensure that the visual reprojection error is more reliable, thereby overcoming the pose change determination error introduced by dynamic features.

[0132] Refer to Figure 5 , Figure 5 which is a structural schematic of a pose determination device provided by an embodiment of the present application. As Figure 5 shown, the pose determination device 500 provided by an embodiment of the present application includes:

[0133] An acquisition module 501, configured to obtain a first image frame and a second image frame captured by an imaging device, where the first image frame and the second image frame are adjacent image frames captured by the imaging device, and the first image frame and the second image frame respectively include dynamic targets;

[0134] A corner extraction module 502, configured to perform corner detection on the first image frame and the second image frame respectively to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame;

[0135] A removal module 503, configured to remove the first corner points in the plurality of first corner points that are in the region of the dynamic target included in the first image frame to obtain the plurality of first corner points after removal;

[0136] removing the second corner points in the plurality of second corner points that are in the region of the dynamic target included in the second image frame to obtain the plurality of second corner points after removal;

[0137] A positioning module 504, configured to determine a pose change of the imaging device when capturing the second image frame relative to when capturing the first image frame according to the multiple first corner points after culling and the multiple second corner points after culling.

[0138] In a possible implementation, the multiple first corner points and the multiple second corner points include one of the following: accelerated segment test feature (FAST) corner points, Harris corner points, and binary robust invariant scalable keypoints (BRISK) corner points.

[0139] In a possible implementation,

[0140] The device is applied to a target vehicle, and the target vehicle is fixedly provided with the imaging device, an inertial measurement unit (IMU), and a wheel speed meter; the acquisition module is configured to acquire motion state data of the target vehicle during the period from capturing the first image frame to capturing the second image frame measured by the IMU;

[0141] Acquire wheel speed data of the target vehicle during the period from capturing the first image frame to capturing the second image frame measured by the wheel speed meter;

[0142] Correspondingly, the positioning module is configured to determine a first pose change of the imaging device when capturing the second image frame relative to when capturing the first image frame according to the parallax of the multiple first corner points after culling and the multiple second corner points after culling in their respective image frames, where the first pose change does not include scale information;

[0143] Determine a second pose change of the imaging device when capturing the second image frame relative to when capturing the first image frame according to the motion state data, the wheel speed data, and the first pose change, where the second pose change includes scale information;

[0144] Perform non-linear optimization on the second pose change to obtain the pose change.

[0145] In a possible implementation, the positioning module is configured to perform non-linear optimization on the second pose change according to a preset optimization function, where the preset optimization function includes a wheel speed meter residual term.

[0146] In a possible implementation, the device further includes:

[0147] A dynamic target detection module, configured to detect dynamic targets in the first image frame and dynamic targets in the second image frame through a pre-trained neural network, so as to obtain the regions where the dynamic targets included in the first image frame are located, and the regions where the dynamic targets included in the second image frame are located.

[0148] Based on the same concept, referring to Figure 6 As shown, an embodiment of the present application provides a pose determination device 600, including a transceiver 610, a processor 620, and a memory 630; the memory 630 is used to store programs, instructions, or codes; the processor 620 is used to execute the programs, instructions, or codes in the memory 630;

[0149] The transceiver 610 is configured to receive a first image frame and a second image frame input by a camera device;;

[0150] The processor 620 is configured to perform corner detection on the first image frame and the second image frame respectively to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame; removing the first corner points in the plurality of first corner points that are in the region where the dynamic target included in the first image frame is located to obtain the plurality of first corner points after removal; removing the second corner points in the plurality of second corner points that are in the region where the dynamic target included in the second image frame is located to obtain the plurality of second corner points after removal; determining the pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the plurality of first corner points after removal and the plurality of second corner points after removal.

[0151] Wherein, the processor 620 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 620 or the instructions in software form. The above-mentioned processor 602 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 630, and the processor 620 reads the information in the memory 630 and combines its hardware to execute the above method steps.

[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0153] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0156] Obviously, those skilled in the art can make various modifications and variations to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. A pose determination method, characterized in that, The method includes: Obtaining a first image frame and a second image frame captured by a camera device, where the first image frame and the second image frame are adjacent image frames captured by the camera device, the first image frame and the second image frame each include a dynamic target, and the camera device is a forward-looking monocular camera; Performing corner detection on the first image frame and the second image frame respectively to obtain a plurality of first corner points of the first image frame and a plurality of second corner points of the second image frame; Removing the first corner points in the area where the dynamic target included in the first image frame is located among the plurality of first corner points to obtain the plurality of first corner points after removal; Removing the second corner points in the area where the dynamic target included in the second image frame is located among the plurality of second corner points to obtain the plurality of second corner points after removal; Determining the pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the plurality of first corner points after removal and the plurality of second corner points after removal.

2. The method according to claim 1, wherein The plurality of first corner points and the plurality of second corner points include one of the following: Fast Angle Features for Segment Testing (FAST) corner points, Harris corner points, and Binary Robust Invariant Scalable Keypoints (BRISK) corner points.

3. The method according to claim 1, wherein The method is applied to a target vehicle, and the target vehicle is fixedly provided with the camera device, an Inertial Measurement Unit (IMU), and a wheel speed meter; the method further includes: Obtaining the motion state data of the target vehicle during the period from shooting the first image frame to shooting the second image frame measured by the IMU; Obtaining the wheel speed data of the target vehicle during the period from shooting the first image frame to shooting the second image frame measured by the wheel speed meter; Correspondingly, the determining the pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the plurality of first corner points after removal and the plurality of second corner points after removal includes: Determining a first pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the disparity of the plurality of first corner points after removal and the plurality of second corner points after removal in their respective image frames, where the first pose change does not include scale information; Determining a second pose change of the camera device when shooting the second image frame relative to when shooting the first image frame according to the motion state data, the wheel speed data, and the first pose change, where the second pose change includes scale information; Performing non-linear optimization on the second pose change to obtain the pose change.

4. The method according to claim 3, wherein The performing non-linear optimization on the second pose change includes: Performing non-linear optimization on the second pose change according to a preset optimization function, where the preset optimization function includes a wheel speed meter residual term.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Detect dynamic objects in the first image frame and the second image frame through a pre-trained neural network to obtain the regions where the dynamic objects included in the first image frame are located and the regions where the dynamic objects included in the second image frame are located.

6. The method according to any one of claims 1 to 4, characterized in that The dynamic object is a vehicle.

7. A pose determination device, characterized in that, The device includes: An acquisition module, configured to acquire a first image frame and a second image frame captured by a camera device, where the first image frame and the second image frame are adjacent image frames captured by the camera device, the first image frame and the second image frame respectively include dynamic objects, and the camera device is a front-view monocular camera; A corner extraction module, configured to obtain a plurality of first corners of the first image frame and a plurality of second corners of the second image frame by performing corner detection on the first image frame and the second image frame respectively; An elimination module, configured to eliminate the first corners in the plurality of first corners that are in the region where the dynamic object included in the first image frame is located to obtain the eliminated plurality of first corners; Eliminate the second corners in the plurality of second corners that are in the region where the dynamic object included in the second image frame is located to obtain the eliminated plurality of second corners; A positioning module, configured to determine the pose change of the camera device when capturing the second image frame relative to when capturing the first image frame according to the eliminated plurality of first corners and the eliminated plurality of second corners.

8. The device according to claim 7, characterized in that, The plurality of first corners and the plurality of second corners include one of the following: accelerated segment test feature (FAST) corners, Harris corners, and binary robust invariant scalable keypoints (BRISK) corners.

9. The device according to claim 7, characterized in that, The device is applied to a target vehicle, and the target vehicle is fixedly provided with the camera device, an inertial measurement unit (IMU), and a wheel speed meter; the acquisition module is configured to acquire the motion state data of the target vehicle during the period from capturing the first image frame to capturing the second image frame measured by the IMU; Acquire the wheel speed data of the target vehicle during the period from capturing the first image frame to capturing the second image frame measured by the wheel speed meter; Correspondingly, the positioning module is configured to determine the first pose change of the camera device when capturing the second image frame relative to when capturing the first image frame according to the disparity of the eliminated plurality of first corners and the eliminated plurality of second corners in their respective image frames, where the first pose change does not include scale information; Determine the second pose change of the camera device when capturing the second image frame relative to when capturing the first image frame according to the motion state data, the wheel speed data, and the first pose change, where the second pose change includes scale information; Perform non-linear optimization on the second pose change to obtain the pose change.

10. The device according to claim 9, characterized in that, The positioning module is configured to perform non-linear optimization on the second pose change according to a preset optimization function, where the preset optimization function includes a wheel speed meter residual term.

11. The device according to any one of claims 7 to 10, characterized in that, The device further includes: A dynamic target detection module, configured to detect dynamic targets in the first image frame and dynamic targets in the second image frame through a pre-trained neural network, so as to obtain regions where the dynamic targets included in the first image frame are located and regions where the dynamic targets included in the second image frame are located.

12. The device according to any one of claims 7 to 10, characterized in that, The dynamic target is a vehicle.

13. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium contains computer instructions for performing the pose determination method according to any one of claims 1 to 6.

14. An operation device, characterized in that, The computing device includes a memory and a processor, and code is stored in the memory, and the processor is configured to obtain the code to perform the pose determination method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Synchronous positioning and map-constructing method for mobile robot facing indoor dynamic environment

    CN109387204A

  • Mobile robot positioning and mapping method

    CN111795686A