Visual obstacle avoidance method and device

By placing a monocular camera on a mobile carrier, collecting front and bottom images and combining feature point fitting and semantic segmentation algorithms, the problem of monocular vision lacking depth of field information is solved, the distance and size of obstacles can be accurately calculated, and the accuracy and application scope of obstacle avoidance are improved.

CN120823573APending Publication Date: 2025-10-21BEIJING EYESTAR TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410434912.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Monocular vision solutions lack depth of field information when avoiding obstacles, resulting in the inability to accurately know the distance and size of the target, limiting their application in obstacle avoidance for vehicles and robots.

Method used

A monocular camera is placed on a mobile carrier to capture images of a limited range in the front and below. A quadratic polynomial surface fitting function is established based on the feature points in the image. Combined with semantic segmentation and Kalman filtering algorithms, the bounding box of obstacles is updated in real time, and the distance and size of the obstacle to the carrier are calculated.

Benefits of technology

It achieves accurate judgment of the distance and size of obstacles in complex environments, improves the accuracy and reliability of obstacle avoidance, and expands the application scope of monocular vision in obstacle avoidance of mobile robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004786898990000031
    Figure BDA0004786898990000031
  • Figure BDA0004786898990000061
    Figure BDA0004786898990000061
  • Figure BDA0004786898990000081
    Figure BDA0004786898990000081
Patent Text Reader

Abstract

The invention relates to a visual obstacle avoidance method and device, and the method comprises the steps: S1, arranging a monocular camera on a moving carrier, and enabling an image collected by the monocular camera to comprise a preset range at the front lower part of the moving carrier; s2, collecting a sample image through the monocular camera and calibrating the sample image to obtain a relational expression of the distance between a target in the image and the mobile carrier; s3, acquiring a formal image through the monocular camera, and identifying an obstacle in the formal image; and S4, determining the distance between the obstacle and the mobile carrier through the relational expression. According to the invention, the monocular camera is arranged on the mobile carrier and only collects the image in the limited range of the front lower part, and the distance information with the target obstacle can be obtained from the collected image, so that the obstacle avoidance judgment is realized, and compared with the existing monocular vision that the depth-of-field information cannot be obtained, so that the obstacle distance cannot be accurately judged, and the obstacle avoidance accuracy is improved. Good application prospects are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and artificial intelligence technology, and in particular to a visual obstacle avoidance method and device. Background Art

[0002] Obstacle avoidance refers to the process in which vehicles, mobile robots, etc., when sensing the presence of static or dynamic obstacles on their planned driving paths through sensors during movement, update the paths in real time according to certain algorithms, thereby bypassing the obstacles and ultimately reaching the target point. It is one of the core functions of autonomous driving.

[0003] Vision is a commonly used obstacle avoidance method and technology. Common computer vision solutions include binocular vision, monocular vision, time-of-flight (ToF)-based depth cameras, structured light-based depth cameras, etc.

[0004] Structured light-based depth cameras utilize an infrared light source. The emitted light is then encoded and projected onto an object. When these patterns are reflected from the object's surface, they deform according to the distance from the object. The image sensor then records the deformed patterns. The sensor then calculates the shape of each pixel in the captured pattern to obtain the corresponding parallax, and thus the depth value. However, the ranging range of structured light solutions is affected by the light spot pattern, resulting in a smaller obstacle avoidance range. Furthermore, the system performs poorly in bright light environments and is easily affected by the lighting.

[0005] The ToF-based depth camera is used by iPad Pro to realize augmented reality gameplay. Its working principle is that an infrared light source emits high-frequency light pulses to an object, then receives the light pulses reflected from the object, and calculates the distance from the object to the camera by detecting the round-trip flight time of the light pulses. The depth camera can obtain RGB images and depth maps at the same time, but whether it is based on ToF or structured light, the effect is not ideal in strong outdoor light environments, so they all need to actively emit light. In addition, the resolution of ToF structured light is relatively low, making it difficult for image information to assist in obstacle avoidance, so it is not suitable for use in fields such as sweeping robots.

[0006] Binocular vision is essentially based on triangulation. Because the two cameras are positioned at different positions and form a baseline of fixed length, there will be different pixel positions when imaging. Therefore, it can better capture and restore the depth of field information of objects and has a clear spatial scale, just like the two eyes of a human.

[0007] Compared to depth cameras, structured light cameras, and binocular vision, monocular vision uses deep learning algorithms to identify objects and then detect their bounding boxes for distance and 3D estimation. These solutions offer advantages such as high resolution, low cost, easy development, and simple deployment, making them widely used in various mobile robotics applications. However, the biggest challenge with monocular vision is that most single photos only provide two-dimensional information, much like a 2D movie, without a direct sense of space. Therefore, we rely on our own experience, such as the notion of objects being obscured and objects appearing larger when closer appearing smaller. Therefore, the information captured by a single camera is extremely limited and cannot achieve the desired spatial effect. Similarly, in machine vision technology, a single camera image cannot capture the distance relationship between each object in the scene and the lens, lacking the third dimension: depth of field. Given the extreme complexity of real-life scenes, the lack of depth of field information from a single camera leads to a high probability of visual miscalculation, potentially miscalculating the actual distance of an object.

[0008] In other words, mobile robots equipped with a single camera (such as sweepers and lawn mowers) recognize all objects as two-dimensional, not three-dimensional. They can only estimate pre-trained objects using a model. However, actual operations and working environments vary greatly, making it impossible to train for all objects that could cause the robot to become stuck. This bottleneck severely limits the application of monocular vision solutions for obstacle avoidance in mobile vehicles, robots, and other vehicles. Summary of the Invention

[0009] Monocular vision recognition, a commonly used computer vision recognition solution, suffers from a major problem when applied to obstacle avoidance: a lack of depth information. Therefore, it is impossible to accurately determine the distance and size of a target through monocular recognition. In light of this, the present invention proposes a visual obstacle avoidance method and device.

[0010] To address the aforementioned issues with existing technologies, the inventors noted that existing monocular cameras focus on an infinitely forward direction. However, if images are captured only within a limited area below and in front of the vehicle, the camera's shooting distance can be determined through actual measurement, further confirming the distance corresponding to each pixel in the image. Consequently, if an obstacle on the road ahead enters the camera's shooting range as the vehicle travels, its distance relative to the vehicle and its own dimensions can be calculated based on the obstacle's bounding box.

[0011] To achieve the above objectives, an embodiment of the present invention provides a visual obstacle avoidance method, including:

[0012] S1, arranging a monocular camera on a mobile carrier, wherein an image captured by the monocular camera includes a predetermined range below and in front of the mobile carrier;

[0013] S2, collecting sample images by the monocular camera and calibrating them to obtain a relationship between the distance between the target in the image and the mobile carrier;

[0014] S3, collecting a formal image through the monocular camera and identifying obstacles in the formal image;

[0015] S4: Determine the distance between the obstacle and the mobile carrier using the relationship in S2.

[0016] Preferably, in S1, the method of making the image captured by the monocular camera include a predetermined range below and in front of the mobile carrier includes:

[0017] The monocular camera is arranged toward the front and lower side of the mobile carrier; or

[0018] The image captured by the monocular camera is cropped.

[0019] Preferably, the S2 specifically includes:

[0020] S21, dividing a number of feature points in the sample image and obtaining pixel coordinates of the feature points;

[0021] S22, determining a plurality of real position points corresponding to the feature points, and actually measuring forward and lateral distances between the real position points and the mobile carrier;

[0022] S23, establishing a quadratic polynomial surface fitting function relationship among the pixel coordinates, the forward distance, and the lateral distance.

[0023] Preferably, the quadratic polynomial surface fitting function in S23 is:

[0024]

[0025] Among them, (a0, a1, a2, a3, a4, a5) and (b0, b1, b2, b3, b4, b5) are fitting coefficients.

[0026] Preferably, in S3, identifying obstacles in the formal image specifically includes:

[0027] The target position is represented by a bounding box. The bounding boxes of the previous and next frames are matched and the bounding box of the same target is predicted and updated in real time using Kalman filtering to obtain the vertex coordinates of the target's bounding box.

[0028] Preferably, the S4 specifically includes:

[0029] Substitute the vertex coordinates of the bounding box into the relationship in S2 to obtain the distance between the target and the mobile carrier.

[0030] On the other hand, an embodiment of the present invention also proposes a device for the above-mentioned visual obstacle avoidance method, comprising: a mobile carrier; a monocular camera, which is arranged on the mobile carrier, and the image captured by the monocular camera includes a predetermined range in front and below the mobile carrier.

[0031] The visual obstacle avoidance method and device of the present invention, by arranging a monocular camera on a mobile carrier, and the monocular camera only captures images within a limited range in front and below, can obtain distance information to the target obstacle from the captured image, thereby realizing obstacle avoidance judgment. Compared with the existing monocular vision that cannot obtain depth of field information and therefore cannot accurately judge the distance to the obstacle, it has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 Schematic diagram of the flow of the visual obstacle avoidance method according to an embodiment of the present invention;

[0034] Figure 2 This is a schematic diagram of the shooting effect of a monocular camera according to an embodiment of the present invention;

[0035] Figure 3 A schematic diagram of selecting image grid point vertices as sampling feature points according to an embodiment of the present invention;

[0036] Figure 4 Schematic diagram of the forward distance and lateral distance involved in the embodiment of the present invention;

[0037] Figure 5 is a schematic diagram of vertices of a bounding box involved in an embodiment of the present invention;

[0038] Figure 6 and Figure 7 Schematic diagram of visual recognition using the visual obstacle avoidance method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The description of the embodiments in this specification should be combined with the corresponding drawings, which should be considered a complete part of this specification. In the drawings, the shapes and thicknesses of the embodiments may be exaggerated and indicated for simplicity or convenience. Furthermore, the various structural components in the drawings will be described separately. It is worth noting that components not shown in the drawings or not described in words are known to those of ordinary skill in the art.

[0040] The description of the embodiments herein and any references to directions and orientations are for ease of description only and are not to be construed as limiting the scope of the present invention. The following description of the preferred embodiments may involve combinations of features, which may exist independently or in combination. The present invention is not specifically limited to the preferred embodiments. The scope of the present invention is defined by the claims.

[0041] like Figure 1 As shown, the visual obstacle avoidance method of the embodiment of the present invention includes the following steps:

[0042] S1, arranging a monocular camera on a mobile carrier, and an image captured by the monocular camera includes a predetermined range below and in front of the mobile carrier.

[0043] In this embodiment, the monocular camera's mounting position and angle on the mobile carrier are adjusted so that it faces the ground at a limited distance below and in front of the carrier, rather than the infinitely distant view. If this cannot be achieved simply by adjusting the monocular camera's mounting position and angle on the carrier, cropping the image captured and output by the camera can also be used.

[0044] Taking an unmanned lawn mower working on the grass as an example, the image captured by the adjusted monocular camera is as follows: Figure 2 As shown in the figure, the entire image range captured by the camera can be measured on the ground to obtain the corresponding real distance. Generally, it is best to adjust the monocular camera so that it can only capture the distance of several meters to tens of meters in front and below the vehicle such as the unmanned lawn mower.

[0045] S2, collects sample images through a monocular camera and calibrates them to obtain the relationship between the target in the image and the mobile carrier.

[0046] Step S2 is the offline calibration stage, which specifically includes:

[0047] S21 , dividing a number of feature points in the sample image, and obtaining pixel coordinates of the feature points.

[0048] In this embodiment, a monocular camera is used to capture a sample image, and n feature points are manually sampled from different distribution positions in the sample image. The pixel coordinates (ui ,v i ).

[0049] Preferably, in order to ensure that the feature points sampled from the sample image are evenly distributed, the sample image can be divided into grids, and then feature points can be sampled from each grid, or the vertices of the grid points can be directly selected as the sampled feature points, such as Figure 3 shown.

[0050] S22, determining a number of real position points corresponding to the feature points, and actually measuring and obtaining the forward distance and the lateral distance between the real position points and the mobile carrier.

[0051] In this embodiment, it is necessary to find the real position point P corresponding to the feature point i on the ground. i , the actual measurement position point P i Forward distance to the mobile carrier fd i and lateral distance sd i .

[0052] Among them, the forward distance refers to the actual position point P on the ground i The vertical distance to the front end of the carrier, the lateral distance is from point P i The vertical distance to the longitudinal axis of the carrier coordinate system, such as Figure 4 shown.

[0053] S23, establishing a quadratic polynomial surface fitting function relationship among the pixel coordinates, the forward distance, and the lateral distance.

[0054] Establish pixel coordinates (u i ,v i ) and forward distance fd i , and the lateral distance sd i The relationship between the quadratic polynomial surface fitting functions is as follows:

[0055]

[0056] Where (a0, a1, a2, a3, a4, a5) and (b0, b1, b2, b3, b4, b5) are fitting coefficients. The values ​​of the fitting coefficients are mainly related to the camera's field of view, the obstacle recognition distance and range designed for the device, and the degree of distortion of the camera lens.

[0057] In specific use, the pixel coordinates of a certain point j in the image are substituted into the above formula to calculate the real space point P corresponding to the pixel point. j Forward distance fd relative to the moving carrier j and the lateral distance sd j .

[0058] S3, collecting a formal image through the monocular camera and identifying obstacles in the formal image.

[0059] Obstacle visual recognition can be defined as a semantic segmentation problem. Traditional methods such as Normalized Cut, Random Forest, and Support Vector Machines (SVM) have low recognition rates and robustness. Machine learning approaches significantly surpass these traditional methods in both accuracy and robustness. Semantic segmentation networks can be broadly categorized into two types: those based on convolutional neural networks (CNNs) and those based on self-attention neural network models (Transformers).

[0060] The basic structure of the semantic segmentation network: use the backbone network to extract image features, downsample through convolution and pooling (Encoder), use deconvolution and depooling upsampling (Decoder) and then perform multi-scale feature fusion to finally obtain a pixel-level segmentation map.

[0061] When training the model, cross-entropy is used as the loss to minimize:

[0062]

[0063] Among them, C represents the number of obstacle categories set, p i is the true value, q i is the predicted value.

[0064] Specifically, in this embodiment, during the online phase, an RGB image (width w and height h, respectively) captured by a camera is inferred through a semantic segmentation network to produce a semantic segmentation result of size w*h*8 bits. The value of each pixel represents its belonging to a predefined category. For example, if the value matches the category corresponding to a tree, the image is considered a tree; if the value matches the category corresponding to a pool, the image is considered a pool.

[0065] After detecting and identifying the obstacle, we need to predict the category and position of the object at the same time, so we need to introduce some concepts related to position. Usually a bounding box (BBox) is used to represent the position of the object. The bounding box is a rectangular box that can just contain the object. In the embodiment of the present invention, the bounding box is used to approximate the size and position of the object. The upper left vertex of the bounding box is A, the upper right vertex is B, the lower left vertex is C, and the lower right vertex is D. Figure 5 shown.

[0066] Preferably, since there is a certain degree of probability and contingency in identifying and marking the obstacle Bounding Box using a single frame image, the position of the Bounding Box will constantly jump or appear intermittently, seriously affecting the tracking and position calculation of the obstacle. Therefore, this embodiment uses the Hungarian assignment algorithm to match the Bounding Boxes from the previous and next frames, and uses Kalman filtering to predict and update the Bounding Box information, thereby ensuring stable tracking of the Bounding Box. The specific algorithm flow is as follows:

[0067] (1) Define the state quantity list svList, which consists of the state quantities of each Bounding Box [x, y, s, r, x', y', s'], where (x, y) is the pixel coordinate of the center point of the Bounding Box, (x', y') is the first-order velocity of the pixel coordinate, s is the area of ​​the Bounding Box, s' is the rate of change of the area, and r is the ratio of the width to the height of the Bounding Box;

[0068] (2) When a new picture is received:

[0069] a) Use the Kalman filter to predict each state variable in the state variable list svList and add 1 to the duration count of the state variable;

[0070] b) Call the image instance segmentation function (Instance Segmentation) in the deep learning algorithm model to obtain the Bounding Box list currBboxList of the current frame;

[0071] c) Convert the items in the state list svList into the lower left and lower right vertex coordinates of the Bounding Box [u A ,v A ,u D ,v D ] After being expressed, the intersection over union (IoU) of the area is calculated with the state quantity in the currBboxList list to obtain the IoU, which is used as the cost of the matching algorithm and forms the corresponding cost matrix.

[0072] d) Substitute the cost matrix into the Hungarian Algorithm to obtain the matching state index idx state and the unmatched bounding box index idx bbox;

[0073] e) For the matched state quantity in svList, find its corresponding bbox value, convert it to [x, y, s, r, x', y', s'] format, call the Kalman filter to update (Update), and reset its duration count to 0;

[0074] f) For the unmatched bbox in currBboxList, fill it into the tail of the state list svList and use it for the next frame image;

[0075] g) For each item in svList, if its duration count exceeds the set value, it will be deleted;

[0076] h) Convert each item in svList to [u A ,v A ,u D ,v D ] format, which is the output result of the filtered Bounding Box vertex pixel coordinates.

[0077] S4: Determine the distance between the obstacle and the mobile carrier using the relationship in S2.

[0078] After obtaining the pixel coordinates of each vertex of the obstacle's bounding box, the distance between the obstacle and the mobile carrier can be obtained by substituting them into the above formula (1).

[0079] Preferably, after obtaining the forward distance and lateral distance from each vertex of the Bounding Box to the carrier, the width W and height H of the corresponding obstacle can be further calculated:

[0080]

[0081] like Figure 6 and Figure 7 FIG. 1 is a schematic diagram of obstacle avoidance identification using the method of an embodiment of the present invention. According to the above method of this embodiment, the information of each target in the image can be obtained as follows:

[0082]

[0083] In summary, the method of the present invention, by arranging a monocular camera on a mobile carrier, and the monocular camera only captures images within a limited range in front and below, can obtain distance information to the target obstacle from the captured image, thereby realizing obstacle avoidance judgment. Compared with the existing monocular vision that cannot obtain depth of field information and therefore cannot accurately judge the distance to the obstacle, it has good application prospects.

[0084] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A visual obstacle avoidance method, characterized in that: The method comprises: S1, arranging a monocular camera on a mobile carrier, wherein an image captured by the monocular camera includes a predetermined range below and in front of the mobile carrier; S2, collecting sample images by the monocular camera and calibrating them to obtain a relationship between the distance between the target in the image and the mobile carrier; S3, collecting a formal image through the monocular camera and identifying obstacles in the formal image; S4: Determine the distance between the obstacle and the mobile carrier using the relationship in S2.

2. The visual obstacle avoidance method according to claim 1, characterized in that: In S1, the method for making the image captured by the monocular camera include a predetermined range below and in front of the mobile carrier includes: The monocular camera is arranged toward the front and lower side of the mobile carrier; or The image captured by the monocular camera is cropped.

3. The visual obstacle avoidance method according to claim 1, characterized in that: The S2 specifically includes: S21, dividing a number of feature points in the sample image and obtaining pixel coordinates of the feature points; S22, determining a plurality of real position points corresponding to the feature points, and actually measuring forward and lateral distances between the real position points and the mobile carrier; S23, establishing a quadratic polynomial surface fitting function relationship among the pixel coordinates, the forward distance, and the lateral distance.

4. The visual obstacle avoidance method according to claim 3, characterized in that: The quadratic polynomial surface fitting function in S23 is expressed as follows: Among them, (a0, a1, a2, a3, a4, a5) and (b0, b1, b2, b3, b4, b5) are fitting coefficients.

5. The visual obstacle avoidance method according to claim 1, characterized in that: In S3, identifying obstacles in the formal image specifically includes: The target position is represented by a bounding box. The bounding boxes of the previous and next frames are matched and the bounding box of the same target is predicted and updated in real time using Kalman filtering to obtain the vertex coordinates of the target's bounding box.

6. The visual obstacle avoidance method according to claim 5, characterized in that: The S4 specifically includes: Substitute the vertex coordinates of the bounding box into the relationship in S2 to obtain the distance between the target and the mobile carrier.

7. A visual obstacle avoidance device, characterized in that: include: Mobile carrier; A monocular camera is provided on the mobile carrier, and an image captured by the monocular camera includes a predetermined range in the front and lower part of the mobile carrier.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which implements the method according to any one of claims 1 to 6 when executed by a processor.

Citation Information

Cited By

  • Unmanned aerial vehicle binocular vision obstacle avoidance system and distance measurement and obstacle avoidance method thereof

    CN121163464A

  • Unmanned aerial vehicle binocular vision obstacle avoidance system and ranging and obstacle avoidance method thereof

    CN121163464B