Near-field image acquisition method and device, equipment and storage medium
By converting the camera image and ranging sensor point cloud data into the image coordinate system on the robot, judging and uploading the near-field image, the problem of invalid data acquisition in the data center is solved, and efficient near-field data acquisition is achieved.
Patent Information
- Application Number
- CN202410333669.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, when robots collect data in full, the data center obtains a large amount of invalid data of non-near-field images, which increases the cost of data storage, uploading and manual screening, and affects the efficiency of obtaining near-field data.
By acquiring the robot's camera image and ranging sensor point cloud data, converting them to the camera's image coordinate system, and using two-dimensional points to determine whether the image is a near-field image, only near-field images are uploaded to the data center to avoid uploading invalid data.
It reduces the cost of data storage and uploading, saves the cost of manual screening, and improves the efficiency of acquiring near-field data.
Smart Images

Figure CN120707622A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotics, and in particular to a near-field image acquisition method, apparatus, device, and storage medium. Background Art
[0002] For unmanned equipment such as mobile robots and autonomous vehicles, acquiring large amounts of scene data is beneficial for improving their navigation performance and their ability to resolve anomalies. When robots need to operate in near-field environments, their visual perception systems must be able to accurately perceive obstacles within close range. When a robot is close to an obstacle, its camera's field of view is mostly filled with obstacles, resulting in insufficient scene information. The robot's visual perception system can easily misidentify obstacles as navigable areas, severely impacting its obstacle avoidance capabilities. Therefore, collecting near-field data from the robot and training its visual perception system based on this data is a key technical approach to improving its near-field perception capabilities.
[0003] In related technologies, robots continuously collect all data during their movements through a full-data acquisition method, and then manually filter out near-field data from the acquired data. However, this full-data acquisition method causes the data center to obtain a large amount of invalid data that is not near-field images. This undoubtedly increases the cost of data storage, upload, and manual screening, and affects the efficiency of acquiring near-field data. Summary of the Invention
[0004] The present application provides a near-field image acquisition method, apparatus, device and storage medium, which solves the problem in the prior art that a data center acquires a large amount of invalid data that is not near-field images, reduces the cost of data storage and uploading, saves the cost of manual screening, and is conducive to improving the efficiency of acquiring near-field data.
[0005] In a first aspect, the present application provides a near-field image acquisition method, comprising:
[0006] Acquire a first image captured by a camera of the robot;
[0007] Acquire point cloud data collected by a ranging sensor of the robot, and convert the point cloud data into an image coordinate system of the camera to obtain a plurality of first two-dimensional points;
[0008] determining, based on the plurality of first two-dimensional points, whether the first image is a near-field image of the robot;
[0009] In the case where the first image is a near-field image, the first image is uploaded.
[0010] Through the above-mentioned technical means, the point cloud data collected by the ranging sensor during the movement of the robot can be projected into the image coordinate system of the first image, so as to accurately determine whether the first image collected by the robot in real time is a near-field image based on the first two-dimensional points obtained by the projection. When it is determined that the near-field image has been collected, the near-field image is uploaded to the data center, avoiding uploading invalid data other than the near-field image to the data center. This solves the problem of the data center obtaining a large amount of invalid data other than the near-field image in the prior art, reduces the cost of data storage and uploading, and saves the cost of manual screening. The data center can directly obtain the near-field image without the need for manual screening of the near-field image, which is conducive to improving the efficiency of obtaining near-field data.
[0011] Optionally, converting the point cloud data into the image coordinate system of the camera to obtain a plurality of first two-dimensional points includes:
[0012] Converting the point cloud data into a camera coordinate system according to a first relative transformation parameter between the camera and the ranging sensor to obtain a plurality of first three-dimensional points;
[0013] The plurality of first three-dimensional points are converted into the image coordinate system of the camera according to the intrinsic parameters of the camera to obtain a plurality of first two-dimensional points.
[0014] Through the above-mentioned technical means, the point cloud data can be accurately projected into the image coordinate system of the first image based on the first relative transformation parameters of the camera and the ranging sensor and the internal parameters of the camera, so that it can be accurately judged whether the first image captures an obstacle close to the robot based on the first two-dimensional point in the image coordinate system, thereby improving the accuracy of acquiring near-field images.
[0015] Optionally, converting the point cloud data into a camera coordinate system according to a first relative transformation parameter between the camera and the ranging sensor to obtain a plurality of first three-dimensional points includes:
[0016] When the timestamps of the first image and the point cloud data are different, converting the point cloud data into a camera coordinate system according to the extrinsic parameters of the camera, the extrinsic parameters of the ranging sensor, the first pose and the second pose to obtain a plurality of first three-dimensional points;
[0017] The first posture is the posture of the robot when the camera captures the first image, and the second posture is the posture of the robot when the ranging sensor collects the point cloud data.
[0018] Through the above technical means, when the timestamps of the first image and the point cloud data are not aligned, the point cloud data is accurately converted to the camera coordinate system when the monocular camera took the first image through the posture of the robot when collecting the first image and point cloud data and the external parameters of the lidar and monocular camera, so as to improve the acquisition accuracy of the near-field image.
[0019] Optionally, determining whether the first image is a near-field image of the robot based on the plurality of first two-dimensional points includes:
[0020] Determine the Z-axis coordinate of the first three-dimensional point as a depth value corresponding to the first two-dimensional point;
[0021] Determine whether the first image is a near-field image of the robot according to the coordinates and depth values of the plurality of first two-dimensional points.
[0022] Through the above technical means, it is possible to accurately judge whether the obstacle captured by the first image is close to the robot based on the depth value and coordinates of the first two-dimensional point, and then accurately determine whether the first image is a near-field image, thereby improving the acquisition accuracy of the near-field image.
[0023] Optionally, determining whether the first image is a near-field image of the robot according to the coordinates and depth values of the plurality of first two-dimensional points includes:
[0024] When the coordinates of at least one of the first two-dimensional points are within the pixel range of the first image and the depth value corresponding to the first two-dimensional point is less than or equal to a preset depth threshold, the first image is determined to be a near-field image of the robot.
[0025] Through the above technical means, when any spatial point of an obstacle captured by the first image is close to the robot, the first image can be determined as a near-field image, so as to obtain a variety of near-field images, which is conducive to improving the robot's near-field perception ability through training.
[0026] Optionally, before converting the point cloud data into a camera coordinate system according to a first relative transformation parameter between the camera and the ranging sensor to obtain a plurality of first three-dimensional points, the method further includes:
[0027] Calibrate the external parameters of the camera and the external parameters of the ranging sensor;
[0028] A first relative transformation parameter between the camera and the ranging sensor is determined based on the extrinsic parameters of the camera and the extrinsic parameters of the ranging sensor.
[0029] Optionally, when the first image is a near-field image, uploading the first image includes:
[0030] If the first image is a near-field image, storing the first image in a local image library;
[0031] When the number of first images in the local image library is equal to a preset number, uploading each first image in the local image library to a data center;
[0032] The first image in the local image library is deleted.
[0033] Through the above-mentioned technical means, the collected near-field images can be temporarily stored in the local image library, so that when the number of images stored in the local image library is equal to the preset number, the near-field images in the local image library can be uniformly uploaded to the data center to avoid the robot from frequently performing the operation of uploading data, which is beneficial to saving the robot's operating power consumption.
[0034] In a second aspect, the present application provides a near-field image acquisition device, comprising:
[0035] an image acquisition module, configured to acquire a first image captured by a camera of the robot;
[0036] a point cloud conversion module, configured to obtain point cloud data collected by the ranging sensor of the robot, and convert the point cloud data into the image coordinate system of the camera to obtain a plurality of first two-dimensional points;
[0037] an image determination module, configured to determine whether the first image is a near-field image of the robot based on the plurality of the first two-dimensional points;
[0038] The image uploading module is configured to upload the first image when the first image is a near-field image.
[0039] In a third aspect, the present application provides a near-field image acquisition device, comprising:
[0040] One or more processors; a memory storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the near-field image acquisition method as described in the first aspect.
[0041] In a fourth aspect, the present application provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the near-field image acquisition method as described in the first aspect.
[0042] In the present application, a first image captured by the robot's camera is obtained; point cloud data captured by the robot's ranging sensor is obtained, and the point cloud data is converted to the camera's image coordinate system to obtain a plurality of first two-dimensional points; based on the plurality of first two-dimensional points, it is determined whether the first image is a near-field image of the robot; if the first image is a near-field image, the first image is uploaded. Through the above technical means, the point cloud data collected by the ranging sensor during the movement of the robot can be projected into the image coordinate system of the first image, so as to accurately determine whether the first image captured by the robot in real time is a near-field image based on the first two-dimensional points obtained by the projection, and upload the near-field image to the data center when it is determined that the near-field image is captured, thereby avoiding uploading invalid data other than the near-field image to the data center, solving the problem of the data center acquiring a large amount of invalid data other than the near-field image in the prior art, reducing the cost of data storage and uploading and saving the cost of manual screening. The data center can directly acquire the near-field image without the need for manual screening of the near-field image, which is conducive to improving the efficiency of acquiring near-field data. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flowchart of a near-field image acquisition method provided by an embodiment of the present application;
[0044] Figure 2 is a schematic diagram of a robot and an obstacle provided in an embodiment of the present application;
[0045] Figure 3 It is a flowchart of converting a three-dimensional point into a first two-dimensional point provided by an embodiment of the application;
[0046] Figure 4 is a flowchart of determining whether a first image is a near-field image according to an embodiment of the present application;
[0047] Figure 5 is a schematic diagram of a first image provided in an embodiment of the present application;
[0048] Figure 6 is a schematic diagram of a projection image corresponding to a first two-dimensional point provided in an embodiment of the present application;
[0049] Figure 7 is a flowchart of uploading a first image to a data center provided by an embodiment of the present application;
[0050] Figure 8 This is a schematic structural diagram of a near-field image acquisition device provided in an embodiment of the present application;
[0051] Figure 9 It is a structural schematic diagram of a near-field image acquisition device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only some, but not all, of the contents related to the present application are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0053] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0054] In the implementation of related technologies, the robot uses a full-volume acquisition method to control the camera in real time while moving to collect image data. All collected image data is uploaded to the data center, and staff manually filter the image data cached in the data center to extract near-field images. However, the full-volume acquisition method collects a large amount of image data, most of which is invalid data other than near-field images. This causes the data center to obtain a large amount of invalid data other than near-field images, increasing the cost of data storage and transmission. The staff also needs more time and energy to screen near-field images, which increases the cost of manual screening and affects the efficiency of near-field image acquisition.
[0055] To solve the above problems, this embodiment provides a near-field image acquisition method to convert the point cloud data collected by the robot's ranging sensor into first two-dimensional points in the camera coordinate system, and judge whether the first image collected by the robot's camera is a near-field image based on the first two-dimensional points, and upload the near-field image to the data center, avoiding uploading non-near-field images to the data center, eliminating the need for manual screening of near-field images, reducing the cost of data storage and uploading, saving the cost of manual screening, and helping to improve the efficiency of acquiring near-field data.
[0056] The near-field image acquisition method provided in this embodiment can be performed by a near-field image acquisition device. The near-field image acquisition device can be implemented through software and / or hardware. The near-field image acquisition device can be composed of two or more physical entities or a single physical entity. In this embodiment, the near-field image acquisition device can be a robot or a processor of a robot.
[0057] The near-field image acquisition device is installed with at least one operating system, including but not limited to Android, Linux, and Windows. The near-field image acquisition device can install at least one application based on the operating system. The application can be a native application of the operating system or an application downloaded from a third-party device or server. In this embodiment, the near-field image acquisition device has at least one application capable of executing the near-field image acquisition method.
[0058] For ease of understanding, this embodiment is described by taking a robot as an example of a subject that executes the near-field image acquisition method.
[0059] Figure 1 This is a flow chart of a near-field image acquisition method provided by an embodiment of the present application. Figure 1 As shown, the steps of the near-field image acquisition method of the server include:
[0060] S110: Acquire a first image captured by a camera of the robot.
[0061] In this embodiment, the first image is an image captured in real time by the robot via a camera. Exemplarily, the robot turns on its camera when it begins moving, controlling the camera to capture the first image at a preset frequency. Alternatively, the robot acquires a preconfigured target area and, based on the target area and its own navigation and positioning information, determines whether the robot has moved into the target area. When the robot moves into the target area, it turns on its camera and controls the camera to capture the first image at a preset frequency. The target area can be understood as an area centered around the obstacle. When the robot moves into the target area, it indicates that the robot is approaching the obstacle, at which point the camera can be turned on to prepare to capture near-field images. When the robot moves outside the target area, it indicates that the robot is moving away from the obstacle, at which point the camera can be turned off to reduce the amount of invalid data captured by the camera in non-near-field images, thereby conserving the robot's energy consumption. A near-field image is an image captured by the robot of the obstacle when it is approaching the obstacle, while a non-near-field image is an image captured by the robot of the obstacle when it is away from the obstacle, or an image that does not contain the obstacle.
[0062] S120 , obtaining point cloud data collected by the ranging sensor of the robot, and converting the point cloud data into the image coordinate system of the camera to obtain a plurality of first two-dimensional points.
[0063] In this embodiment, the ranging sensor is a sensor such as a laser radar or a depth camera that can collect point cloud data of obstacles in front of the robot. Point cloud data is a collection of three-dimensional points scanned by the ranging sensor, and a three-dimensional point is a point formed by three-dimensional coordinates (X, Y, Z). This embodiment is described by taking the ranging sensor as a laser radar as an example. When the laser radar scans an obstacle, it will generate point cloud data corresponding to the obstacle, and the point cloud data can represent the distance between the obstacle and the robot. The shooting direction of the robot's camera is also toward the front of the robot. The obstacle captured by the camera is most likely the obstacle scanned by the laser radar. Therefore, it can be determined based on the point cloud data whether the robot is close to the obstacle, so that when it is determined that the robot is close to the obstacle, the first image captured by the camera is determined to be the near-field image of the robot.
[0064] It should be noted that although the shooting direction of the robot's camera is also in front of the robot, due to the difference in the installation position of the camera and the lidar, the difference in the collection range of the camera and the lidar, and the difference in the collection timestamp of the camera and the lidar, the objects photographed by the camera and the objects scanned by the lidar may be different. Figure 2 Schematic diagram of the robot and obstacles provided in the embodiment of the present application. Figure 2 As shown, obstacle A is within the imaging area 13 of robot 10's camera 11, and both obstacles A and B are within the scanning area 14 of lidar 12. Therefore, camera 11 captures a first image containing obstacle A, and lidar 12 scans point cloud data for obstacles A and B. If robot 10 determines, based on this point cloud data, that obstacle B is close to robot 10, it will identify the first image captured by camera 11 as a near-field image. However, this first image does not contain obstacle B, which is close to robot 11. In other words, this first image is not a near-field image. This results in an incorrect near-field image judgment, affecting the accuracy of near-field image acquisition.
[0065] To improve the accuracy of near-field image acquisition, the point cloud data can be converted to the camera's image coordinate system. Based on the converted point cloud data, it can be used to determine whether the first image captures an obstacle near the robot. For example, the point cloud data includes multiple 3D points. Based on a first relative transformation parameter between the lidar and camera and the camera's intrinsic parameters, each 3D point in the point cloud data can be converted to the camera's image coordinate system to obtain the first 2D point corresponding to the 3D point.
[0066] In this embodiment, Figure 3 This is a flowchart of converting a three-dimensional point into a first two-dimensional point provided by an embodiment of the application. Figure 3 As shown, the step of converting the three-dimensional point into the first two-dimensional point specifically includes S1201-S1202:
[0067] S1201: Convert the point cloud data into a camera coordinate system according to a first relative transformation parameter between the camera and the ranging sensor to obtain a plurality of first three-dimensional points.
[0068] S1202 : Convert the plurality of first three-dimensional points into the image coordinate system of the camera according to the intrinsic parameters of the camera to obtain a plurality of first two-dimensional points.
[0069] The first relative transformation parameter is the transformation matrix used when transforming the laser radar coordinate system to the camera coordinate system. Multiplying the first relative transformation parameter by the 3D point in the point cloud data yields the first 3D point transformed to the camera coordinate system. The calculation expression for the first 3D point is:
[0070]
[0071] Among them, P l c is the first three-dimensional point, x c1 、y c1 and z c1 are the X-axis coordinate, Y-axis coordinate, and Z-axis coordinate of the first three-dimensional point, T l c is the first relative transformation parameter, P l is a three-dimensional point in the point cloud data. In this embodiment, the three-dimensional point in the point cloud data can be converted to the camera coordinate system based on the first relative transformation parameters between the pre-calibrated laser radar and the camera to obtain the first three-dimensional point. However, the operation of directly calibrating the first relative transformation parameters between the laser radar and the camera is relatively complicated. The external parameters of the camera and the external parameters of the laser radar can be calibrated; based on the external parameters of the camera and the external parameters of the laser radar, the first relative transformation parameters between the camera and the laser radar are determined. Among them, the external parameters of the laser radar are the transformation matrix used when the coordinate system of the robot is converted to the coordinate system of the laser radar, and the external parameters of the camera are the transformation matrix used when the coordinate system of the robot is converted to the coordinate system of the camera. Correspondingly, the calculation expression of the first relative transformation parameter is:
[0072]
[0073] in, is the external parameter of the camera, is the external parameter of the laser radar.
[0074] It should be noted that only when the timestamps of the point cloud data collected by the lidar and the first image captured by the camera are aligned can the point cloud data be accurately converted to the camera coordinate system when the camera captured the first image, and then the first two-dimensional point in the same image coordinate system as the first image can be obtained. Otherwise, the three-dimensional points of obstacles not captured by the first image will be mistakenly judged to fall into the image range of the first image when projected to the first two-dimensional points in the image coordinate system, thereby causing misrecognition of the near-field image and affecting the accuracy of near-field image acquisition. To this end, the robot can control the lidar and camera to synchronously collect point cloud data and the first image to ensure that the timestamps of the first image and the point cloud data are aligned.
[0075] Alternatively, when the timestamps of the first image and the point cloud data are different, the robot can convert the point cloud data into the camera coordinate system based on the camera's extrinsic parameters, the range sensor's extrinsic parameters, the first pose, and the second pose to obtain multiple first three-dimensional points. The first pose is the robot's pose when the camera captures the first image, and the second pose is the robot's pose when the range sensor collects the point cloud data. For example, the expression for the first three-dimensional point is:
[0076]
[0077] in, is the transformation matrix corresponding to the first pose, is the transformation matrix corresponding to the second posture. As can be seen from the above expression, the three-dimensional points of the point cloud data can be converted to the robot coordinate system when the laser radar collects the point cloud data through the external parameters of the laser radar, and then the three-dimensional points in the robot coordinate system are converted to the world coordinate system based on the transformation matrix corresponding to the second posture. The three-dimensional points in the world coordinate system are converted to the robot coordinate system when the camera takes the first image through the inverse matrix of the transformation matrix corresponding to the first posture. Finally, the three-dimensional points in the robot coordinate system are converted to the camera coordinate system when the camera takes the first image through the inverse matrix of the external parameters of the camera, so as to ensure that the first two-dimensional points generated corresponding to the point cloud data can be converted to the image coordinate system corresponding to the first image, which is conducive to improving the acquisition accuracy of the near-field image. It should be noted that the posture of the robot includes position and attitude angle. A translation matrix can be generated based on the position and a rotation matrix can be generated based on the attitude angle. The translation matrix and the rotation matrix can be combined to form a transformation matrix that transforms the robot coordinate system to the world coordinate system. This transformation matrix is also the transformation matrix corresponding to the posture of the robot.
[0078] This embodiment improves the accuracy of near-field image acquisition by accurately converting point cloud data into the camera coordinate system when the monocular camera captured the first image, using the robot's posture when collecting the first image and point cloud data and the external parameters of the lidar and monocular camera when the timestamps of the first image and point cloud data are not aligned.
[0079] Furthermore, the intrinsic parameters of the camera include the horizontal focal length, vertical focal length, and the horizontal and vertical pixel numbers that differ between the center pixel coordinates of the image and the pixel coordinates of the image origin. The calculation formula for the first two-dimensional point is:
[0080]
[0081] Among them, x1 and y1 are the coordinates of the first two-dimensional point, f x and f y is the horizontal focal length and vertical focal length, c x and c y is the number of horizontal and vertical pixels.
[0082] This embodiment uses the first relative transformation parameters of the camera and the ranging sensor and the internal parameters of the camera to accurately project the point cloud data into the image coordinate system of the first image, so that it can subsequently accurately determine whether the first image captures an obstacle close to the robot based on the first two-dimensional point in the image coordinate system, thereby improving the accuracy of acquiring near-field images.
[0083] S130: Determine whether the first image is a near-field image of the robot based on the plurality of first two-dimensional points.
[0084] For example, if the obstacle scanned by the lidar is also within the shooting area of the camera, then when the point cloud data corresponding to the obstacle is projected onto the image coordinate system corresponding to the first image, it can be matched to the pixel points corresponding to the obstacle in the first image. Combined with the distance between the obstacle and the robot represented by the point cloud data corresponding to the obstacle, it can be determined whether the first image captures an obstacle close to the robot, that is, whether the first image is a near-field image of the robot.
[0085] In this embodiment, Figure 4 This is a flowchart of determining whether a first image is a near-field image according to an embodiment of the present application. The steps of determining whether the first image is a near-field image specifically include S1301-S1302:
[0086] S1301: Determine the Z-axis coordinate of the first three-dimensional point as a depth value corresponding to the first two-dimensional point.
[0087] S1302: Determine whether the first image is a near-field image of the robot based on the coordinates and depth values of the plurality of first two-dimensional points.
[0088] Exemplarily, when converting the three-dimensional points of the point cloud data to obtain the first three-dimensional points in the camera coordinate system, the Z-axis coordinate of the first three-dimensional points can represent the distance from the spatial points corresponding to the obstacle to the camera. Since the camera and the robot are integrally arranged, that is, the Z-axis coordinate of the first three-dimensional points can represent the distance between the obstacle and the robot. The coordinates of the first two-dimensional points are the pixel coordinates of the projection of the three-dimensional points of the point cloud data onto the image coordinate system of the first image. It can be understood that if the coordinates of the first two-dimensional points fall within the pixel range of the first image, it indicates that the first image captures the spatial points corresponding to the first two-dimensional points. If the coordinates of the first two-dimensional points do not fall within the pixel range of the first image, it indicates that the first image does not capture the spatial points corresponding to the first two-dimensional points. The spatial points corresponding to the first two-dimensional points are a point on the corresponding obstacle. This embodiment combines Figure 5 and Figure 6 to describe the projection images respectively formed by the first image and the first two-dimensional points shown. As Figure 5 shown, the pixel range of the first image is 0 < x < W and 0 < y < H. The coordinates of the first two-dimensional points corresponding to the obstacle B in the projection image all exceed the pixel range of the first image, so it can be determined that the first image does not capture the obstacle B.
[0089] Only when the obstacle captured by the first image is close to the robot can it be determined that the first image is a near-field image. That is, when the coordinates of the first two-dimensional points fall within the pixel range of the first image, it is possible to determine whether the spatial points corresponding to them are close to the robot based on the depth value of the first two-dimensional points. In this embodiment, the depth threshold is preset as the maximum distance between the near-field image of the obstacle captured by the robot and the obstacle. If the depth value of the first two-dimensional points is greater than the preset depth threshold, it indicates that the distance between the current robot and the spatial points corresponding to the first two-dimensional points is not sufficient for the camera to capture the near-field image of the obstacle corresponding to the spatial points. If the depth value of the first two-dimensional points is less than or equal to the preset depth threshold, it indicates that the distance between the current robot and the spatial points corresponding to the first two-dimensional points is sufficient for the camera to capture the near-field image of the obstacle corresponding to the spatial points. Referring to Figure 5 and Figure 6 , only when the first two-dimensional points corresponding to the obstacle A fall within the pixel range of the first image, can it be determined whether the obstacle A is close to the robot based on the comparison result between the depth value of the first two-dimensional points corresponding to the obstacle A and the preset depth threshold. This embodiment can accurately determine whether the obstacle captured by the first image is close to the robot through the depth value and coordinates of the first two-dimensional points, and then accurately determine whether the first image is a near-field image, improving the acquisition accuracy of the near-field image.
[0090] Furthermore, when the coordinates of at least one first two-dimensional point are within the pixel range of the first image and the depth value corresponding to the first two-dimensional point is less than or equal to a preset depth threshold, the first image is determined to be a near-field image of the robot. It can be understood that the spatial point corresponding to the first two-dimensional point is a point on the obstacle. When at least one first two-dimensional point corresponding to the obstacle falls within the pixel range of the first image and the depth value is less than or equal to the depth threshold, it indicates that the first image captures the obstacle and the robot is close to the obstacle, that is, the first image can be determined to be a near-field image captured by the camera close to the obstacle. Figure 5 and Figure 6 If the depth values of the first two-dimensional points corresponding to obstacle A are all greater than a preset depth threshold, obstacle A is determined to be far from the robot, i.e., the first image is determined not to be a near-field image. If the depth value of at least one first two-dimensional point corresponding to obstacle A is less than or equal to the preset depth threshold, obstacle A is determined to be close to the robot, i.e., the first image is determined to be a near-field image. This embodiment determines the first image as a near-field image when any spatial point of an obstacle captured in the first image is close to the robot. This allows for the acquisition of a variety of near-field images, which is beneficial for improving the robot's near-field perception capabilities through training.
[0091] S140: When the first image is a near-field image, upload the first image.
[0092] Exemplarily, when the robot determines that the first image taken by the camera is a near-field image, the robot may upload the first image to the data center based on the communication address of the data center.
[0093] Alternatively, after collecting a certain number of near-field images, the robot can upload the acquired near-field images to the data center in a unified manner to avoid the robot from frequently performing the operation of uploading data. Figure 7 This is a flow chart of uploading the first image to the data center provided by the embodiment of the present application. Figure 7 As shown, the step of uploading the first image to the data center specifically includes S1401-S1403:
[0094] S1401: When the first image is a near-field image, store the first image in a local image library.
[0095] S1402: When the number of first images in the local image library is equal to a preset number, upload each first image in the local image library to the data center.
[0096] S1403: Delete the first image in the local image library.
[0097] Exemplarily, when the robot determines that the first image is a near-field image, it stores the first image in the local image library. When the number of first images stored in the local image library is equal to a preset number, all the first images in the local image library can be packaged into a compressed package and transmitted to the data center to reduce the amount of transmitted data, improve data transmission efficiency, and thereby improve the efficiency of the data center in acquiring near-field images. After the data center receives the compressed package, it decompresses the compressed package to obtain the near-field image collected by the robot. After the robot completely transmits the compressed package to the data center, it can delete the first image stored in the local image library to avoid repeatedly uploading the same near-field image to the data center. This embodiment temporarily stores the collected near-field images in the local image library, so that when the number of images stored in the local image library is equal to the preset number, the near-field images in the local image library are uniformly uploaded to the data center to avoid the robot frequently performing the operation of uploading data, which is beneficial to saving the robot's operating power consumption.
[0098] In summary, the near-field image acquisition method provided by the embodiment of the present application is achieved by acquiring the first image captured by the robot's camera; acquiring the point cloud data captured by the robot's ranging sensor, converting the point cloud data into the camera's image coordinate system, and obtaining a plurality of first two-dimensional points; determining whether the first image is the robot's near-field image based on the plurality of first two-dimensional points; and uploading the first image if the first image is a near-field image. Through the above-mentioned technical means, the point cloud data collected by the ranging sensor during the movement of the robot can be projected into the image coordinate system of the first image, so as to accurately determine whether the first image captured by the robot in real time is a near-field image based on the first two-dimensional points obtained by the projection, and uploading the near-field image to the data center when it is determined that the near-field image is captured, thereby avoiding uploading invalid data other than the near-field image to the data center, solving the problem of the data center acquiring a large amount of invalid data other than the near-field image in the prior art, reducing the cost of data storage and uploading and saving the cost of manual screening. The data center can directly acquire the near-field image without the need for manual screening of the near-field image, which is conducive to improving the efficiency of acquiring near-field data.
[0099] Based on the above embodiments, Figure 8 This is a schematic diagram of the structure of a near-field image acquisition device provided in an embodiment of the present application. Figure 8 The near-field image acquisition device provided in this embodiment specifically includes: an image acquisition module 21, a point cloud conversion module 22, an image judgment module 23 and an image upload module 24.
[0100] The image acquisition module 21 is configured to acquire a first image captured by a camera of the robot;
[0101] a point cloud conversion module 22 configured to obtain point cloud data collected by the ranging sensor of the robot, and convert the point cloud data into the image coordinate system of the camera to obtain a plurality of first two-dimensional points;
[0102] an image determination module 23, configured to determine whether the first image is a near-field image of the robot based on the plurality of the first two-dimensional points;
[0103] The image uploading module 24 is configured to upload the first image when the first image is a near-field image.
[0104] Based on the above embodiment, the point cloud conversion module 22 includes: a first conversion unit, configured to convert the point cloud data into the camera coordinate system according to the first relative transformation parameter between the camera and the ranging sensor, to obtain multiple first three-dimensional points; a second conversion unit, configured to convert the multiple first three-dimensional points into the camera's image coordinate system according to the camera's intrinsic parameters, to obtain multiple first two-dimensional points.
[0105] Based on the above embodiment, the first conversion unit includes: a first conversion subunit, configured to convert the point cloud data into a camera coordinate system according to the external parameters of the camera, the external parameters of the ranging sensor, the first pose and the second pose, to obtain multiple first three-dimensional points when the timestamps of the first image and the point cloud data are different; wherein the first pose is the pose of the robot when the camera captures the first image, and the second pose is the pose of the robot when the ranging sensor collects the point cloud data.
[0106] Based on the above embodiment, the image judgment module 23 includes: a depth value determination unit, configured to determine the Z-axis coordinate of the first three-dimensional point as the depth value corresponding to the first two-dimensional point; a near-field image determination unit, configured to determine whether the first image is a near-field image of the robot based on the coordinates and depth values of multiple first two-dimensional points.
[0107] Based on the above embodiment, the near-field image determination unit includes: a near-field image determination sub-unit, which is configured to determine that the first image is a near-field image of the robot when the coordinates of at least one first two-dimensional point are within the pixel range of the first image and the depth value corresponding to the first two-dimensional point is less than or equal to a preset depth threshold.
[0108] Based on the above embodiment, the near-field image acquisition device further includes: a first calibration module configured to calibrate the extrinsic parameters of the camera and the extrinsic parameters of the ranging sensor; and a change parameter determination module configured to determine a first relative transformation parameter between the camera and the ranging sensor based on the extrinsic parameters of the camera and the extrinsic parameters of the ranging sensor.
[0109] Based on the above embodiment, the image upload module 24 includes: an image storage unit, configured to store the first image in the local image library when the first image is a near-field image; an image upload unit, configured to upload each first image in the local image library to the data center when the number of first images in the local image library is equal to a preset number; and an image deletion unit, configured to delete the first image in the local image library.
[0110] As mentioned above, the near-field image acquisition device provided by the embodiment of the present application obtains a first image captured by the robot's camera; obtains point cloud data collected by the robot's ranging sensor, converts the point cloud data into the camera's image coordinate system, and obtains a plurality of first two-dimensional points; determines whether the first image is the robot's near-field image based on the plurality of first two-dimensional points; and uploads the first image if it is a near-field image. Through the above technical means, the point cloud data collected by the ranging sensor during the movement of the robot can be projected into the image coordinate system of the first image, so as to accurately determine whether the first image collected by the robot in real time is a near-field image based on the first two-dimensional points obtained by the projection, and upload the near-field image to the data center when it is determined that the near-field image is collected, thereby avoiding uploading invalid data other than the near-field image to the data center, solving the problem of the data center obtaining a large amount of invalid data other than the near-field image in the prior art, reducing the cost of data storage and uploading and saving the cost of manual screening. The data center can directly obtain the near-field image without the need for manual screening of the near-field image, which is conducive to improving the efficiency of obtaining near-field data.
[0111] The near-field image acquisition device provided in the embodiment of the present application can be used to execute the near-field image acquisition method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0112] Figure 9 This is a schematic diagram of the structure of a near-field image acquisition device provided in an embodiment of the present application, with reference to Figure 9 The near-field image acquisition device includes: a processor 31, a memory 32, a communication device 33, an input device 34, and an output device 35. The number of processors 31 in the near-field image acquisition device can be one or more, and the number of memories 32 in the near-field image acquisition device can be one or more. The processor 31, memory 32, communication device 33, input device 34, and output device 35 of the near-field image acquisition device can be connected via a bus or other means.
[0113] The memory 32 is a computer-readable storage medium that can be used to store software programs, computer executable programs, and modules, such as the program instructions / modules corresponding to the near-field image acquisition method of any embodiment of the present application (for example, the image acquisition module 21, the point cloud conversion module 22, the image judgment module 23, and the image upload module 24 in the near-field image acquisition device). The memory 32 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the device, etc. In addition, the memory 32 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0114] The communication device 33 is used for data transmission.
[0115] The processor 31 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 32, that is, realizes the above-mentioned near-field image acquisition method.
[0116] The input device 34 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 35 may include a display device such as a display screen.
[0117] The near-field image acquisition device provided above can be used to execute the near-field image acquisition method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0118] An embodiment of the present application also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute a near-field image acquisition method. The near-field image acquisition method includes: acquiring a first image captured by a camera of the robot; acquiring point cloud data captured by a ranging sensor of the robot, converting the point cloud data into an image coordinate system of the camera to obtain a plurality of first two-dimensional points; determining whether the first image is a near-field image of the robot based on the plurality of first two-dimensional points; and uploading the first image if the first image is a near-field image.
[0119] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media, such as CD-ROMs, floppy disks, or tape drives; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. In addition, the storage medium may be located in the first computer system in which the program is executed, or it may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). The storage medium may store program instructions (e.g., embodied as a computer program) that can be executed by one or more processors.
[0120] Of course, the storage medium containing computer-executable instructions provided in an embodiment of the present application is not limited to the above-mentioned near-field image acquisition method, and can also execute related operations in the near-field image acquisition method provided in any embodiment of the present application.
[0121] The near-field image acquisition device, storage medium, and near-field image acquisition equipment provided in the above embodiments can execute the near-field image acquisition method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the near-field image acquisition method provided in any embodiment of the present application.
[0122] The above are only preferred embodiments of the present application and the technical principles employed. The present application is not limited to the specific embodiments described herein, and any obvious changes, readjustments, and substitutions that are apparent to those skilled in the art will not depart from the scope of protection of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the claims.
Claims
1. A near-field image acquisition method, characterized in that: include: Acquire a first image captured by a camera of the robot; Acquire point cloud data collected by a ranging sensor of the robot, and convert the point cloud data into an image coordinate system of the camera to obtain a plurality of first two-dimensional points; determining, based on the plurality of first two-dimensional points, whether the first image is a near-field image of the robot; In the case where the first image is a near-field image, the first image is uploaded.
2. The near-field image acquisition method according to claim 1, characterized in that: The step of converting the point cloud data into the image coordinate system of the camera to obtain a plurality of first two-dimensional points includes: Converting the point cloud data into a camera coordinate system according to a first relative transformation parameter between the camera and the ranging sensor to obtain a plurality of first three-dimensional points; The plurality of first three-dimensional points are converted into the image coordinate system of the camera according to the intrinsic parameters of the camera to obtain a plurality of first two-dimensional points.
3. The near-field image acquisition method according to claim 2, characterized in that: The step of converting the point cloud data into a camera coordinate system according to a first relative transformation parameter between the camera and the ranging sensor to obtain a plurality of first three-dimensional points includes: When the timestamps of the first image and the point cloud data are different, converting the point cloud data into a camera coordinate system according to the extrinsic parameters of the camera, the extrinsic parameters of the ranging sensor, the first pose and the second pose to obtain a plurality of first three-dimensional points; The first posture is the posture of the robot when the camera captures the first image, and the second posture is the posture of the robot when the ranging sensor collects the point cloud data.
4. The near-field image acquisition method according to claim 2, characterized in that: The determining, based on the plurality of the first two-dimensional points, whether the first image is a near-field image of the robot comprises: Determine the Z-axis coordinate of the first three-dimensional point as a depth value corresponding to the first two-dimensional point; Determine whether the first image is a near-field image of the robot according to the coordinates and depth values of the plurality of first two-dimensional points.
5. The near-field image acquisition method according to claim 4, characterized in that: The determining, based on the coordinates and depth values of the plurality of first two-dimensional points, whether the first image is a near-field image of the robot comprises: When the coordinates of at least one of the first two-dimensional points are within the pixel range of the first image and the depth value corresponding to the first two-dimensional point is less than or equal to a preset depth threshold, the first image is determined to be a near-field image of the robot.
6. The near-field image acquisition method according to claim 2, characterized in that: Before converting the point cloud data into a camera coordinate system according to a first relative transformation parameter between the camera and the ranging sensor to obtain a plurality of first three-dimensional points, the method further includes: Calibrate the external parameters of the camera and the external parameters of the ranging sensor; A first relative transformation parameter between the camera and the ranging sensor is determined based on the extrinsic parameters of the camera and the extrinsic parameters of the ranging sensor.
7. The near-field image acquisition method according to claim 1, characterized in that: When the first image is a near-field image, uploading the first image includes: If the first image is a near-field image, storing the first image in a local image library; When the number of first images in the local image library is equal to a preset number, uploading each first image in the local image library to a data center; The first image in the local image library is deleted.
8. A near-field image acquisition device, characterized in that: include: an image acquisition module, configured to acquire a first image captured by a camera of the robot; a point cloud conversion module, configured to obtain point cloud data collected by the ranging sensor of the robot, and convert the point cloud data into the image coordinate system of the camera to obtain a plurality of first two-dimensional points; an image determination module, configured to determine whether the first image is a near-field image of the robot based on the plurality of the first two-dimensional points; The image uploading module is configured to upload the first image when the first image is a near-field image.
9. A near-field image acquisition device, characterized in that: include: one or more processors; A memory stores one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the near-field image acquisition method according to any one of claims 1 to 7.
10. A storage medium containing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, they are used to perform the near-field image acquisition method according to any one of claims 1 to 7.