A method and device for generating a surround view image of a robot, a quadruped robot, equipment and a storage medium
By modeling the real robot and calculating the camera fusion weights, the problems of stitching gaps and artifacts in multi-channel fisheye camera panoramic stitching were solved, and high-quality seamless fusion of panoramic images of the quadruped robot was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHISHEN XINCHUANG (SUZHOU) INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, multi-channel fisheye camera panoramic stitching technology suffers from stitching gaps and artifacts in overlapping areas, making it impossible to achieve high-quality panoramic image fusion.
By modeling a real robot, a virtual robot is generated, and a bowl-shaped curved surface model is determined. The camera fusion weights are calculated using azimuth weights and edge attenuation weights to generate panoramic mesh vertex data, achieving seamless fusion of multiple real-world images.
It achieves seamless fusion of panoramic images, improves image quality, reduces the computational burden on embedded platforms, and is suitable for small-sized quadruped robots and complex environment perception.
Smart Images

Figure CN122435569A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and more specifically, to a method, apparatus, quadruped robot, device, and storage medium for generating panoramic images of a robot. Background Technology
[0002] Quadruped robots are widely used in disaster relief, industrial inspection, and hazardous environment detection. Remote operators need to perceive the robot's 360° surroundings in real time and intuitively to complete tasks such as obstacle avoidance and navigation. Traditional single-camera systems have limited fields of view and cannot provide complete situational awareness. Therefore, multiple fisheye cameras (such as four cameras in the front, back, left, and right, each with a field of view of approximately 190°) are typically mounted on the robot's body, and panoramic images are generated by stitching the images together.
[0003] In multi-channel fisheye camera panoramic stitching technology, there is a large overlap in the fields of view of each camera. During the fusion process, stitching gaps or artifacts often occur in the overlapping areas. Therefore, how to improve the fusion quality of panoramic images is one of the urgent problems to be solved. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus, quadruped robot, device and storage medium for generating a surround view image of a robot, so as to at least solve the problems existing in the related art.
[0005] Specifically, this application is implemented through the following technical solution:
[0006] This application provides a method for generating a surround view image of a robot, including: A virtual robot is obtained by modeling a real robot; the real robot is equipped with multiple real environment cameras, and the viewpoints of adjacent real environment cameras overlap; the virtual robot has virtual environment cameras that correspond one-to-one with the multiple real environment cameras. A bowl-shaped surface model corresponding to the virtual robot is determined; the bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with respect to the virtual robot remain unchanged; For each grid vertex, the azimuth weight and edge attenuation weight of each virtual environment camera relative to the grid vertex are determined respectively; the azimuth weight is used to characterize the degree of deviation of the optical axis center of each virtual environment camera relative to the grid vertex; the edge attenuation weight is used to characterize the image sampling quality at the projection position of the grid vertex on the screen of each virtual camera. For each virtual environment camera, the camera fusion weight of the virtual environment camera at the grid vertex is determined according to the azimuth weight and the edge attenuation weight, and panoramic grid vertex data is generated based on the multiple camera fusion weights corresponding to each grid vertex. Multiple real-world environment images are acquired using the multiple physical environment cameras, and then fused using the panoramic mesh vertex data to obtain a panoramic image.
[0007] This application also provides an apparatus for generating a surround view image of a robot, comprising: The robot modeling module is used to model a real robot to obtain a virtual robot. The real robot is equipped with multiple real environment cameras, and the viewpoints of adjacent real environment cameras overlap. The virtual robot has virtual environment cameras that correspond one-to-one with the multiple real environment cameras. A bowl-shaped model determination module is used to determine a bowl-shaped surface model corresponding to the virtual robot; the bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with respect to the virtual robot remain unchanged; The weight determination module is used to determine the azimuth weight and edge attenuation weight of each virtual environment camera relative to the grid vertex for each grid vertex; the azimuth weight is used to characterize the degree of deviation of the optical axis center of each virtual environment camera relative to the grid vertex; the edge attenuation weight is used to characterize the image sampling quality at the projection position of the grid vertex on the screen of each virtual camera. The weight fusion module is used to determine the camera fusion weight of the virtual environment camera at the grid vertex for each virtual environment camera according to the azimuth weight and the edge attenuation weight, and generate panoramic grid vertex data based on the multiple camera fusion weights corresponding to each grid vertex. The environmental image fusion module is used to acquire multiple real environment images through the multiple physical environment cameras, and to fuse the multiple real environment images using the panoramic mesh vertex data to obtain a panoramic image.
[0008] This application also provides a quadruped robot, which is equipped with multiple real-world environment cameras, and the viewpoints of adjacent environment cameras overlap. The quadruped robot includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the method for generating a surround view image of the robot as described in any of the foregoing embodiments.
[0009] This application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method for generating a surround view image for a robot as described in any of the foregoing embodiments.
[0010] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for generating a surround view image for a robot as described in any of the foregoing embodiments.
[0011] This application also provides a computer program product, including a computer program that, when run by a processor, performs the steps of any of the possible methods for generating a surround view image for a robot described above.
[0012] The technical solutions provided by the embodiments of this application may include the following beneficial effects: In this embodiment, camera fusion weights are determined based on azimuth weights and edge attenuation weights, thereby generating panoramic mesh vertex data. Since the azimuth weight is used to characterize the angular deviation between the optical axis of each virtual environment camera and the mesh vertex, it can reflect the degree of "sovereignty" of the virtual environment camera over that vertex. Therefore, the azimuth weight can be used to achieve a smooth transition in overlapping areas. At the same time, the edge attenuation weight is used to characterize the image sampling quality at the projection position of the mesh vertex on the screen of each virtual camera. Therefore, the edge attenuation weight can be used to suppress areas with severe edge distortion in fisheye images. Furthermore, the panoramic mesh vertex data is used to fuse the real environment images to obtain a panoramic image. In this way, while achieving a seamless fusion effect between real environment images, the quality of the fused image can also be improved.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description
[0014] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present application of a method for generating a surround view image of a robot; Figure 2 This is a schematic diagram illustrating the distribution of real-world cameras for a real robot in an exemplary embodiment of this application; Figure 3 This is a flowchart illustrating the projection process of a unified camera model according to an exemplary embodiment of this application; Figure 4 This is a top view of a bowl-shaped curved surface model illustrated in an exemplary embodiment of this application; Figure 5This is a side view of a bowl-shaped curved surface model illustrated in an exemplary embodiment of this application; Figure 6 This is a top view illustrating an azimuth sector division in an exemplary embodiment of this application; Figure 7 This is a schematic diagram illustrating the weight transition at a sector boundary, as shown in an exemplary embodiment of this application. Figure 8 This is a flowchart illustrating the fusion process of multiple real-world environment images according to an exemplary embodiment of this application; Figure 9 This is a schematic diagram illustrating a viewpoint control according to an exemplary embodiment of this application; Figure 10 This is a schematic diagram illustrating a rendering process according to an exemplary embodiment of this application; Figure 11 This is a schematic diagram of the structure of a device for generating a surround view image of a robot, as illustrated in an exemplary embodiment of this application. Figure 12 This is a schematic diagram of the structure of another device for generating a surround view image of a robot, as illustrated in an exemplary embodiment of this application. Figure 13 This is a hardware structure diagram of a quadruped robot shown in an exemplary embodiment of this application; Figure 14 This is a hardware structure diagram of a computer device illustrated in an exemplary embodiment of this application. Detailed Implementation
[0015] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0016] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0017] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0018] Quadruped robots are widely used in disaster relief, industrial inspection, and hazardous environment detection. Remote operators need to perceive the robot's 360° surroundings in real time and intuitively to complete tasks such as obstacle avoidance and navigation. Traditional single-camera systems have limited fields of view and cannot provide complete situational awareness. Therefore, multiple fisheye cameras (such as four cameras in the front, back, left, and right, each with a field of view of approximately 190°) are typically mounted on the robot's body, and panoramic images are generated by stitching the images together.
[0019] Existing surround view stitching technologies can be mainly classified into the following categories: Bird's Eye View: This method inversely projects a fisheye image onto a ground plane to generate a top-down view. It only works if the ground height is correct; objects above the ground (such as obstacles and pedestrians) will be severely stretched and distorted, and it cannot provide a sense of 3D space.
[0020] Cylindrical projection: Projects the image onto a cylindrical surface centered on the robot and unfolds it. Its bottom is open, so it cannot cover the ground area directly below the robot; near ground details are extremely compressed, making it unsuitable for quadruped robots that need to accurately perceive ground obstacles.
[0021] Spherical projection: The image is projected onto a sphere, with uniform sampling density throughout. However, quadruped robots are more concerned with the nearby ground (requiring high resolution) and the distant environment (which can be appropriately compressed). The uniformity of the sphere leads to wasted near-field resolution and insufficient distant resolution.
[0022] Bowl-shaped curved surface models (such as some automotive surround view patents): These models use a bowl-shaped curved surface that is flat in the middle and curved around the edges, taking into account both the nearby ground and the distant environment. However, existing solutions are designed for the size of cars (the radius of the flat area is usually 3-5 meters), which is not suitable for the small size of quadruped robots (width 0.3-0.5 meters); moreover, they mostly use CPU-side lookup tables to calculate projection in real time, resulting in low frame rates on embedded platforms; they lack support for fisheye camera unified model (UCM); they do not solve the problem of seamless blending of overlapping camera areas during rendering; and they do not address solutions for overlaying the robot's own 3D model onto the surround view image and providing correct occlusion relationships.
[0023] In multi-channel fisheye camera panoramic stitching technology, there is a large overlap in the fields of view of each camera. During the fusion process, stitching gaps or artifacts often occur in the overlapping areas. Therefore, how to improve the fusion quality of panoramic images is one of the urgent problems to be solved.
[0024] Based on the above research, this disclosure provides a method for generating a surround view image of a robot. The method first models a real robot to obtain a virtual robot. The real robot is equipped with multiple real-world cameras, with overlapping views between adjacent real-world cameras. The virtual robot has virtual environment cameras that correspond one-to-one with the multiple real-world cameras. Next, a bowl-shaped surface model corresponding to the virtual robot is determined. The bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with the virtual robot remain unchanged. Then, for each mesh vertex, the azimuth weight and edge attenuation of each virtual environment camera relative to the mesh vertex are determined. The weighting process involves reducing the azimuth weight to characterize the deviation of the optical axis center of each virtual environment camera from the grid vertex; the edge attenuation weight to characterize the image sampling quality at the projection position of the grid vertex on the screen of each virtual camera; then, for each virtual environment camera, the camera fusion weight at the grid vertex is determined based on the azimuth weight and the edge attenuation weight, and panoramic grid vertex data is generated based on the multiple camera fusion weights corresponding to each grid vertex; finally, multiple real environment images are acquired through the multiple physical environment cameras, and the panoramic grid vertex data is used to fuse the multiple real environment images to obtain a panoramic image.
[0025] In this embodiment, camera fusion weights are determined based on azimuth weights and edge attenuation weights, thereby generating panoramic mesh vertex data. Since the azimuth weight is used to characterize the angular deviation between the optical axis of each virtual environment camera and the mesh vertex, it can reflect the degree of "sovereignty" of the virtual environment camera over that vertex. Therefore, the azimuth weight can be used to achieve a smooth transition in overlapping areas. At the same time, the edge attenuation weight is used to characterize the image sampling quality at the projection position of the mesh vertex on the screen of each virtual camera. Therefore, the edge attenuation weight can be used to suppress areas with severe edge distortion in fisheye images. Furthermore, the panoramic mesh vertex data is used to fuse the real environment images to obtain a panoramic image. In this way, while achieving a seamless fusion effect between real environment images, the quality of the fused image can also be improved.
[0026] To facilitate understanding of this embodiment, a method for generating a surround view image of a robot disclosed in this disclosure will first be described in detail. The subject executing the method for generating a surround view image of a robot provided in this disclosure is generally a quadruped robot.
[0027] In some embodiments, the executing entity may also be a computer device, which can be a server. This server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. In other embodiments, the computer device may also be a terminal device, which can be a mobile device, terminal, handheld device, computing device, vehicle-mounted device, etc.
[0028] In other embodiments, the method can also be applied to an implementation environment consisting of computer equipment and servers, or an implementation environment consisting of terminal equipment and servers. Furthermore, this method for generating a surround view image for a robot can also be implemented by a processor calling computer-readable instructions stored in memory.
[0029] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0030] Please see the appendix Figure 1 The above is a flowchart illustrating a method for generating a surround view image of a robot, as shown in an exemplary embodiment of this application. Figure 1 As shown, the method for generating a surround view image of a robot in this embodiment of the present disclosure may include the following steps S101 to S105: S101: Model the real robot to obtain a virtual robot; the real robot is equipped with multiple real environment cameras, and the viewpoints of adjacent real environment cameras overlap; the virtual robot has virtual environment cameras that correspond one-to-one with the multiple real environment cameras.
[0031] In this embodiment, the real robot is a quadruped robot that can operate in a real environment. In other embodiments, the real robot may also be a bipedal robot or a hexaped robot, etc., which is not limited here.
[0032] It is understandable that a real robot is equipped with multiple real-world cameras, each with a different shooting angle, and the perspectives of adjacent real-world cameras overlap.
[0033] In this embodiment, the real-world camera is a fisheye camera. In other embodiments, the real-world camera can also be other types of cameras, such as a wide-angle camera, which is not limited here.
[0034] Here, a virtual robot refers to a three-dimensional digital model of a real robot, which has virtual environment cameras that correspond one-to-one with multiple real environment cameras.
[0035] It should be noted that the number of cameras in the real environment can be set according to actual needs, such as four or six.
[0036] Please see Figure 2 This is a schematic diagram illustrating the distribution of real-world cameras for a real robot, provided as an exemplary embodiment of this application. Figure 2 As shown, the real robot is equipped with four real environment cameras, including cam0 located in front of the robot body (+X-axis direction), cam1 located to the right of the robot body (+Y-axis direction), cam2 located behind the robot body, and cam3 located to the left of the robot body. The camera field of view of each real environment camera is approximately 190 degrees, and there is overlap between the camera fields of view of two adjacent real environment cameras.
[0037] S102: Determine the bowl-shaped surface model corresponding to the virtual robot; the bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with respect to the virtual robot remain unchanged.
[0038] Here, the bowl-shaped curved surface model can refer to a three-dimensional model that wraps around the virtual robot. Since the quadruped robot mainly focuses on the ground under its feet and the surrounding obstacles, and does not need to look at the sky above its head, the model presents a closed bottom, an open top, and an inward concave shape.
[0039] The bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with respect to the virtual robot remain unchanged. Thus, no matter how the virtual robot moves, the bowl-shaped surface model is fixed relative to the virtual robot.
[0040] S103: For each grid vertex, determine the azimuth weight and edge attenuation weight of each virtual environment camera relative to the grid vertex; the azimuth weight is used to characterize the degree of deviation of the optical axis center of each virtual environment camera relative to the grid vertex; the edge attenuation weight is used to characterize the image sampling quality at the projection position of the grid vertex on the screen of each virtual camera.
[0041] In this step, the azimuth weight is used to characterize the degree of deviation of the optical axis center of each virtual environment camera from the grid vertex. That is, if a grid vertex is directly in front of the camera (i.e., the optical axis center), then the pixel of that grid vertex on the camera screen of the virtual environment camera is the clearest and has the highest weight. Conversely, if a grid vertex is at the edge of the camera screen, then the pixel is distorted and the weight is reduced.
[0042] Edge attenuation weights are used to characterize the image sampling quality at the projection position of the grid vertices on each virtual camera screen. In other words, they are used to reduce the impact of areas located at the edge of the camera screen or with poor sampling quality due to occlusion on the panoramic image.
[0043] S104: For each virtual environment camera, determine the camera fusion weight of the virtual environment camera at the grid vertex according to the azimuth weight and the edge attenuation weight, and generate panoramic grid vertex data based on the multiple camera fusion weights corresponding to each grid vertex.
[0044] Here, for each virtual environment camera, the camera fusion weight at the grid vertex is determined based on the azimuth weight and the edge attenuation weight. Specifically, the product of the azimuth weight and the edge attenuation weight can be used as the camera fusion weight at the grid vertex. Then, based on the multiple camera fusion weights corresponding to each grid vertex, panoramic grid vertex data is generated.
[0045] In some implementations, for each mesh vertex, the mesh vertex can be projected onto the camera screen of each virtual environment camera to obtain the camera texture coordinates corresponding to the mesh vertex. Specifically, the Unified Camera Model (UCM) can be used to project the three-dimensional coordinates of each mesh vertex to obtain multiple camera texture coordinates corresponding to the mesh vertex.
[0046] For details, please see Figure 3 The flowchart provided in this application illustrates the projection process of a unified camera model, as follows: Figure 3 As shown, the projection process specifically includes the following steps: for each mesh vertex 3D coordinates Using camera transformation matrices (including rotation matrices) Translation matrix Perform a coordinate transformation on the mesh to obtain the three-dimensional coordinates of the mesh vertex in the camera coordinate system. For details, please refer to formula (1): (1) Then, the 3D coordinates of the mesh vertex in the camera coordinate system are normalized and projected using UCM to obtain the projected coordinates. As shown in formulas (2) to (4): (2) (3) (4) Obtain the projected coordinates Then, radial distortion correction is performed on the projected coordinates to simulate the optical distortion of a fisheye lens, resulting in corrected coordinates. As shown in formulas (5) to (7): (5) (6) (7) Obtain the corrected coordinates Then, the correction coordinates are converted into image pixel coordinates. As shown in formulas (8) to (9): (8) (9) in, , Focal length , These are the coordinates of the image center.
[0047] pixel coordinates Mapped to camera texture coordinates As shown in formulas (10) to (11): (10) (11) in, For the width of the camera screen, This is the height of the camera screen.
[0048] In some implementations, during image fusion, the GPU rasterizer performs hardware linear interpolation of the mesh vertex attributes inside the triangle. If the UV coordinates of a vertex at the boundary abruptly change to (0,0), the interpolation result inside the triangle will be "dragged" towards (0,0), resulting in obvious stretching artifacts. Therefore, in this implementation, after determining the camera texture coordinates, if the camera texture coordinates exceed the camera screen boundary, the camera fusion weight corresponding to that mesh vertex is set to 0, and the camera texture coordinates are clamped to the range of [0,1] to avoid interpolation artifacts.
[0049] Furthermore, after obtaining the multiple camera texture coordinates corresponding to each grid vertex, in step S104, when generating panoramic grid vertex data based on the multiple camera fusion weights corresponding to each grid vertex, vertex data can be generated for each grid vertex based on the three-dimensional coordinates of the grid vertex, the multiple camera fusion weights corresponding to the grid vertex, and the multiple camera texture coordinates corresponding to the grid vertex. Then, the panoramic grid vertex data is generated based on the vertex data of each grid vertex.
[0050] Specifically, please refer to Table 1, which shows the content of the vertex data for each grid vertex in the panoramic mesh vertex data. Each vertex data contains 17 floating-point numbers.
[0051] Table 1
[0052] In this context, the parameterized UV coordinates refer to the position of the mesh vertex on the "structure" of the bowl-shaped surface.
[0053] In some implementations, the above-mentioned panoramic mesh vertex data can be serialized to obtain a binary file, which includes a file header, vertex data of each mesh vertex, index data, file size and other information.
[0054] The file header includes a magic number representing the file type, file version number, total number of grid vertices, total number of triangle indices, number of grid radial sampling points, number of grid azimuth sampling points, and number of environment cameras. Vertex data for each grid vertex is shown in Table 1 and will not be elaborated upon here. The index data includes indices for multiple triangles, stored in row-major order, with each pair of triangles forming a quadrilateral. Since the azimuth direction is a closed loop, modulo wrapping is used on the column indices during index generation.
[0055] Here, the index data can be generated through the following steps: obtain the preset radial sampling point number and preset azimuth sampling point number of the bowl-shaped surface model, where the mesh vertices are stored in row-major order, and the row number is... Corresponding radial direction, column number Iterate through each quadrilateral, and the four vertices of the quadrilateral are located at the i-th vertex and j-th vertex. line, number line, number Column and number If the column has four vertices, then the indices of the four vertices are... , ,in, , , Divide the quadrilateral into two triangles and generate the corresponding triangle indices: First triangle: The vertices are arranged in counter-clockwise order; the second triangle: The vertices are arranged in counterclockwise order, and all triangle indices are stored sequentially to form index data, which is used by the GPU to draw the bowl-shaped surface mesh.
[0056] It should be noted that the generation process of the panoramic mesh vertex data is completed offline. During the actual operation of the real robot, the graphics processing unit (GPU) only performs rasterization, texture sampling, and weighted fusion, which can improve the fusion efficiency and rendering efficiency.
[0057] S105: Acquire multiple real-world environment images using the multiple physical environment cameras, and fuse the multiple real-world environment images using the panoramic mesh vertex data to obtain a panoramic image.
[0058] Once the panoramic mesh vertex data is obtained, it can be used to generate a panoramic image of the physical robot during its operation.
[0059] Specifically, multiple physical environment cameras installed on a physical robot can be used to acquire images of the real environment, and then panoramic mesh vertex data can be used to fuse the multiple real environment images to obtain a panoramic image.
[0060] Here, each fisheye camera input is in NV12 format (4:2:0 chroma subsampling), and is represented on the GPU using two texture planes: Y-plane texture: GL_R8 format, full resolution (1920×1536), storing luminance; UV-plane texture: GL_RG8 format, half resolution (960×768), storing chroma.
[0061] Converting YUV format to RGB format uses the BT.601 coefficient, as shown in formula (12): (12) After acquiring the real-world image, it needs to be converted from NV12 format to RGB format for subsequent fusion processing.
[0062] In this embodiment, since the azimuth weight is used to characterize the angular deviation between the optical axis of each virtual environment camera and the grid vertex, it can reflect the degree of "sovereignty" of the virtual environment camera over that vertex. Therefore, the azimuth weight can be used to achieve a smooth transition in overlapping areas. Meanwhile, the edge attenuation weight is used to characterize the image sampling quality at the projection position of the grid vertex on the screen of each virtual camera. Therefore, the edge attenuation weight can be used to suppress areas with severe edge distortion in fisheye images. Furthermore, the camera fusion weight is determined based on the azimuth weight and the edge attenuation weight, thereby generating panoramic grid vertex data. The panoramic grid vertex data is then used to fuse the real environment images to obtain a panoramic image. In this way, while achieving a seamless fusion effect between real environment images, the quality of the fused image can also be improved.
[0063] Furthermore, in the virtual robot model, the fusion weights (composed of azimuth weights and edge attenuation weights) of each virtual environment camera are pre-calculated for each grid vertex on the bowl-shaped surface to generate "panoramic grid vertex data". When the physical robot is running, it can directly use this pre-calculated data to perform image fusion without the need for real-time weight calculation, which greatly reduces the computational burden of the embedded platform.
[0064] The following section provides a detailed description of the construction process of the bowl-shaped surface model in step S102.
[0065] In this embodiment, the bowl-shaped curved surface model is in the polar coordinate system. The surface is defined as consisting of a flat ground region and a parabolic bowl wall region, and its shape is described by multiple configurable parameters. Specifically, it includes steps (1) to (2): (1) Obtain the model construction parameters of the bowl-shaped surface model; the model construction parameters include the flat radius, the maximum radial distance, the maximum height of the bowl wall and the parabolic curvature coefficient.
[0066] Here, in response to the user's parameter input, the model construction parameters of the bowl-shaped surface model can be obtained, including the flat radius, maximum radial distance, maximum height of the bowl wall, and parabolic curvature coefficient.
[0067] Optionally, the parameter values in the model building parameters can be adapted according to the actual size of the physical robot (the virtual robot's size is the same as its actual size). This allows for flexible adaptation to quadruped robots of different sizes; for example, for small robots, the flat radius can be adjusted. Setting it to 0.1m allows the bowl wall to start climbing earlier, while larger robots can be set to 2.0m to retain a larger flat area.
[0068] The parameters mentioned above are stored in a configuration file (such as a YAML configuration file). When performing parameter operations (such as deletion, modification, or addition) on any parameter, only the content of the configuration file needs to be updated, without modifying the code.
[0069] (2) Establish a polar coordinate system with the center of the virtual robot as the origin, and generate a bowl-shaped surface model with the center of the virtual robot as the origin based on the flat radius, maximum radial distance, maximum height of the bowl wall and parabolic curvature coefficient in the polar coordinate system.
[0070] It can be understood that the bowl-shaped curved surface model consists of a flat ground region and a bowl-wall region, wherein the radial distance of the flat ground region satisfies and , The flat radius is the flat area of the ground region, and the height of the flat ground region is 0. This flat ground region corresponds to the ground directly below and close to the virtual robot, and the zero height is maintained to faithfully reproduce the ground details.
[0071] The bowl wall region (or parabolic bowl wall region) refers to the radial distance. The region satisfies the mapping relationship shown in formula (13): (13) in, alpha It is the parabolic curvature coefficient, used to control the degree of curvature of the bowl wall; z _max is the maximum height of the bowl wall, used to limit the depth of the bowl.
[0072] Based on the above, the Cartesian coordinates of the bowl-shaped surface model can be obtained, as shown in formula (14): (14) in, ,here, Corresponding to the front of the virtual robot (or the +x axis direction). This corresponds to the left side (or +y direction) of the virtual robot.
[0073] like Figure 4 The image shown is a top view of a bowl-shaped curved surface model provided in an exemplary embodiment of this application. Figure 4 As shown, the black square represents the virtual robot, the green circular area is the flat ground area, and the purple area is the bowl wall area. The front of the virtual robot corresponds to the +X axis direction. The virtual robot is directly behind it in the -X axis direction. The virtual robot's left side corresponds to the +Y axis direction ( The virtual robot's right side corresponds to the -Y axis direction ( ).
[0074] like Figure 5 The image shown is a side view of a bowl-shaped curved surface model provided in an exemplary embodiment of this application. Figure 5 As shown, the horizontal axis represents the radial distance. rho (Unit: meters), the vertical axis is the height z (unit: meters). In the figure, the black square represents the virtual robot, the blue curved surface represents the bowl-shaped curve, and half of the green line segment represents the flat radius. (2 meters), the blue double-arrow solid line represents the maximum radial distance ( (6 meters), maximum height of the bowl wall ( (3 meters) and parabolic curvature coefficient ( alpha (0.25).
[0075] In some implementations, to prevent the bowl wall from exceeding its maximum height... z It later transforms into a "flat-topped hat" shape (causing UV flipping and rendering artifacts), and when generating the bowl-shaped curved surface model, the automatic calculation of the bowl wall just reaches the desired shape. The radial distance is shown in formula (15): (15) in, The bowl wall just reaches radial distance, The radical sign is √.
[0076] Furthermore, when generating a bowl-shaped surface model with the virtual robot's center as the origin based on the flat radius, maximum radial distance, maximum bowl wall height, and parabolic curvature coefficient, the maximum radial distance can be compared with a preset radial distance first. If the maximum radial distance is greater than the preset radial distance, then a bowl-shaped surface model with the virtual robot's center as the origin can be generated based on the flat radius, preset radial distance, maximum bowl wall height, and parabolic curvature coefficient.
[0077] The preset radial distance is the distance that the bowl wall in formula (15) just reaches. radial distance It is determined based on the flat radius, the maximum height of the bowl wall, and the parabolic curvature coefficient.
[0078] In other words, if the maximum radial distance configured by the user is greater than the preset radial distance, the maximum radial distance will be automatically clamped to the preset radial distance. In this way, the bowl wall can be a monotonically increasing parabolic segment, and no flat-top area will appear.
[0079] In other implementations, if the maximum radial distance configured by the user is greater than the preset radial distance, a corresponding log can be generated to prompt the user to modify the maximum radial distance.
[0080] In some implementations, step S103, when determining the azimuth weight of each virtual environment camera relative to the grid vertex, includes the following steps (I) to (IV): (I) Determine the vertex azimuth angle of the grid vertex on the two-dimensional plane based on the three-dimensional coordinates of the grid vertex.
[0081] Here, the vertex azimuth angle refers to the orientation angle of the mesh vertex on the two-dimensional horizontal plane (XOY plane) with the virtual robot as the origin, as shown in formula (16): (16) in, The azimuth angle of the grid vertex on the two-dimensional plane.
[0082] (II) For each virtual environment camera, determine the angular deviation of the vertex azimuth angle relative to the optical axis direction of the virtual environment camera.
[0083] The projection direction of the optical axis of each virtual environment camera onto the horizontal plane is denoted as... For example, the front camera is 0 degrees, the right camera is 90 degrees, the rear camera is 180 degrees, and the left camera is 270 degrees.
[0084] As shown in formula (17), the expression for the angle deviation is: (17) in, For angular deviation, To take the absolute value, For virtual environment cameras c In the direction of the optical axis on the horizontal plane, To normalize the included angle to .
[0085] It is understandable that the smaller the angular deviation, the closer the grid vertices are to the front of the camera in the virtual environment (optical axis direction).
[0086] Please see Figure 6 This is a top view of an azimuth sector division provided as an exemplary embodiment of this application. For example... Figure 6 As shown, the center of the circle is the virtual robot, the sector corresponding to each virtual environment camera is 90 degrees, and the line connecting the four dots to the virtual robot is the optical axis direction of the virtual environment camera. In this way, for each grid vertex, the angular deviation of each virtual environment camera relative to that grid vertex can be determined.
[0087] (III) Based on the number of multiple virtual environment cameras, determine the half-sector angle covered by each virtual environment camera, and based on the half-sector angle and the preset feathering transition width, determine the deviation angle threshold.
[0088] Wherein, the half-sector angle is half of the theoretical coverage area of each virtual environment camera, obtained by dividing the virtual environment camera area into 360 degrees evenly, denoted as... For example, when the number of virtual environment cameras is 4, the half-sector angle is 45 degrees.
[0089] The preset feathering transition width refers to the configurable half-width of the feathering band, denoted as... (In this embodiment, the preset feathering transition width is 5 degrees). This is used to create a smooth transition area on both sides of the sector boundary, avoiding hard seams.
[0090] (IV) Determine the azimuth weight based on the angle deviation and the deviation angle threshold.
[0091] The deviation angle threshold includes a first angle threshold and a second angle threshold.
[0092] The first angle threshold is determined based on the difference between the half-sector angle and the feathering transition width, as shown in formula (18): (18) The second angle threshold is determined based on the sum of the half-sector angle and the feathering transition width, as shown in formula (19): (19) The first angle threshold and the second angle threshold constitute the transition angle range.
[0093] If the angle deviation is not greater than the first angle threshold, the azimuth weight is determined to be 1.
[0094] If the angle deviation is greater than the second angle threshold, the azimuth weight is 0.
[0095] If the angle deviation is greater than the first angle threshold and not greater than the second angle threshold (that is, the angle deviation is in the transition angle range), the azimuth weight is determined using the smooth step function based on the angle deviation, the half-sector angle and the preset feathering transition width, as shown in formulas (20) to (21): (20) (twenty one) in, As azimuth weight, As the transition factor, the smoothing step function is: .
[0096] Please see Figure 7 This is a schematic diagram of weight transition at a sector boundary, provided as an exemplary embodiment of this application. Figure 7 As shown, the horizontal axis represents the azimuth angle, which gradually changes from 30 degrees to 60 degrees, covering the sector boundaries of cam0 and cam1. The vertical axis represents the fusion weight, the dashed line represents the sector boundary, and the yellow band represents the transition area.
[0097] The red solid line represents the fusion weight of cam0, indicating the proportion of contribution of the color in the image captured by cam0 to the fused color at the corresponding angle. The green solid line represents the fusion weight of cam1, indicating the proportion of contribution of the color in the image captured by cam1 to the fused color at the corresponding angle. The sum of the fusion weights of cam0 and cam1 is always 1. The transition area shows the change of fusion weights from 40 degrees to 60 degrees. At 45 degrees, the fusion weight is 1. From 40 degrees to 45 degrees, the fusion weight of cam0 gradually decreases from 1 to 0.5, while the fusion weight of cam1 gradually increases from 0 to 0.5. From 45 degrees to 60 degrees, the fusion weight of cam0 rapidly decreases from 0.5 to 0, while the fusion weight of cam1 increases from 0.5 to 1. This avoids the stitching gap caused by direct switching at the sector boundary (45 degrees).
[0098] In some implementations, step S103, when determining the edge attenuation weight of each virtual environment camera relative to the mesh vertex, includes the following steps (i) to (ii): (i) For each virtual environment camera, determine the minimum distance between the camera texture coordinates corresponding to the mesh vertex and each screen boundary of the camera screen based on the camera texture coordinates corresponding to the mesh vertex and the camera screen size of the virtual environment camera.
[0099] Here, camera screen size refers to the image width of the image captured by the camera in the virtual environment. W and height H .
[0100] As shown in formula (22), the expression for determining the minimum distance is: (twenty two) in, The minimum distance.
[0101] It is understandable that the smaller the minimum distance value, the closer the projection point is to the screen boundary, the more severe the fisheye distortion, and the worse the image quality.
[0102] (ii) Determine the edge attenuation weight based on the minimum distance.
[0103] Here, if the minimum distance indicates that the camera texture coordinates are not within the screen boundary, the edge attenuation weight is determined to be 0. That is, when the projection point is located within the camera screen and is far enough away from the boundary, the weight is 1, and the weight gradually decreases to 0 when it is close to the boundary.
[0104] If the minimum distance indicates that the camera texture coordinates are within the camera screen, the edge attenuation weight is determined based on the minimum distance and the preset edge feathering width.
[0105] Specifically, if Then calculate the normalization factor. (The value range is (0,1)), and then a quadratic decay function is used: edge decay weight. Thus, the closer to the boundary ( The smaller the value, the faster the weight decreases, effectively suppressing edge distortion.
[0106] like Then the edge decay weight This indicates that the projected point is located in a high-quality area inside the camera screen, requiring no attenuation.
[0107] It should be noted that, under normal circumstances, the size of the camera screen of the virtual environment camera is the same as the size of the camera image of the virtual environment camera. Therefore, it can be considered that the attenuation of points near the edge of the camera image is more severe.
[0108] Based on the above calculations, the edge attenuation weight of each virtual environment camera relative to each grid vertex is calculated.
[0109] Thus, after obtaining the azimuth weights of the virtual environment camera relative to the grid vertices and the edge attenuation weights, the initial fusion weights can be determined based on the azimuth weights and edge attenuation weights, as shown in formula (23): (twenty three) in, These are the initial fusion weights.
[0110] Then, the initial fusion weights of each virtual environment camera are normalized to obtain the fusion weights of each virtual environment camera at the grid vertex.
[0111] The sum of the fusion weights of each virtual environment camera at the grid vertex is 1.
[0112] Please refer to formula (24): (twenty four) like If the value is greater than 1, then alignment and normalization are performed, as shown in formula (25): (25) in, .
[0113] In this embodiment, the normalization process ensures that regardless of how many cameras cover a certain mesh vertex, the final total color contribution is always 1.0.
[0114] In some implementations, regarding step S105, when fusing the multiple real-world environment images using the panoramic mesh vertex data to obtain a panoramic image, please refer to [link to relevant documentation]. Figure 8 This includes the following steps S1051~S1053: S1051: Obtain the current posture data and current view control parameters of the real robot; the current view control parameters include at least one of azimuth rotation, pitch rotation, translation, and zoom distance.
[0115] In this step, the current attitude data of the real robot can be obtained in real time through sensors such as the inertial measurement unit (IMU) installed on the real robot. The current attitude data includes at least the pitch angle, roll angle and height off the ground of the real robot's torso relative to the horizontal plane. The current attitude data can reflect the current spatial tilt and vertical position of the real robot.
[0116] The current view control parameters can be user-input parameters used to dynamically adjust the pose of the virtual camera (observer) to control the viewing angle and range of the image. Here, the current view control parameters include azimuth rotation (rotating left and right around the vertical axis), pitch rotation (tilting up and down around the horizontal axis), translation (translating the image horizontally or vertically), and zoom distance (such as zooming in or out of the view).
[0117] S1052: Determine the model view projection matrix based on the current posture data of the real robot and the current view control parameters.
[0118] In computer graphics, the core matrix for transforming a 3D object to 2D screen coordinates is obtained by multiplying the model matrix, view matrix, and projection matrix. The model matrix is used to transform the bowl-shaped surface from the local coordinate system to the world coordinate system (following the robot's posture); the view matrix is constructed from viewpoint control parameters, determining the observer's position and orientation; and the projection matrix defines the near and far planes and the field of view of the perspective projection.
[0119] Please refer to formulas (26) to (28) for the expressions of the model-view projection matrix: (26) in, This is the field of view (in degrees), with a default of 15 degrees. This parameter determines the width of the virtual camera's field of view. The smaller the size, the narrower the field of view, and the more the image resembles a telephoto lens; The larger the lens, the wider the field of view, and the more the image resembles a wide-angle lens. The aspect ratio of the screen ( The resolution of the projected image is typically the same as the screen resolution of the display device or the target image being rendered, and is used to ensure that the projected image is not stretched or distorted. This is the distance to the nearest clipping plane (greater than 0). Objects smaller than this distance will be clipped and not included in the rendering. It is usually set to a small positive number (e.g., 0.1 meters) to avoid incorrectly clipping nearby objects. This is the distance to the clipping plane. Objects larger than this distance will be clipped. It should be large enough to encompass the entire bowl-shaped surface and the 3D model.
[0120] (27) in, This is the azimuth rotation amount (around the Z-axis), controlling left and right rotation. This is the pitch angle rotation (around the X-axis), controlling the vertical pitch. , This refers to the XY plane translation amount, which controls the image translation. The scaling distance (or Z-axis distance) controls the scaling range (negative values indicate that the camera is above the bowl).
[0121] (28) Please see Figure 9 This is a schematic diagram illustrating view control as provided in an exemplary embodiment of this application. Figure 9 As shown, a view control coordinate system is established with the origin as the center. The origin is the bottom center of the bowl-shaped surface (i.e., the projection of the robot's center point on the ground, usually coinciding with the robot's center). The X-axis (front) points forward of the robot (positive direction), the Y-axis (left) points to the left of the robot (positive direction), and the Z-axis (up) points vertically upward (positive direction). The observer represents the position of the virtual camera. (Azimuth): Indicates the arc of rotation around the Z-axis; (Pitch angle): Indicates the arc of rotation about the X-axis; (Observation distance): Marked on the line connecting the camera and the center of the bowl bottom, indicating the vertical distance; , This refers to the translation along the X and Y axes.
[0122] S1053: Using the model view projection matrix and the panoramic mesh vertex data, the multiple real environment images are fused to obtain the panoramic image.
[0123] Finally, using the MVP matrix calculated in step S1052 and the pre-stored panoramic mesh vertex data, real-time rendering is performed.
[0124] Specifically, the 3D coordinates of each mesh vertex are first transformed to screen space using the MVP matrix to obtain the screen display coordinates; then the screen coordinates are rasterized to generate multiple fragments, and the texture coordinates and blending weights of each fragment are obtained by interpolation.
[0125] Here, screen space refers to the screen space of the corresponding display device, which is used to display panoramic images, and users can use the display device to control the view and motion (such as moving forward and backward) of the physical robot.
[0126] Screen display coordinates refer to the screen space coordinates obtained after the mesh vertices undergo MVP transformation (usually...). ,in For pixel position, (where the coordinates are depth values) These screen display coordinates determine the pixel positions that the mesh vertices should be drawn on the display device screen.
[0127] Rasterization is the process by which the GPU converts triangular primitives into fragments of screen pixels. Based on the screen display coordinates of the three vertices of the triangle, it determines which screen pixels are covered by the triangle and generates a fragment for each covered pixel. A fragment is a "shading unit" generated by rasterization, containing the corresponding screen pixel coordinates, depth value, and interpolated texture coordinates, blending weights, and other attributes. After processing by the fragment shader, the fragment is finally written to the framebuffer.
[0128] Within a fragment, texture coordinates are obtained by interpolating the barycentric coordinates of the texture coordinates at the three vertices of the triangle. Each fragment has a texture coordinate for each real-world camera, used to sample color from the environment image of that real-world camera.
[0129] The blending weights are also obtained through interpolation. Each fragment has a weight for each real-world camera (with a value range of [0,1], and the sum of the weights of all real-world cameras is 1). This weight determines the proportion of the color of that real-world camera in the final blend.
[0130] Based on the above, the rasterizer acquires triangles formed by screen display coordinates and iterates through all screen pixels covered by these triangles. For each covered pixel, a fragment is generated. Simultaneously, using the texture coordinates of the three vertices of the triangle and the blending weights, multiple texture coordinates (one for each real-world camera) and blending weights (one for each real-world camera) of the fragment are calculated via barycentric coordinate interpolation. This interpolation ensures continuous attribute changes for pixels within the triangle, thus achieving smooth texture mapping and color blending.
[0131] Next, for each fragment, texture sampling is performed on the multiple real environment images according to the texture coordinates corresponding to the fragment to obtain multiple pixel colors, and the multiple pixel colors are fused according to the fusion weights corresponding to the fragment to obtain a fused pixel color.
[0132] After receiving the interpolated texture coordinates and blending weights, the fragment shader performs texture sampling on the environment image captured by each real-world camera: it reads the pixel color from the real-time environment image using the texture coordinates of that camera. Then, it weights and sums the color values sampled from each real-world camera according to their corresponding blending weights to obtain the final blended color of the fragment. If the weight of a real-world camera is 0, sampling can be skipped to improve performance. The final result is an RGB color value representing the correct color of that screen pixel in the panoramic image.
[0133] The blended color calculated by the fragment shader is passed to the rendering pipeline, and the color value is written to the screen pixel position corresponding to the fragment in the frame buffer. After all triangles have been processed, the initial bowl-shaped surface toroidal image stored in the frame buffer is obtained.
[0134] Optionally, when rendering the screen pixels corresponding to the fragment based on the fused pixel color corresponding to each fragment to obtain the surround view image, the screen pixels corresponding to the fragment can be rendered based on the fused pixel color corresponding to each fragment to obtain the initial surround view image.
[0135] This is the output stage of the first rendering pass. After rasterization and fragment shading of all triangles on the bowl-shaped surface are completed, each fragment carries a blended color value. This color value is written to the screen pixel position corresponding to the fragment in the frame buffer, and at the same time, the depth value of each fragment (generated by the MVP transform) is written to the depth buffer.
[0136] Once all triangles have been processed, a complete initial bowl-shaped surface panorama image is formed in the frame buffer. At this point, the image only contains environmental information (ground, obstacles, and distant objects) and does not yet include the robot's own model.
[0137] Next, the three-dimensional model of the virtual robot and its depth information are obtained. Based on the depth information of the three-dimensional model, the three-dimensional model is rendered onto the initial surround view image to obtain a surround view image containing the three-dimensional model and its surrounding environment.
[0138] Here, a pre-built 3D model of a quadruped robot can be loaded, and the vertex data, textures, and lighting parameters required for its rendering can be prepared. The depth information of the 3D model is not stored in advance, but is dynamically generated during the rendering process: after the model undergoes an MVP transformation (sharing the same view and projection matrix as the first pass, but with the model matrix placing the robot model at the center of the bowl-shaped surface) and is rasterized, each model fragment calculates its own depth value. These depth values are used to perform a depth test against the existing bowl-shaped surface depth values in the depth buffer to determine whether the model is above or below the bowl-shaped surface.
[0139] Finally, the second rendering pass is performed, reusing the depth buffer from the first rendering pass (which already stores the depth values of the bowl-shaped surface) to draw the 3D model: For each fragment of the model, the GPU compares its depth value with the existing depth value of the corresponding pixel in the depth buffer. If the model's depth value is smaller (i.e., the model surface is closer to the virtual camera), the fragment's color (calculated by lighting) will override the pixel's original environment color; conversely, if the model's depth value is larger (the model is behind the bowl surface), it is discarded, preserving the original environment color. In this way, without manually handling occlusion relationships, the system automatically achieves the correct visual effect: the robot's legs may "step" on the bowl surface, while its body protrudes above it. Ultimately, the image in the frame buffer becomes a complete panoramic image containing the robot model and its surrounding environment.
[0140] Please see Figure 10 This is a schematic diagram illustrating a rendering process provided in an exemplary embodiment of this application. Figure 10 As shown, the rendering pipeline is implemented based on a framebuffer object (FBO), which includes a color buffer (RGBA) and a depth buffer. The framebuffer is cleared before each rendering pass. The two rendering passes share a depth buffer, thus achieving the correct occlusion relationship between the bowl-shaped curved background and the foreground of the 3D model.
[0141] The first rendering pass is used to implement the panoramic rendering of the bowl-shaped surface, and specifically includes the following process: (1) Status settings: GL_CULL_FACE = ON: Enables face culling, culling back face triangles. In other words, "Normal mode culling back face triangles" means only triangles facing the virtual environment camera are rendered.
[0142] GL_DEPTH_TEST = ON: Enable depth testing to ensure the correct occlusion relationships of the bowl surface (e.g., distant bowl walls are occluded by nearby bowl walls). (2) Vertex shader: u_RenderMode = 1: Used to distinguish the rendering logic of Pass 1 and Pass 2 (e.g., no lighting is needed in Pass 1).
[0143] MVP Projection to Screen: Transforms the mesh vertices of the bowl-shaped curved surface model to screen space using the MVP matrix, outputting screen coordinates and depth values.
[0144] (3) Fragment shader: 4-channel NV12 texture sampling: Image data in NV12 format is acquired from 4 cameras (Y plane and UV plane are separated).
[0145] YUV → RGB + weighted blending: First, the NV12 sample values are converted to RGB colors using BT.601 coefficients.
[0146] Then, based on the pre-calculated and interpolated fusion weights, the RGB colors of each camera are weighted and summed to obtain the final fragment color.
[0147] Output: RGBA colors are written to the color buffer of the FBO, and depth values are written to the depth buffer (reserved for use by the second rendering channel).
[0148] The second rendering pass is used to render the 3D model, and the specific process includes the following steps: (1) State reuse Depth buffer sharing: Pass 2 directly uses the depth buffer output by Pass 1 without needing to clear it.
[0149] Correct occlusion with the bowl surface: This is automatically achieved through depth testing, comparing the fragment depth of the 3D model with the depth of the bowl surface.
[0150] (2) Vertex shader: MVP + Model Matrix: Here, the model matrix transforms the robot model from the local coordinate system to the world coordinate system (usually with the robot's center as the origin, consistent with the bowl-shaped surface). The view matrix and projection matrix are the same as in Pass 1, resulting in the MVP matrix.
[0151] (3) Fragment shader The Blinn-Phong lighting model can specifically include calculating the ambient, diffuse, and specular components, giving the 3D model a sense of depth and realism.
[0152] (4) Output The color of the model fragment (calculated by lighting) is written to the color buffer of the FBO, overwriting the original bowl surface color (at the position where the depth test passed).
[0153] Finally, the FBO stores a complete RGBA image of the bowl-shaped panoramic background and the robot model.
[0154] When sending an image to a display device for display, an RGBA image can be read from the FBO, then the RGBA image can be converted to NV12 format (for easy video encoding), and then pushed to the display device (or remote operation terminal) for display.
[0155] Since the texture coordinates and weights are pre-stored in the panoramic mesh vertex data, only simple texture sampling and weighted summation are required at runtime, resulting in minimal computation and enabling real-time frame rates on embedded platforms.
[0156] In addition, the design of two rendering channels allows the environment background and foreground models to be rendered independently and automatically merged through depth buffering, which not only ensures rendering efficiency but also provides realistic 3D occlusion relationships, providing remote operators with an intuitive reference for the robot's own pose.
[0157] Corresponding to the aforementioned embodiments of the method for generating a surround view image of a robot, this application also provides embodiments of an apparatus for generating a surround view image of a robot.
[0158] Please see Figure 11 This is a schematic diagram illustrating the structure of a device for generating a surround view image of a robot, as shown in an exemplary embodiment of this application. Figure 11 As shown, the device 1100 for generating a surround view image of a robot includes: The robot modeling module 1110 is used to model a real robot to obtain a virtual robot; the real robot is equipped with multiple real environment cameras, and the viewpoints between adjacent real environment cameras overlap; the virtual robot has virtual environment cameras that correspond one-to-one with the multiple real environment cameras. The bowl-shaped model determination module 1120 is used to determine the bowl-shaped surface model corresponding to the virtual robot; the bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with respect to the virtual robot remain unchanged; The weight determination module 1130 is used to determine, for each grid vertex, the azimuth weight and edge attenuation weight of each virtual environment camera relative to the grid vertex; the azimuth weight is used to characterize the degree of deviation of the optical axis center of each virtual environment camera relative to the grid vertex; the edge attenuation weight is used to characterize the image sampling quality at the projection position of the grid vertex on the screen of each virtual camera. The vertex data generation module 1140 is used to determine the camera fusion weight of the virtual environment camera at the grid vertex for each virtual environment camera according to the azimuth weight and the edge attenuation weight, and generate panoramic grid vertex data based on the multiple camera fusion weights corresponding to each grid vertex. The environmental image fusion module 1150 is used to acquire multiple real environment images through the multiple physical environment cameras, and to fuse the multiple real environment images using the panoramic mesh vertex data to obtain a panoramic image.
[0159] Please see Figure 12 This is a schematic diagram illustrating the structure of another apparatus for generating a surround view image of a robot, as shown in an exemplary embodiment of this application. Figure 12 As shown, the device 1100 for generating a surround view image of a robot also includes a texture coordinate generation module 1160; each mesh vertex has three-dimensional coordinates; the texture coordinate generation module 1160 is used for: For each mesh vertex, the mesh vertex is projected onto the camera screen of each virtual environment camera to obtain the camera texture coordinates corresponding to the mesh vertex; The vertex data generation module 1140 is specifically used for: For each mesh vertex, vertex data is generated based on the 3D coordinates of the mesh vertex, the multiple camera fusion weights corresponding to the mesh vertex, and the multiple camera texture coordinates corresponding to the mesh vertex. The panoramic mesh vertex data is generated based on the vertex data of each mesh vertex.
[0160] In some implementations, the weight determination module 1130 determines the azimuth weight of each virtual environment camera relative to the grid vertex through the following steps: Based on the three-dimensional coordinates of the mesh vertices, determine the vertex azimuth angle of the mesh vertices on the two-dimensional plane; For each virtual environment camera, determine the angular deviation of the vertex azimuth angle relative to the optical axis direction of the virtual environment camera; Based on the number of multiple virtual environment cameras, the half-sector angle covered by each virtual environment camera is determined, and based on the half-sector angle and the preset feathering transition width, the deviation angle threshold is determined. The azimuth weight is determined based on the angle deviation and the deviation angle threshold.
[0161] In some embodiments, the deviation angle threshold includes a first angle threshold and a second angle threshold; the weight determination module 1130 is specifically used for: If the angle deviation is not greater than the first angle threshold, the azimuth weight is determined to be 1; the first angle threshold is determined based on the difference between the half-sector angle and the feathering transition width; and / or, If the angle deviation is greater than the first angle threshold but not greater than the second angle threshold, the azimuth weight is determined based on the angle deviation, the half-sector angle, and the preset feathering transition width; the second angle threshold is determined based on the sum of the half-sector angle and the feathering transition width.
[0162] In some implementations, the weight determination module 1130 determines the edge attenuation weight of each virtual environment camera relative to the mesh vertices through the following steps: For each virtual environment camera, the minimum distance between the camera texture coordinates corresponding to the mesh vertex and each screen boundary of the camera screen is determined based on the camera texture coordinates corresponding to the mesh vertex and the camera screen size of the virtual environment camera. In some implementations, the weight determination module 1130 is specifically used for: The edge attenuation weight is determined based on the minimum distance.
[0163] If the minimum distance indicates that the camera texture coordinates are not within the camera image, the edge attenuation weight is determined to be 0; and / or, If the minimum distance indicates that the camera texture coordinates are within the camera image, the edge attenuation weight is determined based on the minimum distance and a preset edge feathering width.
[0164] In some embodiments, the vertex data generation module 1140 is specifically used for: The initial fusion weights are determined based on the azimuth weights and edge attenuation weights. The initial fusion weights of each virtual environment camera are normalized to obtain the fusion weights of each virtual environment camera at the grid vertex; the sum of the fusion weights of each virtual environment camera at the grid vertex is 1.
[0165] In some implementations, the bowl-shaped model determination module 1120 constructs the bowl-shaped surface model through the following steps: Obtain the model construction parameters of the bowl-shaped curved surface model; the model construction parameters include the flat radius, maximum radial distance, maximum height of the bowl wall, and parabolic curvature coefficient; A polar coordinate system is established with the center of the virtual robot as the origin. Under the polar coordinate system, a bowl-shaped surface model with the center of the virtual robot as the origin is generated based on the flat radius, maximum radial distance, maximum height of the bowl wall, and parabolic curvature coefficient.
[0166] In some embodiments, the bowl-shaped model determining module 1120 is specifically used for: If the maximum radial distance is greater than the preset radial distance, a bowl-shaped surface model with the center of the virtual robot as the origin is generated based on the flat radius, the preset radial distance, the maximum height of the bowl wall, and the parabolic curvature coefficient; the preset radial distance is determined based on the flat radius, the maximum height of the bowl wall, and the parabolic curvature coefficient.
[0167] In some embodiments, the environmental image fusion module 1150 is specifically used for: The current posture data and current view control parameters of the real robot are obtained; the current view control parameters include at least one of azimuth rotation, pitch rotation, translation, and zoom distance. The model view projection matrix is determined based on the current posture data of the real robot and the current view control parameters. The multiple real-world images are fused using the model-view projection matrix and the panoramic mesh vertex data to obtain the panoramic image.
[0168] In some embodiments, the environmental image fusion module 1150 is specifically used for: Using the model-view projection matrix, the three-dimensional coordinates of each grid vertex are transformed to screen space to obtain the screen display coordinates; The screen display coordinates are rasterized to obtain multiple fragments, and the texture coordinates and blending weights corresponding to each fragment are determined. For each fragment, texture sampling is performed on the multiple real environment images according to the texture coordinates corresponding to the fragment to obtain multiple pixel colors, and the multiple pixel colors are fused according to the fusion weights corresponding to the fragment to obtain a fused pixel color. Based on the fused pixel color corresponding to each fragment, the screen pixels corresponding to the fragment are rendered to obtain the panoramic image.
[0169] In some embodiments, the environmental image fusion module 1150 is specifically used for: Based on the fused pixel color corresponding to each fragment, the screen pixels corresponding to the fragment are rendered to obtain an initial panoramic image; Obtain the three-dimensional model of the virtual robot and the depth information of the three-dimensional model; Based on the depth information of the 3D model, the 3D model is rendered onto the initial panoramic image to obtain a panoramic image containing the 3D model and its surrounding environment.
[0170] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0171] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0172] Corresponding to the above-described method for generating surround view images of a robot, this disclosure also provides a quadruped robot equipped with multiple real-world environment cameras. The viewpoints of adjacent environment cameras overlap. (See also...) Figure 13 This is a schematic diagram of the structure of the quadruped robot provided in the embodiments of this disclosure, as shown below. Figure 13 As shown, the quadruped robot 1300 includes a processor 1310, a memory 1320, and multiple real-world cameras 1330, and may also include other hardware required for its functions.
[0173] against Figure 13 For a detailed description of the various processors 1310 and memory 1320 shown, please refer to the following sections. Figure 14 The relevant information about processors and memory is already provided in the text, so it will not be repeated here.
[0174] Corresponding to the above-described method for generating a surround view image of a robot, this disclosure also provides a computer device; please refer to [link to relevant documentation]. Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure, such as... Figure 14 As shown, the computer device 1400 includes a processor 1410, an internal bus 1420, memory 1430, a network interface 1440, and non-volatile memory 1450, and may also include other hardware required for its functions. One or more embodiments of this specification can be implemented in software, for example, the processor 1410 reads the corresponding computer program from the non-volatile memory 1450 into the memory 1430 and then runs it. Of course, besides software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0175] The memory 1430, also known as internal memory, is used to temporarily store the computational data in the processor 1410, as well as the data exchanged with non-volatile memory 1450 such as hard disk. The processor 1410 exchanges data with non-volatile memory 1450 through the memory 1430.
[0176] In this embodiment, memory 1430 is specifically used to store application code that executes the solution of this application, and its execution is controlled by processor 1410. That is, when the computer device is running, processor 1410 communicates with network interface 1440, memory 1430 and non-volatile memory 1450 through internal bus 1420, so that processor 1410 executes the application code stored in memory 1430 and non-volatile memory 1450, thereby executing the method for generating a surround view image for a robot as described in the above method embodiment.
[0177] Processor 1410 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware microservices. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0178] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 1400. In other embodiments of this application, the computer device 1400 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0179] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the method for generating a surround view image for a robot as described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.
[0180] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the method for generating a surround view image for a robot in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0181] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0182] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0183] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.
[0184] Computers suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. Basic computer microservices include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0185] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0186] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0187] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and microservices in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program microservices and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0188] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0189] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for generating a surround view image of a robot, characterized in that, include: A virtual robot is obtained by modeling a real robot; the real robot is equipped with multiple real environment cameras, and the viewpoints of adjacent real environment cameras overlap; the virtual robot has virtual environment cameras that correspond one-to-one with the multiple real environment cameras. A bowl-shaped surface model corresponding to the virtual robot is determined; the bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with respect to the virtual robot remain unchanged; For each grid vertex, the azimuth weight and edge attenuation weight of each virtual environment camera relative to the grid vertex are determined respectively; the azimuth weight is used to characterize the degree of deviation of the optical axis center of each virtual environment camera relative to the grid vertex; the edge attenuation weight is used to characterize the image sampling quality at the projection position of the grid vertex on the screen of each virtual camera. For each virtual environment camera, the camera fusion weight of the virtual environment camera at the grid vertex is determined according to the azimuth weight and the edge attenuation weight, and panoramic grid vertex data is generated based on the multiple camera fusion weights corresponding to each grid vertex. Multiple real-world environment images are acquired using the multiple physical environment cameras, and then fused using the panoramic mesh vertex data to obtain a panoramic image.
2. The method according to claim 1, characterized in that, Each grid vertex has three-dimensional coordinates; the method further includes: For each mesh vertex, the mesh vertex is projected onto the camera screen of each virtual environment camera to obtain the camera texture coordinates corresponding to the mesh vertex; The process of generating panoramic mesh vertex data based on the fusion weights of multiple cameras corresponding to each mesh vertex includes: For each mesh vertex, vertex data is generated based on the 3D coordinates of the mesh vertex, the multiple camera fusion weights corresponding to the mesh vertex, and the multiple camera texture coordinates corresponding to the mesh vertex. The panoramic mesh vertex data is generated based on the vertex data of each mesh vertex.
3. The method according to claim 1, characterized in that, The azimuth weight of each virtual environment camera relative to the grid vertices is determined by the following steps: Based on the three-dimensional coordinates of the mesh vertices, determine the vertex azimuth angle of the mesh vertices on the two-dimensional plane; For each virtual environment camera, determine the angular deviation of the vertex azimuth angle relative to the optical axis direction of the virtual environment camera; Based on the number of multiple virtual environment cameras, the half-sector angle covered by each virtual environment camera is determined, and based on the half-sector angle and the preset feathering transition width, the deviation angle threshold is determined. The azimuth weight is determined based on the angle deviation and the deviation angle threshold.
4. The method according to claim 3, characterized in that, The deviation angle threshold includes a first angle threshold and a second angle threshold; determining the azimuth weight based on the angle deviation and the deviation angle threshold includes: If the angle deviation is not greater than the first angle threshold, the azimuth weight is determined to be 1; the first angle threshold is determined based on the difference between the half-sector angle and the feathering transition width; and / or, If the angle deviation is greater than the first angle threshold but not greater than the second angle threshold, the azimuth weight is determined based on the angle deviation, the half-sector angle, and the preset feathering transition width; the second angle threshold is determined based on the sum of the half-sector angle and the feathering transition width.
5. The method according to claim 1, characterized in that, The edge attenuation weight of each virtual environment camera relative to the mesh vertices is determined by the following steps: For each virtual environment camera, the minimum distance between the camera texture coordinates corresponding to the mesh vertex and each screen boundary of the camera screen is determined based on the camera texture coordinates corresponding to the mesh vertex and the camera screen size of the virtual environment camera. The edge attenuation weight is determined based on the minimum distance.
6. The method according to claim 5, characterized in that, Determining the edge attenuation weight based on the minimum distance includes: If the minimum distance indicates that the camera texture coordinates are not within the camera image, the edge attenuation weight is determined to be 0; and / or, If the minimum distance indicates that the camera texture coordinates are within the camera image, the edge attenuation weight is determined based on the minimum distance and a preset edge feathering width.
7. The method according to claim 1, characterized in that, The step of determining the camera fusion weight of the virtual environment camera at the grid vertex based on the azimuth weight and the edge attenuation weight includes: The initial fusion weights are determined based on the azimuth weights and edge attenuation weights. The initial fusion weights of each virtual environment camera are normalized to obtain the fusion weights of each virtual environment camera at the grid vertex; the sum of the fusion weights of each virtual environment camera at the grid vertex is 1.
8. The method according to any one of claims 1-7, characterized in that, The bowl-shaped surface model is constructed through the following steps: Obtain the model construction parameters of the bowl-shaped curved surface model; the model construction parameters include the flat radius, maximum radial distance, maximum height of the bowl wall, and parabolic curvature coefficient; A polar coordinate system is established with the center of the virtual robot as the origin. Under the polar coordinate system, a bowl-shaped surface model with the center of the virtual robot as the origin is generated based on the flat radius, the maximum radial distance, the maximum height of the bowl wall, and the parabolic curvature coefficient.
9. The method according to claim 8, characterized in that, The process of generating a bowl-shaped surface model with the virtual robot's center as the origin, based on the flat radius, the maximum radial distance, the maximum height of the bowl wall, and the parabolic curvature coefficient, includes: If the maximum radial distance is greater than the preset radial distance, the bowl-shaped surface model is generated based on the flat radius, the preset radial distance, the maximum height of the bowl wall, and the parabolic curvature coefficient; the preset radial distance is determined based on the flat radius, the maximum height of the bowl wall, and the parabolic curvature coefficient.
10. The method according to claim 1, characterized in that, The process of fusing the multiple real-world environment images using the panoramic mesh vertex data to obtain a panoramic image includes: The current posture data and current view control parameters of the real robot are obtained; the current view control parameters include at least one of azimuth rotation, pitch rotation, translation, and zoom distance. The model view projection matrix is determined based on the current posture data of the real robot and the current view control parameters. The multiple real-world images are fused using the model-view projection matrix and the panoramic mesh vertex data to obtain the panoramic image.
11. The method according to claim 10, characterized in that, The process of fusing multiple real-world environment images using the model view projection matrix and the panoramic mesh vertex data to obtain the panoramic image includes: Using the model-view projection matrix, the three-dimensional coordinates of each grid vertex are transformed to screen space to obtain the screen display coordinates; The screen display coordinates are rasterized to obtain multiple fragments, and the texture coordinates and blending weights corresponding to each fragment are determined. For each fragment, texture sampling is performed on the multiple real environment images according to the texture coordinates corresponding to the fragment to obtain multiple pixel colors, and the multiple pixel colors are fused according to the fusion weights corresponding to the fragment to obtain a fused pixel color. Based on the fused pixel color corresponding to each fragment, the screen pixels corresponding to the fragment are rendered to obtain the panoramic image.
12. The method according to claim 11, characterized in that, The step of rendering the screen pixels corresponding to each fragment based on the fused pixel color to obtain the surround view image includes: Based on the fused pixel color corresponding to each fragment, the screen pixels corresponding to the fragment are rendered to obtain an initial panoramic image; Obtain the three-dimensional model of the virtual robot and the depth information of the three-dimensional model; Based on the depth information of the 3D model, the 3D model is rendered onto the initial panoramic image to obtain a panoramic image containing the 3D model and its surrounding environment.
13. An apparatus for generating a surround view image of a robot, characterized in that, include: The robot modeling module is used to model a real robot to obtain a virtual robot. The real robot is equipped with multiple real environment cameras, and the viewpoints of adjacent real environment cameras overlap. The virtual robot has virtual environment cameras that correspond one-to-one with the multiple real environment cameras. A bowl-shaped model determination module is used to determine a bowl-shaped surface model corresponding to the virtual robot; the bowl-shaped surface model includes multiple mesh vertices, and the relative positions of each mesh vertex and each virtual environment camera with respect to the virtual robot remain unchanged; The weight determination module is used to determine the azimuth weight and edge attenuation weight of each virtual environment camera relative to the grid vertex for each grid vertex; the azimuth weight is used to characterize the degree of deviation of the optical axis center of each virtual environment camera relative to the grid vertex; The edge attenuation weight is used to characterize the image sampling quality at the projection position of the grid vertex on each virtual camera screen; The weight fusion module is used to determine the camera fusion weight of the virtual environment camera at the grid vertex for each virtual environment camera according to the azimuth weight and the edge attenuation weight, and generate panoramic grid vertex data based on the multiple camera fusion weights corresponding to each grid vertex. The environmental image fusion module is used to acquire multiple real environment images through the multiple physical environment cameras, and to fuse the multiple real environment images using the panoramic mesh vertex data to obtain a panoramic image.
14. A quadruped robot, characterized in that, The quadruped robot is equipped with multiple real-world environment cameras, and the viewpoints of adjacent environment cameras overlap; the quadruped robot includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the program, it implements the steps of the method for generating a panoramic image of the robot as described in any one of claims 1-12.
15. A computer device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the program, it implements the steps of the method for generating a surround view image for a robot as described in any one of claims 1-12.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method for generating a surround view image for a robot as described in any one of claims 1-12.