Image synthesis method and apparatus for robot, device, and medium

By combining implicit radiation field models with rendering methods for preset static scenes and target objects, the problem of high computational resource consumption in explicit rendering is solved, achieving efficient image synthesis and simulation efficiency improvement.

WO2025227595A1PCT designated stage Publication Date: 2025-11-06BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

Patent Information

Application Number
PCT/CN2024/119393
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2024-09-18
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

In existing technologies, explicit rendering methods consume a lot of computational resources and have low simulation efficiency in robot simulation, and cannot effectively utilize the amount of data captured by vision sensors.

Method used

An implicit radiation field model is used to implicitly render a preset static scene that does not require interaction with the robot, while explicit rendering is performed on the target object that requires interaction. The results of the two are then combined to generate a composite image.

Benefits of technology

It reduces the consumption of computing resources by explicit rendering, improves simulation efficiency, and enhances the realism of image synthesis and simulation speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024119393_06112025_PF_FP_ABST
    Figure CN2024119393_06112025_PF_FP_ABST
Patent Text Reader

Abstract

An image synthesis method and apparatus for a robot, a device, and a medium, relating to the technical field of robot vision. The method comprises: on the basis of position information of a real camera on an operation robot in a preset static scene, obtaining an implicit rendering result of the preset static scene by means of an implicit radiance field model pre-trained for the preset static scene; on the basis of the relative placement position of a target object in the preset static scene, generating an explicit rendering result of the target object in a preset simulation environment; and acquiring synthesized image information on the basis of the implicit rendering result and the explicit rendering result, wherein the operation scene of the operation robot comprises the preset static scene and the target object located in the preset static scene, the preset static scene does not need to be in physical contact with the operation robot, and the target object needs to be in physical contact with the operation robot. The present application can reduce the consumption of computing resources by explicit rendering, and improve simulation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Image synthesis method, device and equipment for robot and medium

[0001] The present application claims priority to the Chinese patent application No. 202410535862.8, filed on April 30, 2024, and entitled "Image synthesis method, device and equipment for robot and medium", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the technical field of computer vision, in particular to an image synthesis method, device and equipment for robot and medium. BACKGROUND

[0003] With the development of computer technology, the intelligence level of robots is getting higher and higher. The tasks performed by robots are mostly related to physical interaction with the environment, such as opening and closing doors, visually guiding grasping, and other operations. The images or depth information captured by visual sensors are key parameters for the input of robot control models.

[0004] Due to the limitation of the amount of data captured by visual sensors, in order to obtain a large amount of realistic visual information, explicit rendering methods such as ray tracing are generally used for simulation. The modeling of explicit rendering objects includes three-dimensional modeling entity files, color map files, and material texture maps such as normal and roughness.

[0005] However, in order to obtain good generalization, explicit rendering needs to stack a large number of model files, resulting in excessive storage space consumption, slowing down the simulation speed, and low simulation efficiency.

[0006] SUMMARY

[0007] The present application aims to provide an image synthesis method, device and equipment for robot and medium to reduce the consumption of computing resources by explicit rendering and improve simulation efficiency.

[0008] To achieve the above-mentioned purpose, the technical solutions adopted by the embodiments of the present application are as follows:

[0009] In a first aspect, the embodiments of the present application provide an image synthesis method for a robot, applied to an electronic device, wherein the electronic device is in communication connection with a work robot, a work scene of the work robot includes a preset static scene and a target object located in the preset static scene, the preset static scene does not need to be in physical contact with the work robot, and the target object needs to be in physical contact with the work robot, and the method comprises:

[0010] According to the position information of the real camera on the work robot in the preset static scene, an implicit radiance field model pre-trained for the preset static scene is used to obtain an implicit rendering result of the preset static scene;

[0011] According to the relative placement position of the target object in the preset static scene, an explicit rendering result of the target object is generated in a preset simulation environment;

[0012] According to the implicit rendering result and the explicit rendering result, synthesis image information is obtained.

[0013] Optionally, the implicit rendering result includes an implicit color map and an implicit depth map of the real camera on the work robot in a preset viewport direction;

[0014] The explicit rendering result includes a virtual camera with the same parameters as the real camera on the work robot, an explicit color map and an explicit depth map in the preset viewport direction;

[0015] The implicit rendering result and the explicit rendering result include:

[0016] According to the implicit color map, the implicit depth map, the explicit color map and the explicit depth map, the synthesis image information is obtained.

[0017] Optionally, according to the position information of the real camera on the work robot in the preset static scene, an implicit radiance field model pre-trained for the preset static scene is used to obtain an implicit rendering result of the preset static scene, including:

[0018] The origin coordinates and attitude angles of a plurality of light rays emitted by the real camera in the preset viewport direction are obtained;

[0019] According to the origin coordinates and attitude angles of the plurality of light rays, the implicit radiance field model is used to determine the volume density information and color information of the plurality of light rays;

[0020] According to the volume density information and color information of the plurality of light rays, the implicit color map and the implicit depth map are determined.

[0021] Optionally, the work robot includes a mobile chassis and a mechanical arm, and the origin coordinates and attitude angles of a plurality of light rays emitted by the real camera in the preset viewport direction are obtained, including:

[0022] According to the odometer of the mobile chassis and the forward kinematics equation of the mechanical arm, the camera extrinsic parameters of the real camera are calculated;

[0023] According to the camera extrinsic parameters of the real camera, the origin coordinates and the attitude angle of the plurality of light rays are determined.

[0024] Optionally, the determining the implicit color map and the implicit depth map according to the volume density information and the color information of the plurality of light rays comprises:

[0025] According to the integral of the volume density information and the color information of the plurality of light rays, the implicit color map of the pixel points corresponding to the plurality of light rays is determined.

[0026] According to the jump position of the volume density information of the plurality of light rays, the implicit depth map of the pixel points corresponding to the plurality of light rays is determined.

[0027] Optionally, before the generating the explicit rendering result of the target object in the preset simulation environment according to the relative placement position of the target object in the preset static scene, the method further comprises:

[0028] According to the depth information in the implicit depth map, the relative placement position of the target object in the preset static scene is determined.

[0029] Optionally, the generating the explicit rendering result of the target object in the preset simulation environment according to the relative placement position of the target object in the preset static scene comprises:

[0030] The simple collision body corresponding to the preset static scene is generated in the preset simulation environment.

[0031] According to the relative placement position of the target object in the preset static scene, the target object is rendered on the simple collision body.

[0032] Optionally, the obtaining the composite image information according to the implicit color map, the implicit depth map, the explicit color map and the explicit depth map comprises:

[0033] According to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map, the target depth values of the plurality of pixel points are determined, the target depth values being the depth values in the implicit depth map or the depth values in the explicit depth map.

[0034] According to the color values of the plurality of pixel points in the implicit color map and the explicit color map, the target color values of the plurality of pixel points are determined, the composite image information comprising the target depth values and the target color values of the plurality of pixel points.

[0035] Optionally, the determining the target depth values of the plurality of pixel points according to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map comprises:

[0036] if the depth value of the target pixel point in the explicit depth map is greater than or equal to the depth value in the implicit depth map, determining the target depth value of the target pixel point as the depth value of the target pixel point in the implicit depth map;

[0037] determining target color values of the plurality of pixel points according to color values of the plurality of pixel points in the depth map to which the target depth values of the plurality of pixel points belong, the implicit color map and the explicit color map, comprises:

[0038] if the target depth value of the target pixel point belongs to the implicit depth map, determining the target color value of the target pixel point as the color value of the target pixel point in the implicit color map.

[0039] Optionally, the determining the target depth values of the plurality of pixel points according to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map comprises:

[0040] if the depth value of the target pixel point in the explicit depth map is less than the depth value in the implicit depth map, determining the target depth value of the target pixel point as the depth value of the target pixel point in the explicit depth map;

[0041] determining target color values of the plurality of pixel points according to color values of the plurality of pixel points in the depth map to which the target depth values of the plurality of pixel points belong, the implicit color map and the explicit color map, comprises:

[0042] if the target depth value of the target pixel point belongs to the explicit depth map, determining the target color value of the target pixel point as the color value of the target pixel point in the explicit color map.

[0043] Optionally, after the obtaining the composite image information according to the implicit rendering result and the explicit rendering result, the method further comprises:

[0044] training a visual control model of the work robot according to the composite image information;

[0045] deploying the trained visual control model to the work robot, so as to control the work robot to operate the target object in the preset static scene according to the visual control model.

[0046] In a second aspect, the embodiments of the present application further provide an image synthesis device for a robot, applied to an electronic device, the electronic device being in communication connection with a work robot, a work scene of the work robot including a preset static scene and a target object located in the preset static scene, the preset static scene not needing to be in physical contact with the work robot, and the target object needing to be in physical contact with the work robot, and the device comprising:

[0047] an implicit rendering module, configured to obtain an implicit rendering result of the preset static scene by using an implicit radiance field model pre-trained for the preset static scene according to position information of a real camera on the work robot in the preset static scene;

[0048] an explicit rendering module, configured to generate an explicit rendering result of the target object in a preset simulation environment according to a relative placement position of the target object in the preset static scene;

[0049] an information synthesis module, configured to obtain synthesis image information according to the implicit rendering result and the explicit rendering result.

[0050] Optionally, the implicit rendering result includes an implicit color map and an implicit depth map of the real camera on the work robot in a preset viewport direction.

[0051] The explicit rendering result includes a virtual camera with the same parameters as the real camera on the work robot, an explicit color map and an explicit depth map in the preset viewport direction.

[0052] The information synthesis module is specifically configured to obtain the synthesis image information according to the implicit color map, the implicit depth map, the explicit color map and the explicit depth map.

[0053] Optionally, the implicit rendering module includes:

[0054] a light information obtaining unit, configured to obtain origin coordinates and attitude angles of a plurality of light rays emitted by the real camera in the preset viewport direction;

[0055] an implicit radiance unit, configured to determine volume density information and color information of the plurality of light rays by using the implicit radiance field model according to the origin coordinates and the attitude angles of the plurality of light rays;

[0056] an implicit information generating unit, configured to determine the implicit color map and the implicit depth map according to the volume density information and the color information of the plurality of light rays.

[0057] Optionally, the job robot comprises a mobile chassis and a mechanical arm, and the light ray information acquisition unit is specifically configured to calculate camera extrinsic parameters of the real camera according to an odometer of the mobile chassis and forward kinematics equations of the mechanical arm; and determine origin coordinates and attitude angles of the plurality of light rays according to the camera extrinsic parameters of the real camera.

[0058] Optionally, the implicit information generation unit is specifically configured to determine an implicit color map of the pixel points corresponding to the plurality of light rays according to integrals of the volume density information and the color information of the plurality of light rays; and determine an implicit depth map of the pixel points corresponding to the plurality of light rays according to jump positions of the volume density information of the plurality of light rays.

[0059] Optionally, before the explicit rendering module, the apparatus further comprises:

[0060] a placement position calculation module configured to determine a relative placement position of the target object in the preset static scene according to the depth information in the implicit depth map.

[0061] Optionally, the explicit rendering module is specifically configured to generate a simple collision body corresponding to the preset static scene in the preset simulation environment; and render and generate the target object on the simple collision body according to the relative placement position of the target object in the preset static scene.

[0062] Optionally, the information synthesis module is specifically configured to determine target depth values of a plurality of pixel points according to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map, the target depth values being the depth values in the implicit depth map or the depth values in the explicit depth map; determine target color values of the plurality of pixel points according to the color values of the plurality of pixel points in the depth map to which the target depth values of the plurality of pixel points belong, the implicit color map and the explicit color map; and the synthesized image information comprises the target depth values and the target color values of the plurality of pixel points.

[0063] Optionally, the information synthesis module is specifically configured to determine a target depth value of a target pixel point as a depth value of the target pixel point in the implicit depth map if the depth value of the target pixel point in the explicit depth map is greater than or equal to the depth value in the implicit depth map; and determine a target color value of the target pixel point as a color value of the target pixel point in the implicit color map if the target depth value of the target pixel point belongs to the implicit depth map.

[0064] Optionally, the information synthesizing module is specifically configured to: if the depth value of the target pixel point in the explicit depth map is less than the depth value in the implicit depth map, determining the target depth value of the target pixel point as the depth value of the target pixel point in the explicit depth map; and if the target depth value of the target pixel point belongs to the explicit depth map, determining the target color value of the target pixel point as the color value of the target pixel point in the explicit color map.

[0065] Optionally, after the information synthesizing module, the device further comprises:

[0066] a model training module configured to train a visual control model of the work robot according to the synthesized image information;

[0067] a model deployment module configured to deploy the trained visual control model to the work robot, so as to control the work robot to operate the target object in the preset static scene according to the visual control model.

[0068] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor, a storage medium and a bus, the storage medium stores program instructions executable by the processor, when the electronic device is running, the processor and the storage medium communicate through the bus, and the processor executes the program instructions to perform the steps of the image synthesis method for a robot according to any one of the first aspect.

[0069] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to perform the steps of the image synthesis method for a robot according to any one of the first aspect.

[0070] The present application has the following beneficial effects:

[0071] The image synthesis method, device, equipment and medium for a robot provided by the present application obtain an implicit rendering result by performing implicit rendering on a preset static scene which does not need to interact with a work robot, perform explicit rendering on a target object which needs to interact with the work robot to obtain an explicit rendering result, and synthesize the implicit rendering result and the explicit rendering result to obtain synthesized image information. The method combines explicit rendering and implicit rendering, reduces the content which needs to be explicitly rendered to the target object which needs to interact with the work robot, reduces the consumption of computing resources by explicit rendering, and improves the simulation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0072] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.

[0073] Fig. 1 is an architecture diagram of an image synthesis system provided by the embodiments of the present application;

[0074] Fig. 2 is a flow diagram of an image synthesis method for a robot provided by the embodiments of the present application;

[0075] Fig. 3 is a flow diagram of an image synthesis method for a robot provided by the embodiments of the present application;

[0076] Fig. 4 is a flow diagram of an image synthesis method for a robot provided by the embodiments of the present application;

[0077] Fig. 5 is a flow diagram of an image synthesis method for a robot provided by the embodiments of the present application;

[0078] Fig. 6 is a flow diagram of an image synthesis method for a robot provided by the embodiments of the present application;

[0079] Fig. 7 is an image synthesis diagram provided by the embodiments of the present application;

[0080] Fig. 8 is a flow diagram of an image synthesis method for a robot provided by the embodiments of the present application;

[0081] Fig. 9 is a flow diagram of an image synthesis method for a robot provided by the embodiments of the present application;

[0082] Fig. 10 is a structural diagram of an image synthesis device for a robot provided by the embodiments of the present application;

[0083] Fig. 11 is a schematic diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0084] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, not all the embodiments.

[0085] The detailed description of embodiments of the application in the following enables a person skilled in the art to make or use the application. Numerous modifications and variations within the scope of the present application are possible in light of the above teachings. It is, therefore, to be understood that changes can be made in the particular embodiments of the application disclosed which are within the scope and spirit of this application. With specific reference now to the drawings in

[0086] It is to be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Furthermore, the terms "including," "includes," "comprising," "comprises," "has," "contains" or variants thereof do not preclude the addition of one or more steps, units, features or components to those already provided.

[0087] It should be noted that the features of the embodiments of the application can be combined with each other, without conflict.

[0088] FIG. 1 is an architecture diagram of an image synthesis system provided by an embodiment of the application. As shown in FIG. 1, the image synthesis system includes an electronic device 100 and a job robot 200. The electronic device 100 is in communication connection with the job robot 200. The electronic device 100 performs implicit rendering on a preset static scene in which the job robot 200 performs a job, and performs explicit rendering on a target object in the preset static scene, to obtain synthesis image information according to an implicit rendering result of the preset static scene and an explicit rendering result of the target object.

[0089] In some embodiments, the synthesis image information is sent to the job robot 200 as a visual simulation input of the job robot 200. A visual control model or algorithm in the job robot 200 generates hand-eye coordinated motion control according to the input synthesis image information. According to the motion control action generated by the job robot 200, physical operations such as grabbing, moving, tapping, and switching are performed on the explicitly rendered target object in a preset simulation environment. The synthesis image information and the physical operation result in the preset simulation environment are used as training data to train the visual control model or algorithm of the job robot 200. After the training is completed, the visual control model or algorithm is deployed in the job robot 200. The job robot 200 acquires a collection image including the preset static scene and the target object through an installed camera, and outputs a control instruction according to the collection image by using the visual control model or algorithm, so that the job robot 200 operates the target object according to the control instruction.

[0090] Based on the above image synthesis system, the following describes a specific implementation of an image synthesis method for a robot provided by the application in combination with embodiments.

[0091] FIG. 2 is a flowchart of an image synthesis method for a robot according to an embodiment of the present application. The method can be executed by the electronic device described above, which can be a computer or other device having computing and processing capabilities. The method can include the following steps, as shown in FIG. 2:

[0092] S101: Obtain an implicit rendering result of a real camera arranged on the work robot for the preset static scene according to an implicit radiance field model pre-trained by the work robot for the preset static scene.

[0093] In this embodiment, the work scene of the work robot includes a preset static scene and a target object in the preset static scene. The preset static scene is composed of static objects that cannot be moved, such as large spaces (rooms, factory areas, etc.), large objects, and building structures. The preset static scene can be used as a scene or environmental obstacle for the work robot to work in. During the control of the work robot, the work robot needs to avoid collision with the preset static scene.

[0094] The implicit radiance field model (NeRF) is a light model that can render light information of the preset static scene relative to the work robot. The electronic device communicates with the work robot to obtain position information of a real camera on the work robot in the preset static scene. The implicit radiance field model is used to obtain an implicit rendering result of the preset static scene according to the position information of the real camera in the preset static scene. The implicit rendering result is used to indicate light information of the preset static scene relative to the real camera at the position. The light information includes depth information and color information.

[0095] S102: Generate an explicit rendering result of the target object in the preset simulation environment according to a relative placement position of the target object in the preset static scene.

[0096] In this embodiment, the target object is a physically interactive object in the preset static scene, i.e., an object that can be operated by the work robot in the preset static scene.

[0097] For the preset static scene and the target object, the work robot does not need to physically contact the static objects corresponding to the preset static scene. Therefore, the implicit radiance field model can be used to obtain the implicit rendering result of the preset static scene. The work robot needs to physically contact the target object to operate the target object. Therefore, the target object needs to be rendered in the simulation engine using the explicit rendering method.

[0098] Specifically, by inputting the relative placement position of the target object in the preset static scene and the model file of the target object in the simulation engine, the target object is rendered in the simulation environment constructed by the simulation engine, and then a virtual camera with the same position information as the real camera is used to shoot the target object in the simulation environment to obtain an explicit rendering result, which is used to indicate the light information of the target object relative to the virtual camera at the same position.

[0099] S103: Obtain synthesis image information according to the implicit rendering result and the explicit rendering result.

[0100] In this embodiment, the light information of the preset static scene and the light information of the target object are fused to obtain the synthesis image information. In the fusion process, the occlusion relationship of the light information of the preset static scene and the light information of the target object in each pixel point needs to be determined, and the light information of each pixel point in the synthesis image information is the unoccluded light information of each pixel point in the preset static scene and the target object.

[0101] The image synthesis method for the robot provided in the above embodiment obtains an implicit rendering result by implicitly rendering a preset static scene that does not need to interact with a work robot, obtains an explicit rendering result by explicitly rendering a target object that needs to interact with the work robot, and synthesizes the implicit rendering result and the explicit rendering result to obtain synthesis image information. This method reduces the content that needs to be explicitly rendered to the target object that needs to interact with the work robot by combining explicit rendering and implicit rendering, reduces the consumption of computing resources by explicit rendering, and improves simulation efficiency.

[0102] Furthermore, the implicit rendering result of the preset static scene is obtained by using an implicit radiation field model. On the one hand, the implicit radiation field model is trained by using an image of the preset static scene shot by a real camera, so that the implicit rendering result of the preset static scene obtained by the implicit radiation field model is more realistic. On the other hand, the preset static scene does not need to be modeled in the simulation environment, which can save a lot of time and further improve the image synthesis efficiency.

[0103] In a possible implementation, the implicit rendering result includes an implicit color map and an implicit depth map of the real camera on the work robot in a preset viewport direction; the explicit rendering result includes an explicit color map and an explicit depth map of a virtual camera with the same parameters as the real camera on the work robot in the preset viewport direction; and the process of S103 of obtaining the synthesis image information according to the implicit rendering result and the explicit rendering result can include:

[0104] The synthesis image information is obtained according to the implicit color map, the implicit depth map, the explicit color map, and the explicit depth map.

[0105] In the embodiment, a real camera is installed on the work robot for image acquisition. When generating the implicit rendering result of the preset static scene, instead of using the real camera to capture images of the preset static scene, a preset viewport direction of the real camera on the work robot is determined according to the position and pose of the work robot in the preset static scene, the viewport direction is the view angle direction of the real camera on the work robot relative to the preset static scene, and light information of the preset static scene in the preset viewport direction is obtained by using the implicit radiance field model according to the preset viewport direction. The light information includes an implicit color map and an implicit depth map of the preset static scene relative to the real camera in the preset viewport direction. The color value of each pixel point in the implicit color map is the color value of the corresponding point of the preset static scene relative to the real camera in the preset viewport direction, and the depth value of each pixel point in the implicit depth map is the distance of the corresponding point of the preset static scene relative to the real camera in the preset viewport direction.

[0106] In order to ensure that the implicit rendering result and the explicit rendering result have correct projection relationship, a virtual camera with the same parameters as the real camera is used to capture the target object in the preset simulation environment to obtain light information of the target object in the preset viewport direction. The light information includes an explicit color map and an explicit depth map of the target object relative to the virtual camera in the preset viewport direction. The color value of each pixel point in the explicit color map is the color value of the corresponding point of the target object relative to the virtual camera in the preset viewport direction, and the depth value of each pixel point in the explicit depth map is the distance of the corresponding point of the target object relative to the virtual camera in the preset viewport direction.

[0107] According to the implicit depth map and the explicit depth map, the occlusion relationship of the preset static scene and the target object at each pixel point is determined, the depth value of the pixel point corresponding to the unoccluded part is determined as the target depth value, the color value of the pixel point corresponding to the unoccluded part is determined as the target color value, and the image information includes the target depth value and the target color value.

[0108] The image synthesis method for the robot provided in the above embodiment uses the real camera to generate the implicit color map and the implicit depth map, and uses the virtual camera to generate the explicit color map and the explicit depth map, which have the same camera parameters and viewport direction, so that the preset static scene and the target object have correct projection relationship and occlusion in the synthesized image information, and the realistic effect of the synthesized image information is improved.

[0109] In a possible implementation, FIG. 3 is a flow diagram II of an image synthesis method for a robot according to an embodiment of the present application. As shown in FIG. 3, the process of obtaining the implicit rendering result of a real camera arranged on a work robot for a preset static scene according to an implicit radiance field model pre-trained by the work robot for the preset static scene in S101 can include:

[0110] S201: Obtain the origin coordinates and attitude angles of a plurality of light rays emitted by a real camera in a preset viewport direction.

[0111] In this embodiment, the origin coordinates and attitude angles of a plurality of light rays emitted by the real camera relative to a plurality of pixel points of a captured image are determined according to the size of the captured image and the preset viewport direction, where the origin coordinates are three-dimensional coordinates X (x, y, z) of the light rays in a world coordinate system, and the attitude angles include a horizontal angle θ and a pitch angle φ of the light rays.

[0112] The process of obtaining the origin coordinates of the light rays emitted to the pixel points includes:

[0113] According to the two-dimensional coordinates (u, v) of the pixel point and the camera intrinsic parameters of the real camera, the three-dimensional coordinates of the pixel point in the camera coordinate system are determined, and then the three-dimensional coordinates of the pixel point in the camera coordinate system are rotated and translated according to the camera extrinsic parameters of the real camera to obtain the three-dimensional coordinates of the pixel point in the world coordinate system, so as to represent the origin coordinates of the light ray corresponding to the pixel point by the three-dimensional coordinates.

[0114] For example, the conversion relationship of converting the two-dimensional coordinates of the pixel point into the three-dimensional coordinates in the world coordinate system can be represented as:

[0115] where u and v are two-dimensional coordinates of the pixel point, is the camera depth, f x and f y are the camera intrinsic parameters of the real camera, i.e., the focal lengths in the x direction and the y direction, c x and c y are pixel origin offsets, is the representation of the world coordinate system in the camera coordinate system, which can be represented by the camera extrinsic parameters of the real camera, i.e., a rotation matrix and a translation matrix, is the three-dimensional coordinates of the pixel point in the world coordinate system.

[0116] S202: Determine the volume density information and color information of the plurality of light rays by using the implicit radiance field model according to the origin coordinates and attitude angles of the plurality of light rays.

[0117] In this embodiment, the training process of the implicit radiance field model is as follows:

[0118] The real camera on the operation robot is used to shoot a preset static scene to obtain a plurality of static scene images, and the origin coordinates and attitude angles of the light rays emitted by the real camera relative to each pixel point in each static scene image are determined. The origin coordinates and attitude angles of the light rays are taken as inputs of the model. The model predicts the color value and the volume density information of each pixel point in each static scene image, and predicts the pixel value of each pixel point according to the color value and the volume density information of each pixel point. The model is optimized according to the predicted pixel value of each pixel point and the real pixel value of each pixel point in each static scene image, and an implicit radiance field model is obtained.

[0119] The implicit radiance field model can be expressed as (RGB, σ) = F (X, θ, φ). The input of the implicit radiance field model is the origin coordinates X and the attitude angle (θ, φ) of the light ray, and the output is the color value RGB and the volume density information σ. The volume density information actually represents the opacity of the pixel point.

[0120] For the trained implicit radiance field model, the origin coordinates and the attitude angle of each light ray are input into the implicit radiance field model, and the volume density information and the color information on each light ray are output by the implicit radiance field model.

[0121] S203: Determine an implicit color map and an implicit depth map according to the volume density information and the color information of the plurality of light rays.

[0122] In this embodiment, the color value and the depth value of the pixel point corresponding to each light ray are determined according to the volume density information and the color information of each light ray. The implicit color map is obtained according to the color value of the pixel point corresponding to the plurality of light rays, and the implicit depth map is obtained according to the depth value of the pixel point corresponding to the plurality of light rays.

[0123] In a possible implementation, FIG. 4 is a flowchart of a method for image synthesis of a robot provided by an embodiment of the present application. As shown in FIG. 4, the process of S201 of obtaining the origin coordinates and the attitude angle of the plurality of light rays emitted by the real camera in the preset viewport direction can include the following steps.

[0124] S301: Calculate the camera extrinsic parameter of the real camera according to the odometer of the mobile chassis and the forward kinematics equation of the mechanical arm.

[0125] S302: Determine the origin coordinates and the attitude angle of the plurality of light rays according to the camera extrinsic parameter of the real camera.

[0126] In this embodiment, in order to determine the accurate origin coordinates and the attitude angle of the plurality of light rays emitted by the real camera relative to the preset static scene in the preset viewport direction, the accurate camera extrinsic parameter needs to be determined.

[0127] The working robot is composed of a mobile chassis and a mechanical arm, a odometer is used to determine the change relationship of the position of the mobile chassis of the working robot over time, so as to obtain the position and posture of the mobile chassis, a forward kinematics equation is used to represent the position and posture of the mechanical arm relative to the mobile chassis, the rotation matrix and the translation matrix of the real camera are determined according to the position and posture of the mobile chassis and the position and posture of the mechanical arm relative to the mobile chassis, the pixel points corresponding to the plurality of light rays are subjected to coordinate conversion according to the rotation matrix and the translation matrix of the real camera, the origin coordinates of the plurality of light rays are determined, and the horizontal angle of the light ray is obtained according to the angle between the origin coordinates and the horizontal direction of the world coordinate system, and the pitch angle of the light ray is obtained according to the angle between the origin coordinates and the vertical direction of the world coordinate system.

[0128] In a possible implementation, Fig. 5 is a flow diagram of a method for synthesizing images of a robot according to an embodiment of the present application, as shown in Fig. 5, the process of determining the implicit color map and the implicit depth map according to the volume density information and the color information of the plurality of light rays in S203 can include:

[0129] S401: determining an implicit color map of pixel points corresponding to the plurality of light rays according to the integration of the volume density information and the color information of the plurality of light rays.

[0130] In the embodiment, each light ray is composed of a plurality of continuous points, the implicit color value of the pixel point corresponding to each light ray is obtained by discretizing the plurality of continuous points on each light ray and integrating the volume density value and the color value of the discrete points, and the implicit color values of the pixel points corresponding to the plurality of light rays constitute the implicit color map.

[0131] S402: determining an implicit depth map of pixel points corresponding to the plurality of light rays according to the jump position of the volume density information of the plurality of light rays.

[0132] In the embodiment, the volume density value of the light ray passing through the transparent object is 0, and the transparent object may, for example, be air. For objects made of glass or semi-transparent materials, the volume density value of the light ray thereon is approximately 0. Therefore, for a point i on a light ray L, the volume density value σ i = 0 in the air, when the light ray L encounters a point j on the surface of another solid object, the volume density value σ j ≠ 0, assuming that i and j are two adjacent points in space, the volume density of the light ray L between points i and j changes in space.

[0133] By determining the jump position of the volume density information of each light ray, the distance between the jump position and the imaging plane of the real camera is calculated to obtain the implicit depth value of the pixel point corresponding to each light ray, and the implicit depth values of the pixel points corresponding to the plurality of light rays constitute the implicit depth map.

[0134] It should be noted that the execution sequence of S401 and S402 shown in FIG. 5 is only an example, and in actual application, S401 and S402 do not have a fixed execution sequence.

[0135] The robot-based image synthesis method provided in the above embodiment adopts the implicit radiance field model to obtain the implicit color map and the implicit depth map of the preset static scene. On the one hand, since the implicit radiance field model is trained by using the images of the preset static scene shot by the real camera, the implicit rendering result of the preset static scene obtained through the implicit radiance field model is more realistic. On the other hand, since it is not necessary to model the preset static scene in the simulation environment, a large amount of time can be saved, and the image synthesis efficiency is further improved.

[0136] In a possible implementation, before S102 generates the explicit rendering result of the target object in the preset simulation environment according to the relative placement position of the target object in the preset static scene, the method can further include:

[0137] According to the depth information in the implicit depth map, the relative placement position of the target object in the preset static scene is determined.

[0138] In this embodiment, according to the depth information in the implicit depth map and the camera intrinsic parameter of the real camera, a pixel point cloud constituting a static object in the preset static scene is calculated, the static object is an object in the preset static scene for placing the target object, and the position of the static object in the preset static scene is determined according to the pixel point cloud constituting the static object in the preset static scene.

[0139] Then, according to the position of the static object, a position on the surface of the static object is randomly determined as the relative placement position of the target object in the preset static scene through a random algorithm. The random algorithm makes the relative placement position of the target object and the static object have a certain randomness, increases the simulation richness, and thus increases the number of data sets for training the model of the work robot as a sample.

[0140] It should be noted that since the synthesized image information obtained by the present scheme is sample data for training the visual control model or algorithm of the work robot, it is not necessary to determine the relative placement position of the target object in the preset static scene as the real position of the target object, and the relative placement position of the target object on the static object in the preset static scene can be randomly determined according to the position of the static object in the preset static scene.

[0141] The image synthesis method for the robot provided in the above embodiment determines the relative placement position of the target object in the preset static scene according to the depth information in the implicit depth map, so that the relative placement position of the target object obtained by simulation and the preset static scene conforms to the real scene, ensures that the target object and the preset static scene have accurate projection relationship and occlusion, and improves the realism of the synthesized image information.

[0142] In a possible implementation, Fig. 6 is a flowchart of a method for synthesizing images for a robot according to an embodiment of the present application. As shown in Fig. 6, the process of generating the explicit rendering result of the target object in the preset simulation environment according to the relative placement position of the target object in the preset static scene in S102 can include:

[0143] S501: generating a simple collision body corresponding to the preset static scene in the preset simulation environment.

[0144] S502: rendering the target object on the simple collision body according to the relative placement position of the target object in the preset static scene.

[0145] In this embodiment, since subsequent operations of the work robot on the target object in the real scene need to be simulated in the preset simulation environment, the target object needs to be in the same state of being supported by static objects in the real environment in the preset simulation environment. Therefore, a simple collision body corresponding to the static object in the preset static scene can be generated in the preset simulation environment.

[0146] Specifically, a plurality of point clouds can be extracted according to the implicit depth map of the preset static scene, and a convex hull or convex decomposition algorithm can be used to generate an outer contour of the plurality of point clouds to obtain a simple collision body that can contain the plurality of point clouds. The simple collision body is used to participate in collision explicit dynamics calculation when subsequent dynamics simulation of the target object is performed, and provides a counteracting support force for the target object to overcome gravity. The simple collision volume is in a transparent state in the preset simulation environment and does not participate in explicit rendering.

[0147] After the simple collision body corresponding to the preset static scene is generated, the target object can be rendered on the simple collision body according to the relative placement position of the target object in the preset static scene.

[0148] It should be noted that, in order to perform dynamics simulation on the target object in the preset simulation environment and determine the motion trajectory of the target object based on the operation of the work robot, mass properties and friction properties can be added to the target object in the preset simulation environment. The mass properties can include the mass, center of mass, and moment of inertia of the target object, and the friction properties can include the static friction coefficient and dynamic friction coefficient of the target object relative to the preset static scene.

[0149] The image synthesis method for the robot provided in the above embodiment can generate a simple collision body corresponding to a preset static scene in a preset simulation environment, render a target object on the simple collision body to simulate a state in which the target object is supported in the preset static scene, so that the motion state of the target object can be further simulated in the preset simulation environment.

[0150] For example, FIG. 7 is a schematic diagram of image synthesis provided by an embodiment of the present application. As shown in FIG. 7, according to a preset viewport direction of a real camera on a work robot and a preset static scene, an implicit rendering result of the preset static scene is obtained by using an implicit radiance field model, a physical model of a target object is generated in a preset simulation environment, a simple collision body of the preset static scene is generated according to the implicit rendering result, the physical model of the target object is located on the simple collision body, an explicit rendering result of the target object is obtained by using explicit rendering, and the implicit rendering result and the explicit rendering result are synthesized to obtain synthesis image information. In the preset static scene shown in FIG. 7, a table and a scene in which the table is located are included, and the target object is an apple.

[0151] In a possible implementation, FIG. 8 is a schematic diagram of a flow of a method for image synthesis for a robot provided by an embodiment of the present application. As shown in FIG. 8, the process of obtaining synthesis image information according to an implicit color map, an implicit depth map, an explicit color map and an explicit depth map can include the following steps.

[0152] S601: determining target depth values of a plurality of pixel points according to depth values of the plurality of pixel points in the implicit depth map and the explicit depth map, the target depth values being the depth values in the implicit depth map or the depth values in the explicit depth map.

[0153] In this embodiment, the occlusion relationship of each pixel point on the preset static scene and the target object is determined according to the depth values of each pixel point in the implicit depth map and the explicit depth map, that is, it is determined whether each pixel point is the preset static scene occluding the target object or the target object occluding the preset static scene.

[0154] If the pixel point is the preset static scene occluding the target object, the target depth value of the pixel point is the depth value in the implicit depth map, and if the pixel point is the target object occluding the preset static scene, the target depth value of the pixel point is the depth value in the explicit depth map.

[0155] S602: determining target color values of the plurality of pixel points according to color values of the plurality of pixel points in the depth map to which the target depth values of the plurality of pixel points belong, the implicit color map and the explicit color map, and the synthesis image information includes the target depth values and the target color values of the plurality of pixel points.

[0156] In the embodiment, if the pixel point is a preset static scene occluded target object, the target depth value of the pixel point is a depth value in the implicit depth map, and the target color value of the pixel point is a color value in the implicit color map; if the pixel point is a target object occluded preset static scene, the target depth value of the pixel point is a depth value in the explicit depth map, and the target color value of the pixel point is a color value in the explicit color map.

[0157] In some embodiments, the process of determining the target depth value of the plurality of pixel points according to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map in S601 can include:

[0158] If the depth value of the target pixel point in the explicit depth map is greater than or equal to the depth value in the implicit depth map, the target depth value of the target pixel point is determined as the depth value of the target pixel point in the implicit depth map.

[0159] The process of determining the target color value of the plurality of pixel points according to the color values of the plurality of pixel points in the depth map to which the target depth value of the plurality of pixel points belongs, the implicit color map and the explicit color map in S602 can include:

[0160] If the target depth value of the target pixel point belongs to the implicit depth map, the target color value of the target pixel point is determined as the color value of the target pixel point in the implicit color map.

[0161] In the embodiment, the occlusion relationship of each pixel point on the preset static scene and the target object is determined according to the depth value of each pixel point in the implicit depth map and the explicit depth map, including: if the depth value of the target pixel point in the explicit depth map is greater than or equal to the depth value in the implicit depth map, that is, the distance between the preset static scene corresponding to the target pixel point and the real camera is closer than the distance between the target object and the real camera, it is determined that the target pixel point is a preset static scene occluded target object, the target depth value of the target pixel point is a depth value in the implicit depth map, and the target color value of the target pixel point is a color value in the implicit color map.

[0162] In other embodiments, the process of determining the target depth value of the plurality of pixel points according to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map in S601 can include:

[0163] If the depth value of the target pixel point in the explicit depth map is less than the depth value in the implicit depth map, the target depth value of the target pixel point is determined as the depth value of the target pixel point in the explicit depth map.

[0164] The process of determining the target color value of the plurality of pixel points according to the color values of the plurality of pixel points in the depth map, the implicit color map and the explicit color map to which the target depth value of the plurality of pixel points belongs in S602 can include:

[0165] If the target depth value of the target pixel point belongs to the explicit depth map, the target color value of the target pixel point is determined as the color value of the target pixel point in the explicit color map.

[0166] In the embodiment, the occlusion relationship of each pixel point on the preset static scene and the target object is determined according to the depth values of each pixel point in the implicit depth map and the explicit depth map, including: if the depth value of the target pixel point in the explicit depth map is less than the depth value in the implicit depth map, that is, the distance between the target object corresponding to the target pixel point and the real camera is closer than the distance between the preset static scene and the real camera, it is determined that the target pixel point is the target object occluding the preset static scene, and the target depth value of the target pixel point is the depth value in the explicit depth map, and the target color value of the target pixel point is the color value in the explicit color map.

[0167] For example, the code for obtaining the synthesized image information according to the implicit color map, the implicit depth map, the explicit color map and the explicit depth map is as follows:

[0168] The image synthesis method for the robot provided in the above embodiment determines the target depth value of the plurality of pixel points according to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map, determines the target color value of the plurality of pixel points according to the color values of the plurality of pixel points in the depth map, the implicit color map and the explicit color map to which the target depth value of the plurality of pixel points belongs, so that the obtained synthesized image information has a correct occlusion relationship.

[0169] In a possible implementation, FIG. 9 is a flowchart of the image synthesis method for the robot provided in the embodiment of the application, as shown in FIG. 9, after the synthesized image information is obtained according to the implicit rendering result and the explicit rendering result in S103, the method can further include:

[0170] S701: training a visual control model of the work robot according to the synthesized image information.

[0171] S702: deploying the trained visual control module to the work robot to control the work robot to operate the target object in the preset static scene according to the visual control module.

[0172] In the embodiment, the synthesized image information is sent to the work robot as visual simulation input of the work robot, a visual control model or algorithm in the work robot generates hand-eye coordination motion control according to the input synthesized image information, physical operations such as grabbing, moving, clicking, and switching of the target object for explicit rendering are generated in the preset simulation environment according to the motion control action generated by the work robot, and the synthesized image information and the physical operation result in the preset simulation environment are used as training data to train the visual control model or algorithm of the work robot. After the training is completed, the visual control model or algorithm is deployed in the work robot.

[0173] The work robot acquires a collection image including a preset static scene and a target object through a real camera, and outputs a control instruction according to the collection image by using the visual control model or algorithm, so that the work robot operates the target object according to the control instruction.

[0174] The image synthesis method for the robot provided in the above embodiment trains the visual control model of the work robot by using the synthesized image information, can increase the sample data amount for training the visual control model, improves the training effect of the visual control model, and guarantees the work accuracy of the work robot.

[0175] On the basis of the method embodiment, the embodiment of the present application further provides an image synthesis device for a robot. FIG. 10 is a structural schematic diagram of the image synthesis device for the robot provided by the embodiment of the present application. As shown in FIG. 10, the device can include:

[0176] The implicit rendering module 801 is configured to acquire an implicit rendering result of a real camera arranged on the work robot for the preset static scene according to an implicit radiance field model pre-trained by the work robot for the preset static scene.

[0177] The explicit rendering module 802 is configured to generate an explicit rendering result of the target object in the preset simulation environment according to a relative placement position of the target object in the preset static scene.

[0178] The information synthesis module 803 is configured to acquire synthesized image information according to the implicit rendering result and the explicit rendering result.

[0179] Optionally, the implicit rendering result includes an implicit color map and an implicit depth map of the real camera on the work robot in a preset viewport direction.

[0180] The explicit rendering result includes an explicit color map and an explicit depth map of a virtual camera having the same parameters as the real camera on the work robot in the preset viewport direction.

[0181] The information synthesis module 803 is specifically configured to acquire the synthesized image information according to the implicit color map, the implicit depth map, the explicit color map and the explicit depth map.

[0182] Optionally, the implicit rendering module 801 comprises:

[0183] The light information acquisition unit is configured to acquire the origin coordinates and the attitude angle of a plurality of light rays emitted by the real camera in the preset viewport direction.

[0184] The implicit radiation unit is configured to determine the body density information and the color information of the plurality of light rays by using an implicit radiation field model according to the origin coordinates and the attitude angle of the plurality of light rays.

[0185] The implicit information generation unit is configured to determine the implicit color map and the implicit depth map according to the body density information and the color information of the plurality of light rays.

[0186] Optionally, the work robot comprises a mobile chassis and a mechanical arm, and the light information acquisition unit is specifically configured to calculate the camera extrinsic parameter of the real camera according to the odometer of the mobile chassis and the forward kinematics equation of the mechanical arm, and determine the origin coordinates and the attitude angle of the plurality of light rays according to the camera extrinsic parameter of the real camera.

[0187] Optionally, the implicit information generation unit is specifically configured to determine the implicit color map of the pixel points corresponding to the plurality of light rays according to the integration of the body density information and the color information of the plurality of light rays, and determine the implicit depth map of the pixel points corresponding to the plurality of light rays according to the jumping position of the body density information of the plurality of light rays.

[0188] Optionally, before the explicit rendering module 802, the apparatus can further comprise:

[0189] The placement position calculation module is configured to determine the relative placement position of the target object in the preset static scene according to the depth information in the implicit depth map.

[0190] Optionally, the explicit rendering module 802 is specifically configured to generate a simple collision body corresponding to the preset static scene in the preset simulation environment, and render and generate the target object on the simple collision body according to the relative placement position of the target object in the preset static scene.

[0191] Optionally, the information synthesis module 803 is specifically configured to determine the target depth value of a plurality of pixel points according to the depth values of the plurality of pixel points in the implicit depth map and the explicit depth map, the target depth value being the depth value in the implicit depth map or the depth value in the explicit depth map, determine the target color value of the plurality of pixel points according to the color values of the plurality of pixel points in the depth map to which the target depth value of the plurality of pixel points belongs, the implicit color map and the explicit color map, and the synthesized image information comprises the target depth value and the target color value of the plurality of pixel points.

[0192] Optionally, the information synthesizing module 803 is specifically configured to: if the depth value of the target pixel point in the explicit depth map is greater than or equal to the depth value in the implicit depth map, determine the target depth value of the target pixel point as the depth value of the target pixel point in the implicit depth map; and if the target depth value of the target pixel point belongs to the implicit depth map, determine the target color value of the target pixel point as the color value of the target pixel point in the implicit color map.

[0193] Optionally, the information synthesizing module 803 is specifically configured to: if the depth value of the target pixel point in the explicit depth map is less than the depth value in the implicit depth map, determine the target depth value of the target pixel point as the depth value of the target pixel point in the explicit depth map; and if the target depth value of the target pixel point belongs to the explicit depth map, determine the target color value of the target pixel point as the color value of the target pixel point in the explicit color map.

[0194] Optionally, after the information synthesizing module 803, the apparatus can further include:

[0195] a model training module configured to train a visual control model of the work robot according to the synthesized image information;

[0196] a model deployment module configured to deploy the trained visual control model to the work robot, so as to control the work robot to operate the target object in the preset static scene according to the visual control model.

[0197] The apparatus is used for executing the method provided in the foregoing embodiments, and has similar implementation principles and technical effects, which will not be described here.

[0198] The above modules can be one or more integrated circuits configured to implement the above method, for example, one or more application specific integrated circuits (ASICs), or one or more microprocessors, or one or more field programmable gate arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can invoke program code. For another example, the modules can be integrated together to be implemented in the form of a system on a chip (SOC).

[0199] FIG. 11 is a schematic diagram of an electronic device according to an embodiment of the present application. As shown in FIG. 11, the electronic device 100 can include a processor 1001, a storage medium 1002, and a bus. The storage medium 1002 stores program instructions executable by the processor 1001. When the electronic device 100 is running, the processor 1001 communicates with the storage medium 1002 through the bus. The processor 1001 executes the program instructions to perform the above method embodiments. The specific implementation and technical effects are similar, and will not be repeated here.

[0200] Optionally, the present application further provides a computer readable storage medium, and the storage medium stores a computer program. The computer program is executed by a processor to perform the above method embodiments.

[0201] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units explicitly or implicitly shown or discussed can be indirect coupling or communication connection through some interface, apparatus or unit, and can be electrical, mechanical or other forms.

[0202] The units described as separate components can or can not be physically separated, and the components explicitly shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.

[0203] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.

[0204] The integrated unit implemented in the form of the software function unit can be stored in a computer readable storage medium. The software function unit is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute part of steps of the method described in various embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media capable of storing program codes.

[0205] The above merely describes specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image synthesis method for a robot, characterized by, A working scene of a working robot includes a preset static scene and a target object in the preset static scene, the preset static scene does not need to be in physical contact with the working robot, and the target object needs to be in physical contact with the working robot, and the method comprises: According to the position of the real camera on the working robot in the preset static scene, an implicit radiance field model pre-trained for the preset static scene is used to obtain an implicit rendering result of the preset static scene; According to the relative placement position of the target object in the preset static scene, an explicit rendering result of the target object is generated in a preset simulation environment; According to the implicit rendering result and the explicit rendering result, synthetic image information is obtained.

2. The method of claim 1, wherein, The implicit rendering result includes an implicit color map and an implicit depth map of the real camera on the working robot in a preset viewport direction; The explicit rendering result includes a virtual camera with the same parameters as the real camera on the working robot, an explicit color map and an explicit depth map in the preset viewport direction; The implicit rendering result and the explicit rendering result include: According to the implicit color map, the implicit depth map, the explicit color map and the explicit depth map, the synthetic image information is obtained.

3. The method of claim 2, wherein, According to the position of the real camera on the working robot in the preset static scene, an implicit radiance field model pre-trained for the preset static scene is used to obtain an implicit rendering result of the preset static scene, including: Obtaining the origin coordinates and attitude angles of a plurality of light rays emitted by the real camera in the preset viewport direction; According to the origin coordinates and attitude angles of the plurality of light rays, the implicit radiance field model is used to determine the volume density information and color information of the plurality of light rays; According to the volume density information and color information of the plurality of light rays, the implicit color map and The implicit depth map is determined.

4. The method of claim 3, wherein, The working robot includes a mobile chassis and a mechanical arm, and the origin coordinates and attitude angles of the plurality of light rays emitted by the real camera in the preset viewport direction are obtained, including: According to the odometer of the mobile chassis and the forward kinematics equation of the mechanical arm, the camera extrinsic parameters of the real camera are calculated; According to the camera extrinsic parameters of the real camera, the origin coordinates and attitude angles of the plurality of light rays are determined.

5. The method of claim 3, wherein, According to the volume density information and color information of the plurality of light rays, the implicit color map and the implicit depth map are determined, including: According to the integration of the volume density information and color information of the plurality of light rays, the implicit color map of the pixel points corresponding to the plurality of light rays is determined; According to the jump position of the volume density information of the plurality of light rays, the implicit depth map of the pixel points corresponding to the plurality of light rays is determined.

6. The method of claim 2, wherein, Before the explicit rendering result of the target object is generated in the preset simulation environment according to the relative placement position of the target object in the preset static scene, the method further comprises: According to the depth information in the implicit depth map, the relative placement position of the target object in the preset static scene is determined.

7. The method of claim 1, wherein, The method further includes: generating an explicit rendering result of the target object in the preset simulation environment according to the relative placement position of the target object in the preset static scene, including: generating a simple collision body corresponding to the preset static scene in the preset simulation environment; 8. The method of claim 2, wherein, rendering the target object on the simple collision body according to the relative placement position of the target object in the preset static scene. The method further includes: acquiring the composite image information according to the implicit color map, the implicit depth map, the explicit color map and the explicit depth map, including: determining a target depth value of each of the plurality of pixel points according to the depth value of each of the plurality of pixel points in the implicit depth map and the explicit depth map, the target depth value being the depth value in the implicit depth map or the depth value in the explicit depth map; 9. The method of claim 8, wherein, determining a target color value of each of the plurality of pixel points according to the depth map to which the target depth value of each of the plurality of pixel points belongs, the color value of each of the plurality of pixel points in the implicit color map and the explicit color map, the composite image information including the target depth value and the target color value of each of the plurality of pixel points. The method further includes: if the depth value of the target pixel point in the explicit depth map is greater than or equal to the depth value in the implicit depth map, determining the target depth value of the target pixel point as the depth value of the target pixel point in the implicit depth map. The method further includes:

10. The method of claim 8, wherein, if the target depth value of the target pixel point belongs to the implicit depth map, determining the target color value of the target pixel point as the color value of the target pixel point in the implicit color map. The method further includes: if the depth value of the target pixel point in the explicit depth map is less than the depth value in the implicit depth map, determining the target depth value of the target pixel point as the depth value of the target pixel point in the explicit depth map. The method further includes:

11. The method of claim 1, wherein, if the target depth value of the target pixel point belongs to the explicit depth map, determining the target color value of the target pixel point as the color value of the target pixel point in the explicit color map. The method further includes: training a visual control model of the work robot according to the composite image information; deploying the trained visual control model to the work robot to control the work robot to operate the target object in the preset static scene according to the visual control model.

12. An image synthesizing apparatus for a robot, characterized by comprising: A working scene of a working robot includes a preset static scene and a target object in the preset static scene, the preset static scene does not need to be in physical contact with the working robot, the target object needs to be in physical contact with the working robot, and the device includes: An implicit rendering module is configured to obtain an implicit rendering result of the preset static scene by using an implicit radiation field model pre-trained for the preset static scene according to a position of a real camera on the working robot in the preset static scene; An explicit rendering module is configured to generate an explicit rendering result of the target object in a preset simulation environment according to a relative placement position of the target object in the preset static scene; An information synthesis module is configured to obtain synthesis image information according to the implicit rendering result and the explicit rendering result.

13. An electronic device, comprising: The device includes: A processor, a storage medium, and a bus, the storage medium stores program instructions executable by the processor, when the electronic device is running, the processor and the storage medium communicate through the bus, and the processor executes the program instructions to perform the steps of the image synthesis method for the robot according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is executed by the processor to perform the steps of the image synthesis method for the robot according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Three-dimensional scene rendering method and device and storage medium

    CN114119849A

  • Image processing method and image processing device

    CN117237514A

  • Image synthesis method and device for robot, equipment and medium

    CN118229863A

  • High resolution neural rendering

    US20220301257A1

Cited By

  • Space intelligent visual physical process inference method based on implicit physical large model

    CN121706990A