Model training method and computing device cluster

By using environment mapping and randomization of the application scenario during object grasping model training to generate image data, the robustness and generalization problems of the object grasping model in real-world scenarios are solved, achieving more accurate grasping pose determination and environmental adaptation.

CN121010692APending Publication Date: 2025-11-25HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411147458.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2024-08-20
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

The object grasping model has low robustness and generalization in real-world applications, mainly due to the difference between the training data and the actual application scenarios.

Method used

By providing environment textures for the application scenario, images containing objects are generated, and an object grasping model is trained based on these images. This includes random adjustments to the 3D model and randomization of environmental parameters to improve the model's adaptability.

Benefits of technology

It improves the robustness and generalization of the object grasping model in specific application scenarios, enabling it to more accurately determine the grasping pose and adapt to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010692A_ABST
    Figure CN121010692A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and a computing device cluster. The method comprises the steps that a first interface is provided, an environment map of an application scene is obtained from the first interface, and the environment map comprises grabbing equipment and transmission equipment in the application scene; a first image is obtained through rendering according to the environment map and a first three-dimensional model, the first three-dimensional model is used for representing a first object, and the first image comprises an image of the first object on the conveying device; and updating parameters of the object capturing model according to training data, wherein the training data comprises the first image. According to the method, the image for training the object grabbing model is generated according to the environment map provided by the user, the robustness and generalization of the object grabbing model can be improved, and therefore the grabbing requirement of the user for the actual application scene can be better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a model training method and a cluster of computing devices. Background Technology

[0002] With the development of industrial technology, the efficiency and accuracy of industrial systems are becoming increasingly important. In some industrial applications, object grasping models typically provide mechanized equipment with the grasping pose for grasping objects, and the equipment then grasps the objects in the application scenario based on this grasping pose.

[0003] In related technologies, object grasping models can be trained based on training data from different application scenarios. However, there are significant differences between the training data and the data from actual application scenarios, resulting in low robustness and generalization ability of object grasping models in real-world applications. Summary of the Invention

[0004] This application provides a model training method and a computing device cluster, which trains an object grasping model based on training data generated for the application scenario, thereby improving the robustness and generalization of the object grasping model in the application scenario.

[0005] In a first aspect, this application provides a model training method. The method includes: providing a first interface; obtaining an environment map of an application scenario from the first interface, the environment map including a grasping device and a transmission device in the application scenario; rendering a first image based on the environment map and a first three-dimensional model, the first three-dimensional model representing a first object, the first image including an image of the first object on the transmission device; updating parameters of an object grasping model based on training data, the training data including the first image, the object grasping model deployed on the grasping device, the object grasping model being used to determine the grasping pose of the object grasped by the grasping device based on the image.

[0006] In the above solution, an image containing objects is generated based on the environment texture provided by the user, and the object grasping model is trained based on this image. This allows the object grasping model to provide a more accurate grasping pose for the grasping device. In other words, this solution generates relevant training data based on the actual application scenario to train the object grasping model, which can improve the robustness and generalization of the object grasping model in that application scenario.

[0007] In one possible implementation, the method further includes: receiving a model determination request from the first interface, the model determination request being used to determine a second three-dimensional model, the second three-dimensional model being used to represent a second object; and randomly adjusting the geometric parameters and / or reflection parameters of the second three-dimensional model to obtain the first three-dimensional model.

[0008] In the above solution, the first three-dimensional model is obtained by randomly adjusting the second three-dimensional model provided by the user, which can generalize the object grasping model from the second object to the first object, thereby improving the generalization ability of the object grasping model to objects in the application scenario.

[0009] In one possible implementation, the first image comprises multiple images of multiple objects moving in the application scene at multiple moments, including the first object. The step of rendering the first image based on the environment map and the first 3D model includes: determining the poses of the multiple objects at multiple moments as they move in the application scene based on the initial poses of the multiple objects, multiple 3D models, and arrangement parameters. The multiple 3D models represent the multiple objects, including the first 3D model, and the arrangement parameters indicate the arrangement of the multiple objects. The multiple 3D models are then rendered based on the environment map, the material parameters of the multiple objects, and the poses of the multiple objects at multiple moments as they move in the application scene, to obtain multiple images of the multiple objects moving in the application scene at multiple moments.

[0010] In one possible implementation, before rendering the first image based on the environment map and the first 3D model, the method further includes: randomly generating the initial poses of the plurality of objects, the material parameters of the plurality of objects, and the arrangement parameters.

[0011] In one possible implementation, before rendering the first image based on the environment map and the first 3D model, the method further includes: randomly generating values ​​for the environment parameters according to the range of values ​​for the environment parameters; and adjusting the values ​​of the environment parameters in the environment map to the randomly generated values ​​of the environment parameters.

[0012] In the above solution, randomizing the environmental parameters of the user-provided environment texture can enable the object grasping model to better adapt to the environmental changes in the application scenario, thereby improving the generalization of the object grasping model to the environment in the application scenario.

[0013] In one possible implementation, the training data further includes the label value corresponding to the first image, which includes the grasping pose of the grasping device when grasping the first object.

[0014] The grasping pose for the first object can include the coordinates of the grasping point on the surface of the first object and the direction of the normal vector. Specifically, each point on each surface of the 3D model of the first object includes coordinates and a normal vector. Multiple points on each surface of the 3D model of the first object can be evaluated to obtain a grasping score for each point. Then, the grasping point of the first object is determined based on the grasping scores of multiple points on each surface. The normal vector of each point can include the normal vector of the plane containing the point, and the normal vector of the plane is perpendicular to the plane.

[0015] The grasping pose of the first object can also be determined based on the pose of the first object at multiple moments as it moves in the application scenario.

[0016] In one possible implementation, the training data further includes a second image and a label value corresponding to the second image. The second image includes an image of the second object on the transmission device. The second image is rendered based on the environment map and the second 3D model. The label value corresponding to the second image includes the grasping pose of the grasping device grasping the second object.

[0017] In one possible implementation, before updating the parameters of the object grasping model based on training data, the method further includes: sending a plurality of third images to a user through a second interface, the plurality of third images including the first image; and receiving an image determination request from the second interface, the image determination request being used to determine the first image.

[0018] Secondly, this application also provides a model training apparatus. The apparatus includes a display module, a generation module, and a training module.

[0019] The display module is used to provide a first interface and obtain an environment map of the application scenario from the first interface. The environment map includes the grabbing device and the transmission device in the application scenario.

[0020] The generation module is used to render a first image based on the environment texture and the first three-dimensional model. The first three-dimensional model is used to represent a first object, and the first image includes an image of the first object on the transmission device.

[0021] The training module is used to update the parameters of the object grasping model based on the training data, which includes the first image. The object grasping model is deployed on the grasping device and is used to determine the grasping pose of the object grasped by the grasping device based on the image.

[0022] In one possible implementation, the display module is further configured to receive a model determination request from the first interface, the model determination request being used to determine a second three-dimensional model, the second three-dimensional model being used to represent a second object. The generation module is further configured to randomly adjust the geometric parameters and / or reflection parameters of the second three-dimensional model to obtain the first three-dimensional model.

[0023] In one possible implementation, the first image comprises multiple images of multiple objects moving in the application scene at multiple moments, including the first object. The generation module is specifically configured to: determine the poses of the multiple objects at multiple moments as they move in the application scene based on the initial poses of the multiple objects, multiple 3D models, and arrangement parameters, wherein the multiple 3D models represent the multiple objects, including the first 3D model, and the arrangement parameters indicate the arrangement of the multiple objects; and render the multiple 3D models based on the environment map, the material parameters of the multiple objects, and the poses of the multiple objects moving in the application scene at multiple moments to obtain multiple images of the multiple objects moving in the application scene at multiple moments.

[0024] In one possible implementation, before the first image is rendered based on the environment map and the first 3D model, the generation module is further configured to: randomly generate the initial poses of the plurality of objects, the material parameters of the plurality of objects, and the arrangement parameters.

[0025] In one possible implementation, before rendering the first image based on the environment map and the first 3D model, the generation module is further configured to: randomly generate the values ​​of the environment parameters according to the range of values ​​of the environment parameters; and adjust the values ​​of the environment parameters in the environment map to the randomly generated values ​​of the environment parameters.

[0026] In one possible implementation, the training data further includes the label value corresponding to the first image, which includes the grasping pose of the grasping device when grasping the first object.

[0027] In one possible implementation, the training data further includes a second image and a label value corresponding to the second image. The second image includes an image of the second object on the transmission device. The second image is rendered based on the environment map and the second 3D model. The label value corresponding to the second image includes the grasping pose of the grasping device grasping the second object.

[0028] In one possible implementation, before updating the parameters of the object grasping model based on training data, the display module is further configured to send a plurality of third images to the user through a second interface, and receive an image determination request from the second interface, the plurality of third images including the first image, the image determination request being used to determine the first image.

[0029] Thirdly, this application also provides a computing device. The computing device includes a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the model training method provided by the first aspect or any possible implementation thereof.

[0030] Fourthly, this application also provides a computing device cluster. This computing device cluster includes multiple computing devices as provided in the third aspect.

[0031] Fifthly, this application also provides a computer-readable storage medium. This computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the model training method provided by the first aspect or any possible implementation thereof.

[0032] Sixthly, this application also provides a computer program product containing instructions. When the computer program product is run on a computer, it causes the computer to execute the model training method provided by the first aspect or any possible implementation thereof.

[0033] Any of the devices, computer storage media, or computer program products provided above are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects of the corresponding solutions in the corresponding methods provided above, and will not be repeated here. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of an object grasping application scenario provided in an embodiment of this application;

[0035] Figure 2 This is a flowchart of a model training method provided in an embodiment of this application;

[0036] Figure 3 This is a schematic diagram of data interaction during model training on a computing device in a cloud platform, provided in an embodiment of this application.

[0037] Figure 4 This is a schematic diagram of a computing device training model provided in an embodiment of this application;

[0038] Figure 5 This is a data interaction diagram of a crawling device training a model in an application scenario provided by an embodiment of this application;

[0039] Figure 6 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;

[0040] Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0041] Figure 8 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0042] Figure 9 This application provides an embodiment of a computing device cluster deployment Figure 6 A schematic diagram of the device shown. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings.

[0044] In the description of the embodiments of this application, the words "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the words "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a specific manner.

[0045] In the description of the embodiments in this application, the term "and / or" is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, B existing alone, and A and B existing simultaneously. Furthermore, unless otherwise stated, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals.

[0046] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0047] Before introducing the embodiments of this application, the terms used in the embodiments of this application will be explained below.

[0048] A 3D model is a digital representation of an object's shape, size, location, and surface properties in three-dimensional space. 3D models can be created in various ways, including manual modeling, scanning physical objects, and reconstruction from 2D images or videos. In practical applications, 3D models can take different representations, such as, but not limited to, point cloud models, mesh models, and models obtained through the truncated signed distance function (TSDF).

[0049] A texture image is a two-dimensional image that can imbue a three-dimensional model with rich surface visual properties. These properties can include, but are not limited to, color, texture details, material texture, and lighting effects, making the three-dimensional model more visually realistic and vivid. Texture images not only enhance the model's visual appeal but also allow users to better understand the model's shape, structure, and purpose by simulating real-world materials and textures.

[0050] Coordinate mapping is used to define how a two-dimensional texture image is mapped onto the surface of a three-dimensional model. For example, a UV map establishes a precise correspondence between points on the surface of a 3D model and pixels in the texture image based on a set of two-dimensional coordinates (U and V). In a UV map, the "U" and "V" axes represent the horizontal and vertical positions of the two-dimensional texture image, similar to the "X" and "Y" axes in a Cartesian coordinate system. UV mapping enables a natural fit of textures on the surface of a 3D model, thereby enhancing the visual effect of the 3D model and providing users with a more realistic experience.

[0051] Simulators are tools used to model 3D models of objects and simulate their movement. In object grasping scenarios, simulators can be used to simulate the movement of objects on conveyor devices, allowing the simulation to obtain the object's pose at different times.

[0052] A renderer is a tool used to convert 3D models into 2D images and is an important component of computer graphics. In object-grabbing scenarios, renderers can be used to render 3D models based on environmental information, resulting in 2D images that include both the objects and the environment.

[0053] Pose refers to the position and orientation of an object, which can be represented by three-dimensional coordinates and three-dimensional rotation angles.

[0054] Figure 1 This is a schematic diagram illustrating an object grasping application scenario provided in an embodiment of this application. This application scenario can be a logistics sorting scenario, which may include a grasping system, an object, and a conveying device. The grasping system can be used to grasp the object as it moves on the conveying device, thereby allowing the object to enter the next processing step.

[0055] Specifically, the crawling system may include, for example: Figure 1The diagram shows a photographing device 101 and a grasping device 102. The photographing device 101 can take pictures of the transmission device 103 at regular intervals and input the generated images into the grasping device 102. An object grasping model can be deployed in the grasping device 102. The object grasping model can process the input images to determine the grasping pose of the grasping device 102 when grasping an object. The grasping device 102 may include a robot capable of grasping. In some embodiments, the photographing device 101 and the grasping device 102 can be integrated into a single device, such as a robot with a photographing function.

[0056] In related technologies, object grasping models are trained based on images containing objects, enabling the model to predict the grasping pose of an object from an input image. However, this approach cannot generalize to unseen objects, exhibiting poor generalization. Furthermore, this approach suffers from overfitting, meaning the training data deviates significantly from the data in real-world applications, affecting the robustness of the object grasping model in real-world scenarios.

[0057] Therefore, this application provides a model training method that can solve the above-mentioned problems.

[0058] The method provided in this application embodiment uses an environment map of an application scenario provided by the user, then generates an image based on the environment map and a 3D model of an object, and updates the parameters of the object grasping model based on the image and its corresponding label values. In this method, generating an image based on the user-provided environment map of the application scenario, and then updating the parameters of the object grasping model using this image, can improve the robustness of the object grasping model to the application scenario, thereby better meeting the user's needs.

[0059] The model training method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0060] Figure 2 This is a flowchart of a model training method provided in an embodiment of this application. Figure 2 As shown, the method may include steps S201-S203 as follows. This method can be applied to a training device to train the object grasping model described above. The training device may include a computing device 104 in a cloud platform, the grasping device 102 described above, or the terminal device 103 described above. The terminal device 103 may include, but is not limited to, user devices such as smartphones, laptops, and tablets.

[0061] The following uses a cloud platform as an example. Figure 2 The steps shown are explained below.

[0062] In S201, a first interface is provided, from which the environment texture of the application scene is obtained.

[0063] A cloud platform client can be deployed on the user's terminal device 103. Once the user runs the client on terminal device 103, the client can display a primary interface. For example... Figure 3 As shown, users can send the environment texture of the object grasping model's application scenario to the computing device 104 in the cloud platform through this first interface. In this way, the computing device 104 can obtain the environment texture through this first interface.

[0064] This environment map can include Figure 1 The application scenario shown includes a grasping device 102 and a conveying device 103. The environment map can include a panoramic image of the application scenario captured by a user using a panoramic camera. This environment map can be a standard image or a high dynamic range (HDR) image.

[0065] In S202, a first image is rendered based on the environment map and the first 3D model. The first 3D model is used to represent the first object.

[0066] In this step, the computing device 104 can input the environment map and the information of the first object into the renderer, thereby rendering a first image of the first object moving in the application scene. This first image includes the image of the first object on the transmission device 103. The information of the first object may include texture images, UV maps, material parameters of the first object, the pose of the first object at one or more moments while moving in the application scene, and information such as the first 3D model.

[0067] Texture images and / or UV maps may also include texture images and UV maps from the cloud platform. Texture images and / or UV maps may also be provided by the user. Specifically, such as... Figure 3 As shown, the computing device 104 can also obtain the aforementioned texture image and / or UV map through the first interface.

[0068] The first 3D model may include a user-provided 3D model. In other words, the aforementioned object information may also include this first 3D model.

[0069] The first 3D model may also include 3D models from a 3D model library on a user-specified cloud platform. Specifically, prior to this step, the computing device 104 may also receive a model determination request from the user via a first interface. The model determination request is used to determine the first 3D model.

[0070] The first 3D model may further include a 3D model obtained by adjusting a second 3D model provided by the user or a second 3D model from a user-specified 3D model library. Specifically, prior to this step, the computing device 104 may receive a model determination request from the user via a first interface. The model determination request is used to determine a second 3D model, which represents a second object. Then, the computing device 104 may randomly adjust the geometric parameters and / or reflection parameters of the second 3D model to obtain the first 3D model. By randomizing the user-provided second 3D model, the generalization ability of the object grasping model to objects can be improved.

[0071] The aforementioned first image may also include multiple images taken at multiple moments as multiple objects move within the application scenario. In other words, the first image consists of multiple images captured while multiple simulated objects move on the transmission device, and each image may include one or more of the multiple objects. These multiple objects may all be the first object, or they may include the first object and other objects.

[0072] Specifically, the computing device 104 can randomly generate initial poses, material parameters, and arrangement parameters for multiple objects, and then simulate and render the images of these objects. The computing device 104 can utilize a pre-built simulator for simulation. For example... Figure 4 As shown, the computing device 104 can input information such as the initial poses of multiple objects, multiple 3D models, arrangement parameters, and movement speeds into the simulator. The simulator then simulates the movement of multiple objects on a transport device in an application scenario, thereby determining the poses of the multiple objects at multiple moments during their movement within the application scenario. Among the multiple 3D models input into the simulator is a first 3D model representing the first object. The computing device 104 can also generate 3D models of the multiple objects at multiple moments during their movement within the application scenario using the simulator. It should be noted that the 3D model of each object is identical at different moments. Furthermore, the computing device 104 can input environment maps into the simulator, thereby enabling a more realistic simulation of the movement of objects within the application scenario.

[0073] The simulator described above may include simulation software with simulation capabilities. The computing device 104 can call the application programming interface (API) provided by the simulation software to implement the simulation process described above. This simulation software may run on the computing device 104 or on other devices. Other devices may include, but are not limited to, devices on a cloud platform.

[0074] After the simulation, the computing device 104 can perform rendering using a pre-built-in renderer. For example... Figure 4As shown, the computing device 104 can input environment maps and information about multiple objects into a renderer, thereby rendering multiple 3D models to obtain multiple images of the objects moving within the application scene at multiple moments. The information for each object may include a texture image, UV map, material parameters, pose of the object at multiple moments during its movement within the application scene, and information about the 3D model. The texture images and / or UV maps of the multiple objects may be the same or different. These texture images and / or UV maps may be from a cloud platform or provided by the user.

[0075] The renderer described above may include rendering software with rendering capabilities. Similarly, the computing device 104 may call the application programming interface provided by the rendering software to implement the rendering process described above. Furthermore, the rendering software may run on the computing device 104 or on other devices. Other devices may include, but are not limited to, devices in a cloud platform.

[0076] Before rendering the first image, the computing device 104 can also randomly adjust the values ​​of the environment parameters in the environment map, thereby improving the generalization of the object grasping model to the environment. Specifically, the computing device 104 can randomly generate the values ​​of the environment parameters according to the range of values ​​of the environment parameters, and then adjust the values ​​of the environment parameters in the environment map to the randomly generated values ​​of the environment parameters. The range of values ​​of the environment parameters can be set by the user through interaction. The environment parameters may include one or more parameters such as light intensity, light direction, background color temperature and color difference. Among them, the offset angle can be changed by adjusting the z-axis rotation angle of the environment map, thereby changing the light direction.

[0077] In S203, the parameters of the object grasping model are updated based on the training data.

[0078] In this step, the computing device 104 can train the object grasping model using either an unsupervised learning method or a supervised learning method. This application embodiment does not impose specific limitations on the particular unsupervised or supervised learning methods used.

[0079] In the case of using an unsupervised learning method, the training data may include the first image described above. The computing device 104 can update the parameters of the object grasping model based on the first image. Furthermore, the training data may also include a second image. The second image includes the image of the second object on the conveying device. The second image can be obtained using the same generation method as the first image; that is, the second image can be generated based on the environment texture and the second 3D model provided by the user. The specific process will not be elaborated here.

[0080] In the case of using a supervised learning method, the training data includes a first image and its corresponding label value. The computing device 104 can update the parameters of the object grasping model based on the first image and its corresponding label value. The label value corresponding to the first image includes the grasping pose of the grasping device 102 when grasping the first object. The grasping pose of the grasping device 102 when grasping the first object can include the coordinates of the grasping points on the surface of the first object and the direction of the normal vector. Specifically, the computing device 104 can determine the coordinates and normal vectors of points on each surface of the three-dimensional model of the first object. Furthermore, the computing device 104 can evaluate multiple points on each surface of the three-dimensional model of the first object, obtain a grasping score for each point, and then determine the grasping point of the first object based on the grasping scores of multiple points on each surface. The normal vector of each point can include the normal vector of the plane containing the point, and the normal vector of the plane is perpendicular to the plane.

[0081] Specifically, the computing device 104 can input the first image into the object grasping model to obtain the grasping pose prediction value output by the object grasping model. Then, based on the grasping pose prediction value and the aforementioned label value, the loss value of the object grasping model is determined, and the parameters of the object grasping model are adjusted based on the loss value. It should be noted that this application embodiment does not impose specific limitations on the specific type and structure of the object grasping model, or the loss function, etc. The object grasping model can employ AI networks, including but not limited to neural networks.

[0082] In supervised learning, the training data may further include a second image and its corresponding label values. The second image can be generated based on a user-provided environment texture and a second 3D model. The specific generation process of the second image can be referred to the above description of the generation of the first image, and will not be repeated here.

[0083] In this embodiment, the computing device 104 can also send multiple third images to the user through a second interface, allowing the user to determine the first image from among the multiple third images. The multiple third images include the first image. These multiple third images can be generated according to the process described in S202. Specifically, the computing device 104 can receive an image determination request from the second interface, which is used to determine the first image. That is, the computing device 104 can determine, based on the image determination request, that the user instructs the training of the object grasping model based on the first image. By interacting with the user and allowing the user to select from multiple images, an image more suitable for the actual application scenario can be determined, thus further improving the robustness and generalization of the object grasping model.

[0084] In this embodiment, after the object grasping model has been trained to a certain extent, it can be tested using a test dataset. The test results are determined based on the model's output, and then randomization is performed again to generate images based on the test results. The user then selects training data again and trains the model. This process is repeated, ultimately enabling the object grasping model to achieve targeted robustness and generalization at a relatively low training cost. For example, if the object grasping model's output is inaccurate under a certain lighting intensity value, images generated based on environment maps matching that value can be added to the training data.

[0085] It should be noted that the above Figure 2 The steps shown are not limited to being performed by a single computing device in the cloud platform. In some embodiments, the steps can also be performed by multiple computing devices in the cloud platform. For example, the cloud platform includes computing devices 1041 to 1043. Computing devices 1041 to 1043 can each execute steps S201 to S203. Figure 2 The simulation and rendering steps in step S202 shown can also be performed by multiple computing devices. For example, computing device 1042 simulates and reconstructs the three-dimensional model of the object, and computing device 1043 renders the three-dimensional model.

[0086] The above Figure 2 In the illustrated embodiment, after the user uploads data through the data upload interface, the user-provided data is randomized, and training data for the object grasping model is generated based on the randomized data. Updating the parameters of the object grasping model based on the training data can improve the generalization ability of the object grasping model.

[0087] In some embodiments, the above Figure 2 The steps shown can also be applied to the above. Figure 1 In the gripping device 102 shown. That is to say, the gripping device 102 can be based on Figure 2 The steps shown train the object grasping model. Figure 5 As shown, the grasping device 102 can interact with the user, obtain an environment map provided by the user, generate a first image based on the environment map and the first 3D model, and update the parameters of the object grasping model based on the first image. The specific execution process of the grasping device 102 can be referred to the above. Figure 2 The description of the execution process of the computing device 104 in the illustrated embodiment will not be repeated here.

[0088] based on Figure 2 The method embodiments shown in this application also provide a model training device.

[0089] Figure 6This is a schematic diagram of the structure of a model training device 600 provided in this application. Figure 6 As shown, the model training device 600 includes a display module 601, a generation module 602, and a training module 603.

[0090] The display module 601 is used to provide a first interface and obtain an environment map of the application scene from the first interface. The environment map includes the grabbing device and the transmission device in the application scene.

[0091] The generation module 602 is used to render a first image based on the environment texture and the first three-dimensional model, wherein the first three-dimensional model is used to represent a first object, and the first image includes an image of the first object on the transmission device.

[0092] The training module 603 is used to update the parameters of the object grasping model based on the training data, which includes the first image. The object grasping model is deployed on the grasping device and is used to determine the grasping pose of the object grasped by the grasping device based on the image.

[0093] It should be noted that, Figure 6 The model training device 600 provided in the illustrated embodiment, when executing the model training method, is only illustrated by the division of the above-described functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the model training device provided in the above embodiment and... Figure 2 The model training method embodiments shown belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0094] Figure 6 The model training device 600 shown can be applied to the computing devices in the aforementioned cloud platform (e.g., the computing device 104 described above) to perform the aforementioned... Figure 2 The steps shown.

[0095] The display module 601, generation module 602, and training module 603 can all be implemented in software or in hardware. For example, the implementation of the display module 601 will be described below. Similarly, the implementation methods of the generation module 602 and training module 603 can be the same as those of the display module 601.

[0096] As an example of a software functional unit, display module 601 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, display module 601 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0097] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0098] As an example of a hardware functional unit, the display module 601 may include at least one computing device, such as a server. Alternatively, the display module 601 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0099] The multiple computing devices included in the display module 601 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the display module 601 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the display module 601 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0100] In some embodiments, Figure 6 The model training device 600 shown can also be applied to the grasping device 102 described above, for performing the above-mentioned tasks. Figure 2 The steps shown are as follows. Among them, the display module 601, the generation module 602, and the training module 603 can all be implemented by software.

[0101] Figure 7 This is a schematic diagram of the hardware structure of a computing device 700 provided in an embodiment of this application.

[0102] The computing device 700 can be the aforementioned computing device 104, the aforementioned terminal device 103, or the grasping device 102. See also Figure 7 The computing device 700 includes a processor 701, a memory 702, a communication interface 703, and a bus 704. The processor 701, memory 702, and communication interface 703 communicate via the bus 704. The processor 701, memory 702, and communication interface 703 may also communicate using other connection methods besides the bus 704. It should be understood that this application does not limit the number of processors and memories in the computing device 700.

[0103] Processor 701 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0104] The memory 702 may include volatile memory, such as random access memory (RAM). The processor 701 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD). The memory 702 stores executable program code, which the processor 701 executes to implement the functions of the aforementioned display module 601, generation module 602, and training module 603, thereby achieving… Figure 2 The model training method shown is as follows. That is, memory 702 stores the data used for execution. Figure 2 The instructions for the model training method shown. Alternatively, executable code is stored in memory 702, and processor 701 executes the executable code to implement the functions of the aforementioned model training device 600, thereby achieving... Figure 2 The model training method shown is as follows. That is, memory 702 stores the data used for execution. Figure 2 The instructions for the model training method are shown.

[0105] The communication interface 703 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 700 and other devices or communication networks.

[0106] The 704 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus 104 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 104 may include a path for transmitting information between various components of the computing device 700 (e.g., memory 702, processor 701, communication interface 703).

[0107] The aforementioned devices can be disposed on separate chips, or at least partially or entirely on the same chip. Whether to dispose of the devices independently on different chips or integrate them on one or more chips often depends on the needs of the product design. This application does not limit the specific implementation of the aforementioned devices.

[0108] Figure 7The computing device 700 shown is merely exemplary. In the implementation process, the computing device 700 may also include other components, which will not be listed one by one in this article.

[0109] based on Figure 2 The method embodiments shown in this application also provide a computing device cluster.

[0110] Figure 8 This application provides a computing device cluster 800.

[0111] The computing device cluster 800 may include the cloud platform described above. The computing device cluster includes at least one computing device. This computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0112] like Figure 8 As shown, the computing device cluster includes at least one computing device 700. The memory 702 of one or more computing devices 700 in the computing device cluster may store the same memory for executing... Figure 2 The instructions for the method shown.

[0113] In some possible implementations, the memory 702 of one or more computing devices 700 in the computing device cluster may also store memory for execution. Figure 2 The instructions of the method shown are partial. In other words, a combination of one or more computing devices 700 can jointly execute instructions for performing... Figure 2 The instructions for the method shown.

[0114] It should be noted that the memories 702 in different computing devices 700 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the model training device 600. That is, the instructions stored in the memories 702 of different computing devices 700 can implement the functions of one or more modules among the display module 601, the generation module 602, and the training module 603.

[0115] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 9 One possible implementation is shown. For example... Figure 9As shown, two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 702 in computing device 700A stores instructions for executing the functions of display module 601. Simultaneously, the memory 702 in computing device 700B stores instructions for executing the functions of generation module 602 and training module 603. These include display module 601, generation module 602, and training module 603.

[0116] Figure 9 The connection method between the computing device clusters shown can be such that, considering the model training method provided in this application requires a large amount of storage and computing resources (e.g., large amounts of data storage), the functions implemented by the generation module 602 and the training module 603 are delegated to the computing device 700B. It should be understood that... Figure 9 The functions of the computing device 700A shown can also be performed by multiple computing devices 700. Similarly, the functions of the computing device 700B can also be performed by multiple computing devices 700.

[0117] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 8 and Figure 9 The connection method of the computing device cluster. The difference is that the memory 702 of one or more computing devices 700 in this computing device cluster can store the same information for execution. Figure 2 The instructions for training the model shown.

[0118] In some possible implementations, the memory 702 of one or more computing devices 700 in the computing device cluster may also store memory for execution. Figure 2 The instructions for the model training method shown are partial. In other words, a combination of one or more computing devices 700 can jointly execute instructions for training the model. Figure 2 The instructions for training the model shown.

[0119] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform... Figure 2 The model training method shown.

[0120] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute... Figure 2 The model training method shown, or instructions to the computing device to perform... Figure 2 The model training method shown.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A model training method, characterized in that, The method includes: A first interface is provided, from which an environment map of the application scenario is obtained. The environment map includes the grasping device and the transmission device in the application scenario. A first image is obtained by rendering the environment map and the first 3D model, wherein the first 3D model is used to represent the first object, and the first image includes an image of the first object on the transmission device. The parameters of the object grasping model are updated using training data, which includes the first image. The object grasping model is deployed on the grasping device and is used to determine the grasping pose of the object grasped by the grasping device based on the image.

2. The method according to claim 1, characterized in that, The method further includes: A model determination request is received from the first interface. The model determination request is used to determine a second three-dimensional model, which is used to represent a second object. The geometric parameters and / or reflection parameters of the second three-dimensional model are randomly adjusted to obtain the first three-dimensional model.

3. The method according to claim 1 or 2, characterized in that, The first image comprises multiple images of multiple objects moving within the application scenario at multiple moments, including the first object. The process of rendering the first image based on the environment texture and the first 3D model includes: Based on the initial poses of the multiple objects, multiple 3D models, and arrangement parameters, the poses of the multiple objects at multiple moments when they move in the application scenario are determined. The multiple 3D models are used to represent the multiple objects, and the multiple 3D models include the first 3D model. The arrangement parameters are used to indicate the arrangement of the multiple objects. Based on the environment texture, the material parameters of the multiple objects, and the poses of the multiple objects at multiple moments when they move in the application scene, the multiple 3D models are rendered to obtain multiple images of the multiple objects at multiple moments when they move in the application scene.

4. The method according to claim 3, characterized in that, Before rendering the first image based on the environment map and the first 3D model, the method further includes: The initial poses, material parameters, and arrangement parameters of the multiple objects are randomly generated.

5. The method according to any one of claims 1-4, characterized in that, Before rendering the first image based on the environment map and the first 3D model, the method further includes: The values ​​of the environmental parameters are randomly generated based on their range. The values ​​of the environment parameters in the environment map are adjusted to randomly generated values.

6. The method according to any one of claims 1-5, characterized in that, The training data also includes the label value corresponding to the first image, which includes the grasping pose of the grasping device when grasping the first object.

7. The method according to any one of claims 2-6, characterized in that, The training data also includes a second image and a label value corresponding to the second image. The second image includes an image of the second object on the transmission device. The second image is rendered based on the environment map and the second 3D model. The label value corresponding to the second image includes the grasping pose of the grasping device when grasping the second object.

8. The method according to any one of claims 1-7, characterized in that, Before updating the parameters of the object grasping model based on the training data, the method further includes: Multiple third images are sent to the user through a second interface, the multiple third images including the first image; The second interface receives an image determination request, which is used to determine the first image.

9. A model training device, characterized in that, The device includes: The display module is used to provide a first interface and obtain an environment map of the application scenario from the first interface. The environment map includes the grabbing device and the transmission device in the application scenario. A generation module is used to render a first image using the environment texture and the first three-dimensional model, wherein the first three-dimensional model is used to represent a first object, and the first image includes an image of the first object on the transmission device. A training module is used to update the parameters of an object grasping model using training data, the training data including the first image. The object grasping model is deployed on the grasping device and is used to determine the grasping pose of the object grasped by the grasping device based on the image.

10. The apparatus according to claim 9, characterized in that, The display module is also configured to receive a model determination request from the first interface, the model determination request being used to determine a second three-dimensional model, the second three-dimensional model being used to represent a second object; The generation module is also used to randomly adjust the geometric parameters and / or reflection parameters of the second three-dimensional model to obtain the first three-dimensional model.

11. The apparatus according to claim 9 or 10, characterized in that, The first image includes multiple images of multiple objects moving in the application scenario at multiple moments, wherein the multiple objects include the first object, and the generation module is specifically used for: Based on the initial poses of the multiple objects, multiple 3D models, and arrangement parameters, the poses of the multiple objects at multiple moments when they move in the application scenario are determined. The multiple 3D models are used to represent the multiple objects, and the multiple 3D models include the first 3D model. The arrangement parameters are used to indicate the arrangement of the multiple objects. Based on the environment texture, the material parameters of the multiple objects, and the poses of the multiple objects at multiple moments when they move in the application scene, the multiple 3D models are rendered to obtain multiple images of the multiple objects at multiple moments when they move in the application scene.

12. The apparatus according to claim 11, characterized in that, Before rendering the first image based on the environment texture and the first 3D model, the generation module is further configured to: The initial poses, material parameters, and arrangement parameters of the multiple objects are randomly generated.

13. The apparatus according to any one of claims 9-12, characterized in that, Before rendering the first image based on the environment texture and the first 3D model, the generation module is further configured to: The values ​​of the environmental parameters are randomly generated based on their range. Adjust the values ​​of the environment parameters in the environment map to randomly generated values.

14. The apparatus according to any one of claims 9-13, characterized in that, The training data also includes the label value corresponding to the first image, which includes the grasping pose of the grasping device when grasping the first object.

15. The apparatus according to any one of claims 10-14, characterized in that, The training data also includes a second image and a label value corresponding to the second image. The second image includes an image of the second object on the transmission device. The second image is rendered based on the environment map and the second 3D model. The label value corresponding to the second image includes the grasping pose of the grasping device when grasping the second object.

16. The apparatus according to any one of claims 9-15, characterized in that, Before updating the parameters of the object grasping model based on the training data, the display module is also used to send multiple third images to the user through the second interface, and to receive an image determination request from the second interface, wherein the multiple third images include the first image, and the image determination request is used to determine the first image.

17. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method according to any one of claims 1 to 8.

19. A computer program product, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method according to any one of claims 1 to 8.