Training data generating device, training data generating method using the same, and robotic arm system

By generating virtual scene simulation and machine learning technology for training data, the occlusion problem when the robotic arm picks up objects is solved, and high-precision actual object picking operations are achieved.

CN116117828BActive Publication Date: 2025-09-16IND TECH RES INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111536541.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-12
Filing Date
2021-12-15
Publication Date
2025-09-16
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

When the robotic arm is picking up objects, it is easy for the object to be blocked, resulting in material jamming, object damage, and failure to pick up the object.

Method used

Through the training data generation device, training data is generated using the virtual scene simulation unit, the vertical projection virtual camera unit and the perspective projection virtual camera unit. Combined with machine learning technologies such as Mask R-CNN, the occlusion status of the actual object can be accurately judged and the robotic arm can be controlled to accurately pick up objects.

Benefits of technology

The success rate of the robotic arm in actual object retrieval has been improved, and problems such as object damage and material jamming have been reduced. The average accuracy rate has been increased to 70-95%, which is significantly better than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116117828B_ABST
    Figure CN116117828B_ABST
Patent Text Reader

Abstract

The training data generation device includes a virtual scene simulation unit, a vertical projection virtual camera unit, an object occlusion determination unit, a perspective projection virtual camera unit, and a training data generation unit. The virtual scene simulation unit is used to generate a virtual scene containing multiple objects. The vertical projection virtual camera unit is used to extract a vertical projection image of the virtual scene. The object occlusion determination unit is used to mark the occlusion status of each object based on the vertical projection image. The perspective projection virtual camera unit is used to extract a perspective projection image of the virtual scene. The training data generation unit is used to generate training data for the virtual scene based on the perspective projection image and the occlusion status of each object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a training data generating device, a training data generating method using the same, and a robotic arm system. Background Art

[0002] In the field of object retrieval, when a robotic arm retrieves an obscured object (covered by another object), it can easily cause the object to become stuck, damaged, and / or fail to retrieve. Therefore, improving the accuracy of object retrieval is one of the goals of technicians in this field. Summary of the Invention

[0003] The present invention discloses a training data generating device and a training data generating method using the same.

[0004] According to one embodiment of the present disclosure, a training data generating device is proposed. The training data generating device includes a virtual scene simulation unit, a vertical projection virtual camera unit, an object occlusion judgment unit, a perspective projection virtual camera unit, and a training data generating unit. The virtual scene simulation unit is used to generate a virtual scene, and the virtual scene includes multiple objects. The vertical projection virtual camera unit is used to extract a vertical projection image of the virtual scene. The object occlusion judgment unit is used to mark the occlusion status of each object based on the vertical projection image. The perspective projection virtual camera unit is used to extract a perspective projection image of the virtual scene. The training data generating unit is used to generate training data for the virtual scene based on the perspective projection image and the occlusion status of each object.

[0005] According to another embodiment of the present disclosure, a training data generation method is provided. The training data generation method includes the following steps: generating a virtual scene, the virtual scene including multiple objects; extracting a vertical projection of the virtual scene; marking the occlusion status of each object based on the vertical projection; extracting a perspective projection of the virtual scene; and generating training data for the virtual scene based on the perspective projection and the occlusion status of each object.

[0006] According to another embodiment of the present disclosure, a robotic arm system is provided. The robotic arm system includes a robotic arm; a perspective projection camera disposed on the robotic arm and configured to capture images of a plurality of real objects; and a controller configured to generate a learning model based on the aforementioned training data generation method, analyze the real object images using the learning model, and thereby determine the actual occlusion state of each real object; and control the robotic arm to extract real objects whose actual occlusion state is unobstructed.

[0007] In order to better understand the above and other aspects of the present disclosure, the following embodiments are specifically described in detail with reference to the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1AFIG. 4 is a functional block diagram of a training data generating device according to an embodiment of the present disclosure.

[0009] Figure 1B Schematic diagram of a robotic arm system according to another embodiment of the present disclosure.

[0010] Figure 2 Schematic diagram of the virtual scene in Figure 1.

[0011] Figure 3A This is a schematic diagram of the perspective projection virtual camera unit in FIG1 extracting a perspective projection image.

[0012] Figure 3B for Figure 2 Schematic diagram of the perspective projection diagram of the virtual scene.

[0013] Figure 4A 1 is a schematic diagram of extracting a vertical projection image by the vertical projection virtual camera unit.

[0014] Figure 4B for Figure 2 Schematic diagram of the vertical projection of the virtual scene.

[0015] Figure 5A for Figure 2 Schematic diagram of a judgment scenario where the subject is the judge.

[0016] Figure 5B for Figure 5A Schematic diagram of determining the vertically projected object image of a scene.

[0017] Figure 5C for Figure 4B Schematic diagram of the vertical projection object image.

[0018] Figure 6 For application Figure 1A Flowchart of a training data generating method of a training data generating device.

[0019] Figure 7 for Figure 6 Flowchart of the method for determining the shielding state of the step.

[0020] [Description of Reference Numerals]

[0021] 1: Robotic arm system

[0022] 10: Virtual Scene

[0023] 100: Training data generating device

[0024] 10e: Perspective Projection

[0025] 10v: Vertical projection

[0026] 11-14: Subject

[0027] 110: Virtual scene simulation unit

[0028] 11e~14e:Viewing angle projection object image

[0029] 11v~14v:Vertical projection of the target image

[0030] 12': Subject of judgment scene

[0031] 120: Vertical projection virtual camera unit

[0032] 130: Object occlusion judgment unit

[0033] 140: Perspective Projection Virtual Camera Unit

[0034] 15: Container

[0035] 150: Training data generation unit

[0036] 160: Machine Learning Unit

[0037] 22: Actual Object

[0038] G1: bearing surface

[0039] M1: Learning Model

[0040] ST: shielded state

[0041] TR: Training data

[0042] V1: Vertical line of sight DETAILED DESCRIPTION

[0043] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0044] Please refer to Figure 1A ~5, Figure 1A is a functional block diagram of a training data generating apparatus 100 according to an embodiment of the present disclosure. Figure 1B is a schematic diagram of a robotic arm system 1 according to another embodiment of the present disclosure. Figure 2 is a schematic diagram of the virtual scene 10 in FIG1 , Figure 3A 1 is a schematic diagram of the perspective projection virtual camera unit 140 extracting the perspective projection diagram 10e, Figure 3B for Figure 2 FIG10e is a schematic diagram of a perspective projection of a virtual scene 10, Figure 4A 1 is a schematic diagram of the vertical projection virtual camera unit 120 extracting the vertical projection image 10v, and Figure 4B for Figure 2FIG10 v is a schematic diagram of a vertical projection of a virtual scene 10. ...

[0045] like Figure 1A As shown, the training data generation device 100 includes a virtual scene simulation unit 110, an orthographic virtual camera unit 120, an object occlusion determination unit 130, a perspective virtual camera unit 140, and a training data generation unit 150. At least one of the virtual scene simulation unit 110, the orthographic virtual camera unit 120, the object occlusion determination unit 130, the perspective virtual camera unit 140, and the training data generation unit 150 may be, for example, software, firmware, or a physical circuit formed by a semiconductor process. At least one of the virtual scene simulation unit 110, the orthographic virtual camera unit 120, the object occlusion determination unit 130, the perspective virtual camera unit 140, and the training data generation unit 150 may be integrated into a single unit or integrated into a processor or controller.

[0046] like Figure 1A As shown, the virtual scene simulation unit 110 is used to generate a virtual scene 10, and the virtual scene 10 includes a plurality of objects 11 to 14. The vertical projection virtual camera unit 120 is used to extract a vertical projection image 10v of the virtual scene 10 (the vertical projection image 10v is in FIG. Figure 4B The object occlusion determination unit 130 is used to mark each object 11 to 14 (objects 11 to 14 are in the vertical projection image 10v). Figure 2 ) is occluded in the state ST. The perspective projection virtual camera unit 140 is used to extract the perspective projection image 10e of the virtual scene 10 (objects 11 to 14 are in the Figure 3B ). The training data generating unit 150 is used to generate the training data TR of the virtual scene 10 based on the perspective projection image 10e and the occlusion status ST of each object 10. For example, each piece of training data TR includes the perspective projection image 10e and the occlusion status ST of each object 10. Compared with the perspective projection image 10e, the vertical projection image 10v can better reflect the actual occlusion status of the objects 11 to 14, so that the occlusion status ST determined based on the vertical projection image 10v can better conform to the actual occlusion status of the object. In this way, the learning model M1 generated by subsequent training based on the training data TR (the learning model M1 in Figure 1A ) can be used to more accurately determine the actual shading state of the actual object. The aforementioned shading state ST, for example, includes multiple categories, such as two categories such as "shaded" and "unshaded".

[0047] The "projected virtual camera unit" herein, for example, utilizes computer vision technology to project a three-dimensional (3D) image onto a reference surface (e.g., a plane) to create a two-dimensional (2D) image. Furthermore, the aforementioned virtual scene 10, perspective projection image 10e, and vertical projection image 10v may be data generated during the unit's operation and are not necessarily actual images.

[0048] like Figure 1A As shown, the training data generating device 100 can generate multiple virtual scenes 10 and generate training data TR for each virtual scene 10. The machine learning unit 160 can learn the occlusion judgment of the object based on the multiple training data TR and produce a learning model M1. The machine learning unit 160 can be a subcomponent of the training data generating device 100, or it can be configured independently of the training data generating device 100, or the machine learning unit 160 can be configured in the robot arm system 1 (the robot arm system 1 is in Figure 1B In addition, the learning model M1 generated by the machine learning unit 160 is obtained by, for example, using Mask R-CNN (object segmentation), Faster-R-CNN (object detection) or other types of machine learning technologies or algorithms.

[0049] like Figure 1B As shown, the robotic arm system 1 includes a robotic arm 1A, a controller 1B and a perspective projection camera 1C. The perspective projection camera 1C is an actual (physical) camera. The robotic arm 1A can extract the actual object 22, and the extraction method is, for example, clamping or sucking (vacuum suction or magnetic suction). In actual object-picking applications, the perspective projection camera 1C can extract the actual object images P1 of multiple actual objects 22 below. The controller 1B is electrically connected to the robotic arm 1A and the perspective projection camera 1C, and can judge the actual shading status of these actual objects 22 based on (or analyze) the actual object image P1, and control the robotic arm 1A to extract the actual objects 22 below accordingly. The controller 1B can be based on (or loaded with) the learning model M1 completed by the above training. Since the learning model M1 generated by training according to the training data TR (the learning model M1 is then Figure 1A ) can be used to more accurately determine the actual obstruction state of the actual object 22. This allows the controller 1B of the robotic arm system 1 to accurately determine the actual obstruction state of the actual object 22 in real-world scenarios based on the learned model M1, and further control the robotic arm 1A to accurately extract the actual object 22, such as extracting a completely unobstructed actual object 22 (e.g., an actual object with an actual obstruction state of "unobstructed"). In this way, in actual object-grabbing applications, the success rate of object grazing can be increased and / or problems such as object damage and material jamming can be reduced.

[0050] In addition, if Figure 1BAs shown, the perspective projection camera 1C used in the robotic arm system 1 captures the actual object image P1 using the perspective projection principle or method. The actual object image P1 captured by the perspective projection camera 1C is based on the same perspective projection principle as the perspective projection diagram 10e learned by the learning model M1. Therefore, the controller 1B can accurately determine the object's occlusion status.

[0051] like Figure 1A and 2 As shown, the virtual scene simulation unit 110 uses, for example, V-REP (Virtual Robot Experiment Platform) technology to generate virtual scenes 10. Each virtual scene 10 is, for example, a randomly generated three-dimensional (3D) image. Figure 2 For example, virtual scene 10 includes virtual objects 11-14 and a virtual container 15, with objects 11-14 located within container 15. The stacking, arrangement, and / or number of the multiple objects in each virtual scene 10 are randomly generated by virtual scene simulation unit 110, such as scattered stacking, random placement, and / or close placement. Therefore, the stacking, arrangement, and / or number of the multiple objects in each virtual scene 10 are not exactly the same. Furthermore, the objects in virtual scene 10 can be any type of object, such as hand tools, stationery, plastic bottles, food ingredients, mechanical parts, and other objects used in factories, such as machinery processing plants, food processing plants, stationery factories, resource recycling plants, electronic device processing plants, and electronic device assembly plants. The 3D models of the objects and / or container 15 can be pre-created using modeling software, for example. The virtual scene simulation unit 110 generates the virtual scene 10 based on these pre-modeled 3D models of the objects and / or container 15.

[0052] like Figure 3A As shown, the perspective projection virtual camera unit 140 shoots the virtual scene 10 as if it were a general perspective projection camera (with a cone-shaped shooting perspective A1), and extracts the perspective projection image 10e (the perspective projection image 10e is in Figure 3B ).like Figure 3B As shown, the perspective projection diagram 10e is, for example, a perspective projection 2D image of the virtual scene 10 (3D image), which has a corresponding Figure 2 11e-14e of the objects 11-14 are projected from perspective. From perspective projection image 10e, perspective projection object images 11e-13e overlap. For example, perspective projection object image 12e is obscured by perspective projection object image 11e. In other words, the object obscuration state of perspective projection object image 12e, as determined from perspective projection image 10e, is "obscured." However, this is not the actual obscuration state of perspective projection object image 12e in virtual scene 10.

[0053] like Figure 4AAs shown, the vertical projection virtual camera unit 120 shoots the virtual scene 10 with a vertical line of sight V1, and extracts a vertical projection image 10v (the vertical projection image 10v is shown in FIG. Figure 4B As shown). Figure 4B As shown, the vertical projection image 10v is, for example, a vertically projected two-dimensional image of the virtual scene 10 (three-dimensional image), which has a corresponding Figure 2 11v-14v of objects 11-14 are shown. From vertical projection image 10v, vertical projection object image 12v is not obscured by vertical projection object image 11v. In other words, the actual obscuration status of vertical projection object image 12v, as determined from vertical projection image 10v, is "unobstructed." Because the obscuration status ST of an object in the disclosed embodiment is determined from vertical projection image 10v, it can reflect the actual obscuration status of each object in virtual scene 10.

[0054] like Figure 4A As shown, the objects 11 to 14 (or container 15) are placed relative to the supporting surface G1, and the vertical projection 10v is an image generated by viewing the virtual scene 10 with a vertical line of sight V1. The vertical line of sight V1 described herein is, for example, perpendicular to the supporting surface G1. In this embodiment, the supporting surface G1 is, for example, a horizontal plane; however, in another embodiment, the supporting surface G1 can be an inclined plane, and the vertical line of sight V1 is also perpendicular to the inclined supporting surface G1. Alternatively, the direction of the vertical line of sight V1 can be defined to be consistent with or parallel to the object-picking direction of the robotic arm (not shown).

[0055] In addition, the embodiments of the present disclosure may utilize multiple methods to determine the shielding status of an object, one of which is described below.

[0056] Please refer to Figures 5A to 5C , Figure 5A for Figure 2 The object 12 is a schematic diagram of a judged subject scene 12', Figure 5B for Figure 5A The vertical projection object image 12v' of the judged object scene 12' is shown in FIG. Figure 5C Draw Figure 4B Schematic diagram of the vertically projected object image 12v.

[0057] The object occlusion judgment unit 130 is also used to: hide objects other than the object being judged of these objects in the virtual scene, and correspondingly generate an object scene of the object being judged. The vertical projection virtual camera unit 120 is also used to: obtain a vertical projection object image of the object scene of the object being judged. The object occlusion judgment unit 130 is also used to: obtain a difference value between the vertical projection object image and the vertical projection object image of the object being judged in the vertical projection diagram; and, determine the occlusion state ST based on this difference value. The object occlusion judgment unit 130 is used to: determine whether the difference value is less than a default value; when the difference value is less than the default value, determine the occlusion state of the object being judged to be "unoccluded"; and, when the difference value is not less than the default value, determine the occlusion state of the object being judged to be "occluded". The unit of the above-mentioned "default value" is, for example, the number of pixels. The embodiment of the present disclosure does not limit the numerical range of the "default value", which may depend on the actual situation.

[0058] For example, taking object 12 as the subject of judgment, Figure 5A As shown, the object shielding judgment unit 130 hides Figure 2 For example, the object occlusion judgment unit 130 sets the object 12 to the "visible state", but sets the remaining objects 11 and 13 to 14 to the "invisible state". Then, the object occlusion judgment unit 130 generates a judged object scene 12' for the object 12. Figure 5B As shown, the vertical projection virtual camera unit 120 extracts the vertical projection object image 12v' of the subject object scene 12'. The object occlusion judgment unit 130 obtains the vertical projection object image 12v' of the subject object scene 12' and the vertical projection object image 12v of the object 12 in the vertical projection image 10v (the vertical projection object image 12v is in FIG. Figure 5C ); and determining the obstruction state ST based on this difference value. If the difference value between the vertically projected object image 12v and the vertically projected object image 12v' is substantially zero or less than a default value, it indicates that the object 12 is not obstructed, and the object obstruction determination unit 130 may accordingly determine the obstruction state ST of the object 12 as "unobstructed." Conversely, if the difference value between the vertically projected object image 12v and the vertically projected object image 12v' is not equal to zero or greater than the default value, it indicates that the object 12 is obstructed, and the object obstruction determination unit 130 may accordingly determine the obstruction state ST of the object 12 as "obstructed." Furthermore, the aforementioned default value is greater than zero, such as 50, but may also be any integer less than or greater than 50.

[0059] The same method can be used to determine the occlusion status of each object in the virtual scene 10. Figure 2For example, the training data generating device 100 can use the same method to determine the occlusion states ST of the objects 11, 13, and 14 in the virtual scene 10 as "occluded", "unoccluded", and "unoccluded", respectively. Figure 4B shown.

[0060] Please refer to Figure 6 , which is in Figure 1A Flowchart of the training data generating method of the training data generating device 100.

[0061] In step S110, Figure 2 As shown, the virtual scene simulation unit 110 uses, for example, V-REP technology to generate a virtual scene 10 , which includes a plurality of virtual objects 11 - 14 and a virtual container 15 .

[0062] In step S120, Figure 4B As shown, the vertical projection virtual camera unit 120 utilizes, for example, computer vision technology to extract a vertical projection image 10v of the virtual scene 10. This extraction method, for example, utilizes computer vision technology to convert the three-dimensional virtual scene 10 into a two-dimensional vertical projection image 10v. However, as long as the three-dimensional virtual scene 10 can be converted into a two-dimensional vertical projection image 10v, the present disclosure is not limited to the processing technology employed; it may even be any suitable existing technology.

[0063] In step S130, Figure 1B As shown, the object occlusion determination unit 130 marks the occlusion status ST of each object 11 to 14 according to the vertical projection image 10v. The method for determining the occlusion status ST of the object has been described above and will not be repeated here.

[0064] In step S140, Figure 3B As shown, the perspective projection virtual camera unit 140 uses, for example, computer vision technology to extract a perspective projection image 10e of the virtual scene 10. This extraction method, for example, uses computer vision technology to convert the three-dimensional virtual scene 10 into a two-dimensional perspective projection image 10e. However, as long as the three-dimensional virtual scene 10 can be converted into a two-dimensional perspective projection image 10e, the present disclosure is not limited to the processing technology used; it can even be any suitable existing technology.

[0065] In step S150, Figure 1B As shown, the training data generation unit 150 generates training data TR for the virtual scene 10 based on the view projection image 10e and the occlusion status ST of each object 11-14. For example, the training data generation unit 150 integrates the data of the view projection image 10e and the data of the occlusion status ST of each object 11-14 into a piece of training data TR.

[0066] After obtaining a set of training data TR, the training data generation device 100 repeats steps S110-S150 to generate the next set of training data TR for the virtual scene 10, which includes the perspective projection image 10e of the next virtual scene 10 and the occlusion status ST of each object. Based on this principle, the training data generation device 100 can obtain N sets of training data TR. The present disclosure does not limit the value of N; it can be a positive integer equal to or greater than 2, such as 10, 100, 1000, 10,000, or even more or less.

[0067] Please refer to Figure 7 , which is Figure 6 The following is a flowchart of the method for determining the shielding state in step 130. Figure 2 The object 12 is taken as an example to be judged. The judgment method of the shielding state ST of other objects is the same and will not be repeated here.

[0068] In step S131, the object occlusion determination unit 130 hides the objects 11 and 13-14 other than the object 12 in the virtual scene 10, and correspondingly generates a determined object scene 12' of the object 12, such as Figure 5A shown.

[0069] In step S132, the vertical projection virtual camera unit 120 uses computer vision technology, for example, to extract the vertical projection object image 12v' of the subject object scene 12', such as Figure 5B The method of extracting the vertical projection object image 12v' is the same as the method of extracting the vertical projection image 10v.

[0070] In step S133, the object occlusion determination unit 130 obtains the vertically projected object image 12v' (the vertically projected object image 12v' is Figure 5B ) and vertical projection 10v (vertical projection 10v in Figure 4B ) of the vertical projection object image 12v (the vertical projection object image 12v is Figure 5C For example, the object occlusion determination unit 130 performs a difference operation between the vertically projected object image 12v′ and the vertically projected object image 12v, and uses the absolute value of the difference as the difference value.

[0071] In step S134, the object occlusion determination unit 130 determines whether the difference value is less than a default value. If the difference value is less than the default value, the process proceeds to step S135, where the object occlusion determination unit 130 determines the occlusion state ST of the object 12 to be "unoccluded." If the difference value is not less than the default value, the process proceeds to step S136, where the object occlusion determination unit 130 determines the occlusion state ST of the object 12 to be "occluded."

[0072] In an experiment, 2000 virtual scenes 10 and mask-r-cnn machine learning technology were used to obtain a learning model M1. The controller 1C judged the actual occlusion status of the actual object for the actual object images of multiple actual scenes (each scene contains different arrangements of actual objects) based on the learning model M1. According to the experimental results, the mean average precision (mAP) of the training data generation method using the embodiment of the present disclosure is between 70% and 95%, or even higher, while the mean average precision of the existing training data generation method generally does not exceed 50%. The above average precision is generated under the premise that the IoU (Intersection Over Union) is 0.5. Compared with the existing training data generation method, the training data generation method of the embodiment of the present disclosure does have a higher average precision.

[0073] In summary, the embodiments of the present disclosure propose a training data generating device and a training data generating method using the same, which can generate at least one set of training data. The training data includes the occlusion status of each object in the virtual scene and the perspective projection map of the virtual scene. The occlusion status of each object is obtained by extracting the vertical projection map of the virtual scene based on the vertical projection virtual camera unit, so it can reflect the actual occlusion status of the object, and the perspective projection map is the same as the perspective projection principle of the general perspective projection camera used in actual object picking, so the mechanical system can accurately judge the occlusion status of the actual object. In this way, the learning model generated based on (or learning) the above training data can improve the average accuracy of the robotic arm system in actual object picking.

[0074] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A training data generating device, comprising: A virtual scene simulation unit, used to generate a virtual scene, wherein the virtual scene includes a plurality of objects; A vertical projection virtual camera unit, used to extract a vertical projection image of the virtual scene; an object occlusion determination unit, configured to mark an occlusion status of each object according to the vertical projection image; A perspective projection virtual camera unit, used to extract a perspective projection image of the virtual scene; and The training data generating unit is used for generating the training data of the virtual scene according to the viewing angle projection image and the shielding state of each object. 2 . The training data generating apparatus according to claim 1 , wherein the masking state includes masked and unmasked.

3. The training data generating apparatus according to claim 1 , wherein the object occlusion determination unit is further configured to conceal the objects other than the subject of the judgment in the virtual scene and correspondingly generate a subject object scene for the subject of the judgment; and the vertical projection virtual camera unit is further configured to obtain a vertically projected object image of the subject object scene; in, The object occlusion determination unit is further configured to obtain a difference between the vertically projected object image of the object scene of the person being determined and the vertically projected object image of the person being determined in the vertical projection image; and The shielding state is determined according to the difference value.

4. The training data generating apparatus according to claim 3, wherein the object occlusion determination unit is further configured to determine whether the difference value is less than a default value; If the difference value is less than the default value, determining that the shielding status of the person being judged is not shielded; and If the difference value is not less than the default value, it is determined that the shielding state of the person being judged is shielded.

5. The training data generating device according to claim 1, wherein the objects are placed relative to the supporting surface; the perspective projection virtual camera unit also uses a vertical line of sight to extract the perspective projection image; the vertical line of sight is perpendicular to the supporting surface.

6. The training data generating device according to claim 1, wherein the virtual scene includes a container, the objects are placed in the container, the container is placed relative to the supporting surface, and the vertical projection image is an image generated by viewing the virtual scene with a vertical line of sight, and the vertical line of sight is perpendicular to the supporting surface.

7. A method for generating training data, comprising: generating a virtual scene, wherein the virtual scene includes a plurality of objects; Extracting a vertical projection image of the virtual scene; Marking the shading status of each object according to the vertical projection image; Extracting a perspective projection image of the virtual scene; and Training data of the virtual scene is generated according to the viewing angle projection image and the shielding status of each object.

8. The training data generation method according to claim 7, wherein the step of marking the occlusion state of each object according to the vertical projection image comprises: The shielding status of each object is marked as shielded or not shielded according to the vertical projection image.

9. The training data generating method according to claim 7, further comprising: Hiding the objects in the virtual scene except for the objects being judged; Correspondingly, a judged object scene of the judged person is generated; Obtaining a vertically projected object image of the subject's object scene; Obtaining a difference value between the vertically projected object image of the subject's object scene and the vertically projected object image of the subject in the vertical projection image; and The shielding state is determined according to the difference value.

10. The training data generating method according to claim 9, further comprising: Determine whether the difference value is less than the default value; If the difference value is less than the default value, determining that the shielding state of the person being judged is not shielded; as well as If the difference value is not less than the default value, it is determined that the shielding state of the person being judged is shielded.

11. The training data generation method according to claim 7, wherein in the step of generating the virtual scene, the objects are placed relative to a supporting surface; the training data generation method further comprises: The viewing angle projection image is extracted with a vertical sight line, wherein the vertical sight line is perpendicular to the supporting surface.

12. The training data generation method according to claim 7, wherein in the step of generating the virtual scene, the virtual scene includes a container, the objects are placed in the container, and the container is placed relative to the supporting surface; the training data generation method further comprises: The viewing angle projection image is extracted with a vertical sight line, wherein the vertical sight line is perpendicular to the supporting surface.

13. A robotic arm system comprising: robotic arms; a perspective projection camera, disposed on the robotic arm and used to capture actual object images of a plurality of actual objects; as well as A controller is used to generate a learning model according to the training data generation method described in claim 7, and analyze the actual object image with the learning model to determine the actual shading state of each actual object; and control the robotic arm to extract the actual object whose actual shading state is not shading.

Citation Information

Patent Citations

  • Apparatus for taking out bulk stored articles by robot

    US20140039679A1

  • Real-Time Determination of Object Metrics for Trajectory Planning

    US20160136808A1

  • Generating Synthetic Image Data for Machine Learning

    US20200342652A1

  • Bin-picking system for randomly positioned objects

    US7313464B1