Information processing device and information processing method
The information processing device generates virtual three-dimensional spaces with human models to create accurate training data, enhancing object detection in camera images and reducing reliance on costly 3D sensors.
Patent Information
- Application Number
- JP2022180876
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-11-11
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2042-11-11
AI Technical Summary
Conventional methods face difficulties in generating appropriate data for machine learning to infer the three-dimensional position of an object from camera images.
An information processing device and method that generates virtual three-dimensional spaces with human models, calculates their areas and distances, and outputs images with associated data for machine learning.
Enables effective generation of data for inferring the three-dimensional position of objects, improving detection accuracy in various environments and reducing the need for expensive 3D sensors.
Smart Images

Figure 0007750214000001 
Figure 0007750214000002 
Figure 0007750214000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device and an information processing method. [Background technology]
[0002] Conventionally, an approach has been known in which a large amount of data is generated by simulation and used for learning (see, for example, Patent Document 1). Patent Document 1 discloses a technology for acquiring center data by simulating multiple 3D range image sensors when a human agent is moving.
[0003] In recent years, technology that uses deep learning and other techniques to detect objects from two-dimensional images taken with a camera has also been put to practical use. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2020-163509 Summary of the Invention [Problem to be solved by the invention]
[0005] However, with conventional technologies, for example, when performing supervised learning to infer the three-dimensional position of an object from an image captured by a camera, it can be difficult to properly obtain data for machine learning.
[0006] An object of the present disclosure is to provide a technology that can appropriately generate data for machine learning to infer the three-dimensional position of an object that is a subject from an image captured by a camera. [Means for solving the problem]
[0007] In a first aspect of the present disclosure, there is provided an information processing device having: a generation unit that generates information of a virtual three-dimensional space in which a human model is placed and generates an image of the virtual three-dimensional space captured by a virtual camera; a calculation unit that calculates information indicating an area of the human model in the image and information indicating a distance from the virtual camera to the human model in the virtual three-dimensional space; and an output unit that outputs the image and the information calculated by the calculation unit in association with each other.
[0008] In addition, a second aspect of the present disclosure provides an information processing method in which an information processing device executes the following processes: generating information of a virtual three-dimensional space in which a human model is placed; generating an image of the virtual three-dimensional space captured by a virtual camera; calculating information indicating the area of the human model in the image and information indicating the distance from the virtual camera to the human model in the virtual three-dimensional space; and outputting the image in association with the calculated information. [Effects of the Invention]
[0009] According to one aspect, it is possible to appropriately generate data for machine learning to infer the three-dimensional position of an object that is the subject from an image captured by a camera. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a configuration of an information processing apparatus according to an embodiment. [Figure 2] 10 is a flowchart illustrating an example of processing by the information processing apparatus according to the embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a virtual three-dimensional space according to the embodiment. [Figure 4] 10A and 10B are diagrams illustrating an example of learning images and correct answer data according to the embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of a learning DB (database) according to the embodiment. [Figure 6] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing apparatus according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] The principles of the present disclosure will be described with reference to some exemplary embodiments. It should be understood that these embodiments are set forth for illustrative purposes only, to aid those skilled in the art in understanding and practicing the present disclosure, without implying any limitation on the scope of the disclosure. The disclosure described herein may be implemented in various ways other than those described below. In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0012] <Configuration> The configuration of an information processing device 10 according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the configuration of the information processing device 10 according to an embodiment. In the example of Fig. 1, the information processing device 10 includes an acquisition unit 11, a generation unit 12, a calculation unit 13, and an output unit 14. Each of these units may be realized by cooperation between one or more programs installed in the information processing device 10 and hardware such as a processor and memory of the information processing device 10.
[0013] The acquisition unit 11 acquires various information from a storage unit within the information processing device 10 or an external device. The acquisition unit 11 acquires, for example, a person model (three-dimensional data of a person). The generation unit 12 generates information on a virtual three-dimensional space in which the person model is placed, and generates an image of the virtual three-dimensional space captured by a virtual camera.
[0014] The calculation unit 13 calculates information indicating the area of the human model in the image generated by the generation unit 12 and information indicating the distance from the virtual camera to the human model in the virtual three-dimensional space generated by the generation unit 12. The output unit 14 outputs the image generated by the generation unit 12 and the information calculated by the calculation unit 13 in association with each other.
[0015] <Processing> Next, an example of processing of the information processing device 10 according to the embodiment will be described with reference to Fig. 2 to Fig. 5. Fig. 2 is a flowchart showing an example of processing of the information processing device 10 according to the embodiment. Fig. 3 is a diagram showing an example of a virtual three-dimensional space according to the embodiment. Fig. 4 is a diagram showing an example of a learning image and correct answer data according to the embodiment. Fig. 5 is a diagram showing an example of a learning DB (database) 501 according to the embodiment.
[0016] The process in Fig. 2 may be executed when a predetermined operation is performed by a user (an operator of the information processing device 10), for example. The information processing device 10 may also repeatedly execute the process in Fig. 2 until the number of images designated by the user or the like is generated. Note that the processes in Fig. 2 may be executed in a different order as long as there is no contradiction.
[0017] In step S1, the acquisition unit 11 acquires one or more human models (three-dimensional data of a person). Here, the acquisition unit 11 may acquire, for example, a three-dimensional (3D) model of a person created using Computer Aided Design (CAD). Alternatively, the acquisition unit 11 may acquire, for example, a three-dimensional model of a human shape that combines multiple three-dimensional shapes (for example, a cube, a rectangular parallelepiped, a sphere, and a cylinder).
[0018] The acquisition unit 11 may acquire each person model via the Internet using a search engine, for example, or may acquire each person model from a specific database.
[0019] Next, the acquisition unit 11 acquires posture (pose) data of each human model (step S2). Here, the acquisition unit 11 may acquire posture data indicating the three-dimensional angles (roll, pitch, and yaw) of each bone (armature) of a human, for example. Alternatively, the acquisition unit 11 may acquire posture data indicating the three-dimensional angle of each bone at any point in time from time-series data of the three-dimensional angle of each bone during a specific human behavior (e.g., walking, lifting an object, dancing).
[0020] The acquiring unit 11 may acquire the posture data of each person via the Internet using a search engine, for example, or may acquire the posture data of each person from a specific database.
[0021] Next, the acquisition unit 11 acquires texture data of each human model (step S3). Here, the acquisition unit 11 may acquire, for example, the pattern of each clothing part on the surface of the human model, and the texture of an image or the like.
[0022] The acquiring unit 11 may acquire the texture data of each person via the Internet using a search engine, for example, or may acquire the texture data of each person from a specific database.
[0023] Next, the generation unit 12 generates (changes, modifies) each person model (step S4). Here, the generation unit 12 may generate a model for each person by setting the person model acquired in step S1 to the posture acquired in step S2. This makes it possible to generate learning data that can detect people performing various movements, for example. Furthermore, the generation unit 12 may randomly change the angle of each joint for each person, and then calculate (determine) the position of the extremities and the like from the position and angle of each joint using forward kinematics.
[0024] Then, the generation unit 12 may apply (paste) the texture acquired in step S3 to the surface of the model for each person, thereby generating learning data that can detect people wearing clothes of various colors and patterns, for example.
[0025] The generation unit 12 may also change the size of each human model. This makes it possible to generate learning data that can detect people of various heights, for example. In this case, the generation unit 12 may determine the height of each human model based on, for example, the average height and standard deviation specified by the user. The generation unit 12 may then enlarge or reduce the size of each human model so that the height of each human model is equal to the determined height. This makes it possible to perform machine learning based on images of human models with heights that correspond to, for example, the average height of workers working in a certain factory, thereby improving the performance of detecting workers in the factory.
[0026] As a result, for example, when multiple images are generated by executing the process of Figure 2 multiple times, the generation unit 12 generates an image in which the size of the human model in the virtual three-dimensional space is a first size, and an image in which the size of the human model in the virtual three-dimensional space is a second size different from the first size.
[0027] Next, the generation unit 12 places each human model in a virtual three-dimensional space (step S5). Here, the generation unit 12 may place each human model in a random position and orientation on the ground (e.g., floor) in the virtual three-dimensional space, for example.
[0028] Next, the generation unit 12 determines lighting conditions in the virtual three-dimensional space (step S6). Here, the generation unit 12 may, for example, determine the position of one or more lights in the virtual three-dimensional space. In this case, the generation unit 12 may, for example, determine random positions on the ceiling surface of the virtual three-dimensional space, which is an indoor space, as the positions of lighting fixtures.
[0029] Furthermore, the generation unit 12 may, for example, randomly determine the brightness (lumens) of each of one or more lights in the virtual three-dimensional space. This makes it possible to generate learning data that can more appropriately detect information about a person, even in a relatively dark room. In this case, the generation unit 12 may, for example, determine the brightness of each light based on the average value and standard deviation of the brightness specified by the user. This allows machine learning to be performed based on images in a lighting environment that corresponds to, for example, the average value of the brightness of lighting fixtures installed in a certain factory, thereby improving the performance of detecting workers in the factory.
[0030] As a result, for example, when multiple images are generated by executing the process of Figure 2 multiple times, the generation unit 12 generates an image in which the brightness of the lighting in the virtual three-dimensional space is a first brightness, and an image in which the brightness of the lighting in the virtual three-dimensional space is a second brightness that is different from the first brightness.
[0031] Furthermore, the generating unit 12 may, for example, randomly determine the illumination angle of each of one or more lights in the virtual three-dimensional space, thereby generating learning data that can more appropriately detect information about a person within the area illuminated by the spotlight.
[0032] Next, the generation unit 12 determines the shooting conditions in the virtual three-dimensional space (step S7). Here, the generation unit 12 may, for example, determine the position and angle of view of a virtual camera in the virtual three-dimensional space. In this case, the generation unit 12 may, for example, determine at least one of the height of the virtual camera from the ground in the virtual three-dimensional space and the angle of view of the virtual camera based on information specified by a user. This allows machine learning to be performed based on images captured from a height from the ground of a camera provided on a robot moving within a factory, for example, thereby improving the performance of detecting workers in the factory.
[0033] In the example of FIG. 3, a virtual camera 302, a person model 303, a person model 304, a person model 305, a virtual light 306, and a virtual light 307 are arranged in a virtual three-dimensional space 301.
[0034] Next, the generation unit 12 generates an image of the virtual three-dimensional space photographed by the virtual camera (step S8). Here, the generation unit 12 may generate, by simulation, an image in which a subject such as each human model is photographed under the light of virtual lighting, for example, from the position and angle of view of the virtual camera.
[0035] Next, the calculation unit 13 calculates correct answer data for machine learning for each person who is a subject of the image (step S9). Here, the calculation unit 13 may calculate, as the correct answer data, information indicating the area of each person model in the image and information indicating the distance from the virtual camera to each person model in the virtual three-dimensional space.
[0036] In this case, the calculation unit 13 may calculate the area of each human model in the image based on, for example, the position and angle of view of the virtual camera in the virtual three-dimensional space and the position of each human model in the virtual three-dimensional space. The information indicating the area of the human model in the image may be, for example, a bounding box, which is a rectangle (rectangle or square) surrounding the human model depicted in the image. Furthermore, the information indicating the area of the human model in the image may be, for example, segmentation data, which is a group of pixels of the human model depicted in the image.
[0037] The calculation unit 13 may calculate the distance from the virtual camera to each human model in the virtual three-dimensional space (depth as seen from the virtual camera) based on, for example, the position of the virtual camera in the virtual three-dimensional space and the position of each human model in the virtual three-dimensional space.
[0038] Furthermore, the calculation unit 13 may calculate, for example, a label of a human model in the image as the correct answer data. In this case, the label may include information indicating the gender, age, height, etc. that has been assigned to the human model in advance.
[0039] Fig. 4 shows an image 401 captured by the virtual camera 302 of the virtual three-dimensional space 301 of Fig. 3, and examples of each correct answer data in the image. The example of Fig. 4 shows an example in which information 411 indicating the area of a person model 303, information 412 indicating the distance from the virtual camera 302 to the person model 303 in the virtual three-dimensional space 301, and a label (gender, age, height) 413 of the person model 303 are superimposed and displayed. Also shown is an example in which information 421 indicating the area of a person model 304, information 422 indicating the distance from the virtual camera 302 to the person model 304 in the virtual three-dimensional space 301, and a label 423 of the person model 304 are superimposed and displayed.
[0040] Next, the output unit 14 associates the image with correct answer data for machine learning for each person who is a subject of the image, and outputs the associated data to the training DB 501 (step S10). Note that the training DB 501 may be recorded in a storage device inside the information processing device 10, or may be recorded in a storage device external to the information processing device 10 (for example, a cloud server).
[0041] 5, data (records) of combinations of images with one or more pieces of correct answer data are recorded in the learning DB 501. Note that the output unit 14 may output a single image file in which the correct answer data is added to the image data as image metadata.
[0042] 2, the information processing device 10 may generate a second image and supervised data by changing at least one of the illumination conditions in step S6 and the shooting conditions in step S7, thereby enabling faster generation of a data set of a combination of images and supervised data.
[0043] <Other> Conventionally, when preparing training data to generate a trained model that detects the position of a person based on images taken with a single camera, collecting images and assigning correct answer data to the images requires a huge amount of work. Also, differences in lighting and shooting conditions between the shooting environment of the prepared images and the shooting environment when inference is performed can make it difficult to improve detection accuracy.
[0044] By performing machine learning using training data generated by the technology disclosed herein, it is possible to reduce the possibility of a robot or the like moving around in various facilities, such as a factory, home, office, or restaurant, colliding with a person. Furthermore, it is possible for a robot or the like working collaboratively with a human to more appropriately act on a person (e.g., handing an object to a human, receiving an object from a human, etc.). Furthermore, by performing machine learning using training data generated by the technology disclosed herein, it is possible to detect the distance to a person, etc., using a monocular camera as an inference sensor. Therefore, it is possible to detect the distance to a person, etc., using a relatively inexpensive device compared to using a 3D distance image sensor such as a stereo camera or LiDAR (Light Detection and Ranging).
[0045] <Hardware configuration> Fig. 6 is a diagram showing an example of the hardware configuration of an information processing device 10 according to an embodiment. In the example of Fig. 6, the information processing device 10 (computer 100) includes a processor 101, a memory 102, and a communication interface 103. These components may be connected via a bus or the like. The memory 102 stores at least a part of a program 104. The communication interface 103 includes an interface required for communication with other network elements.
[0046] When the program 104 is executed by the processor 101, memory 102, and the like in cooperation with each other, the computer 100 performs at least some of the processing of the embodiments of the present disclosure. The memory 102 may be of any type. As a non-limiting example, the memory 102 may be a non-transitory computer-readable storage medium. The memory 102 may also be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. Although only one memory 102 is shown in the computer 100, several physically different memory modules may exist in the computer 100. The processor 101 may be of any type. The processor 101 may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture, as a non-limiting example. The computer 100 may have multiple processors, such as application-specific integrated circuit chips that are time-slaved to a clock that synchronizes the main processor.
[0047] Embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device.
[0048] The present disclosure also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, that execute on a target real or virtual processor or device to perform the processes or methods of the present disclosure. Program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or divided among program modules as desired in various embodiments. The machine-executable instructions of the program modules may be executed in local or distributed devices. In a distributed device, the program modules may be located in both local and remote storage media.
[0049] The program code for executing the methods of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus. When the program code is executed by the processor or controller, the functions / acts in the flowcharts and / or implementing block diagrams are performed. The program code may be executed entirely on the machine, partly on the machine, as a standalone software package, partly on the machine and partly on a remote machine, or entirely on a remote machine or server.
[0050] The program can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible recording media. Examples of non-transitory computer-readable media include magnetic recording media, magneto-optical recording media, optical disk media, and semiconductor memory. Magnetic recording media include, for example, flexible disks, magnetic tapes, and hard disk drives. Magneto-optical recording media include, for example, magneto-optical disks. Optical disk media include, for example, Blu-ray discs, CD (Compact Disc)-ROMs (Read Only Memory), CD-Rs (Recordable), and CD-RWs (Rewritable). Semiconductor memory includes, for example, solid-state drives, mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory). The program may also be provided to a computer by various types of temporary computer-readable media. Examples of temporary computer-readable media include electrical signals, optical signals, and electromagnetic waves. The temporary computer-readable medium can supply the program to the computer via a wired communication path such as an electric wire or an optical fiber, or via a wireless communication path.
[0051] <Modification> The information processing device 10 may be a device contained in a single housing, but the information processing device 10 of the present disclosure is not limited to this. Each unit of the information processing device 10 may be realized by cloud computing configured by one or more computers, for example.
[0052] The present invention is not limited to the above-described embodiment, and can be modified as appropriate within the scope of the invention. [Explanation of symbols]
[0053] 10. Information processing equipment 11 Acquisition Department 12 Generation part 13 Calculation section 14 Output section
Claims
1. A generation unit that generates information about a virtual three-dimensional space in which a human model is placed, the human model having information indicating at least one of gender and age, and generates an image of the virtual three-dimensional space captured by a virtual camera; a calculation unit that calculates information indicating an area of the human model in the image, information indicating a distance from the virtual camera to the human model in the virtual three-dimensional space, and a label including information indicating at least one of the gender and the age of the human model; an output unit that outputs the image and the information calculated by the calculation unit in association with each other; An information processing device having the above.
2. the generating unit generates a first image in which the brightness of lighting in the virtual three-dimensional space is set to a first brightness, and a second image in which the brightness of lighting in the virtual three-dimensional space is set to a second brightness different from the first brightness. The information processing device according to claim 1 .
3. the generation unit generates a third image in which a size of the human model in the virtual three-dimensional space is a first size, and a fourth image in which a size of the human model in the virtual three-dimensional space is a second size different from the first size.
3. The information processing device according to claim 1.
4. the generation unit determines at least one of a height of the virtual camera from the ground in the virtual three-dimensional space and an angle of view of the virtual camera based on information specified by a user.
3. The information processing device according to claim 1.
5. Generating information about a virtual three-dimensional space in which a human model is placed, the human model having information indicating at least one of gender and age, and generating an image of the virtual three-dimensional space captured by a virtual camera; calculating information indicating an area of the human model in the image, information indicating a distance from the virtual camera to the human model in the virtual three-dimensional space, and a label including information indicating at least one of the gender and the age of the human model; The image and the calculated information are output in association with each other. An information processing method in which processing is performed by an information processing device.
Citation Information
Patent Citations
Data processing device and data processing method
JP2019191874A
Rule generation device, rule generation method, and rule generation program
JP2020077343A
Simulation system, simulation program and learning device
JP2020163509A
Training using rendered images
US20220351427A1