Information processing apparatus
The information processing apparatus addresses the challenge of constructing appropriate learning data for object detection by using 3D models and CG images from multiple viewpoints, resulting in improved detection accuracy and versatility.
Patent Information
- Application Number
- JP2023203423
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-11
AI Technical Summary
Existing object detection methods face challenges in constructing appropriate learning data, especially when using 3D models, due to restrictions on obtaining on-site images and the need for high detection accuracy.
An information processing apparatus that uses a three-dimensional model of an object and learns from image data of three-dimensional computer-generated (CG) images from at least two viewpoints of an imaging device to construct appropriate learning data for object detection.
The proposed solution enables the construction of effective learning data for object detection, improving detection accuracy and versatility by considering multiple viewpoints and physical constraints.
Smart Images

Figure 2025088614000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus that detects an object in video data obtained by a photographing apparatus.
Background Art
[0002] There is a need to visualize the production site in order to grasp the manufacturing process, and the demand for object (for example, cart) management is increasing. By performing cart tracking and work time measurement from the video of an existing surveillance camera (photographing apparatus) and managing them, costs can be reduced, leading to the visualization of problems that have not been visible until now.
[0003] In Patent Document 1, Faster R-CNN, which is an object detection method, is adopted, and learning is performed using a dataset in which actual objects are annotated before object detection. In addition, in order to consider improving the accuracy of object detection, a combination with a dataset created using a 3D model in addition to actual objects is disclosed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the learning process of object detection in Patent Document 1, it has been shown that it is preferable to use annotation teaching data combining an actual model and a 3D model. However, there are various restrictions on obtaining learning data of on-site images using an actual model, and there is room for consideration as to what kind of learning data should be constructed to improve detection accuracy when using a 3D model.
[0006] The present invention is an invention for solving the above problems, and an object thereof is to provide an information processing apparatus capable of constructing appropriate learning data for object detection.
Means for Solving the Problems
[0007] To achieve the above object, an information processing apparatus of the present invention is an information processing apparatus for detecting an object of video data obtained by an imaging device, and when performing learning of object detection, the information processing apparatus uses a three-dimensional model of the object and learns using image data of three-dimensional CG from at least two viewpoints of the imaging device. Other aspects of the present invention will be described in the embodiments described later.
Effects of the Invention
[0008] According to the present invention, appropriate learning data for object detection can be constructed.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Embodiments for Carrying Out the Invention
[0010] First, the outline of this embodiment will be explained. In this embodiment, a method of constructing an object detection image dataset for transfer learning prepared according to a detection target using only 3D CG images for object detection using deep learning is proposed. When constructing an object detection image dataset, not only is it costly in terms of time and money, but there are also cases where restrictions are imposed to protect intellectual property rights and confidentiality obligations. Through the case of constructing an object detection image dataset using 3D CG images and detecting an object (a cart in this embodiment) from a camera image using an object detection model Faster R-CNN using deep learning, the effectiveness of the constructed object detection image dataset is evaluated.
[0011] When creating an image dataset for object detection, in this embodiment, by using at least two viewpoints of virtual cameras, it contributed to the improvement of detection accuracy and versatility. Specifically, it was realized by preparing two patterns of tilt angles under the camera viewpoint conditions.
[0012] In terms of physical constraints, by observing the constraints of grounding and up-and-down directions and arranging objects without interference (not encroaching on other objects), it contributed to the improvement of detection accuracy. It is realized by physical calculations, coordinate settings, and interference checks in the virtual space. Note that the virtual camera refers to the camera in the three-dimensional space when creating the image dataset.
[0013] That is, the information processing apparatus of this embodiment is an information processing apparatus that detects an object in video data obtained by a photographing apparatus (real camera). When the information processing apparatus performs learning of object detection, when using a three-dimensional model of the object, by considering arranging a virtual camera at the position in the three-dimensional CG where the photographing apparatus (real camera) is assumed to be arranged, it is characterized by learning using the image data of the three-dimensional CG from at least two viewpoints where the photographing apparatus (real camera) can be arranged.
[0014] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. FIG. 1 is a block diagram showing an information processing apparatus 100 according to an embodiment. The information processing apparatus 100 includes an image generation unit 10 that creates an image dataset for object detection, a learning unit 20 that causes Faster R-CNN to learn the dataset created by the image generation unit 10, a detection / evaluation unit 30 that detects and evaluates an object using the learning model learned by the learning unit 20, a storage unit 40, an input unit 51, an output unit 52, a communication unit 53, and the like. The storage unit 40 includes a learning data holding unit 41, a learning model group holding unit 42, a detection result holding unit 43, and the like.
[0015] The image generation unit 10 includes a three-dimensional model creation unit 11 (see FIG. 3), a three-dimensional model setting unit 12 (see FIG. 5), a virtual camera setting unit 13 (see FIG. 6), and a CG image capturing unit 14 (see FIG. 8).
[0016] The learning unit 20 (see FIG. 1) includes an image input unit 21 that inputs the image captured by the CG image capturing unit 14, and a learning model generation unit 22 that generates a learning model using the image data in the learning data holding unit 41.
[0017] The detection / evaluation unit 30 (see FIG. 1) includes an object detection unit 31 that detects an object using the learning models in the learning model group holding unit 42, an evaluation value calculation unit 32 that evaluates the detection results in the detection result holding unit 43, and a work time measurement unit 33 that determines in which defined area of which work area the object detected by the object detection unit 31 exists, and integrates the determined frames to measure the work time of each work area.
[0018] In the work time measurement unit 33, the area for which the work time is to be measured in advance is defined as a bounding box (BBOX). This area is called a defined work area (defined area). Note that the bounding box is a rectangular frame line (rectangle) that surrounds an image or the like.
[0019] The information processing apparatus 100 includes a memory, a processor, a storage device such as an HD (Hard Disk), a communication unit such as a NIC (Network Interface Card), a user interface unit, and the like. As an example of the processor, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) can be considered, but other semiconductor devices may be used as long as they are the main body for executing predetermined processing.
[0020] Then, the program stored in the storage device is loaded into the memory, and the loaded program is executed by the processor. As a result, the functions of the image generation unit 10, the learning unit 20, and the detection / evaluation unit 30 shown in FIG. 1 are realized. The information processing apparatus 100 has a user interface unit, and includes a display as an output unit 52 and a mouse, a keyboard, and a touch panel as an input unit 51.
[0021] Figure 2 is a diagram showing examples of images taken in the real space and the virtual space. As a result of using the video of the actual factory where the cart is shown as learning data for detection and confirming undetected and misdetected cases, it was confirmed that learning data with variations is necessary to improve the detection accuracy. However, considering actual use, it is difficult to take videos inside the factory every time due to time and cost limitations, restrictions for protecting intellectual property and confidential information, the mental burden on workers, and the trouble of taking pictures. Therefore, using 3D CG software, we consider adding to the learning of the cart detector that does not require shooting in the real space by operating and shooting the cart in a 3D virtual space. In this embodiment, cart video shooting and dataset construction using Unity are performed. Blender is 3D CG software equipped with functions such as modeling, simulation, and rendering. Since it is suitable for creating 3D carts and spaces, it is applied to the learning data and verification data of this embodiment. The process of generating a CG image using Unity will be described. First, as a preliminary preparation, a sample real space is set. Next, using Unity, a 3D model of the cart and a 3D model of the sample real space are created. Figure 2 shows image 2A in the real space and image 2B in the virtual space.
[0022] Figure 3 is a flowchart showing the process S110 of the 3D model creation unit 11. The process S110 is composed of an object representation process for detection (S111 to S113), a background representation process (S114, S115), and an illumination representation process (S116, S117).
[0023] The 3D model creation unit 11 creates a rough shape using a mesh, and the size of the circumscribed rectangular parallelepiped is (S Ox , S Oy , S OzSet it (S111). Specifically, the size of the cart is determined. Next, the 3D model creation unit 11 expresses the details using fillet processing (S112). In that case, let the fillet radius be r. A fillet is a process of rounding the corners in the field of mechanical engineering. Then, the 3D model creation unit 11 sets the pasting of textures and materials for the target object (S113). A texture generally refers to the overall characteristics, material sense, and effects that capture the partial changes such as the visual color and brightness uniformity of the surface of the material, and the unevenness that feels the strength of the tactile force.
[0024] The 3D model creation unit 11 sets the background structure (S114). Specifically, it is the size of the circumscribed rectangular parallelepiped of the virtual space (S Bx , S By , S Bz ) is set. Then, the 3D model creation unit 11 sets the pasting of textures and materials for the target virtual space (S115). Here, the texture refers to a single color, pattern, photo, etc., and the material refers to the three primary colors, material, reflectivity, etc.
[0025] The 3D model creation unit 11 selects the type of lighting (S116) and sets the luminous intensity and illuminance (S117).
[0026] Figure 4 is a diagram showing an example of the arrangement of the background / detection object. In Figure 4, for the world coordinate system Σ W with respect to the background circumscribed rectangular parallelepiped coordinate system Σ B , the detection object coordinate system Σ O , and the lighting coordinate system Σ L are arranged. The cart of the detection object is arranged in the virtual space.
[0027] Figure 5 is a flowchart showing the process S120 of the 3D model setting unit 12. The process S120 is composed of a background / detection object arrangement process (S121~S124) and a lighting arrangement process (S125, S126).
[0028] The three-dimensional model setting unit 12 sets the background circumscribed rectangular parallelepiped coordinate system Σ B (S121), sets the detection object coordinate system Σ O (S122), sets the detection object posture (S123), and arranges the detection object within the background circumscribed rectangular parallelepiped (S124). When arranging this detection object, it is advisable to consider the physical constraints in the real space. For example, the detection object is grounded on the ground within the background.
[0029] The three-dimensional model setting unit 12 sets the illumination coordinate system Σ L (S125) and sets the illumination posture (S126). The illumination coordinate system Σ L is set within the coordinate range of the background circumscribed rectangular parallelepiped. When setting the illumination posture, it is advisable to reproduce the illumination and shadows in the real space.
[0030] FIG. 6 is a flowchart showing the process S130 of the virtual camera setting unit 13. As described above, the virtual camera is a camera in the three-dimensional space when creating the image data set. The virtual camera setting unit 13 sets the distance d between the virtual camera and the detection object (S131), sets the inclination angle range from the horizontal plane as (0deg ≦ θ < 90deg) or (-90deg ≦ θ < 0deg), and sets at least two patterns at intervals of Δθ within the range (S132). The virtual camera is installed on the spherical surface with a radius d centered on the detection object. The range of d is 0 to D (predetermined value). The virtual camera optical axis passes through the center point coordinates P OC (1 / 2·S OX , 1 / 2·S OY , 0) of the xy plane of the detection object.
[0031] The virtual camera setting unit 13 sets the rotation angle around the vertical axis as (0deg ≦ ψ < 360deg) and sets it within the range (S133), and sets the resolution of the rotation angle as Δψdeg (S134). The imaging viewpoints of the virtual camera depend on the inclination angle range, the rotation angle range, and the distance d. When setting the range of the inclination angle and the rotation angle, it is advisable to consider the imaging viewpoints in the real space, the variations of the learning image data set, and the number of images.
[0032] FIG. 7 is a diagram showing an example of parameters of a virtual camera. The parameters related to the virtual camera in FIG. 7 are as follows. P OC : Point through which the camera optical axis passes (center point of the xy plane of the detection object) θ: Tilt angle (with respect to the horizontal plane) Δθ: Resolution of the tilt angle φ: Rotation angle (around the vertical axis) Δφ: Resolution of the rotation angle d: Distance to the detection object
[0033] FIG. 8 is a flowchart showing the process S140 of the CG image capturing unit 14. The CG image capturing unit 14 prepares virtual cameras for the number of tilt angle patterns and sets each camera to have a different tilt angle (S141), sets the minimum value of the rotation angle range as the initial value of the rotation angle (S142), and captures a CG image from the current virtual camera position and orientation (S143).
[0034] Then, the CG image capturing unit 14 outputs the BBOX (Bounding BOX) coordinates of the detection object from the virtual camera video as automatic annotation (S144), and adds Δφ deg of the resolution to the current rotation angle (S145). Then, the CG image capturing unit 14 determines whether the rotation angle is within the set range (S146). If the rotation angle is within the set range (S146, Yes), it proceeds to S143. If the rotation angle is not within the set range (S146, No), the process ends. The captured CG image etc. are stored in the learning data holding unit 41 via the image input unit 21.
[0035] FIG. 9 is a flowchart showing the process S220 of the learning unit 20. The learning unit 20 acquires the captured CG image and the BBOX information of the detection object from the learning data storage unit 41 (S221), and introduces a pre-trained model (S222). For object detection from the learning image and the input image, it is advisable to use the object detection method of the deep learning model pre-trained with the large-scale database MSCOCO in advance. Not limited to MSCOCO, pre-trained models of other large-scale databases may also be used.
[0036] The learning unit 20 sets the learning conditions of the object detector (S223). Specifically, there are the number of epochs, batch size, and learning rate. The batch size refers to (the number of pieces processed at one time) = (the number randomly selected). The number of epochs is the number that determines how many times this routine is repeated by dividing the dataset into N subsets according to the batch size and training each subset N times for learning.
[0037] The learning unit 20 inputs the learning image dataset (S224), learns using Faster R-CNN (S225), and generates a detection model (S226). The generated detection model is stored in the learning model group storage unit 42.
[0038] FIG. 10 is a flowchart showing the process S300 of the detection / evaluation unit 30. The detection / evaluation unit 30 performs object detection on the input image (real image) based on the learning model (S301), and stores the detection result in the detection result storage unit 43. The detection / evaluation unit 30 calculates evaluation values (precision, recall) from the detection result (S302), and displays the calculation result on the display unit (output unit 52) and the like.
[0039] The evaluation of the proposed method in this embodiment was examined using precision and recall. First, True Positive (TP), False Negative (FN), False Positive (FP), and True Negative (TN) are defined. A predicted value of 1 indicates that an object for detection is predicted in a certain working area of a certain frame, and a correct value of 1 indicates that the correct object in that frame is the object for detection. TP represents the number of frames where the predicted value is 1 and the correct value is also 1, FN represents the number of frames where the predicted value is 0 and the correct value is 1, FP represents the number of frames where the predicted value is 1 and the correct value is 0, and TN represents the number of frames where the predicted value is 0 and the correct value is 0.
[0040] Precision is shown in Equation (1), and recall is shown in Equation (2). Precision = TP / (TP + FP) ··· Equation (1) Recall = TP / (TP + FN) ··· Equation (2)
[0041] Precision represents the proportion of the objects predicted by the system as objects for detection that are correctly the objects for detection. When the precision is high in Equation (1), it means that there are few false detections of being an object for detection.
[0042] Recall represents the proportion of the objects that are actually objects for detection and are correctly detected by the system. When the recall is high in Equation (2), it means that there are few misses of being an object for detection.
[0043] Figure 11 is a diagram showing the object detection process (Faster R-CNN). In this embodiment, Faster R-CNN was adopted for the object detection process. Faster R-CNN is a general object detection model composed of two stages: an RPN (Region Proposal Network) that estimates object candidate regions from a feature map and a network that estimates the class label and rectangular position of the objects existing in the object candidate regions.
[0044] In Backborn CNN71, a feature map is generated for the input image (input image 70). Next, in RPN72, the generated feature map is used as input to estimate object candidate regions. The estimated object candidate regions are converted into fixed-length feature vectors by ROI Pooling73. Finally, FC74,75 outputs the class probability (class 76) of what kind of object each object candidate region is and the BBOX (BB77) of the object position. Note that FC is the abbreviation of Fully connected layer, and BBOX is the abbreviation of Bounding Box.
[0045] Specifically, the fixed-length feature vector converted by ROI Pooling73 is input into two-layer FC74,75 to obtain an intermediate feature vector. Then, the intermediate feature vector is input into FC75A for outputting the object class probability and FC75B for the deviation of the object position, and the object class probability and the accurate rectangular coordinates of the object are output. Here, if the class probability of the object class is 0.9 or more, it is regarded as correctly identified as an object, and a rectangle is drawn.
[0046] When the information processing apparatus 100 is used as a monitoring system, video data is captured from a photographing apparatus (not shown) installed in the area to be monitored via the communication unit 53 (see FIG. 1). The information processing apparatus 100 includes an object detection unit 31 that detects an object in a frame of the video data, and a working time measurement unit 33 that determines in which defined area of which work area the object detected by the object detection unit 31 exists, and measures the working time of each work area by integrating the determined frames.
[0047] The object detection unit 31 detects an object in the frame of the video data as a rectangle, and the working time measurement unit 33 calculates an evaluation index indicating the degree of overlap between the detection rectangle detected by the object detection unit 31 and the rectangle of the defined area of each work area, and determines that an object exists in the area where the evaluation index is equal to or greater than a predetermined threshold value and the evaluation index is the largest.
[0048] (Example of settings of the virtual camera setting unit) As camera viewpoint conditions, set the range of the rotation angle ψ and the tilt angle θ of the camera around a single cart.
[0049] FIG. 12 is a top view showing an image of the rotation angle ψ. FIG. 12 is an example of setting the rotation angle ψ in FIG. 7. The distance from the detection object is set to d = 5 m, and the virtual camera's orbit rotates 360 degrees clockwise. That is, as shown in FIG. 12, the camera is installed around a stationary cart, the camera optical axis is directed at the cart, and shooting is performed on a circumference with a constant radius.
[0050] FIG. 13 is a side view showing a shooting image of the tilt angle θ. As shown in FIG. 13, six virtual cameras (cameras C1, C2, C3, C4, C5, C6) are prepared, and a stationary cart is photographed at different tilt angles. In this embodiment, the tilt angle θ has six patterns of 10 deg, 20 deg, ···, 60 deg. As described above, by setting the rotation angle ψ and the tilt angle θ, it becomes possible to operate and shoot the camera viewpoint conditions. That is, 3D CG images from other viewpoints can be created for the rotation angle ψ and the tilt angle θ.
[0051] (Example of learning data) FIG. 14 is a diagram showing an example of a CG image of learning data for one viewpoint. The learning data can arbitrarily set the rotation angle ψ and the tilt angle θ, and the tilt angle of the camera can be changed for each learning data. FIG. 14 is an example of learning data for one viewpoint. Learning data D1 is an example where the tilt angle θ is 10 deg, learning data D2 is an example where the tilt angle θ is 20 deg, learning data D3 is an example where the tilt angle θ is 30 deg, learning data D4 is an example where the tilt angle θ is 40 deg, learning data D5 is an example where the tilt angle θ is 50 deg, and learning data D6 is an example where the tilt angle θ is 60 deg. For each tilt angle, the resolution Δφ of the rotation angle range is 1 deg, and one image is generated for each 1 deg of the rotation angle. Therefore, each of the learning data D1 to D6 generated 360 images. That is, 3D CG images from one viewpoint were generated for the tilt angle θ, and 3D CG images from other viewpoints were generated for the rotation angle ψ.
[0052] FIG. 15 is a diagram showing an example of a CG image of learning data from two viewpoints. The learning data D1+D6 is an example of a combination of inclination angles θ of 10 deg and 60 deg, the learning data D2+D6 is an example of a combination of inclination angles θ of 20 deg and 60 deg, the learning data D3+D6 is an example of a combination of inclination angles θ of 30 deg and 60 deg, the learning data D4+D6 is an example of a combination of inclination angles θ of 40 deg and 60 deg, and the learning data D5+D1 is an example of a combination of inclination angles θ of 50 deg and 10 deg. For each inclination angle, the resolution Δφ of the rotation angle range is 1 deg, and one image is generated for each 1 deg of the rotation angle. Therefore, each of the learning data D1+D6, D2+D6, ···, D5+D1 generated 360 images each. That is, 3D CG images from two viewpoints were generated for the inclination angle θ, and 3D CG images from other viewpoints were generated for the rotation angle ψ. Note that the learning conditions are 360 images each in FIG. 14, 720 images in FIG. 15, and the number of learning times is 50 times.
[0053] (Verification data example) FIG. 16 is a diagram showing an example of an image of verification data. An image of one cart taken in the real space is used as the verification data. In this embodiment, as the verification data, three videos with the inclination angles fixed at 10 deg, 20 deg, and 30 deg are taken. The three taken videos are frame-divided to construct 360 verification data each. FIG. 16 shows an example of an inclination angle of 30 deg. Here, the resolution Δψ of the rotation angle range is 1 deg. Also, it is assumed that one image is generated for each 1 deg of the rotation angle. Therefore, the verification data was generated 360 each for each inclination angle.
[0054] (Transfer learning) As described above, in this embodiment, Faster R-CNN, which is an object detection model, is used. By inputting learning data into Faster R-CNN, object detection becomes possible. The larger the amount of learning data, the more improvement in detection accuracy is expected. However, the number of image sheets of the learning data used in this embodiment is insufficient, and there is concern about a decrease in detection accuracy. Therefore, in this embodiment, an improvement in detection accuracy using transfer learning is attempted. Transfer learning is a method of applying a learning device used for learning any learning data as a learning device for different learning data.
[0055] FIG. 17 is a diagram showing an image of transfer learning. Transfer learning is roughly classified into pre-training and additional learning. In pre-training, MSCOCO, a large-scale database, is learned. In additional learning, the learning device after pre-training is utilized as the learning task of the learning data of this embodiment. Therefore, in the embodiment, additional learning of the CG image of the cart is performed on Faster R-CNN pre-trained with MSCOCO.
[0056] (Evaluation value calculation result) FIG. 18 is a diagram showing the calculation results of the recall rate at one viewpoint. The calculation results in FIG. 18 are the results for the verification data of (1) tilt angle 10 deg, (2) tilt angle 20 deg, and (3) tilt angle 30 deg with respect to the learning data D1, D2, ···, D6. From the calculation results, it was confirmed that the most effective learning data for the verification data with a tilt angle of 10 deg is the tilt angle of 10 deg. Similarly, it was confirmed that the most effective learning data for the verification data with a tilt angle of 30 deg is the tilt angle of 30 deg. Also, it was confirmed that the learning data with a tilt angle of 60 deg has the largest difference in tilt angle of the verification data and the detection accuracy has decreased.
[0057] Figure 19 is a diagram showing the calculation results of the reproduction rates for two viewpoints. The calculation results in Figure 19 are the results for the verification data at (1) tilt angle 10 deg, (2) tilt angle 20 deg, and (3) tilt angle 30 deg for the training data D1+D6, D2+D6, ···, D5+D1. From the calculation results, it was confirmed that the training data at the tilt angle for two viewpoints had a higher reproduction rate than the training data at the tilt angle for one viewpoint. In particular, it was confirmed that the reproduction rate of all the training data for the verification data at the tilt angle of 10 deg was 100%. Similarly, it was confirmed that the reproduction rate of all the training data for the verification data at the tilt angle of 30 deg was 100%. In the case of two viewpoints, it was found that a stable reproduction rate can be obtained even if the camera viewpoints in the real space and the virtual space do not match. Note that although the results for two viewpoints are shown in this embodiment, the same is expected for three or more viewpoints.
[0058] As a result, it was confirmed that the tilt angle of the camera during CG image generation is an important condition for improving the cart detection accuracy in the real space. Also, it was confirmed that by setting the tilt angle in the construction of the training data to at least two viewpoints, not only the cart detection accuracy but also the robustness is improved. From the above, the training data construction method in this embodiment is effective for object detection.
[0059] (Influence of the background of the training data set) The influence of changes in the background of the training data set on object detection was verified. Figure 20 is a diagram showing a 3D CG image regarding the background conditions of the virtual camera. In Unity, carts (trolley models) were randomly arranged, and 500 data sets were constructed respectively in an environment with a colored background (refer to symbol B1), a white background (refer to symbol B2), and no background (refer to symbol B3) using automatic annotation. In Figure 20, two types of carts described later were arranged. The constructed data sets were trained with Faster R-CNN with pre-training.
[0060] The colored background is the case where an "appropriate (= similar to / conforming to the real environment)" background structure, texture material, and lighting position are set. The white background refers to the case where an "appropriate" background structure and lighting position are set. Note that the texture material is set to "none". The backgroundless case refers to the case where an "appropriate" lighting position is set. Note that the background structure and texture material are set to "none".
[0061] As a result, it was found that the learning dataset constructed with a color background is more effective than the white background and the backgroundless case. Using a learning dataset automatically constructed in 3D CG with varying background conditions and the object detection model Faster R-CNN using deep learning, cart detection was performed. As a result, it was confirmed that the mean precision of the object detector trained with the learning dataset constructed using the 3D CG color background created along the real space is 91% and the mean recall is 99%. From the above results, it was also confirmed that cart detection in the real space can be handled even by learning only 3D CG images without using real space images. In particular, for the image data, it is preferable to use 3D CG image data that reproduces lighting similar to or conforming to the real environment.
[0062] (Influence of the detail level of the 3D model) What represents the detail of the 3D model is not limited to the number of polygons, and various concepts are possible. The number of polygons is just an example. The number of polygons refers to the number of faces that make up the 3D object. The larger the number of polygons, the denser the object becomes, and thus it is represented in more detail. The influence of the number of polygons of the 3D model was examined.
[0063] FIG. 21 is a diagram showing a 3D CG image related to the number of polygons of a 3D model. FIG. 21 shows 3D model images of carriages with different numbers of polygons. As described above, a polygon is the minimum unit that constitutes the surface of a 3D model, and the more polygons there are, the more detailed the expression can be. The carriage is three-dimensionally represented by a collection of polygons. In FIG. 21, a carriage model composed only of cubes and having 20 or fewer polygons (polygon number ≤ 20) is defined as level 1. A carriage model having more than 20 and 100 or fewer polygons (20 < polygon number ≤ 100) formed by a combination of a plurality of geometric shapes is defined as level 2. A carriage model created using LiDAR and having more than 100 polygons (polygon number > 100) is defined as level 3. Two different carriages are shown in FIG. 21.
[0064] That is, Level 1: Outline of a carriage created by deforming a cube Level 2: Model of a carriage with fillets added to reproduce details Level 3: Model of a carriage with textures and materials similar to the real thing added It can be said. Note that a fillet is a process of rounding the corners in the field of mechanical engineering. A texture refers to the feel and texture of an object.
[0065] Datasets of 300 images each were constructed for levels 1, 2, and 3. Using a learning dataset automatically constructed in 3D CG and an object detection model Faster R-CNN using deep learning, carriage detection was performed. As a result, it was found that the higher the level of detail of the 3D model used when constructing the learning dataset, the higher the precision and recall rates of carriage detection tended to be.
Explanation of Signs
[0066] 10 Image generation unit 11 3D model creation unit 12 3D model setting unit 13 Virtual camera setting unit 14 CG image capturing unit 20 Learning unit 21 Image Input Unit 22 Learning Model Generation Unit 30 Detection and Evaluation Unit 31 Object Detection Unit 32 Evaluation Value Calculation Unit 33 Working Time Measurement Unit 40 Memory Unit 41 Learning Data Holding Unit 42 Learning Model Group Holding Unit 43 Detection Result Holding Unit 51 Input Unit 52 Output Unit 53 Communication Unit 70 Input Image 71 Backborn CNN 72 RPN 73 ROI Pooling 74,75 FC (Fully connected layer) 76 Class 77 BB (Bounding Box) 100 Information Processing Device
Claims
1. An information processing apparatus for detecting an object in video data obtained by a photographing device, wherein when performing learning of object detection, the information processing apparatus uses a three-dimensional model of the object and learns using image data of three-dimensional CG from at least two viewpoints of the photographing device. An information processing apparatus characterized by the above.
2. The image data is image data in which a background structure, texture, material, and illumination position are set. The information processing apparatus according to claim 1, characterized by the above.
3. The image data uses a three-dimensional model of an object in which fillets are added to reproduce details. The information processing apparatus according to claim 1, characterized by the above.
4. The image data uses a three-dimensional model of an object reproduced with settings of texture and material added. The information processing apparatus according to claim 1, characterized by the above.
5. The image data uses image data of three-dimensional CG that reproduces illumination similar to or conforming to an actual environment. The information processing apparatus according to claim 1, characterized by the above.
6. The image data uses image data of three-dimensional CG that reproduces an object based on physical constraints. The information processing apparatus according to claim 1, characterized by the above.
7. The information processing apparatus includes an object detection unit that detects an object in a frame of the video data, and a working time measurement unit that determines in which defined area of which working area the object detected by the object detection unit exists, and integrates the determined frames to measure the working time of each working area. The information processing apparatus according to any one of claims 1 to 6, characterized by the above.
8. The object detection unit detects an object in a frame of the video data as a rectangle, and the working time measurement unit calculates an evaluation index indicating the degree of overlap between the detection rectangle detected by the object detection unit and the rectangle of the defined area of each working area, and determines that an object exists in the area where the evaluation index is equal to or greater than a predetermined threshold value and the evaluation index is the largest. The information processing apparatus according to claim 7, characterized by the above.
Citation Information
Patent Citations
Monitoring system and monitoring method
JP2022180238A