System for generating virtual images, system control method, and system program
The system generates virtual images of moving objects from multiple directions using cameras with fisheye lenses or PTZ functions, addressing the limitations of fixed-position inspection and complex camera setups, ensuring thorough defect detection.
Patent Information
- Application Number
- JP2024026250
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-09-05
AI Technical Summary
Existing visual inspection methods for moving objects on conveyor belts struggle with overlooking defects due to the need for manual observation from fixed positions or multiple camera setups, which can be complex and prone to interference.
A system that generates multiple virtual images of a moving object from various directions using cameras with fisheye lenses or PTZ functions, setting virtual viewpoints, and generating images that appear to rotate, allowing for comprehensive visual inspection without additional mechanical complexity.
Enables efficient and comprehensive visual inspection of moving objects by generating virtual images that allow for easy observation from multiple angles, reducing the likelihood of defect overlooks and simplifying the inspection process.
Smart Images

Figure 2025129547000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to techniques for generating virtual images, and more particularly to techniques for generating multiple virtual images for viewing a moving object from multiple directions. [Background technology]
[0002] On factory production lines, etc., visual inspections are performed on products or workpieces transported by conveyor belts. The main visual inspection methods are visual inspection, in which the object is checked directly with the naked eye, and image inspection, in which images captured by a camera are checked. Each method has the following problems. First, when performing a visual inspection, the inspection staff must closely observe the moving object, which makes it easy for them to overlook something.
[0003] Furthermore, when observing images captured by a camera tracking an object from one direction, inspection staff only need to observe the object displayed at a fixed position on the screen, making it less likely that defects will be overlooked. However, there is a possibility that defects in parts of the object that are not visible in the image may be overlooked. Furthermore, when performing visual inspections using images of an object captured by a camera from multiple directions, inspection staff must sequentially view the images from multiple cameras. This may result in defects being overlooked. Furthermore, interference with other devices may make it difficult to properly position multiple cameras around the production line. Another option is to use a single camera to capture images of the object from multiple directions while rotating the object itself. However, this method requires the production line to include a mechanism for rotating the object, which increases the complexity of the production line. Therefore, a technology is needed to more simply observe objects moving on conveyor belts, etc.
[0004] Regarding technology for performing visual inspection of an object using multiple cameras, for example, Japanese Patent Application Laid-Open No. 2023-69348 (Patent Document 1) discloses an appearance inspection device that "includes a lighting device that irradiates illumination light onto the workpiece in a photography room having a work set section where the workpiece is placed, and multiple photography devices that capture images of the workpiece from multiple directions, an image processing device processes the image information captured by the photography devices and detects positional deviation from a reference image of the workpiece, and a photography condition adjustment device adjusts the conditions for photography in accordance with the positional deviation of the workpiece detected by the image processing device" ("Abstract"). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2023-69348 Summary of the Invention [Problem to be solved by the invention]
[0006] The technology disclosed in Patent Document 1 may be capable of visually inspecting a workpiece. However, the technology disclosed in Patent Document 1 assumes that the camera, lighting device, and workpiece are fixed. Therefore, the technology disclosed in Patent Document 1 may not be able to observe an object moving on a belt conveyor or the like. Therefore, a technology is needed that can observe a moving object from multiple directions.
[0007] The present disclosure has been made in view of the above-described background, and an object in one aspect is to generate a plurality of virtual images for observing a moving object from a plurality of directions. [Means for solving the problem]
[0008] According to one embodiment, a system is provided. The system includes an object detection unit that detects the position of a moving object, a capture unit that captures images of the object using one or more cameras, a virtual viewpoint setting unit that sets a virtual viewpoint, and a virtual image generation unit that generates a virtual image that appears to be captured of the object from the virtual viewpoint. The capture unit captures images of the object using one or more cameras at each of a plurality of capture positions along the movement path of the object. The virtual viewpoint setting unit sets each of a plurality of virtual viewpoints corresponding to each of the plurality of capture positions. The virtual image generation unit generates each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints from one or more images captured at each of the plurality of capture positions.
[0009] In one aspect, the system further includes an output unit that outputs each of a plurality of virtual images corresponding to each of a plurality of virtual viewpoints so that the object appears to rotate.
[0010] In one aspect, one or more cameras have a fisheye lens, and photographing the object includes obtaining one or more images by flattening at least a portion of the fisheye image.
[0011] In one aspect, generating each of a plurality of virtual images corresponding to each of a plurality of virtual viewpoints includes inputting a chronologically ordered set of images including one or more images taken at each of a plurality of shooting positions into an inference model, and generating each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints.
[0012] In one aspect, setting each of a plurality of virtual viewpoints corresponding to each of a plurality of shooting positions includes, when the one or more cameras are two cameras, setting each of the plurality of virtual viewpoints so that it moves on a line connecting the two cameras.
[0013] In one aspect, setting each of a plurality of virtual viewpoints corresponding to each of a plurality of shooting positions includes setting each of the plurality of virtual viewpoints so as to generate a virtual image in which the size, rotation axis, and rotation speed of the object always appear constant.
[0014] In one aspect, setting each of the plurality of virtual viewpoints to generate a virtual image in which the size, rotation axis, and rotation speed of the object appear constant includes setting a virtual viewpoint at a position away from the object by a predetermined second distance each time the object moves by a predetermined first distance. The virtual viewpoint is a position rotated by a predetermined angle around the object relative to the previous virtual viewpoint.
[0015] In one aspect, the object detection unit detects the position of the object using a position estimation model.
[0016] In one aspect, setting each of a plurality of virtual viewpoints corresponding to each of a plurality of shooting positions includes setting each of the plurality of virtual viewpoints in an output unit so that the object can be displayed from a predetermined direction.
[0017] In one aspect, setting each of a plurality of virtual viewpoints corresponding to each of a plurality of shooting positions includes setting each of a plurality of virtual viewpoints for each of a plurality of objects so that, when displaying the plurality of objects in the output unit, each of the plurality of objects can be displayed from a predetermined direction.
[0018] In one aspect, the virtual viewpoint setting unit selects one or more cameras to be used for shooting from among the one or more cameras, based on the positional relationship between the one or more cameras, the positions of the object, and the virtual viewpoint.
[0019] In one aspect, the system further includes an analysis unit that determines abnormalities in the object by analyzing multiple virtual images using image analysis AI.
[0020] In one aspect, the system further includes an illumination unit that illuminates the object with light to improve inspection accuracy of the object when photographing the object.
[0021] According to another embodiment, there is provided a control method for a system, the control method including: photographing an object using one or more cameras at each of a plurality of photographing positions along a movement path of the object; setting each of a plurality of virtual viewpoints corresponding to each of the plurality of photographing positions; and generating each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints from one or more images photographed at each of the plurality of photographing positions.
[0022] According to yet another embodiment, there is provided a system program that causes the system to photograph an object using one or more cameras at each of a plurality of photographing positions along a movement path of the object, set each of a plurality of virtual viewpoints corresponding to each of the plurality of photographing positions, and generate each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints from one or more images photographed at each of the plurality of photographing positions. [Effects of the Invention]
[0023] According to one embodiment, multiple virtual images can be generated for viewing a moving object from multiple directions.
[0024] The above and other objects, features, aspects and advantages of the present disclosure will become apparent from the following detailed description of the disclosure taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a diagram showing an example of the appearance and operation overview of a system 100 according to the present embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a relationship between each of a plurality of virtual viewpoints and each of a plurality of virtual images. [Figure 3] FIG. 1 illustrates an example of a functional configuration of a system 100 according to the present embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a device 110. [Figure 5]It is a diagram showing an example of the procedure until a virtual image is generated by the system 100. [Figure 6] It is a diagram showing a modified example of the virtual viewpoint setting procedure. [Figure 7] It is a diagram showing an example of a state where virtual viewpoints are simultaneously set for each of a plurality of objects. [Figure 8] It is a diagram showing an example of a state where inspection images of a plurality of objects are simultaneously output. [Figure 9] It is a diagram showing an example of the switching function of a plurality of cameras. [Figure 10] It is a diagram showing an example of a function of setting a plurality of virtual viewpoints for each shooting position. [Figure 11] It is a diagram showing an example of a state of inspection using the light source L300.
Embodiments for Carrying Out the Invention
[0026] Hereinafter, embodiments of the technical idea according to the present disclosure will be described while referring to the drawings. In the following description, the same parts are denoted by the same reference signs. Their names and functions are also the same. Therefore, detailed descriptions thereof will not be repeated. Also, each embodiment, each modification example, each software configuration, each hardware configuration, each function, and each process, etc. may be selectively combined as appropriate.
[0027] <A. Application Examples of the System>
[0028] An application example of the system 100 according to the present embodiment will be described with reference to FIGS. 1 and 2 . Typically, the system 100 can be used as an inspection system for workpieces or products in a production line or the like. Hereinafter, the system 100 will be described as an inspection system for workpieces in a production line, but this is merely an example. The system according to the present embodiment can also be used as a surveillance system for use in airports, stores, or the like. Furthermore, the system according to the present embodiment can also be used as a recording system. For example, the system according to the present embodiment can be used to monitor people passing through a security check area at an airport or suspicious individuals walking through a store from all angles. For another example, the system according to the present embodiment can generate and store virtual images of people, vehicles, or the like passing through a certain location. That is, the system according to the present embodiment can generate multiple virtual images of any object passing through a predetermined area for any application. The system according to the present embodiment also has functions such as analyzing the generated virtual images and presenting them to a user.
[0029] FIG. 1 is a diagram illustrating an example of the appearance and operation overview of a system 100 according to the present embodiment. The system 100 generates multiple virtual images of an object W101 moving on a movement area 150. Inspection staff can easily visually inspect the entire object W101 by checking the multiple generated virtual images. For the purposes of the following explanation, the movement area 150 is assumed to be the movement path of the object W101 formed by a belt conveyor. The object W101 is assumed to be a product or the like moving on the belt conveyor. Furthermore, in this specification, "movement of an object" includes both the object moving under its own power and the object moving due to the force of another object.
[0030] The system 100 includes cameras C101 and C102, a device 110, and a display device 120. The cameras C101 and C102, the device 110, and the display device 120 are configured to be able to communicate with each other via cables, a network, or the like. The shooting area 160 is an area in which the cameras C101 and C102 shoot an object W101. The shooting area 160 is divided into a first half shooting area 160A and a second half shooting area 160B by a line 130 connecting the cameras C101 and C102. The object W101 moves at a constant speed on a moving area 150, which is a belt conveyor, in the direction of an arrow 170. A scene 180A shows the object W101 passing through the shooting area 160A. A scene 180B shows the object W101 passing through the shooting area 160B.
[0031] As an example, the cameras C101 and C102 are arranged to sandwich the movement area 150. A line 130 connects the cameras C101 and C102. As an example, the cameras C101 and C102 may be arranged so that the line 130 is perpendicular to the traveling direction of the object W101.
[0032] The cameras C101 and C102 capture images of the object W101 at each of a plurality of capture positions on the movement area 150 or the movement path. More specifically, while the object W101 remains in the capture area 160, the cameras C101 and C102 capture images of the object W101 each time the object W101 moves a certain distance or after a certain period of time has elapsed. First, the cameras C101 and C102 capture images of the object W101 multiple times while the object W101 is in the capture area 160A. This allows the system 100 to acquire one or more images of the object W101 captured from the right side of the figure. Next, the cameras C101 and C102 capture images of the object W101 multiple times while the object W101 is in the capture area 160B. This allows the system 100 to acquire one or more images of the object W101 captured from the left side of the figure. In the example of FIG. 1, it is assumed that the object W101 is photographed at six photographing positions P1, P2, P3, P4, P5, and P6 that are arranged at equal intervals.
[0033] As another example, the cameras C101 and C102 may continue to capture images of the object W101 while the object W101 remains within the imaging area 160. In this case, the cameras C101 and C102 output images of the object W101. The system 100 acquires images from each of the cameras C101 and C102. The system 100 may then extract frames (i.e., images) from each image at the time when the object W101 passes through each of the imaging positions P1, P2, P3, P4, P5, and P6.
[0034] In one aspect, the cameras C101 and C102 may be cameras equipped with fisheye lenses (also referred to as fisheye lens cameras). In this case, the system 100 acquires fisheye images from the cameras C101 and C102. The system 100 generates a virtual image using frames (i.e., images) obtained by flattening the fisheye images. Alternatively, the cameras C101 and C102 may have a function for flattening captured images. In another aspect, the cameras C101 and C102 may be cameras equipped with normal lenses. In this case, the cameras C101 and C102 have a PTZ (Pan, Tilt, Zoom) function for controlling the attitude and zoom of the cameras C101 and C102. In the example of FIG. 1, the cameras C101 and C102 are fisheye lens cameras with a 180° angle of view and are arranged facing each other. This allows the system 100 to perform PTZ operations on the cameras C101 and C102 through digital processing without changing the orientation of the lenses.
[0035] In some aspects, system 100 may include three or more cameras. In this case, system 100 may select one or more of the three or more cameras to use for capturing images. In other aspects, the cameras included in system 100 may include a fisheye lens, a normal lens, a wide-angle lens, or any other lens. As described above, system 100 is designed to use a fisheye lens camera. However, system 100 may also use a camera with a normal lens, etc., as long as the camera can reproduce an angle of view equivalent to that of a fisheye lens camera using a PTZ function or the like.
[0036] The device 110 has four functions: a function for controlling the cameras C101 and C102, a function for setting a virtual viewpoint, a function for generating a virtual image, and a function for outputting various information to the display device 120. Each function will be explained in turn.
[0037] First, the control function of cameras C101 and C102 will be described. Device 110 may be integrated with a programmable logic controller (PLC) that controls the line, or may be configured to be able to communicate with the PLC. The PLC is a control device that controls various actuators that make up the line and collects signals from various sensors arranged on the line. Device 110 can track the current position of object W101 by referencing variable values of a control program stored in the PLC. Device 110 transmits commands to cameras C101 and C102 to capture an image of object W101 every time object W101 moves a certain distance. In one aspect, device 110 may transmit commands to cameras C101 and C102 to capture an image of object W101 every time a certain period of time elapses.
[0038] Next, the virtual viewpoint setting function will be described. A "virtual viewpoint" is a point where a camera is virtually positioned. An image that appears to be captured from the virtual viewpoint is called a "virtual image." A virtual image is generated from one or more actually captured images. The device 110 sets multiple virtual viewpoints. Each virtual viewpoint is positioned at any position on the line 130 connecting the cameras C101 and C102. Each virtual viewpoint may be positioned between the cameras C101 and C102, or at another position. For example, one or more virtual viewpoints may be positioned outside the camera C101 or C102 on the line 130 as viewed from the movement area 150. Each of the multiple virtual viewpoints can be associated with a respective one of the multiple capture positions. Using the scene 180A as an example, each of the virtual viewpoints V101 to V103 is associated with each of the capture positions P1, P2, and P3. Using scene 180B as an example, virtual viewpoints V104 to V106 are each associated with a corresponding one of shooting positions P4, P5, and P6. The relationship between each virtual viewpoint and each shooting position will be described using virtual viewpoint V101 and shooting position P1 as an example. Cameras C101 and C102 capture an image of object W101 located at shooting position P1. System 100 acquires an image captured by camera C101 and a fisheye image captured by camera C102, and flattens these images. From the flattened image of object W101 located at shooting position P1, system 100 generates a virtual image that appears to have been captured of object W101 located at shooting position P1 from virtual viewpoint V101.
[0039] Next, the virtual image generation function will be described. The device 110 generates a virtual image corresponding to each of the multiple virtual viewpoints that have been set. More specifically, the device 110 may generate a virtual image from one or more images acquired at shooting positions corresponding to each virtual viewpoint. As an example, the device 110 generates a virtual image VI201 (see FIG. 2) that appears to be of an object W101 photographed from a virtual viewpoint V101 from a first image photographed by a camera C101 and a second image photographed by a camera C102 at shooting position P1. Similarly, the device 110 generates virtual images VI202 to VI206 (see FIG. 2) from images photographed by cameras C101 and C102 at shooting positions P2 to P6, respectively. In one aspect, the device 110 may generate a virtual image from two or more images using techniques such as stereo matching and homography transformation. In another aspect, the device 110 may generate a virtual image from one or more images using a machine learning inference model or the like. To explain the operation of the inference model using FIG. 1 as an example, the object W101 is photographed sequentially (i.e., in chronological order) at each of the photographing positions P1 to P6. That is, one or more images photographed at each photographing position can be considered to be images of the object W101 at a certain point in time in the time series. Therefore, the system 100 can acquire a chronologically ordered set of images by grouping one or more images photographed at each photographing position. The system 100 can acquire a virtual image corresponding to each photographing position by inputting the acquired image set into the inference model. The chronologically ordered set of images includes images of the object W101 photographed from various directions in chronological order. The inference model can infer and generate a virtual image from these images of the object W101 photographed from various directions in chronological order. The device 110 may store programs for performing stereo matching, homography transformation, etc. in the secondary storage device 430 (see FIG. 4). The device 110 may also store an inference model in the secondary storage device 430. The device 110 may reference and execute these programs or inference models as needed.
[0040] In the above description, one or more cameras included in system 100 are described as outputting images of object W101, but this is merely an example. In one aspect, one or more cameras included in system 100 may output a video of object W101. In this case, system 100 acquires one or more videos from each of the one or more cameras. System 100 then extracts a frame (i.e., an image) at each time from each of the one or more videos and flattens each of the obtained one or more images. System 100 then generates a virtual image from the one or more flattened images. In the following description, an "image" may refer to an image captured by one or more cameras, or may refer to a frame (i.e., an image) at a certain time extracted from videos captured by one or more cameras.
[0041] Next, a function of outputting various information to the display device 120 will be described. The device 110 outputs the generated multiple virtual images to the display device 120. The device 110 can determine the display order of the multiple virtual images, etc., so that the inspection staff can easily check the entire object W101. As an example, the device 110 may cause the display device 120 to switch between the multiple virtual images and display them so that the object W101 appears to rotate. Alternatively, the device 110 may generate an image from the multiple virtual images that makes the object W101 appear to rotate, and output the image to the display device 120.
[0042] Display device 120 outputs the multiple virtual images input from device 110. In one aspect, display device 120 may display the multiple virtual images sequentially or simultaneously. In another aspect, display device 120 may display a video generated based on the multiple virtual images.
[0043] FIG. 2 illustrates an example of the relationship between each of a plurality of virtual viewpoints and each of a plurality of virtual images. The virtual viewpoint setting 200 indicates the orientation of each of the plurality of virtual viewpoints V101 to V106 relative to the object W101. According to the virtual viewpoint setting 200, it can be seen that the virtual viewpoints V101 to V106 are set to surround the object W101. Furthermore, the output information 210 illustrates an example of information output to the display device 120. For example, the display device 120 may sequentially display virtual images VI201, VI202, VI203, VI204, VI205, and VI206 corresponding to the virtual viewpoints V101 to V106. This allows the inspection staff to perform a visual inspection of the entire periphery of the object W101 simply by checking the display device 120.
[0044] As described with reference to FIGS. 1 and 2, system 100 may set a plurality of virtual viewpoints corresponding to a plurality of shooting positions. Setting a plurality of virtual viewpoints corresponding to a plurality of shooting positions includes setting each of the plurality of virtual viewpoints so that object W101 can be displayed from a predetermined direction on output unit 360 (see FIG. 3). That is, display device 120 may first display object W101 photographed from a predetermined direction and then rotate and display it in a fixed orientation. In the example of FIG. 2, display device 120 always first displays virtual image VI201, which appears to have been photographed from virtual viewpoint V101. Then, display device 120 displays virtual images VI201 to VI206 in order.
[0045] As described above, the system 100 uses multiple cameras to capture multiple images of the moving object W101. The system 100 then sets a virtual viewpoint for each capture position and generates multiple virtual images corresponding to each of the multiple virtual viewpoints. These multiple virtual images are images of the object W101 that appear to be captured from different directions. The system 100 can also switch between the multiple generated virtual images and display them on the display device 120. This allows the inspection staff to easily perform a visual inspection of the entire periphery of the object W101.
[0046] In one aspect, the system 100 may be used for any purpose, such as a surveillance system for a security area or a store, a recording system for accidents at intersections, etc. The system 100 may capture any object passing through a predetermined area.
[0047] 1 and 2, six virtual viewpoints have been described, but this is merely an example. In one aspect, the system 100 may set any number of virtual viewpoints depending on the application of the system 100.
[0048] 1 and 2, device 110 outputs a plurality of virtual images to display device 120, but this is merely an example. In one aspect, device 110 may input the generated plurality of virtual images to a trained model for visual inspection (which may also be read as an inspection AI (Artificial Intelligence)). In this case, the trained model outputs the inspection results of the visual inspection. In another aspect, device 110 may output the generated plurality of virtual images to display device 120 and input them into the trained model for visual inspection. When system 100 is used for purposes other than visual inspection of products, system 100 may use different trained models for each purpose.
[0049] Furthermore, in the descriptions of FIGS. 1 and 2, although the system 100 is constituted by one device 110, this is merely an example. In certain aspects, the system 100 is constituted by a combination of multiple devices 110. Also, the device 110 may include a personal computer, a workstation, a server, a tablet, or a smartphone. Further, the device 110 may include a System-on-a-chip (SoC) and a System-on-Module (SoM). Additionally, the device 110 may include any peripheral devices such as switches, routers, displays, keyboards, and mice. Moreover, the device 110 may include virtual machines and instances constructed on a cloud environment. In certain aspects, the system 100 may be connected to input / output devices such as a display and a keyboard and used as a stand-alone device. In other aspects, the system 100 may provide various functions as a service or a web application via a network. In this case, a user may use the functions of the system 100 via a browser or client software installed on their own terminal.
[0050] <B. Configuration of the System>
[0051] Next, referring to FIGS. 2 and 3, the functional configuration and hardware configuration of the system 100 will be described. Additionally, the roles of each configuration will also be explained. In certain aspects, a part of each configuration shown in FIG. 2 may be realized as a program. In this case, by various hardware shown in FIG. 3 collaborating to execute the program, the functions of each configuration shown in FIG. 2 can be realized. Also, a part of the functions shown in FIG. 2 may be realized as dedicated hardware.
[0052] FIG. 3 is a diagram showing an example of the functional configuration of system 100 according to the present embodiment. System 100 includes device 110, multiple cameras C100, one or more light sources L300, and display device 120. Device 110 includes object detection unit 310, virtual viewpoint setting unit 320, illumination unit 330, imaging unit 340, virtual image generation unit 350, output unit 360, and analysis unit 370. Hereinafter, when referring to multiple cameras collectively, they will be referred to as camera C100. When referring to individual cameras, each camera will be referred to as cameras C101, C102 to C10n. Similarly, when referring to one or more light sources collectively, they will be referred to as light source L300. When referring to individual light sources, each light source will be referred to as light sources L301, L302 to L30n.
[0053] The object detection unit 310 detects the position of the object W101. In one aspect, the object detection unit 310 may detect the position of the object W101 based on information acquired from a PLC. For example, the PLC may have information such as the rotation amount of various actuators and signals acquired from sensors installed along the conveyor belt. In another aspect, the object detection unit 310 may detect the position, orientation, or both of the object W101 by analyzing images or videos acquired from the camera C100 using a position estimation model or position estimation AI. In another aspect, the object detection unit 310 may detect the position of the object W101 using one or more LiDAR (Light Detection and Ranging) devices or the like. The LiDAR devices may be disposed near the camera C100 or at other locations. Furthermore, the object detection unit 310 may detect when the object W101 enters the shooting area 160 and when the object W101 leaves the shooting area 160.
[0054] The virtual viewpoint setting unit 320 sets a plurality of virtual viewpoints corresponding to a plurality of imaging positions along the movement path of the object W101. Setting a plurality of virtual viewpoints may alternatively be referred to as determining the positions of the plurality of virtual viewpoints. Using FIG. 1 as an example, the virtual viewpoint setting unit 320 sets virtual viewpoints V101 to V106 corresponding to imaging positions P1 to P6, respectively. The virtual viewpoint setting unit 320 sets each virtual viewpoint so that a plurality of virtual images for observing the moving object W101 from various directions can be generated. Using FIG. 2 as an example, the virtual viewpoint setting unit 320 sets virtual viewpoints V101 to V106. By checking virtual images VI201 to VI206 corresponding to these virtual viewpoints, the inspection staff can visually inspect the object W101 from approximately all directions.
[0055] The illumination unit 330 controls one or more light sources L300 as necessary to illuminate the object W101 with light. The illumination unit 330 may use one light source L300 to illuminate the object W101 with light, or may use multiple light sources L300 to illuminate the object W101 with light. The illumination unit 330 may switch between light sources L300 depending on the application. As an example, when inspecting the surface of the object W101 for defects, a light source L300 that illuminates a lattice-shaped light as shown in FIG. 11 may be used. Illuminating the surface of the object W101 with lattice-shaped light makes it easier to detect distortions and the like on the surface of the object W101. If the light source L300 is unnecessary, the system 100 may not use or be provided with the light source L300.
[0056] The photographing unit 340 photographs the object W101 using one or more cameras C100. In the example of FIG. 1, the photographing unit 340 photographs the object W101 multiple times using cameras C101 and C102. The one or more images obtained by photographing are used to generate a virtual image that appears to have been photographed from each of multiple virtual viewpoints. Alternatively, the photographing unit 340 may acquire video of the object W101 using one or more cameras C100. In this case, the photographing unit 340 extracts images of the object W101 at each photographing position from the video.
[0057] The virtual image generation unit 350 generates a virtual image that appears to have been captured from a set virtual viewpoint from one or more images acquired from the image capture unit 340. In the example of FIG. 1, cameras C101 and C102 capture images of the object W101 at each of the image capture positions P1 to P6. That is, the system 100 may capture two or more images of the object W101 at each of the image capture positions P1 to P6. Furthermore, corresponding virtual viewpoints V101 to V106 are set for each of the image capture positions P1 to P6. As an example, the virtual image generation unit 350 may generate a virtual image VI201 from one or more images of the object W101 captured at the image capture position P1. Similarly, the virtual image generation unit 350 may generate virtual images VI203 to VI206 from one or more images of the object W101 captured at each of the image capture positions P2 to P6. As described above, each of virtual images VI203-VI206 is an image of object W101 that appears to have been captured from each of virtual viewpoints V101-V106. In one aspect, virtual image generation unit 350 may generate a virtual image that appears to have been captured from a set virtual viewpoint from one or more images extracted from one or more videos that have actually been captured. In another aspect, virtual image generation unit 350 may generate a virtual image from two or more images using stereo matching and homography transformation. Furthermore, in another aspect, virtual image generation unit 350 may generate a virtual image from one or more images using a machine learning inference model or the like.
[0058] In the examples of FIGS. 1 and 2, the number of shooting positions and virtual viewpoints is six, but this is merely an example. System 100 may set any number of shooting positions and virtual viewpoints. In some aspects, system 100 may be configured to adjust the number of shooting positions and virtual viewpoints. Reducing the number of shooting positions and virtual viewpoints reduces the load on system 100. Conversely, increasing the number of shooting positions and virtual viewpoints can be expected to improve the accuracy of visual inspection. For example, output unit 360 may connect multiple virtual images to generate and output a smooth video. Furthermore, analysis unit 370 may analyze many virtual images to perform a highly accurate visual inspection of object W101. Additionally, system 100 may adjust the number of virtual viewpoints or the interval between virtual viewpoints so that the generated video is displayed smoothly at 30 or 60 frames per second (fps), for example.
[0059] The output unit 360 outputs the generated virtual image to the display device 120. For example, when an inspection staff member visually inspects the appearance of the object W101, the output unit 360 may output the virtual image to the display device 120. In one aspect, the output unit 360 may output multiple virtual images to the display device 120 in sequence while switching between them. In another aspect, the output unit 360 may connect the multiple virtual images to convert them into a video and output the video to the display device 120. In either case, the display device 120 displays the image or video so that the object W101 appears to rotate. In another aspect, the output unit 360 may transmit the generated virtual image to another device. For example, the output unit 360 may transmit multiple virtual images or a video generated from the multiple virtual images to a terminal held by the inspection staff member. In yet another aspect, the output unit 360 may output multiple virtual images to the analysis unit 370 instead of the display device 120. The analysis unit 370 has an appearance inspection function using AI or the like, enabling unmanned appearance inspection of the object W101. If visual inspection by a human is not required, the output unit 360 does not need to output a virtual image to the display device 120. Alternatively, if visual inspection and inspection by AI or the like are used in combination, the output unit 360 may output multiple virtual images to both the display device 120 and the analysis unit 370.
[0060] The analysis unit 370 determines whether or not there is an abnormality in the object W101. When the system 100 is applied to the production line shown in FIG. 1, the analysis unit 370 determines whether or not there is a scratch on the surface of the object W101 and outputs the determination result. The analysis unit 370 may display the determination result on the display device 120 or may transmit the determination result to a user's terminal or the like. In a certain aspect, the analysis unit 370 may include an inspection model or an inspection AI for the object W101. The inspection model or the inspection AI may be, for example, a trained model. Furthermore, the analysis unit 370 may be configured to be able to set an arbitrary inspection model or an inspection AI for each application of the system 100.
[0061] As described with reference to FIGS. 1 to 3, the system 100 includes an object detection unit 310 that detects the position of a moving object W101. The system 100 also includes an imaging unit 340 that captures images of the object W101 using one or more cameras C100. The system 100 also includes a virtual viewpoint setting unit 320 that sets a virtual viewpoint. The system 100 also includes a virtual image generation unit 350 that generates a virtual image that appears to be captured of the object W101 from the virtual viewpoint. The imaging unit 340 captures images of the object W101 at each of a plurality of imaging positions along the movement path of the object W101 using one or more cameras C100. The virtual viewpoint setting unit 320 sets a plurality of virtual viewpoints V101 to V106 that correspond to the plurality of imaging positions P1 to P6, respectively. The virtual image generating section 350 generates a plurality of virtual images VI201 to VI206 corresponding to a plurality of virtual viewpoints V101 to V106, respectively, from one or more images captured at a plurality of imaging positions P1 to P6.
[0062] System 100 also includes output unit 360 that outputs a plurality of virtual images VI201-VI206 corresponding to a plurality of virtual viewpoints V101-V106, respectively, so that object W101 appears to be rotating. In one aspect, output unit 360 may output virtual images VI201-VI206 to display device 120 while switching between them in order. In another aspect, output unit 360 may generate an image from virtual images VI201-VI206 that makes object W101 appear to be rotating, and output the image to display device 120.
[0063] One or more of the cameras C100 may also have a fisheye lens. Photographing the object W101 may include obtaining one or more images by flattening at least a portion of the fisheye image.
[0064] Furthermore, generating each of the multiple virtual images corresponding to each of the multiple virtual viewpoints may include inputting a chronologically ordered image set including one or more images captured at each of the multiple shooting positions into an inference model and generating each of the multiple virtual images corresponding to each of the multiple virtual viewpoints. The object W101 is photographed sequentially (i.e., in chronological order) at each of the multiple shooting positions. The system 100 may acquire the chronologically ordered image set by grouping one or more images captured at each shooting position. Furthermore, the system 100 may acquire the virtual image corresponding to each shooting position by inputting the acquired image set into an inference model.
[0065] Furthermore, setting each of the multiple virtual viewpoints V101 to V106 corresponding to each of the multiple shooting positions P1 to P6 includes, when the one or more cameras C100 are two cameras C100, setting each of the multiple virtual viewpoints V101 to V106 so that it moves on a line connecting the two cameras C100.
[0066] The system 100 also includes an analysis unit 370 that determines whether an abnormality exists in the object W101 by analyzing the virtual images VI201 to VI206 using an image analysis AI. The analysis unit 370 may be configured to be able to change the trained model used depending on the application. As an example, the image analysis AI may be a trained model for detecting an abnormality in the object W101.
[0067] As described above, the system 100 can take advantage of the fact that the object W101 is moving. Therefore, the system 100 can use a fixed camera to generate a virtual image that allows the object W101 to be observed from various angles. This allows the user to visually inspect the entire periphery of the object W101 simply by checking the virtual image displayed on the display device 120.
[0068] 4 is a diagram illustrating an example of a hardware configuration of device 110. Device 110 includes processor 410, primary storage device 420, secondary storage device 430, external device interface 440, input interface 450, output interface 460, and communication interface 470. In one aspect, device 110 may include hardware as a PLC. In another aspect, device 110 may be configured to be able to cooperate with an external PLC via external device interface 440 or communication interface 470.
[0069] The processor 410 may execute programs for implementing various functions of the device 110. The processor 410 may be configured, for example, with at least one integrated circuit. According to an embodiment, the integrated circuit may include at least one central processing unit (CPU), at least one graphics processing unit (GPU), at least one field programmable gate array (FPGA), at least one application specific integrated circuit (ASIC), at least one artificial intelligence (AI) chip, or a combination thereof.
[0070] The primary storage device 420 functions as a workspace for the processor 410. The primary storage device 420 stores programs executed by the processor 410 and data referenced by the processor 410. In one aspect, the primary storage device 420 may be realized by a dynamic random access memory (DRAM), a static random access memory (SRAM), or the like.
[0071] Secondary storage device 430 is a non-volatile memory and may store programs executed by processor 410 and data referenced by processor 410. In this case, processor 410 references data read from secondary storage device 430 to primary storage device 420. In one aspect, secondary storage device 430 may be realized by a hard disk drive (HDD), a solid state drive (SDD), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), a flash memory, or the like.
[0072] External device interface 440 can be connected to any external device such as a printer, a scanner, an external HDD, etc. In one aspect, external device interface 440 may be realized by a USB (Universal Serial Bus) terminal or the like.
[0073] Input interface 450 can be connected to any input device such as a keyboard, a mouse, a touchpad, or a gamepad. In one aspect, input interface 450 may be realized by a USB terminal, a PS / 2 terminal, a Bluetooth (registered trademark) module, or the like.
[0074] Output interface 460 can be connected to any output device such as a cathode ray tube display, a liquid crystal display, an organic electroluminescence (EL) display, etc. In one aspect, output interface 460 can be realized by a USB terminal, a D-sub terminal, a DVI (Digital Visual Interface) terminal, an HDMI (registered trademark) (High-Definition Multimedia Interface) terminal, a DisplayPort terminal, etc.
[0075] Communication interface 470 is connected to a wired or wireless network device. In one aspect, communication interface 470 may be implemented by a wired local area network (LAN) port, a Wi-Fi (registered trademark) (Wireless Fidelity) module, or the like. In another aspect, communication interface 470 may transmit and receive data using a communication protocol such as TCP / IP (Transmission Control Protocol / Internet Protocol), UDP (User Datagram Protocol), or the like.
[0076] <C.フローチャート>
[0077] 5 is a diagram showing an example of a procedure for generating a virtual image by system 100. In one aspect, processor 410 may load a program for performing the processing of FIG. 5 from secondary storage device 430 to primary storage device 420 and execute the program. In another aspect, some or all of the processing may be realized as a combination of circuit elements configured to perform the processing. Furthermore, in another aspect, the following steps may be executed in a reverse order.
[0078] In step S510, the system 100 determines whether the object W101 has entered the shooting area 160. In one aspect, the system 100 may determine whether the object W101 has entered the shooting area 160 based on position information of the object W101 held by the PLC, etc. In another aspect, the system 100 may analyze images captured by the cameras C101, C102, etc. to determine whether the object W101 has entered the shooting area 160. If the system 100 determines that the object W101 has entered the shooting area 160 (YES in step S510), the system 100 transfers control to step S520. If not (NO in step S510), the system 100 transfers control to step S510.
[0079] In step S520, the system 100 detects that the object W101 has arrived at the photographing position. Using Fig. 1 as an example, the system 100 detects that the object W101 is at one of the photographing positions P1 to P6. Similar to step S510, the system 100 can detect that the object W101 is at one of the photographing positions P1 to P6 by using information held by the PLC or image analysis technology.
[0080] In step S530, the system 100 captures an image of the object W101. More specifically, the system 100 captures an image of the object W101 using two or more cameras. Using FIG. 1 as an example, the system 100 captures an image of the object W101 using cameras C101 and C102. As a result, the system 100 obtains two or more images of the object W101.
[0081] In step S540, system 100 sets a virtual viewpoint. The virtual viewpoint corresponds to the shooting position in step S520. In one aspect, each of the multiple virtual viewpoints set by system 100 may be associated in advance with each of the multiple shooting positions.
[0082] In step S550, system 100 generates a virtual image. More specifically, system 100 generates the virtual image from one or more images obtained in step S530. The virtual image is an image that appears to be of object W101 photographed from the virtual viewpoint set in step S540.
[0083] In step S560, system 100 determines whether the object W101 has exited the imaging area 160. If system 100 determines that the object W101 has exited the imaging area 160 (YES in step S560), it ends the process. Otherwise (NO in step S560), system 100 transfers control to step S520. While the object W101 exists in the imaging area 160, system 100 can generate a plurality of virtual images by repeating the processes of steps S520 to S550. Taking FIGS. 1 and 2 as examples, system 100 can generate virtual images VI201 to VI206 by photographing the object W101 from different directions respectively.
[0084] As described with reference to FIG. 5, system 100 can execute the method defined by the control program by executing the control program. More specifically, system 100 can photograph the object W101 at each of a plurality of imaging positions in the movement path of the object W101 using one or more cameras C100, set each of a plurality of virtual viewpoints corresponding to each of the plurality of imaging positions, and generate each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints from one or more images photographed at each of the plurality of imaging positions.
[0085] <D. Variant 1>
[0086] In Variant 1, system 100 sets the virtual viewpoint so that the distance from the virtual viewpoint to the object is always constant. Since there are no changes in the functional configuration and hardware configuration, the descriptions thereof will not be repeated.
[0087] Fig. 6 is a diagram showing a modified example of the procedure for setting a virtual viewpoint. In the example of Fig. 1, each of the virtual viewpoints V101 to V106 was located on the line 130 connecting the cameras C101 and C102. In contrast, in this modified example, each of the virtual viewpoints V601 to V611 is located so that the distance from the object W101 is always constant. More specifically, each of the virtual viewpoints V601 to V611 is set at a position a predetermined distance 630 away from the object W101.
[0088] In the example of FIG. 6, cameras C101 and C102 capture images of object W101 at eleven shooting positions Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10, and Q11. Furthermore, system 100 sets virtual viewpoints V601, V602, V603, V604, V605, V606, V607, V608, V609, V610, and V611. Each of virtual viewpoints V601 to V611 corresponds to a corresponding one of shooting positions Q1 to Q11. Furthermore, each virtual viewpoint is set so that the distance from each virtual viewpoint to each shooting position is always constant. As an example, the distance from virtual viewpoint V601 to shooting position Q1 is 630. Similarly, the distance from each of the virtual viewpoints V602 to V611 to each of the shooting positions Q2 to Q11 is also a distance 630.
[0089] Virtual viewpoint arrangement 650 shows the arrangement of virtual viewpoints V601-V611 around the production line. As can be seen from virtual viewpoint arrangement 650, virtual viewpoints V601-V611 are arranged so as to form a complex trajectory between cameras C101 and C102. Positional relationship 660 of each virtual viewpoint shows the positional relationship of virtual viewpoints V601-V611 as seen from object W101. As can be seen from positional relationship 660 of each virtual viewpoint, virtual viewpoints V601-V611 are arranged at equal intervals on a circumference 670 centered on object W101. As a result, the size of object W101 captured in the virtual image corresponding to each virtual viewpoint is always constant. In other words, system 100 can obtain an image that appears as if object W101 were photographed while the camera was moving around circumference 670.
[0090] When the object W101 is placed at a fixed position, such as on a production line, the object detection unit 310 detects the position of the object W101 from information held by a PLC, etc. Conversely, when the position of the object W101 is unknown each time, the object detection unit 310 may detect the position of the object W101 by analyzing images or videos from the cameras C101 and C102. Alternatively, the object detection unit 310 may detect the position of the object W101 using a time-of-flight (TOF) sensor, a light detection and ranging (LiDAR) sensor, etc.
[0091] For each shooting position, the virtual viewpoint setting unit 320 can determine the position of the next virtual viewpoint by rotating the virtual viewpoint as viewed from the object W101 by a predetermined angle 640 from the previous virtual viewpoint. As an example, the virtual viewpoint V602 is tilted by an angle 640 more than the virtual viewpoint V601 as viewed from the object W101. Similarly, with each shooting position, the virtual viewpoint setting unit 320 tilts the virtual viewpoint by an angle 640 as viewed from the object W101. That is, the virtual viewpoint setting unit 320 changes the coordinates of the object W101 from the previous shooting position to the current shooting position. Then, the virtual viewpoint setting unit 320 rotates the virtual viewpoint as viewed from the object W101 by a predetermined angle 640 from the previous virtual viewpoint.
[0092] The virtual image generation unit 350 generates virtual images that appear to have been photographed from each of the virtual viewpoints V601 to V611. These virtual images appear to have been photographed of the object W101 from a fixed distance at different angles. The output unit 360 outputs the virtual viewpoints V601 to V611 to the display device 120. The display device 120 displays an image or video in which the object W101 photographed from a fixed distance appears to be rotating. In reality, the interval between each photographing position (distance 630) is extremely short, and it is desirable to set the virtual viewpoints frequently (for example, 30 or 60 times per second).
[0093] As described with reference to FIG. 6, setting each of the plurality of virtual viewpoints corresponding to each of the plurality of imaging positions includes setting each of the plurality of virtual viewpoints so as to generate a virtual image in which the size, rotation axis, and rotation speed of the object W101 always appear constant.
[0094] Furthermore, setting each of the plurality of virtual viewpoints so as to generate a virtual image in which the size, rotation axis, and rotation speed of the object W101 always appear constant includes setting a virtual viewpoint at a position separated from the object W101 by a predetermined second distance (distance 630) every time the object W101 moves a predetermined first distance. The movement of the first distance is the distance from the previous imaging position to the current imaging position (for example, the distance from imaging position Q1 to imaging position Q2). The virtual viewpoint is a position rotated by a predetermined angle (angle 640) around the object W101 with respect to the previous virtual viewpoint.
[0095] As described above, in the first modification, the system 100 generates a virtual image that appears to be captured from a constant distance at all times. Thereby, the inspection staff can always observe the object W101 from a constant distance, and can easily perform an appearance inspection of the object W101.
[0096] <E. Second Modification>
[0097] In the second modification, the system 100 generates virtual images of each of the plurality of objects, enabling the plurality of objects to be inspected simultaneously. Since there are no changes in the functional configuration and hardware configuration, the description thereof will not be repeated.
[0098] 7 is a diagram showing an example of how virtual viewpoints are set simultaneously for multiple objects. In the example of scene 710, objects W101A, W101B, W101C, W101D, and W101E are moving in the direction of arrow 170 on a conveyor belt, which is movement area 150. These objects are collectively referred to as object W101. Furthermore, five shooting positions R1, R2, R3, R4, and R5 are set within shooting area 160. Each of shooting positions R1, R2, R3, R4, and R5 is associated with a respective virtual viewpoint V701, V702, V703, V704, and V705.
[0099] The system 100 can simultaneously capture images of multiple objects W101 in the moving area 150. The system 100 can also simultaneously generate virtual images of each object W101. As an example, the system 100 captures an image of the object W101A at the image capture position R1. Based on the image obtained from the image capture, the system 100 generates a virtual image of the object W101A that appears to have been captured from a virtual viewpoint V701. Concurrently with this series of processes, the system 100 captures an image of the object W101B at the image capture position R2. The system 100 also generates a virtual image of the object W101B that appears to have been captured from a virtual viewpoint V702. Similarly, the system 100 captures images of the objects W101C to W101E at the image capture positions R3 to R5. Furthermore, system 100 generates virtual images that appear as if each of objects W101C-W101E were photographed from each of virtual viewpoints V703-V705. In this manner, system 100 can simultaneously generate virtual images of each of objects W101C-W101E in parallel. This allows system 100 to efficiently generate virtual images of multiple objects W101 present in shooting area 160.
[0100] Positional relationship 760 of each virtual viewpoint indicates the positional relationship of each virtual viewpoint V701-V705 as seen from the object W101. In the example of Fig. 7, five virtual viewpoints V701-V705 are evenly arranged clockwise around the object W101. The arrangement of the virtual viewpoints shown in Fig. 7 is merely an example, and any number of virtual viewpoints may be arranged at any position as seen from the object W101.
[0101] Depending on the application of the system 100, the orientation, position, and movement direction of the object W101 may be random. For example, suppose the system 100 is used in a manufacturing line that includes manual processes, a store surveillance camera, or other applications. In such cases, the orientation, position, and movement direction of the object W101 (workpiece or person) may be random. Therefore, the system 100 may dynamically change the virtual viewpoint for each capture position. In this way, the system 100 can always generate virtual images that appear to have been captured of the object W101 from each predetermined virtual viewpoint. For example, comparing scene 720 with scene 710, the orientations of the objects W101A to W101E are different. In such cases, the system 100 may adjust the positions of the virtual viewpoints V701 to V705 at each of the capture positions R1 to R5. In this way, system 100 can stably generate virtual images that appear to be taken of object W101 from each predetermined virtual viewpoint, regardless of, for example, a positional shift of object W101. In the example of Fig. 7, system 100 can set virtual viewpoints as shown in positional relationship 760 of each virtual viewpoint, regardless of, for example, a positional shift of object W101.
[0102] FIG. 8 is a diagram illustrating an example of simultaneous output of inspection images of multiple objects. System 100 may simultaneously generate virtual images of multiple objects W101 as shown in FIG. 7 . System 100 may also simultaneously display the virtual images of the multiple objects W101 on display device 120. In one aspect, system 100 may create an image for each of the multiple objects W101 from the multiple virtual images and display the image on display device 120. In the example of FIG. 8 , system 100 displays virtual images of four objects W101A, W101B, W101C, and W101D on display device 120, each of which has the same size and orientation. In another aspect, display device 120 may display any number of objects W101 greater than or equal to two. System 100 also aligns the rotation axes of objects W101A, W101B, W101C, and W101D. The system 100 can rotate the multiple objects W101 on the display device 120 in the same manner by switching between virtual images of the multiple objects W101 at the same time.
[0103] 7 and 8, setting each of the multiple virtual viewpoints corresponding to each of the multiple shooting positions includes setting each of the multiple virtual viewpoints for each of the multiple objects W101 so that each of the multiple objects W101 can be displayed from a predetermined direction when displaying the multiple objects W101 in output unit 360. That is, display device 120 can first display the multiple objects W101A-W101D photographed from a predetermined direction, and then rotate and display them in a fixed orientation.
[0104] When the system 100 is used for visual inspection of products on a production line, the inspection staff can observe multiple objects W101 of the same size and orientation. As another example, when the system 100 is used for monitoring people in a security area, the monitoring staff can monitor multiple people from the same direction. In either case, the efficiency of inspecting the objects W101 can be improved by making it possible to check multiple objects W101 of the same size and orientation.
[0105] In a certain situation, the system 100 may gradually shift the orientation of the virtual images of the plurality of objects W101 on the display device 120. By doing so, after visually inspecting the object W101A from the direction of the virtual viewpoint V701, the inspection staff can visually inspect the object W101B from the direction of the virtual viewpoint V701. That is, the inspection staff can visually inspect each object W101 from the same direction by utilizing the time difference.
[0106] <F. Other functions>
[0107] In addition to the functions described so far, the system 100 includes functions such as switching the camera to be used based on the position of the object W101, setting a plurality of virtual viewpoints for each shooting position, and using lighting. These functions can be used in combination with typical examples or other variations.
[0108] FIG. 9 is a diagram showing an example of the function of switching between a plurality of cameras. The system 100 may include three or more cameras C100. Also, the system 100 can change the combination of cameras C100 to be used according to the position of the object W101 or the shooting position.
[0109] Depending on the positional relationship among the plurality of cameras, the object W101, and the virtual viewpoint, it may be difficult to generate a virtual image. For example, assume that the object W101 is located near the line connecting two cameras and the virtual viewpoint is set near the center of the line. In this case, when performing projective transformation on the images obtained from the two cameras, the images are greatly distorted, so the generated virtual image may also be distorted. Therefore, the system 100 may select the camera to be used for each combination of the position of the camera C100, the position of the object W101, and the virtual viewpoint. As an example, the system 100 may select a combination of cameras with less distortion in projective transformation from among the plurality of cameras C100.
[0110] In the example of FIG. 9, there are shooting positions S1, S2, and S3. Furthermore, virtual viewpoints V901, V902, and V903 are associated with the shooting positions S1, S2, and S3, respectively. The system 100 selects the combination of cameras C101B and C102B for shooting at the shooting position S1. Similarly, the system 100 selects the combination of cameras C102A and C102B and the combination of cameras C101A and C102A for shooting at the shooting positions S2 and S3, respectively. In this way, the system 100 can select a combination from among the multiple cameras C100 that is appropriate for the combination of the position of camera C100, the position of object W101, and the virtual viewpoint. This allows the system 100 to generate a virtual image with minimal distortion.
[0111] As explained with reference to Figure 9, the virtual viewpoint setting unit 320 can select one or more cameras C100 to be used for shooting from one or more cameras C100 based on the positional relationship between one or more cameras, the position of the object W101, and the virtual viewpoint.
[0112] FIG. 10 is a diagram illustrating an example of a function for setting multiple virtual viewpoints for each shooting position. As described above, the system 100 may include an analysis unit 370. The analysis unit 370 can use an inspection model or an inspection AI, and therefore can instantly analyze more images than a human. Therefore, the system 100 may set multiple virtual viewpoints for each shooting position to improve the accuracy of the inspection. In this case, the system 100 generates virtual images that appear to be taken of the object W101 from each of the multiple virtual viewpoints for each shooting position.
[0113] In the example of scene 1010, system 100 sets virtual viewpoints V1001A and V1001B corresponding to shooting position T1. In this case, system 100 generates virtual images corresponding to each of virtual viewpoints V1001A and V1001B. Similarly, in the example of scene 1020, system 100 sets virtual viewpoints V1002A and V1002B corresponding to shooting position T2 and generates virtual images corresponding to each virtual viewpoint. In some aspects, system 100 may generate three or more virtual viewpoints for each shooting position. In other aspects, system 100 may set multiple virtual viewpoints on line 130 so that they are symmetrical when viewed from the center of line 130. In this way, system 100 can improve the accuracy of inspection using analysis unit 370 by setting multiple virtual viewpoints for each shooting position and generating more virtual images.
[0114] FIG. 11 illustrates an example of an inspection using a light source L300. The system 100 may illuminate the object W101 with light from one or more light sources L300. As an example, the system 100 may illuminate the object W101 with a grid-like light. The system 100 generates a virtual image of the illuminated object W101. Conventional techniques set the illumination pattern of the light source (i.e., the orientation of the grid, etc.) to make it easier to find flaws on the object W101 from video or images captured by a physical camera. In contrast, the system 100 may set a virtual viewpoint based on a predetermined illumination pattern to make it easier to find flaws on the object W101. Using the example of FIG. 11, given the relative positions of the object W101 and the light source L300, it is easier to find flaws on the object W101 from virtual viewpoint V1101 than from virtual viewpoint V1102. In this case, the system 100 selects the virtual viewpoint V1101. In one aspect, system 100 may select, from among a plurality of light sources L300, a light source L300 that has a light irradiation pattern that makes it easy to find flaws on object W101 from a predetermined virtual viewpoint.
[0115] Suppose there is an abnormality 1120, such as a scratch, on the surface of the object W101. In this case, the grid-like light irradiated onto the abnormality 1120 is distorted. As a result, the grid displayed in the virtual image is also distorted. Virtual image VI1130 is a virtual image of a normal portion of the object W101. Virtual image VI1140 is a virtual image of a portion of the object W101 where the abnormality 1120 is present. Comparing virtual image VI1130 and virtual image VI1140, it can be seen that the shape of the grid on the surface of the object W101 is different. The inspection staff or analysis unit 370 can easily detect the abnormality in the object W101 based on the change in the shape of the grid.
[0116] As described with reference to FIG. 11, the system 100 includes an illumination unit 330 that illuminates the object W101 with light for improving the inspection accuracy of the object W101 when photographing the object W101.
[0117] In some aspects, the various configurations and processes described with reference to Figures 1 to 11 may be used in appropriate combination. For example, some or all of the functions shown in Figures 7 to 11 may be applied to the processes shown in Figures 1 and 2.
[0118] <G.まとめ>
[0119] As described above, the system 100 according to the present embodiment can set virtual viewpoints for each of a plurality of shooting positions for the moving object W101. The system 100 can also generate virtual images corresponding to each virtual viewpoint and output these virtual images. This allows the user or the inspection AI to easily check the object W101 from a plurality of directions.
[0120] Furthermore, the system 100 can be used to inspect, monitor, or record any moving object W101, and therefore can be used for a wide range of applications, such as inspecting products, monitoring specific areas, and recording the passage of pedestrians or vehicles.
[0121] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present disclosure is defined by the claims, not by the above description, and is intended to include all modifications within the meaning and scope equivalent to the claims. Furthermore, the disclosures described in the embodiments and each modification are intended to be implemented, as far as possible, either alone or in combination. [Explanation of symbols]
[0122] 100 system, 110 device, 120 display device, 130 line, 150 movement area, 160, 160A, 160B shooting area, 170 arrow, 180A, 180B, 710, 720, 1010, 1020 scene, 200 virtual viewpoint setting, 210 output information, 310 object detection unit, 320 virtual viewpoint setting unit, 330 irradiation unit, 340 shooting unit, 350 virtual image generation unit, 360 output unit, 370 analysis unit, 410 processor, 420 primary storage device, 430 secondary storage device, 440 external device interface, 450 input interface, 460 output interface, 470 communication interface, 630 distance, 640 angle, 650 arrangement, 660, 760 positional relationship, 670 circumference, 1120 Abnormal, C100,C101,C102 Camera, L300 Light source, P1,P2,P3,P4,P5,P6,Q1,Q2,Q3,Q4,Q5,Q6,Q7,Q8,Q9,Q10,Q11,R1,R2,R3,R4,R5,S1,S2,S3,T1,T2 Shooting position,V101,V102,V103,V104,V106,V601,V602,V603,V604,V605,V606,V607,V608,V609,V610,V611,V701,V702,V703,V705,V701,V702,V703,V704,V705,V901,V902,V903,V1001A,V1001B,V1002A,V1002B,V1101,V1102 Virtual viewpoint,VI201,VI202,VI203,VI204,VI205,VI206,VI1130,VI114 Virtual image,W101 Object.
Claims
1. an object detection unit that detects the position of a moving object; an imaging unit that images the object using one or more cameras; a virtual viewpoint setting unit that sets a virtual viewpoint; a virtual image generating unit that generates a virtual image that appears to be taken of the object from the virtual viewpoint, the photographing unit photographs the object at each of a plurality of photographing positions along a moving path of the object using the one or more cameras; the virtual viewpoint setting unit sets a plurality of virtual viewpoints corresponding to the plurality of shooting positions, The virtual image generation unit generates each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints from one or more images captured at each of the plurality of shooting positions.
2. The system of claim 1 , further comprising an output unit that outputs each of the plurality of virtual images corresponding to each of the plurality of virtual viewpoints so that the object appears to rotate.
3. the one or more cameras have a fisheye lens; The system of claim 1 , wherein photographing the object comprises obtaining the one or more images by flattening at least a portion of a fisheye image.
4. The system of claim 1, wherein generating each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints includes inputting a time-sequential image set including the one or more images taken at each of the plurality of shooting positions into an inference model and generating each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints.
5. A system described in any one of claims 1 to 4, wherein setting each of the plurality of virtual viewpoints corresponding to each of the plurality of shooting positions includes, when the one or more cameras are two cameras, setting each of the plurality of virtual viewpoints so that it moves on a line connecting the two cameras.
6. The system described in any one of claims 1 to 4, wherein setting each of the plurality of virtual viewpoints corresponding to each of the plurality of shooting positions includes setting each of the plurality of virtual viewpoints so as to generate the virtual image in which the size, rotation axis, and rotation speed of the object always appear constant.
7. setting each of the plurality of virtual viewpoints so as to generate the virtual image in which the size, the rotation axis, and the rotation speed of the object appear to be constant includes setting the virtual viewpoint at a position away from the object by a predetermined second distance each time the object moves by a predetermined first distance; The system according to claim 6 , wherein the virtual viewpoint is a position rotated by a predetermined angle around the object relative to a previous virtual viewpoint.
8. The system according to any one of claims 1 to 4, wherein the object detection unit detects the position of the object using a position estimation model.
9. 3. The system of claim 2, wherein setting each of the plurality of virtual viewpoints corresponding to each of the plurality of shooting positions includes setting each of the plurality of virtual viewpoints so that the object can be displayed from a predetermined direction in the output unit.
10. Setting each of the plurality of virtual viewpoints corresponding to each of the plurality of shooting positions includes: The system according to claim 9, further comprising: when displaying a plurality of objects in the output unit, setting each of the plurality of virtual viewpoints for each of the plurality of objects so that each of the plurality of objects can be displayed from a predetermined direction.
11. The system described in any one of claims 1 to 4, wherein the virtual viewpoint setting unit selects one or more cameras to be used for shooting from the one or more cameras based on the positional relationship between the one or more cameras, the position of the object, and the virtual viewpoint.
12. The system according to any one of claims 1 to 4, further comprising an analysis unit that determines abnormalities in the object by analyzing the plurality of virtual images using image analysis AI.
13. The system according to claim 12 , further comprising an illumination unit that illuminates the object with light for improving inspection accuracy of the object when photographing the object.
14. 1. A method for controlling a system, comprising: photographing an object using one or more cameras at each of a plurality of photographing positions along a path of movement of the object; setting a plurality of virtual viewpoints corresponding to the plurality of shooting positions; generating a plurality of virtual images corresponding to each of the plurality of virtual viewpoints from one or more images captured at each of the plurality of shooting positions.
15. A program for a system, photographing an object using one or more cameras at each of a plurality of photographing positions along a path of movement of the object; setting a plurality of virtual viewpoints corresponding to the plurality of shooting positions; and generating each of a plurality of virtual images corresponding to each of the plurality of virtual viewpoints from one or more images taken at each of the plurality of shooting positions.
Citation Information
Patent Citations
Appearance inspection device and appearance inspection method
JP2023069348A