Control system, control method, and program
The control system optimizes data transmission by using distance information and selective data trimming to reduce bandwidth requirements for three-dimensional model generation, ensuring accurate model creation with reduced costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-17
- Publication Date
- 2026-03-30
AI Technical Summary
Existing techniques for generating three-dimensional models require large transmission bandwidths due to the large-sized data of imaging and distance images, leading to increased system costs.
A control system that acquires distance information and controls the transmission of data based on this information, including methods to trim and select data for transmission, such as using mask images and selecting cameras with overlapping imaging ranges, to reduce the amount of data sent.
This approach reduces the amount of data transmitted while maintaining or improving the accuracy of three-dimensional model generation.
Smart Images

Figure 2026054940000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a technique for controlling data transmission.
Background Art
[0002] There is a technique for synchronously imaging an object such as a subject with each imaging device set at a plurality of different positions to obtain imaging images and distance images from a plurality of viewpoints, and generating three-dimensional shape data (three-dimensional model data) of an object such as the object based on the imaging images and distance images from the plurality of viewpoints. Also, Patent Document 1 discloses a technique for generating a three-dimensional model based on a plurality of camera images (imaging images) with different viewpoints and a depth map (distance image) obtained from a depth camera.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the technique disclosed in Patent Document 1, a three-dimensional model is generated using a plurality of imaging images (camera images) and distance images (depth maps). Therefore, when transmitting large-sized data such as imaging images and distance images to a generation device that generates a three-dimensional model, a large transmission bandwidth is required, resulting in a problem of increased system cost.
[0005] Therefore, an object of this disclosure is to reduce the amount of transmission.
Means for Solving the Problems
[0006] The control system disclosed herein includes acquisition means for acquiring distance information relating to the distance from an imaging device to a subject, and control means for controlling the transmission of the distance information based on the distance information. [Effects of the Invention]
[0007] According to this disclosure, the amount of data transmitted can be reduced. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an example of the configuration of a system for generating three-dimensional models. [Figure 2] This figure shows an example of the internal configuration of an image processing device. [Figure 3] This is a diagram showing an example of the internal configuration of a control device. [Figure 4] This figure shows an example of overlapping imaging areas by cameras. [Figure 5] This figure shows an example of the internal configuration of a three-dimensional model generation device. [Figure 6] This is a control flowchart for the control device and the image processing device. [Figure 7] This figure shows an example of the hardware configuration of a control device. [Modes for carrying out the invention]
[0009] Hereafter, embodiments relating to this disclosure will be described with reference to the drawings. The following embodiments are not limiting to this disclosure, and not all combinations of features described in these embodiments are essential to the solutions of this disclosure. The configuration of the embodiments may be modified or changed as appropriate depending on the specifications and various conditions (usage conditions, usage environment, etc.) of the apparatus to which this disclosure applies. Furthermore, the embodiments described hereafter may be combined as appropriate. In the following embodiments, the same or similar configurations and processing steps will be denoted by the same reference numerals, and redundant descriptions will be omitted.
[0010] The transmission control device of the control system according to this embodiment performs control to transmit data including an image captured by the imaging device, a depth image, and information representing the state of the imaging device at the time of imaging, to a three-dimensional model generation device that generates a three-dimensional model. <First Embodiment> Figure 1 is a block diagram showing an example configuration of a control system according to the first embodiment. The control system according to the first embodiment includes, for example, a plurality of cameras 1 to 9, each of which is an imaging device, image processing devices 10 to 18 provided in each of the cameras 1 to 9, a control device 20, and a three-dimensional model generation device 30. The image processing devices 10 to 18 may be built into each of the cameras 1 to 9, or they may be separate devices connected in a one-to-one correspondence with each of the cameras 1 to 9. In this embodiment, an example in which each camera is equipped with an image processing device will be described. The image processing devices 10 to 18 are connected to the control device 20 and the three-dimensional model generation device 30 via a HUB 40. In the first embodiment, an example is given in which the control device 20 of the control system shown in Figure 1 realizes the function of a transmission control device.
[0011] Cameras 1 through 9 are imaging devices, each capable of acquiring both captured images and depth images, and are positioned at different locations. Cameras 1 through 9 perform synchronized imaging frame by frame, thereby acquiring captured images and depth images from multiple different viewpoints. Each camera also outputs information representing the camera's state at the time of imaging (hereinafter referred to as camera information), such as camera identification information, position, orientation, and direction at the time of imaging, zoom value (magnification) at the time of imaging, and depth of field at the time of imaging. Note that the camera information only needs to include at least one of the following: camera position, orientation, and direction, camera zoom value, and camera depth of field.
[0012] In this embodiment, the captured image is assumed to be an RGB image such as 4096 x 2160 (4K compatible resolution), but it is not limited to this example. The depth image is a depth image in which the distance value in the depth direction from the camera for each pixel is represented by shades of gray. For example, the closer the distance from the camera to the subject, the closer the value approaches black, and conversely, the further the distance from the camera to the subject, the closer the value approaches white. In other words, the depth image is a depth image that expresses the distance (depth) from the camera to the subject in the depth direction of three-dimensional space using shades of grayscale. The depth image is assumed to be an image such as 4096 x 2160 (4K compatible resolution), but it is not limited to this example. Thus, cameras 1 to 9 in this embodiment are cameras that can acquire both captured images (RGB images) and depth images. Examples of cameras that can acquire depth images include depth cameras equipped with sensors that can measure the distance from the camera to the subject, and TOF (Time Of Flight) cameras that use infrared technology to measure distance. Cameras 1-9 perform synchronized imaging for each frame, acquiring both the captured image and the depth image at the same time, and assigning the same acquisition time to each of these captured images and depth images.
[0013] Image processing devices 10 to 18 perform predetermined processing using the captured images and depth images acquired through synchronous imaging by their respective cameras 1 to 9. As will be described in detail later, the predetermined processing performed by the image processing devices includes extracting a foreground image from the captured images acquired by the corresponding cameras to create a mask image, and extracting distance information from the depth image based on the mask image to obtain depth information data. As will also be described in detail later, the image processing devices also perform the processing of transmitting the mask image, depth information data, and the captured images from the corresponding cameras in accordance with data transmission control commands received from the control device 20, which will be described later. Details of the configuration examples of image processing devices 10 to 18 and data transmission control based on data transmission control commands received from the control device 20 will be described later. The data output from these image processing devices 10 to 18, along with the camera information of each camera 1 to 9, is sent to the control device 20 and the three-dimensional model generation device 30 via the HUB 40.
[0014] The control device 20 transmits control commands for data transmission in each image processing device 10 to 18 based on the camera information for each camera 1 to 9 sent from each image processing device 10 to 18. A configuration example of the control device 20 and details of the transmission of data transmission control commands will be described later.
[0015] The three-dimensional model generation device 30 generates a three-dimensional model (three-dimensional shape data) corresponding to an object such as an object of a subject using the data transmitted from the image processing devices 10 to 18. The configuration example of the three-dimensional model generation device 30 and the details of the generation process of the three-dimensional model will be described later. Note that the three-dimensional model generation device 30 is also assumed to have a function of generating an image (virtual viewpoint image) corresponding to any one of a plurality of virtual viewpoints based on the generated three-dimensional model and the captured image. The three-dimensional shape data (three-dimensional model) generated by the three-dimensional model generation device 30 is represented by a point cloud, and each point that is a component of the three-dimensional shape data is represented by coordinate values of x, y, and z. Also, the coloring of the three-dimensional shape data is performed by assigning RGB values to those points. In the present embodiment, the three-dimensional shape data is represented by a point cloud, but it is not limited to this, and for example, it may be configured by a mesh representing a certain plane.
[0016] FIG. 2 is a diagram showing a schematic configuration example inside the image processing device 10, which is representative of the image processing devices 10 to 18 shown in FIG. 1. Note that FIG. 2 also shows the camera 1 and the HUB 40 connected to the image processing device 10. The image processing device 10 has, for example, a reception unit 101, a foreground extraction unit 102, a distance information extraction unit 103, a transmission unit 104, and a command reception unit 105.
[0017] The reception unit 101 receives the captured image (RGB image), the distance image (depth image), and the camera information from the corresponding camera 1 described above. The reception unit 101 outputs the captured image to the foreground extraction unit 102 and outputs the distance image to the distance information extraction unit 103. Note that in FIG. 2, the path of the camera information is omitted, but the camera information is sent to the transmission unit 104 described later, and is sent from the transmission unit 104 to the control device 20 and the three-dimensional model generation device 30 through the HUB 40.
[0018] The foreground extraction unit 102 compares a background image, which is a pre-captured imaging image without a foreground, with the imaging image captured by the camera 1, extracts an area with a difference greater than or equal to a predetermined value in these images as a foreground area, and obtains an image containing only the extracted foreground area as a foreground image. That is, in the foreground extraction unit 102, an image area of an object such as a subject that was not shown in the background image is extracted as the foreground image.
[0019] There are various methods for extracting the foreground area from the imaging image. In this embodiment, for example, a method according to the so-called background difference method is used. For example, the foreground extraction unit 102 determines that a pixel whose RGB value at the corresponding pixel position of the imaging image by the camera 1 and the background image is greater than or equal to a preset determination threshold value is a foreground pixel. The foreground image has the same image size as the imaging image, and each pixel represented by 1 Bit with the foreground pixels being white and the background pixels being black. Of course, the method for extracting the foreground area is not limited to this example, and any method that can compare images may be used. Also, the foreground area only needs to be information that can indicate that it is a foreground for each pixel, and the method is not particularly limited. Further, as the foreground extraction method, in addition to the background difference method described above, for example, a frame difference method, a foreground extraction method using artificial intelligence, or other methods capable of extracting the foreground may be used.
[0020] In the case of this embodiment, since the foreground image generated by the foreground extraction unit 102 is used as a mask image when extracting distance information from the distance image in the distance information extraction unit 103 described later, in the following description, the foreground image will be referred to as a mask image. The foreground extraction unit 102 outputs the mask image to the distance information extraction unit 103 and the transmission unit 104, and also outputs the imaging image (RGB image) captured by the camera 1 to the transmission unit 104.
[0021] The distance information extraction unit 103 receives the distance image from the receiving unit 101 and the mask image from the foreground extraction unit 102. The distance information extraction unit 103 then trims the distance image based on the mask image. In this embodiment, the trimming of the distance image using the mask image is performed to reduce the amount of data transmitted when the distance image is transmitted from the subsequent transmission unit 104. The distance information extraction unit 103 trims a rectangular area from the distance image input from the receiving unit 101 that corresponds to the foreground area of the mask image (the foreground area where pixels are white), and generates a distance image consisting only of the distance information of the trimmed rectangular area. In other words, the distance information extraction unit 103 generates a distance image after trimming from the aforementioned 4096 × 2160 (4K compatible resolution) distance image to extract only the distance information of each pixel within the rectangular area corresponding to the foreground area of the mask image. The distance information extraction unit 103 then combines the distance information of each pixel after trimming according to the foreground region of the mask image with the coordinate information of each pixel on the original distance image to generate depth information data, and outputs this depth information data to the transmission unit 104.
[0022] Furthermore, the distance information extracted by the distance information extraction unit 103 for each pixel represents absolute distance. For example, the distance information extraction unit 103 holds the median of the distance information and treats the difference from that median as distance information representing absolute distance. By using the difference from the median as distance information in this way, it is possible to reduce the bit length required to represent the distance information and reduce the amount of data transmitted when transmitting distance information. In addition, by using the distance information extracted by the distance information extraction unit 103 as information of the difference from the median, the control device 20 and the three-dimensional model generation device 30 that receive the transmitted data of the distance information can reconstruct the original distance information from the distance information of the median and each difference.
[0023] The command receiving unit 105 receives control commands that control the ON / OFF status of various data transmissions, which have been generated and transmitted by the control device 20 as described later, and sends these control commands to the transmission unit 104. The transmission unit 104 transmits the captured image captured by camera 1, the aforementioned mask image, and depth information data to HUB 40. In this embodiment, the transmission unit 104 controls the transmission of various data ON / OFF based on control commands that the command receiving unit 105 receives to determine whether to ON or OFF the transmission of various data. For example, in the initial state when the transmission unit 104 has not received a data transmission control command from the command receiving unit 105, it transmits the captured image, mask image, and depth information data together. Also, for example, if it receives a control command to turn OFF the transmission of depth information data, the transmission unit 104 transmits only the captured image and mask image, and does not transmit the depth information data. Also, for example, if it receives a control command to turn OFF the transmission of the mask image, the transmission unit 104 transmits only the captured image, and does not transmit the mask image or depth information data.
[0024] Figure 2 illustrates the configuration and operation of image processing device 10 as a representative example, but the other image processing devices 11 to 18 have the same configuration and operate similarly as image processing device 10. Therefore, in the other image processing devices 11 to 18 as well, the transmission unit 104 controls the ON / OFF transmission of various data based on the data transmission control commands received from the control device 20. In other words, in the system shown in Figure 1, the amount of data transmitted from image processing devices 10 to 18 is controlled by the data transmission control commands from the control device 20.
[0025] Figure 3 shows an example of the internal configuration of the control device 20 shown in Figure 1. Each processing unit shown in Figure 3 is implemented using the hardware configuration and program described later in Figure 7. As shown in Figure 3, the control device 20 includes an information holding unit 201, a selection unit 202, a state acquisition unit 203, a UI generation unit 204, a display unit 205, an input unit 206, a control unit 207, and a communication unit 208. Figure 3 also shows a HUB 40.
[0026] The information storage unit 201 stores camera information for each of the cameras 1 to 9. As mentioned above, the camera information includes camera identification information, camera parameters indicating the camera's position, orientation, and direction, camera zoom value, and camera depth of field information. In this embodiment, the information storage unit 201 is assumed to store camera information for nine cameras, as described above, but the information storage unit 201 can also store and manage camera information for, for example, around 100 cameras.
[0027] The status acquisition unit 203 acquires information indicating the current status of each camera from 1 to 9 and the status of the image processing devices 10 to 18, and outputs it to the UI generation unit 204. The UI generation unit 204 generates UI (user interface) images representing the system and the state of each camera 1 to 9 in Figure 1, based on the camera information held and managed in the information holding unit 201 and the information indicating the current state of each camera and image processing device acquired by the state acquisition unit 203. The display unit 205 includes a frame buffer and a display panel. The display unit 205 stores (overwrites) the display image of the UI image generated by the UI generation unit 204 in the frame buffer. Then, the display unit 205 reads the display image of the UI image stored in the frame buffer at a predetermined refresh rate and displays it on the display panel. The display panel is, for example, a liquid crystal panel or an organic EL panel. The input unit 206 receives operation information from a controller (not shown) entered by the user and outputs it to the control unit 207. The controller may be, for example, a keyboard, mouse, touch panel, or joystick.
[0028] The selection unit 202 selects an image processing device to transmit depth information data based on the camera information stored and managed in the information holding unit 201. In this embodiment, if the image processing device is built into the camera, the selection unit 202 will select the camera to transmit depth information data based on the camera information. The selection unit 202 then requests the control unit 207 to control the image processing device of the selected camera to transmit depth information data. In other words, the selection unit 202 selects an image processing device of a camera among the image processing devices 10 to 18 of cameras 1 to 9 that will not transmit depth information data, and requests the control unit 207 to control the selected image processing device so that it does not transmit depth information data. Details of the selection process in the selection unit 202 will be described later.
[0029] The control unit 207 transmits a control command to turn on data transmission via the communication unit 208, which enables the transmission of depth information data from the image processing device of the camera selected for data transmission. In other words, the control unit 207 transmits a control command to turn off data transmission via the communication unit 208, which enables the transmission of depth information data from the image processing device of the camera selected for non-data transmission. In addition, the control unit 207 also gives control instructions to each part of the control device 20 based on operation information input from the user via the input unit 206.
[0030] In the following section, the selection unit 202 will describe the method for selecting a camera image processing device that transmits depth information data, or in other words, the method for selecting a camera image processing device that does not transmit depth information data. In this embodiment, the selection unit 202 uses, for example, one of the three selection methods from the first to the third selection methods, or a selection method obtained by combining two or more of the first to third selection methods with predetermined weights.
[0031] The first selection method is one that selects based on the degree of overlap in the imaging ranges of each camera. In this embodiment, the imaging range of a camera is the range to be imaged, determined based on camera parameters indicating the position, orientation, and direction of the camera at the time of imaging, the focal length (zoom value) of the camera, and the angle of view determined by the image sensor. Furthermore, overlap of the imaging ranges of each camera means that the imaging range of one camera overlaps with the imaging range of another camera, and the degree of overlap of the imaging ranges is the ratio of the imaging range of one camera to the imaging range of another camera.
[0032] In the first selection method, the selection unit 202 identifies the imaging range for each camera based on the camera parameters of each camera (position, orientation, and direction) and the focal length and field of view determined by the image sensor of each camera. The selection unit 202 then selects the image processing devices of cameras with a greater degree of overlap in their imaging ranges as image processing devices that do not transmit camera information data. For example, the selection unit 202 counts the number of pixels in the area that overlaps with the imaging range of other cameras for all cameras. The selection unit 202 then selects the image processing devices of cameras with a larger number of pixels in the overlapping imaging range area as image processing devices that do not transmit depth information data.
[0033] Figure 4 shows an example of an image 50 captured by camera 1. In the example in Figure 4, it is assumed that the imaging ranges of camera 9 and camera 2 overlap in part with the imaging range of camera 1. In the example in Figure 4, region 51 represents the region where the imaging range of camera 9 overlaps with the imaging range of camera 1, while region 52 represents the region where the imaging range of camera 2 overlaps with the imaging range of camera 1. Furthermore, it is assumed that the region in the image 50 captured by camera 1 where the imaging ranges of camera 9 and camera 2 overlap accounts for approximately 30% of the entire image 50. In other words, the degree of overlap between the imaging range of camera 1 and the imaging ranges of the other cameras is approximately 30%.
[0034] The selection unit 202 calculates the degree of overlap in imaging ranges for all cameras and selects the image processing unit of the camera with the greatest degree of overlap with the imaging ranges of other cameras as the image processing unit that does not transmit camera information data. In other words, the selection unit 202 excludes the image processing unit of the camera with the greatest degree of overlap with the imaging ranges of other cameras from the selection of image processing units that transmit camera information data. For example, if the selection unit 202 wants to exclude two camera image processing units from the selection of image processing units 10 to 18 for cameras 1 to 9, it first excludes the camera image processing unit with the greatest degree of overlap in imaging ranges. Then, the selection unit 202 recalculates the degree of overlap in imaging ranges for each of the remaining camera image processing units and excludes the camera image processing unit with the greatest degree of overlap in imaging ranges from the selection, as described above. In this way, the selection unit 202 completes the selection process when it has excluded a predetermined number of camera image processing units (two in the above example) from the selection.
[0035] Furthermore, if, for example, there are two cameras with the greatest overlap in their imaging ranges, the selection unit 202 will exclude the image processing device of one of those two cameras, for example, the one with the higher camera identification number, from the selection. For example, if the cameras with the greatest overlap in their imaging ranges are camera 9 and camera 1, the image processing device 18 of camera 9, which has the highest identification number, will first be excluded from the selection, and then the image processing device 10 of camera 1 will be excluded from the selection.
[0036] For example, if the imaging directions of cameras 1 to 9 are different, the distance information extracted from the same region on each distance image acquired by each camera 1 to 9 will be different, and the surfaces of objects etc. corresponding to that region will also be different. For this reason, the selection unit 202 limits the cameras to those for which the degree of overlap in imaging ranges is determined, to those whose angle between the directions of each camera (i.e., the angle between the optical axes of each camera) is within a predetermined angle threshold (for example, within 30°). In other words, the selection unit 202 does not include cameras whose optical axis orientation is greater than the predetermined angle threshold in the determination of the degree of overlap in imaging ranges, and determines that there is no overlap in the imaging ranges of those cameras. Here, 30° is given as an example of the predetermined angle threshold, but it is not limited to this, and an angle threshold other than 30° may be used.
[0037] In the first selection method described above, depth information data from the image processing devices of cameras with a large degree of overlap in their imaging ranges is not transmitted, while depth information data from the image processing devices of cameras with no or small degree of overlap in their imaging ranges is transmitted. In other words, the first selection method reduces the amount of data transmitted by not transmitting depth information data from the image processing devices of cameras with a large degree of overlap in their imaging ranges. When generating a three-dimensional model with the three-dimensional model generation device 30 described later, a highly accurate three-dimensional model can be generated by using depth information data from as many areas as possible. According to the first selection method, by transmitting depth information data from the image processing devices of cameras with no or small degree of overlap in their imaging ranges, it becomes possible to transmit depth information data from a larger area overall, thus enabling the generation of a highly accurate three-dimensional model.
[0038] The second selection method is one that selects based on the camera's zoom value. Here, the larger the camera's zoom value, the larger the subject or object will appear in the captured image. Conversely, the smaller the camera's zoom value, the smaller the subject or object will appear.
[0039] For example, the larger the camera's zoom value, the more the depth information data in the camera's image processing device corresponds to an image in which the subject or other object is captured larger, resulting in data that indicates distance with higher accuracy than when the zoom value is small. Furthermore, when depth information data from a camera's image processing device with a larger zoom value is transmitted, the three-dimensional model generation device 30, described later, can generate a more accurate three-dimensional model.
[0040] On the other hand, the smaller the camera's zoom value, the wider the range of depth information data that the camera's image processing device will have compared to when the zoom value is large. In other words, when depth information data is transmitted to an image processing device of a camera with a smaller zoom value, the three-dimensional model generation device 30, described later, can generate three-dimensional models of objects and other objects over a wider range.
[0041] Therefore, in the second selection method, the selection unit 202 prioritizes selecting the image processing device of the camera with a larger zoom value as the image processing device for transmitting depth information data. Alternatively, the selection unit 202 prioritizes selecting the image processing device of the camera with a smaller zoom value as the image processing device for transmitting depth information data.
[0042] In addition, as a second selection method, the selection unit 202 may select, for example, a camera image processing device with a zoom value within a predetermined range as the image processing device for transmitting depth information data. That is, the selection unit 202 may select, for example, a camera image processing device that can capture a subject with a reasonably wide angle and a certain size as the image processing device for transmitting depth information data. In this case, the three-dimensional model generation device 30, which will be described later, will have the advantage of being able to use depth information data that achieves both accuracy and a wide range when generating the three-dimensional model.
[0043] The third selection method is one that is based on the camera's depth of field. Here, the depth of field in a camera is determined by the zoom value (focal length) and the f-number. The larger the zoom value, the shallower the depth of field, and the smaller the f-number, the shallower the depth of field. The shallower the depth of field, the narrower the range in which the image is in focus. When the depth of field is shallow, the extraction accuracy (separation accuracy of separating the subject as the foreground) when extracting subjects that are out of focus as foreground background decreases. On the other hand, in order to improve the accuracy of the three-dimensional model generated by the three-dimensional model generation device 30 described later, processing such as not using the mask image of the out-of-focus area in the three-dimensional model generation may be performed.
[0044] Therefore, as a third selection method, the selection unit 202 preferentially selects the image processing device of a camera with a shallower depth of field as the image processing device for transmitting depth information data. As a result, depth information data obtained from the image processing device of the camera with a shallow depth of field is preferentially transmitted, making it possible to cover areas where the mask image cannot be used with depth information data and further improve the accuracy of the generated three-dimensional model.
[0045] Figure 5 shows an example of the internal configuration of the three-dimensional model generation device 30. The three-dimensional model generation device 30 includes a receiving unit 301, a storage control unit 302, a storage unit 303, and a model generation unit 304. Figure 5 also shows a HUB 40 connected to the three-dimensional model generation device 30, and as explained in Figure 1, the three-dimensional model generation device 30 is connected to the image processing devices 10 to 18 of each camera 1 to 9 via the HUB 40.
[0046] The receiving unit 301 receives data transmitted from the image processing devices 10 to 18. The receiving unit 301 outputs the mask image or depth information data from the received data to the model generation unit 304 and outputs the captured image to the storage control unit 302. The storage control unit 302 outputs the captured image input from the receiving unit 301 to the storage unit 303. The model generation unit 304 generates three-dimensional shape data of an object, such as a subject, based on a mask image or depth information data, and outputs it to the storage unit 303. The storage unit 303 stores the captured images sent from the storage control unit 302, and also stores the three-dimensional shape data sent from the model generation unit 304. The three-dimensional model generation device 30 also has the function of generating virtual viewpoint images and the like from an arbitrary virtual viewpoint using the captured images and three-dimensional shape data stored in the storage unit 303.
[0047] The following describes how the model generation unit 304 generates three-dimensional models of objects such as bodies. In this embodiment, the model generation unit 304 generates a three-dimensional model of an object using a mask image representing the foreground region. There are various methods for generating a three-dimensional model of an object using a mask image of the foreground region, but in this embodiment, we will describe an example in which the three-dimensional model of an object is extracted and generated using the so-called viewing volume cross-section method. In the view volume cross-eyed method, the target area from which the three-dimensional model is extracted is divided into small rectangular prisms (hereinafter referred to as voxels). The number of pixels for each voxel when it appears in multiple captured images is calculated using three-dimensional computation, and it is determined whether or not that voxel corresponds to a foreground pixel. If a voxel becomes a foreground pixel in all the images captured by the cameras, that voxel is identified as a voxel that constitutes an object in the target area.
[0048] The model generation unit 304 performs processing using this viewing volume cross-section method, leaving only the voxels identified as the foreground of all camera images and deleting all other voxels. The remaining voxels are the voxels that constitute the objects present in the target area, and three-dimensional shape data of the objects is generated based on these. Each of these voxels corresponds to a point in the point cloud described above, and each voxel is a point with x, y, and z coordinates that constitute the three-dimensional shape data.
[0049] After generating a three-dimensional model as described above, the model generation unit 304 corrects the three-dimensional model using depth information data. At this time, the model generation unit 304 compares the three-dimensional model with the depth information data to check whether the distance between each voxel of the three-dimensional model and the camera matches the distance information contained in the depth information data. If the distances match, the model generation unit 304 considers the voxel to be correct. On the other hand, if the distances do not match, the model generation unit 304 performs further processing, such as trimming the voxel, based on the distance information in the depth information data.
[0050] The model generation unit 304 corrects the three-dimensional model by repeating this process. In this embodiment, as described above, the image processing devices 10 to 18 of cameras 1 to 9 are selected to either transmit depth information data or not. Therefore, when correcting the three-dimensional model, correction can be performed using depth information data, which can efficiently improve the accuracy of the three-dimensional model.
[0051] Next, we will explain the generation and transmission of data transmission ON / OFF control commands by the control device 20, and the operation of receiving control commands and setting the transmission of depth information data in the image processing device 10. Figure 6 is a flowchart showing the process flow in which the control device 20 generates a data transmission ON / OFF control command and sends it to the image processing devices 10-18, and the process flow in which the image processing devices 10-18 perform data transmission setting processing in response to the control command. Here, the control device 20 generates a control command that instructs each camera's image processing device to transmit depth information data ON / OFF based on camera information, and each image processing device sets the transmission of depth information data ON / OFF in response to that control command. In the flowchart of Figure 6, the symbol S represents a processing step (processing step). Also, in the flowchart of Figure 6, S601-S605 are processes performed in the control device 20, and S611-S614 are processes performed in the image processing devices 10-18.
[0052] In the control device 20, the first step in the process of S601 is for the selection unit 202 to acquire camera information from the information holding unit 201. Next, in the S602 process, the selection unit 202 selects a camera image processing device that transmits depth information data based on the camera information. The selection unit 202 makes a selection using one of the first to third selection methods described above, or a selection method that combines two or more of the first to third selection methods with predetermined weights. Next, in the process of S603, the selection unit 202 transmits the selection result of the camera's image processing device to the control unit 207. Next, in the S604 process, the control unit 207 sends a control command to the image processing device of the camera that was selected, for example, not to transmit depth information data (transmission OFF), based on the selection result in the selection unit 202, instructing it not to transmit depth information data.
[0053] In an image processing device selected by the control device 20 to turn off the transmission of depth information data, the first step in the process of S611 is for the command receiving unit 105 to receive a control command from the control device 20. Next, in the process of S612, the command receiving unit 105 transmits a control command to the transmission unit 104. Next, in the S613 process, the transmission unit 104 sets the transmission of depth information data to OFF according to the control command, and then transmits the completion of this setting to the command receiving unit 105. Next, in the process of S614, the command receiving unit 105 notifies the control device 20 that the setting is complete. On the other hand, when the control device 20 receives notification from the image processing device that the settings have been completed, as part of the process in S605, the control unit 207 completes the processing for controlling the transmission of depth information data.
[0054] As described above, the control device 20 of the first embodiment selects an image processing device for a camera that will transmit data (or will not transmit data) based on the camera information of each of the cameras 1 to 9, and generates and transmits a data transmission control command corresponding to that selection. The image processing devices 10 to 18 set the ON / OFF status of data transmission based on the control command received from the control device 20. This makes it possible to reduce the amount of data transmitted. When selecting an image processing device that will transmit data, one of the first to third selection methods described above, or a selection method that combines two or more of the first to third selection methods with predetermined weights, is used. The first to third selection methods are selection methods based on the camera's position, orientation, direction, camera zoom value, and degree of overlap of the camera's imaging range. By using the depth information data transmitted from the image processing device selected based on these selection methods, it becomes possible to generate a highly accurate three-dimensional model.
[0055] <Second Embodiment> In the second embodiment, an example is described in which the functions of the transmission control device are realized in the image processing devices 10 to 18, which correspond to cameras 1 to 9 in Figure 1, respectively. In the image processing device according to the second embodiment, the distance information extraction unit 103 extracts distance information from the distance image acquired by the corresponding camera, and the transmission unit 104 decides whether or not to transmit depth information data based on that distance information. In other words, in the case of the second embodiment, the processing performed by the transmission unit 104 of the image processing device 10 shown in Figure 2 is different from that of the first embodiment. In the case of the second embodiment, the control device 20 may or may not perform ON / OFF control of the transmission of depth information data to the image processing device of the camera described in the first embodiment. The example configuration of the image processing device 10 in Figure 2 described above will also be used in the description of the second embodiment.
[0056] In the second embodiment, the transmission unit 104 of the image processing device 10 determines whether or not to transmit depth information data acquired frame by frame by the corresponding camera 1, based on the distance information as described in the first embodiment. In this second embodiment, the transmission unit 104 uses one of the following five decision methods, from the first to the fifth, or a combination of two or more of the first to fifth decision methods, as a method for deciding whether or not to transmit depth information data. The first to third decision methods are methods that make a decision based on information about a region where the distance information changes. The fourth and fifth decision methods are methods that make a decision according to the distance value of the depth information data.
[0057] The first determination method is one that determines the region based on the total size of the area where the distance in the depth image changes. The region where the distance information changes can be identified, for example, as the region where the distance changes in a depth image acquired by a camera, even if the depth image itself contains no objects such as a subject. For example, if the total size of the region where the distance information changes is large, it may indicate that there is some kind of obstacle in the vicinity of the camera, preventing accurate data from being obtained. On the other hand, if the total size of the region where the distance information changes is small, it may indicate the presence of some kind of noise.
[0058] Therefore, as a first decision method, the transmission unit 104 decides not to transmit depth information data if the total size of the region where distance information changes is greater than a first size threshold predetermined as a significantly large size. Also as a first decision method, the transmission unit 104 decides not to transmit depth information data if the total size of the region where distance information changes is smaller than a second size threshold predetermined as a significantly small size. Note that the first size threshold and the second size threshold are different thresholds. According to the first decision method, the amount of transmitted data can be reduced by not transmitting depth information data when the total size of the region where distance information changes is significantly large or significantly small. Furthermore, in the case of the first decision method, the three-dimensional model generation device 30 can generate a highly accurate three-dimensional model based on depth information data when the total size of the region where distance information changes is not significantly large or significantly small.
[0059] The second decision method is one that determines based on the individual size of regions where distance information changes. For example, regions where the change in distance information is extremely small are likely to be noise. Therefore, in the second decision method, the transmission unit 104 decides not to transmit depth information data corresponding to each region where there is a change in distance information if the size of that region is smaller than a predetermined third size threshold. The size of each region can be determined by the number of pixels in that region, and if the number of pixels in that region is smaller than the number of pixels predetermined as the third size threshold, the transmission unit 104 decides not to transmit depth information data corresponding to that region. According to the second decision method, the amount of transmitted data can be reduced by not transmitting depth information data corresponding to regions where the change in distance information is small. In addition, in the second decision method, regions where the change in distance information is extremely small can be treated as noise, and the three-dimensional model generation device 30 can generate a highly accurate three-dimensional model based on depth information data from which noisy regions have been excluded.
[0060] The third decision method is one that determines the location based on the coordinates of areas where distance information changes. Devices such as cameras may have characteristics such as poor accuracy in areas near the edges of the image. Therefore, the third decision method determines whether or not to send depth information data depending on the coordinates of the region where the distance information has changed. According to the third decision method, the amount of transmitted data can be reduced by not transmitting depth information data corresponding to the region where the distance information has changed. In addition, in the case of the third decision method, processing is performed according to the characteristics of the device, such as the accuracy of regions near the edge of the image being reduced, so the three-dimensional model generation device 30 can generate a highly accurate three-dimensional model based on depth information data that excludes regions with reduced accuracy.
[0061] The fourth determination method is a method that determines based on the magnitude of the unevenness of an object such as a subject. When using the fourth determination method, the transmission unit 104 checks the depth information data in the region where the distance information changes and determines whether the unevenness is large. Here, the magnitude of the unevenness can be determined, for example, by the distance from the camera to the object in the region where the distance has changed in the distance image acquired by the imaging device, compared to a distance image when there is no object such as a subject. For example, the transmission unit 104 determines that the unevenness is large if, for pixels in the same region where the distance information has changed, the difference between the average distance of the top predetermined percentage of pixels that are farther away and the average distance of the top predetermined percentage of pixels that are closer to the object is greater than or equal to a predetermined difference threshold. The top predetermined percentage can be 10%, for example.
[0062] In the fourth decision method, the transmission unit 104 decides to transmit depth information data for areas with large irregularities. According to the fourth decision method, depth information data for areas other than those with large irregularities is not transmitted, thus reducing the amount of data transmitted. Also, in the fourth decision method, since depth information data for areas with large irregularities is transmitted, the three-dimensional model generation device 30 can generate a highly accurate three-dimensional model even in areas with large irregularities.
[0063] The fifth decision method is one that determines the distance from the camera to an object such as a subject. For example, the transmission unit 104 checks the distance information in a region where the distance information changes, and decides whether or not to transmit based on whether the distance information is within a predetermined distance range threshold. The distance range threshold is set in advance, for example, based on the distance measurement accuracy information of the depth sensor equipped in the camera. Furthermore, if the depth sensor has the characteristic that its distance measurement accuracy is high between 10m and 20m and low in other ranges, the transmission unit 104 decides to transmit depth distance data if the distance is within the range of 10m to 20m, and not transmit it otherwise. According to the fifth decision method, depth information data is not transmitted when the distance information is not within a predetermined distance range threshold, thus reducing the amount of transmitted data. In addition, with the fifth decision method, depth information data containing distance information in the range where the depth sensor's distance measurement accuracy is high is transmitted, so the three-dimensional model generation device 30 can generate a highly accurate three-dimensional model.
[0064] The first to fifth decision methods described above can be used individually or in combination. Here, a combination means sending the depth information data only when all five decision methods (first to fifth) determine that it should be transmitted. Alternatively, for example, if the first decision method has a set upper limit for the transmission size, the depth information data determined by the fourth decision method may be given priority for transmission.
[0065] As described above, according to the second embodiment, the image processing device determines depth information data to be transmitted on a frame-by-frame basis based on distance information, thereby enabling the transmission of data for improving the accuracy of the three-dimensional model more efficiently.
[0066] Figure 7 shows an example of the hardware configuration of the control device 20 according to this embodiment. The control device 20 according to this embodiment includes a CPU 222, ROM 223, RAM 224, large-capacity storage unit 225, NIC (Network Interface Card) 226, input unit 227, display unit 228, etc. The CPU 222 is the central processing unit and controls the entire computer using computer programs and data stored in the ROM 223 and RAM 224. In other words, the CPU 222 functions as each of the processing units shown in Figure 3. The ROM 223 stores configuration data for the control unit 20 and boot programs, etc. The RAM 224 has an area for temporarily storing computer programs and data loaded from the ROM 223, data acquired from external sources via the NIC 226, etc. Furthermore, the RAM 224 has a work area used by the CPU 222 when executing various processes. That is, the RAM 224 can be allocated as frame memory, for example, or various other areas can be provided as appropriate.
[0067] The input unit 227 consists of, for example, a keyboard or mouse, and receives various instructions from the user and inputs them to the CPU 222. The display unit 228 consists of, for example, a liquid crystal display, and displays the processing results from the CPU 222. The large-capacity storage unit 225 has, for example, an HDD or SSD. The large-capacity storage unit 225 stores the OS (operating system), computer programs that enable the CPU 222 to implement the functions of each processing unit shown in Figure 3, etc. Furthermore, the large-capacity storage unit 225 may also store image data to be processed. The computer programs and data stored in the large-capacity storage unit 225 are loaded into the RAM 224 as appropriate according to the control of the CPU 222 and become the target of processing by the CPU 222. The NIC 226 can be connected to networks such as LANs and the Internet, and various other devices such as projection devices and display devices, and the control device 20 can acquire and send various information via this NIC 226. The system bus 221 is a bus that connects the above-mentioned parts. The operation of each of the above-mentioned configurations is mainly controlled by the CPU 222.
[0068] Furthermore, the hardware configurations of the image processing units 10 to 18, even when configured differently from cameras 1 to 9, can generally be realized with configurations similar to those in Figure 7. When the image processing unit is realized with the configuration of Figure 7, the large-capacity storage unit 225 stores, for example, computer programs that enable the CPU 222 to implement the functions of each processing unit shown in Figure 2. In other words, the CPU 222 implements the functions of each processing unit shown in Figure 2 by executing these computer programs.
[0069] This disclosure can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions. The embodiments described above are merely examples of concrete implementations for carrying out this disclosure, and the technical scope of this disclosure should not be limited by them. In other words, this disclosure can be implemented in various ways without departing from its technical concept or its main features.
[0070] Each embodiment of the disclosure includes the following configurations, methods, and programs. (Composition 1) An acquisition means for acquiring distance information regarding the distance from the imaging device to the subject, A control means for controlling the transmission of the distance information based on the distance information, A control system characterized by having the following features. (Configuration 2) The acquisition means acquires distance information and information representing the state of the imaging device at the time of imaging from a plurality of imaging devices. The control means is The imaging device that transmits the distance information is selected based on the acquired information representing the state of the imaging device. The control system according to Configuration 1, characterized in that it sends an instruction to the selected imaging device for the imaging device to transmit the distance information, and does not send the instruction to an imaging device other than the selected imaging device. (Composition 3) The control system according to configuration 2, characterized in that the information representing the state of the imaging device includes at least one of the following: information indicating the position, orientation, and direction of the imaging device at the time of imaging; the zoom value of the imaging device at the time of imaging; and the depth of field of the imaging device at the time of imaging. (Composition 4) The control system according to configuration 3, characterized in that the control means acquires the degree of overlap of the imaging ranges of the multiple imaging devices based on information indicating the position, orientation and direction of the imaging devices at the time of imaging, and selects the imaging device based on the degree of overlap. (Composition 5) The control means is characterized in that it excludes the imaging devices from the selection target in order of the degree of overlap of the imaging ranges of the imaging devices. (Composition 6) The control system according to configuration 4, characterized in that the control means limits the imaging devices to which the degree of overlap of the imaging range is to be acquired to imaging devices in which the angle between the directions of the imaging devices at the time of imaging is within a predetermined angle threshold. (Composition 7) The control means is characterized in that it selects the imaging device based on the zoom value of the imaging device during imaging, as described in any one of configurations 3 to 6. (Composition 8) The control system according to configuration 7, characterized in that the control means selects an imaging device whose zoom value is greater than that of other imaging devices, or an imaging device whose zoom value is smaller than that of other imaging devices. (Composition 9) The control system according to configuration 7 or 8, characterized in that the control means selects an imaging device in which the zoom value is within a predetermined range. (Composition 10) The control means is characterized in that it selects the imaging device based on the depth of field of the imaging device during imaging, as described in any one of configurations 3 to 9. (Composition 11) The control system according to configuration 10, characterized in that the control means selects the imaging device prioritizing the imaging device having a shallower depth of field. (Composition 12) The acquisition means acquires the distance information from the distance image acquired in response to imaging by the imaging device, The control means is characterized in that it determines whether or not to transmit the distance information based on the total size of the region in the distance image acquired by the imaging device in the absence of a subject. This is the control system according to any one of configurations 1 to 11. (Composition 13) The control system according to configuration 12, characterized in that the control means decides not to transmit the distance information if the total size of the region where the distance has changed is greater than a first threshold, or less than a second threshold different from the first threshold. (Composition 14) The acquisition means acquires the distance information from the distance image acquired in response to imaging by the imaging device, The control system according to any one of configurations 1 to 13, characterized in that the control means determines whether or not to transmit the distance information based on the individual sizes of regions where the distance has changed in the distance image acquired by the imaging device, with respect to the distance image acquired when there is no subject. (Composition 15) The control system according to configuration 14, characterized in that the control means decides not to transmit the distance information if the individual sizes of the regions where the distance has changed are smaller than a third threshold. (Composition 16) The control means is characterized in that it determines whether or not to transmit the distance information based on the coordinates of the region where the distance has changed in the distance image acquired by the imaging device, for a distance image acquired when there is no subject. This is the control system according to any one of configurations 1 to 15. (Composition 17) The control system according to any one of configurations 1 to 16, characterized in that the control means determines whether or not to transmit the distance information to a distance image acquired when there is no subject, based on the distance from the imaging device to the subject in the region where the distance has changed in the distance image acquired by the imaging device. (Composition 18) The control system according to configuration 17, characterized in that the control means determines to transmit the distance information when the difference between the average of the upper predetermined proportion of distance information where the distance from the imaging device to the subject is farther away in the region where the distance has changed and the average of the upper predetermined proportion of distance information where the distance from the imaging device to the subject is closer is greater than or equal to a predetermined difference threshold. (Composition 19) The control system according to configuration 17 or 18, characterized in that the control means decides to transmit the distance information if the distance from the imaging device to the subject in the region where the distance has changed is within a predetermined distance range threshold. (Method 1) An acquisition process to acquire distance information regarding the distance from the imaging device to the subject, A control step that controls the transmission of the distance information based on the distance information, A control method characterized by having the following features. (Program 1) A program to cause a computer to perform the control method described in Method 1. [Explanation of symbols]
[0071] 10: Image processing device, 20: Control device, 30: Model generation device, 103: Distance calculation unit, 104, 704: Transmission unit, 105: Command receiving unit, 201: Information holding unit, 202: Selection unit, 207: Control unit
Claims
1. An acquisition means for acquiring distance information regarding the distance from the imaging device to the subject, A control means for controlling the transmission of the distance information based on the distance information, A control system characterized by having the following features.
2. The acquisition means acquires distance information and information representing the state of the imaging device at the time of imaging from a plurality of imaging devices. The control means is The imaging device that transmits the distance information is selected based on the acquired information representing the state of the imaging device. The control system according to claim 1, characterized in that it sends an instruction to the selected imaging device for the imaging device to transmit the distance information, and does not send the instruction to an imaging device other than the selected imaging device.
3. The control system according to claim 2, characterized in that the information representing the state of the imaging device includes at least one of the following: information indicating the position, orientation and direction of the imaging device at the time of imaging; the zoom value of the imaging device at the time of imaging; and the depth of field of the imaging device at the time of imaging.
4. The control means is characterized in that it obtains the degree of overlap of the imaging ranges of the plurality of imaging devices based on information indicating the position, orientation and direction of the imaging devices at the time of imaging, and selects the imaging device based on the degree of overlap, as described in claim 3.
5. The control means is characterized in that it excludes the imaging devices from the selection target in order of the degree of overlap of the imaging ranges of the imaging devices.
6. The control system according to claim 4, characterized in that the control means limits the imaging devices to which the degree of overlap of the imaging range is to be acquired to imaging devices in which the angle between the directions of the imaging devices at the time of imaging is within a predetermined angle threshold.
7. The control means is characterized in that it selects the imaging device based on the zoom value of the imaging device during imaging, as described in claim 3.
8. The control system according to claim 7, characterized in that the control means selects an imaging device whose zoom value is greater than that of other imaging devices, or an imaging device whose zoom value is smaller than that of other imaging devices.
9. The control system according to claim 7, characterized in that the control means selects an imaging device in which the zoom value is within a predetermined range.
10. The control means is characterized in that it selects the imaging device based on the depth of field of the imaging device during imaging, as described in claim 3.
11. The control means is characterized in that it selects the imaging device prioritizing the imaging device having a shallower depth of field, as described in claim 10.
12. The acquisition means acquires the distance information from the distance image acquired in response to imaging by the imaging device, The control means is characterized in that it determines whether or not to transmit the distance information based on the total size of the region in the distance image acquired by the imaging device in the absence of a subject, with respect to the distance image acquired in the absence of a subject.
13. The control system according to claim 12, characterized in that the control means decides not to transmit the distance information if the total size of the region where the distance has changed is greater than a first threshold, or less than a second threshold different from the first threshold.
14. The acquisition means acquires the distance information from the distance image acquired in response to imaging by the imaging device, The control system according to claim 1, characterized in that the control means determines whether or not to transmit the distance information to a distance image acquired without a subject, based on the individual sizes of regions where the distance has changed in the distance image acquired by the imaging device.
15. The control system according to claim 14, characterized in that the control means decides not to transmit the distance information if the individual sizes of the regions where the distance has changed are smaller than a third threshold.
16. The control system according to claim 1, characterized in that the control means determines whether or not to transmit the distance information based on the coordinates of the region where the distance has changed in the distance image acquired by the imaging device, with respect to the distance image acquired when there is no subject.
17. The control system according to claim 1, characterized in that the control means determines whether or not to transmit the distance information to the distance image acquired by the imaging device in a distance image without a subject, based on the distance from the imaging device to the subject in a region where the distance has changed in the distance image acquired by the imaging device.
18. The control system according to claim 17, characterized in that the control means determines to transmit the distance information when the difference between the average of the upper predetermined proportion of distance information where the distance from the imaging device to the subject is farther away and the average of the upper predetermined proportion of distance information where the distance from the imaging device to the subject is closer is greater than or equal to a predetermined difference threshold.
19. The control system according to claim 17, characterized in that the control means determines to transmit the distance information when the distance from the imaging device to the subject in the region where the distance has changed is within a predetermined distance range threshold.
20. An acquisition process to acquire distance information regarding the distance from the imaging device to the subject, A control step that controls the transmission of the distance information based on the distance information, A control method characterized by having the following features.
21. A program for causing a computer to execute the control method described in claim 20.
Citation Information
Patent Citations
3D model generation apparatus, method and program
JP2023094430A