Image processing method, information processing device, and computer program
The image processing method efficiently reconstructs moving images by generating and aligning three-dimensional model data from multiple captures, addressing inefficiencies in existing technologies by accurately aligning with pre-configured templates for realistic video representation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2026-04-15
AI Technical Summary
Existing technologies face inefficiencies in reconstructing moving images from photographed videos of objects, particularly in generating three-dimensional model data and aligning it with pre-configured video templates for accurate representation.
An image processing method involving capturing images from multiple directions to generate first three-dimensional model data, correcting orientation and position using markers, and replacing this data with pre-configured template data to reconstruct moving images, while adjusting size and position for accurate alignment.
Enables efficient reconstruction of moving images by aligning actual three-dimensional model data with pre-configured templates, ensuring accurate and realistic video representation.
Smart Images

Figure 0007846652000001 
Figure 0007846652000002 
Figure 0007846652000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing method, an information processing apparatus, and a computer program.
Background Art
[0002] Patent Document 1 describes a technique for generating a virtual space image visible from a virtual camera in a virtual space by mapping a texture generated from a group of captured images obtained by capturing an object from a plurality of shooting directions by a shooting unit to a primitive in the virtual space.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the above-described technology, when an object is photographed and a moving image is reconstructed using the photographed video of the object, it is expected that the reconstruction process is efficiently performed.
[0005] Therefore, when an object is photographed and a moving image is reconstructed using the photographed video of the object, it is possible to efficiently perform the reconstruction process.
Means for Solving the Problems
[0006] One embodiment for solving the above problem is an image processing method for reconstructing a moving image, comprising: a step of a processing unit generating a plurality of images from which images of the exterior of the model are captured from multiple directions by a shooting unit, the processing unit generating first three-dimensional model data of the model included in the images; a step of the processing unit reconstructing the moving image by replacing the second three-dimensional model data with the first three-dimensional model data in a moving image template data which has been configured in advance to include second three-dimensional model data; and a step of an output unit outputting information indicating the acquisition location of the reconstructed moving image. The present invention also relates to an image processing method for reconstructing moving images, wherein the appearance of a model is reconstructed by an imaging unit. of The process includes: a step of generating first three-dimensional model data of the model included in the images from multiple images generated by taking images from multiple directions; a step of reconstructing the moving image by replacing the second three-dimensional model data with the first three-dimensional model data in a moving image template data that has been configured in advance to include second three-dimensional model data; a step of the processing unit accepting the selection of the moving image template data; and a step of displaying a recommended pose for the moving image on a display device in response to the acceptance of the selection of the moving image template data. [Effects of the Invention]
[0007] When photographing an object and reconstructing a video using the captured image of the object, the reconstruction process can be carried out efficiently. [Brief explanation of the drawing]
[0008] [Figure 1] A diagram showing an example configuration of an image processing system corresponding to the embodiment, and a diagram showing an example of the hardware configuration of the information processing device 100. [Figure 2] A diagram showing an example of the appearance of a model corresponding to an embodiment. [Figure 3] A figure showing an implementation example of an image processing system corresponding to the embodiment, and an example of an AR marker. [Figure 4]A diagram showing an example of the data structure of a data table corresponding to an embodiment. [Figure 5] A flowchart illustrating an example of the process for generating PV data corresponding to the embodiment. [Figure 6] A flowchart corresponding to an example of the imaging process corresponding to the embodiment. [Figure 7] A flowchart showing an example of the PV reconstruction process corresponding to the embodiment. [Figure 8] A diagram illustrating the shooting environment when taking images of a model corresponding to an embodiment. [Figure 9] This figure shows a comparative example of the output results when the PV template conversion corresponding to the embodiment is performed on temporary data and when it is performed on the actual data. [Modes for carrying out the invention]
[0009] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims, and not all combinations of features described in the embodiments are essential to the invention. Two or more of the features described in the embodiments may be arbitrarily combined. Also, identical or similar configurations will be given the same reference numeral, and redundant descriptions will be omitted. Furthermore, in each figure, the top, bottom, left, right, front, and back directions relative to the paper will be used as the top, bottom, left, right, front, and back directions of the parts (or components) in this embodiment, and will be used in the descriptions within the text.
[0010] First, the configuration of the image processing system corresponding to this embodiment will be described. Figure 1(A) is a diagram showing an example of the configuration of the image processing system 10 corresponding to this embodiment. The image processing system 10 is configured by connecting an information processing device 100 to an imaging unit 110, a support arm 120, a model support device 130, a display device 140, a printing device 150, etc. Note that the system configuration is not limited to that shown in Figure 1(A), and the information processing device 100 may be further connected to an external server, cloud server, etc. via a network. Such external server, etc., can execute at least some of the processing of the embodiment described below.
[0011] The information processing device 100 controls the operation of at least one of the shooting unit 110, the support arm 120, and the model support device 130 to photograph the object to be photographed from any angle to generate multiple images, and generates three-dimensional model data (main data) from these multiple images. Then, using a template for a moving image (promotional video: PV data) (PV template) that includes pre-prepared temporary three-dimensional model data (temporary data), the PV data can be reconstructed by replacing the temporary data in the PV template with the main data and converting it. Furthermore, the information processing device 100 can also function as an image processing device that generates images to be displayed in a virtual space by using the generated main data as data for a virtual space. An image processing device may be prepared separately from the information processing device 100. In this embodiment, the object to be photographed is an assembleable plastic model, an action figure (a figure with movable joints), a toy, a doll, etc., and these are collectively referred to as "models". Furthermore, in this embodiment, a promotional video (PV) is not limited to a video intended for advertising, promoting, or selling the model being filmed, but also refers to any attractive moving image (short movie, movie, or demo video, etc., regardless of the format or type of moving image) using the model being filmed. In the following explanation, "PV" or "promotional video" will be used as a general term to include these moving images.
[0012] Next, the imaging unit 110 is an imaging device that, in accordance with the control of the information processing device 100, captures (scans) the three-dimensional shape of the model to be photographed and outputs an image of the model. The imaging unit 110 can, for example, utilize a smartphone with a camera and an application for capturing three-dimensional shapes installed. Alternatively, the imaging unit 110 may be configured as a three-dimensional scanner device capable of outputting three-dimensional shape and color information. For example, the Space Spider manufactured by Artec can be used as such a three-dimensional scanner device. For example, by acquiring approximately 500 to 800 scan images, 3D model data of the entire model can be obtained.
[0013] The support arm 120 is a position and orientation control device that moves the imaging unit 110 to a predetermined imaging position and orientation according to the control of the information processing device 100. The support arm 120 may be configured to allow manual changes to the imaging position and orientation, and to maintain the changed position and orientation in a fixed position, or it may be configured to be controllable by the information processing device 100. When the support arm 120 is configured to be controllable by the information processing device 100, for example, the xArm 7 manufactured by UFACTORY can be used. The xArm 7 consists of seven joints and can move in a manner similar to a human arm. Alternatively, the imaging unit 110 may be positioned manually instead of using the support arm 120.
[0014] The model support device 130 is a support base that supports a model with a fixed pose. The model support device 130 may be configured to be rotatable, for example, when the model is installed on the support base or at the tip of a support rod. In this embodiment, after the imaging unit 110 is positioned at an arbitrary imaging position and imaging angle by the support arm 120, the model support device 130 is rotated one full turn to perform imaging. By performing this at a plurality of imaging positions and imaging angles, an image of the entire model can be obtained. Here, by driving the support arm 120 and the model support device 130 in synchronization, the imaging process can be performed more simply and with high accuracy. Alternatively, instead of the model support device 130, the imaging unit 110 may be manually moved around the model to perform imaging from an arbitrary imaging position and imaging angle.
[0015] The display device 140 is a display device such as a liquid crystal display (LCD), and displays the image acquired by the imaging unit 110, or displays the original data of the three-dimensional model data generated from the captured image, or displays the PV data reconstructed using the original data. The display device 140 can also display a two-dimensional barcode (QR code) indicating the download destination link on the screen. The printing device 150 is a printing device that prints a QR code indicating the download destination link on a paper medium or the like for the user to acquire and view the PV data after the PV data is reconstructed.
[0016] FIG. 1(B) shows an example of the hardware configuration of the information processing device 100. The CPU 101 is a device that performs overall control of the information processing device 100 and calculates, processes, and manages data. For example, it controls the imaging timing and the number of imaging shots in the imaging unit 110, and controls the joints of the arm of the support arm 120 to position the imaging unit 110 at an arbitrary imaging position and imaging angle. Further, after the imaging position and imaging angle of the imaging unit 110 are determined, the model support device 130 can be rotated to perform an imaging operation by the imaging unit 110. The CPU 101 can also function as an image processing unit that processes the image output from the imaging unit 110.
[0017] The RAM 102 is a volatile memory and is used as a main memory of the CPU 101 and a temporary storage area such as a work area. The ROM 103 is a non-volatile memory, and image data, other data, various programs for the operation of the CPU 101, etc. are stored in respective predetermined areas. The CPU 101 controls each part of the information processing apparatus 100 by using the RAM 102 as a work memory according to, for example, a program stored in the ROM 103. Note that the program for the operation of the CPU 101 is not limited to being stored in the ROM 103 and may be stored in the storage device 104.
[0018] The storage device 104 is constituted by, for example, a magnetic disk such as an HDD or a flash memory. In the storage device 104, application programs, an OS, control programs, related programs, game programs, etc. are stored. Further, data of a PV template to be reconstructed is included. The PV template includes temporary data, and the size and pose of the three-dimensional model may be different for each PV template. The storage device 104 can read and write data based on the control of the CPU 101. The storage device 104 may be used instead of the RAM 102 or the ROM 103.
[0019] The communication device 105 is a communication interface for communicating with the imaging unit 110, the support arm 120, and the model support device 130 based on the control of the CPU 101. The communication device 105 may be further configured to be able to communicate with an external server or the like. The communication device 105 can include a wireless communication module, and the module can include a well-known circuit mechanism including an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identification module card, a memory, etc. Here, the communication between the information processing apparatus 100 and the imaging unit 110, the support arm 120, and the model support device 130 may be performed by wireless communication.
[0020] Furthermore, the communication device 105 may also include a wired communication module for wired connection. The wired communication module enables communication with other devices, including the display device 140, via one or more external ports. It may also include various software components for processing data. The external ports connect to other devices directly or indirectly via a network, such as via Ethernet, USB, or IEEE 1394. It should be noted that software that achieves equivalent functionality to the above devices can also be used as a substitute for the hardware devices.
[0021] The operation unit 106 consists of, for example, buttons, a keyboard, a touch panel, etc., and accepts user input. The display control unit 107 functions as an interface for displaying information on the display device 140 connected to the information processing device 100, and controls the operation of the display device 140. Some of the functions of the operation unit 106 may be provided on the display device 140. That is, the display device 140 may be configured as a device equipped with a touch panel, such as a tablet terminal.
[0022] Next, with reference to Figure 2, an example of a model to be photographed in this embodiment will be described. Model 200 is a model having the appearance of a human (robot or human). The model may be assembled and painted as, for example, a plastic model. Alternatively, it may be a pre-assembled model, such as an action figure with movable joints. The model in Figure 2 is merely an example for illustrative purposes, and the shape of the model is not limited to those with a human-like appearance; it can be a model of any shape, such as a general vehicle, a racing car, a military vehicle, an aircraft, a ship, an animal, a virtual life form, etc. It should be noted that the item to be photographed is not limited to a model, as long as the three-dimensional shape of the item can be photographed by the photography unit 110.
[0023] Model 200 consists of the following parts: head 201, chest 202, right arm 203, left arm 204, right torso 205, left torso 206, right leg 207, left leg 208, right foot 209, and left foot 210, which are joined together to form the model. At least a portion of each individual part 201-210 is supported so as to be rotatable (or swingable) relative to an adjacent part. For example, the head 201 is rotatably supported relative to the chest 202, and the right arm 203 and left arm 204 are rotatably supported relative to the chest 202. In this way, each part of model 200 is provided with a joint structure, so that model 200 can assume any posture.
[0024] Next, an example of the implementation of the image processing system 10 in this embodiment will be described with reference to Figure 3. Figure 3(A) shows an implementation example in which PV data is reconstructed from the main data generated by photographing a model, and a download link for viewing the reconstructed PV data can be issued by a printer. Figures 3(B) and (C) show an example of the configuration of the AR (Augmented Reality) marker in this embodiment.
[0025] In Figure 3(A), the case 301 contains the information processing device 100, a drive system for driving the support arm 302, a drive system for driving the turntable 306, and the like. The support arm 302 can also be manually adjusted for shooting direction and position. The surface of the case 301 is flat so that promotional posters can be attached.
[0026] The support arm 302 corresponds to the support arm 120 and can support and fix the position of the terminal 303, which functions as an imaging unit 110, either under the control of the information processing device 100 or manually. The support arm 302 can also operate to control the tilt of the terminal 303.
[0027] Terminal 303 is a touch-panel terminal with a built-in camera, and can be used as a smartphone, tablet, digital camera, etc. Terminal 303 can take pictures of the model 200 and transmit them to the information processing device 100. The ring light 304 is a lighting device used when terminal 303 takes pictures of the model 200, and it can evenly illuminate the model 200 to minimize shadows. In addition to the ring light 304, auxiliary lights such as a top light and lights on the left, right, and bottom may be installed as additional light sources.
[0028] The background sheet 305 is a background sheet for photography, and a white sheet can be used, for example. The turntable 306 can mount the model 200 and rotate the model. Multiple AR markers 310 are arranged on the turntable 306, as shown in Figure 3(B). The AR markers 310 are used to adjust the orientation and position of the photographed model. As shown in Figures 3(B) and 3(C), three-dimensional coordinates can be identified from the AR markers 310, and the inclination of the X, Y, and Z axes can be identified. For example, if the three-dimensional coordinates identified from the AR markers 310 shown in Figure 3(B) are used as a reference, the orientation and inclination of the data can be correctly corrected by matching the three-dimensional coordinates identified from the AR markers 310 shown in Figure 3(C) to that reference.
[0029] In Figure 3(A), the model is placed on a translucent (transparent) base, but other support devices, such as those called "action bases," may also be used. An action base has a support column on a base that is bent into a "V" shape, and the model can be attached to the end of the column. An AR marker may also be placed on the end of the column. The method of supporting the model can be changed depending on the pose. For example, in an upright position, the model can be placed on a transparent base for shooting. On the other hand, if you want to photograph the soles of the feet, such as in a flying pose, you can use an action base. Note that an action base may also be used to photograph the upright position.
[0030] The display device 307 is a device that corresponds to the display device 140 and may have a touch panel function. The display device 140 can, for example, display icon images as options for the PV that the user wishes to reconfigure, and accept the user's selection. The printer 308 is, for example, a thermal printer and can print a QR code indicating the download URL after the PV data reconfiguration is complete. Alternatively, the QR code may be displayed on the display device 307 so that the user can scan it with their smartphone or other device.
[0031] Next, with reference to Figure 4, the data structure of the various tables stored in the storage device 104 of the information processing device 100 will be described. In the PV management table 400 in Figure 4(A), the model type information 402 that identifies the model type of the model, information on temporary data included in the PV 403, and the storage location 404 for the PV template data are registered in association with the PV name 401.
[0032] First, the PV name 401 contains identification information to uniquely identify the template data of the promotional video (PV) that will be the source of the moving image to be reconstructed in this embodiment. The model type information 402 is information to identify the type of model to which the PV corresponds. In this embodiment, PVs may be prepared for each type of model. For example, a first shape may be a normal shape (for example, an external shape that is a scaled-down version of a life-size model), and a second shape may be a deformed shape (for example, an external shape that is a deformed version of a life-size model, such as a 2-head or 3-head proportion model). The model type information 402 indicates in M1 that the corresponding PV corresponds to a model of the first shape, and in M2 that it corresponds to a model of the second shape.
[0033] Temporary data information 403 registers information about the temporary data included in the PV. In this embodiment, the PV template data is constructed using temporary data, and the PV can be reconstructed by replacing this temporary data with the actual data generated from the captured image. Storage location 404 indicates information about the storage location of the PV template data.
[0034] Table 400 can be used to register other information related to the PV. This information may include, for example, icon images for PV selection, sample movies, and information about recommended poses in the PV.
[0035] Next, Figure 4(B) shows an example configuration of the user information table 410, which registers information of users who use the image processing system 10. Table 410 registers the username 411, PV name 412, captured image data 413, three-dimensional model 414, and URL 415. The username 411 is information to uniquely identify the user (or player) who owns the model captured by the imaging unit 110. The PV name 412 is where the PV name selected by the user is registered. The PV name 412 corresponds to one of the pieces of information registered as PV name 401 in table 400.
[0036] The captured image data 413 contains information about the captured image obtained when the camera unit 110 photographs a model owned by the user registered under username 411. The three-dimensional model 414 contains the main data generated based on the captured image data 413. The URL 415 is link information indicating the storage location of the PV data reconstructed using the main data of the three-dimensional model 414, and is a URL indicated by a QR code printed by the printing device 150.
[0037] Next, with reference to Figure 5, an example of processing performed by the information processing device 100 corresponding to this embodiment will be described. At least a part of the processing corresponding to the flowchart is realized by the CPU 101 of the information processing device 100 executing a program stored in the ROM 103 or the storage device 104.
[0038] First, in S501, CPU 101 accepts user registration. It accepts input of the user's name and contact information. Each user is assigned a user identifier to uniquely identify them. This user identifier corresponds to username 411 in Figure 4(B). CPU 101 stores the entered user information in storage device 104, associating it with the time the input was received and the user identifier.
[0039] In the subsequent S502, the CPU 101 displays icons for selectable promotional videos (PVs) on the display device 140 via the display control unit 107, and accepts the user's selection of a PV they wish to create. Any number of PV icons exist, and each PV contains different scenes. A sample movie may be displayed at this time. In addition, PVs may be prepared for each type of model. For example, a first shape may be a normal shape (e.g., an external shape that is simply a scaled-down version of a life-size model), and a second shape may be a deformed shape (e.g., an external shape that is a deformed version of a life-size model). The user can determine whether the shape of the model they own is a normal shape or a deformed shape, and select the PV corresponding to that shape.
[0040] In the subsequent S503, the CPU 101 accepts the user's selection of a PV. Each PV is pre-assigned identification information, and the identification information of the selected PV is linked with the user information received in S501 and stored in the table 410 of the storage device 104.
[0041] In the subsequent S504, the CPU 101 displays the recommended pose for the model in the selected PV on the display device 140 via the display control unit 107. In this embodiment, the model's pose differs for each PV, and by capturing the model's appearance in a pose that matches the selected PV, it is possible to reconstruct a more realistic video. Poses for the model in this embodiment include, for example, an upright posture, a flying posture, and an attacking posture. Since PVs are prepared for the upright posture, the flying posture, and the attacking posture, it is important to take the pose corresponding to the content of each PV in order to create a PV that does not look unnatural.
[0042] The user checks the recommended pose of the model displayed on the display device 140, decides on the pose of their own model, and sets it on the model support device 130. The CPU 101 can determine whether the model has been set on the model support device 130 based on the image captured by the imaging unit 110. Alternatively, a switch that turns on when a model is set on the model support device 130 may be provided, and the CPU 101 may detect the signal from the switch to make the determination. Alternatively, a button for accepting operation when the model has been set may be displayed on the display device 140, and the CPU 101 may detect whether an operation in response to the button operation has been received. In S505, the CPU 101 detects that the model has been set on the model support device 130 by one of the above methods. In response to this detection, the process proceeds to S506.
[0043] In S506, the imaging unit 110 performs imaging processing to generate an image. The captured image is transmitted to the information processing device 100, and the CPU 101 stores it in the storage device 104's table 410 as captured image data 413, associated with a user identifier, etc. In the subsequent S507, the CPU 101 performs PV reconstruction processing based on the captured image data 413. The CPU 101 stores the reconstructed PV data in the storage device 104, associated with the user information registered in S501.
[0044] In the subsequent S508, a URL and QR code are issued for the user to access the PV data. The URL and QR code may be sent to the user's email address or displayed on the display device 140. Once the user obtains the URL or QR code, they can access the storage location of the PV data and download the reconstructed PV data in S509.
[0045] Next, with reference to Figure 6, the details of the shooting process in S506 will be explained. In this embodiment, the drive information of the support arm 120 and the number of times the shooting unit 110 takes a picture at the shooting position are registered for each shooting position. When taking a picture at each shooting position, the model support device 130 is rotated to perform the shooting.
[0046] First, in S601, the CPU 101 controls the support arm 120 to move the imaging unit 110 to one of the registered imaging positions. Alternatively, the support arm 120 may be manually controlled to move to one of the imaging positions. The CPU 101 also controls the model support device 130 to move it to the rotation start position. In the subsequent S602, the CPU 101 rotates the model support device 130 at the imaging position while the imaging unit 110 takes a picture, thereby acquiring an image of the model 200. This allows for the acquisition of multiple images with different imaging angles. For example, if the imaging angle is changed by 15 degrees, 24 images can be taken, and if it is changed by 10 degrees, 36 images can be taken. By making the range of change in the imaging angle smaller, the accuracy of the main data generated in the subsequent stage is improved.
[0047] Alternatively, the model support device 130 may be stopped at a specific rotational position, and the model may be photographed while moving the model to the imaging unit 110 vertically by controlling the support arm 120, or by the photographer manually changing the angle of the imaging unit 110. This allows the model to be photographed from above or below. In particular, photographing from below makes it possible to photograph the soles of the model's feet. In this embodiment, the shooting direction and shooting angle are selected to cover the entire circumference of the model.
[0048] In this embodiment, the model can be suspended and photographed by setting it on a transparent support rod extending from the model support device 130. Alternatively, it can be photographed standing upright directly on the turntable of the model support device 130. In the former shooting method, the soles of the model's feet can also be photographed, making it suitable for photographing a flying posture. On the other hand, in the latter shooting method, the soles of the feet cannot be photographed, but it is suitable for photographing an upright posture. Note that the upright posture may also be photographed in the former shooting method.
[0049] Furthermore, AR markers are placed on the support rods and turntable of the model support device 130, and when taking images, these AR markers are included in the image generation. The AR markers are used to correct the orientation and position of the object being photographed.
[0050] In the following step S603, the CPU 101 determines whether there are any unselected shooting positions among the pre-set shooting positions. If there are unselected shooting positions, it returns to S601 and continues processing. On the other hand, if there are no unselected shooting positions, it proceeds to S604. In S604, the shooting unit 110 transmits the image obtained above to the information processing device 100, and the CPU 101 associates the received image with user information, etc., and stores it as captured image data 413 in the table 410 of the storage device 104.
[0051] Next, referring to Figure 7, the details of the process in S507, which reconstructs the PV based on the captured images obtained in S506, will be explained. First, in S701, the CPU 101 acquires multiple images taken by the imaging unit 110, which are stored in the storage device 104 as captured image data 413. In the following S702, the CPU 101 generates 3D model data of the model, which will become the main data, from the multiple images of the acquired model taken from different angles.
[0052] In the method for generating the 3D model data, for example, a 3D model corresponding to the shape of the model can be determined using a known viewing volume cross-section method, with the model image obtained from multiple images. In this viewing volume cross-section method, multiple model images are used to create a cone (viewing volume) with the camera position as the vertex and the silhouette as the cross-section. The silhouette is then back-projected into 3D space, and the 3D shape is reconstructed as the intersection (common part) of these cones.
[0053] In S703, CPU101 applies pre-set regulations regarding the position and size of the model and performs resizing of the data generated in S702. This resizing process adjusts the size of the model so that it is not cut off if the model exceeds the field of view of the planned PV. The size of the model is predetermined for each PV, with length, width, height, and depth. If any of these exceeds the specified size, it is resized to fit within the specified size. Note that for PVs for the first shape, the resizing process is performed on the model of the first shape, but not on the model of the second shape. Similarly, for PVs for the second shape, the resizing process is performed on the model of the second shape, but not on the model of the first shape.
[0054] In the subsequent S704, CPU101 adjusts the orientation and position of the model that has been resized according to the regulations. This process may be performed before S703. The three-dimensional model contains AR markers, and the orientation and position of the front of the three-dimensional model are determined according to the orientation of these AR markers. The AR markers indicate how much the three-dimensional model is tilted or shifted in the XYZ three-dimensional coordinate system, so the tilt and shift are reset and the three-dimensional coordinate system is readjusted. Alternatively, correction values in the original three-dimensional coordinate system may be calculated, and then these correction values may be used to adjust the orientation and position of the three-dimensional model.
[0055] Since the temporary data included in the PV template has already been corrected for direction and position, correcting the main data to be generated will prevent any shifts in the direction or position of the model in the PV, even when replacing the temporary data. After processing in S704, AR markers can be removed from the main data.
[0056] Furthermore, in S704, the origin of this data is determined. In the case of a model in an upright posture, it will be displayed placed on the ground or other surface, so the origin is set at the bottom of this data, and the direction and position of this data are adjusted accordingly. In the case of an upright posture, the soles of the feet become the origin, so any model can be placed on the ground. Also, in the case of a flying posture rather than an upright posture, the origin may be set at the center of gravity of the model instead of the bottom. Note that if the model has weapons or other equipment, the center of gravity may shift due to their influence, so when setting the center of gravity, it may be better to define the center of gravity of the model body only, excluding weapons and equipment. The origin set in this data serves as the basis for determining the placement of this data when replacing temporary data in the PV template.
[0057] In the subsequent S705 step, the main data generated in the above steps is replaced with the temporary data included in the PV template and output as PV video data. The PV template sets the movement of the background and objects, the camera angle, as well as the movement and camera angle (viewpoint position) of the temporary data, and the template data can be converted into viewable video data and output. Here, by replacing the temporary data with the main data generated in S702 and converting it into video data, it becomes possible to reconstruct PV data using any main data. The PV template can be generated using, for example, UNREAL ENGINE 5 (registered trademark).
[0058] More specifically, the PV template contains information such as the position of objects and models in 3D space, the starting position of the objects and models and their movement in 3D space (position changes over time), camera angle information, and lighting and other effect information, all aligned with the PV's timeline. In this template data, temporary data can be replaced with the actual data generated by S702, and the objects and models can be made to move and operate in 3D space. By extracting these as 2D images on a frame-by-frame basis and compressing them using any video encoding scheme such as H264, video data can be generated as PV data.
[0059] Alternatively, the camera angle of the 3D model data may be specified for each frame that makes up the PV data. Textures of the 3D model corresponding to each camera angle may be generated, and the PV may be reconstructed by replacing the temporary data textures on a frame-by-frame basis. In the subsequent S706, the CPU 101 saves the PV data reconstructed in S705 to the storage device 104.
[0060] The above description explains the case where the imaging unit 110 transmits multiple images to the information processing device 100 to generate this data. However, the process of generating this data from multiple images may utilize an external server or a cloud server. In that case, the information processing device 100 can transmit the multiple images to the external server and obtain this data as a processing result. Similarly, the PV reconstruction process in S705 may also utilize an external server or a cloud server.
[0061] Figure 8 illustrates the lighting when photographing the model 200 with the camera unit 110. Figure 8(A) shows an example of lighting arrangement viewed from above the model 200, with a ring light 304 positioned in front of the model 200 being photographed. In this case, if the model 200 is illuminated by the ring light 304 alone, the light L0 emitted from the ring light 304 alone will cast a shadow on the back of the model 200. Therefore, auxiliary lights 801 and 802 are placed to the side of the ring light 304, and the light L1 and L2 emitted from them are used to prevent shadows from being included in the shooting range.
[0062] Next, Figure 8(B) shows an example of lighting arrangement when viewed from the side of the model 200. The top light 803 is positioned above the ring light 304 in front of the subject 200, and the auxiliary light 804 is positioned below it. This arrangement uses the top light 803 to illuminate the head of the model 200, enabling line-based image acquisition of the head, and also to cancel out the shadows cast by the illumination light L3 of the top light 803 with the illumination light L4 of the auxiliary light 804. The illumination light L3 of the top light 803 is also canceled out by the illumination lights L1 and L2 of the auxiliary lights 801 and 802 mentioned above.
[0063] The auxiliary lights 801, 802, 804 and the top light 803 shown in Figures 8(A) and 8(B) do not all need to be used; at least one of them can be used to counteract the shadow cast by the ring light 304.
[0064] Figure 9 shows a comparison of the output results when a PV template conversion is performed using temporary data versus when it is performed using the actual data. In Figure 9(A), the output result when the PV template was converted using temporary data 901 is shown for scene 900. In contrast, the output result when the PV template was converted using the actual data 902 is shown for the reconstructed scene 900. The actual data 902 is placed in the same position as the temporary data 901, indicating that a replacement has occurred. Figure 9(B) shows a different scene 910 in the same PV, where the temporary data 911 has been replaced with the actual data 912.
[0065] In this way, in this embodiment, by using the PV template generated for the temporary data, it becomes possible to efficiently reconstruct similar PV data simply by replacing the temporary data with the actual data.
[0066] In the above description of the embodiment, we described a case in which a promotional video is generated that displays three-dimensional model data, obtained from images taken of a model, in a virtual space. However, the embodiment is not limited to the generation of PV data; the three-dimensional model data can also be used as game data for virtual space. Furthermore, it can be broadly applied to processes that make characters move in virtual space. For example, the character can appear in events, concerts, sports, online meetings, etc., held in virtual space. Moreover, the technology of this embodiment can also be applied to video technologies that merge the real world and the virtual world, such as cross-reality (XR), making it possible to perceive things that do not exist in the real world.
[0067] In this embodiment, three-dimensional model data can be generated from images of the model's appearance and then featured in the user's preferred promotional video. Some models, such as assembly-type plastic models, are painted or otherwise finished by the user to create their own unique works. By reflecting these individual characteristics in the video or character representation in the virtual space, the appeal can be significantly enhanced.
[0068] <Summary of Embodiments> The above embodiments disclose at least the following image processing method, information processing device, and computer program.
[0069] (1) An image processing method for reconstructing a moving image, The process involves a processing unit capturing images of the model's exterior from multiple directions using a photography unit, and from these images, generating a first three-dimensional model data of the model included in the images. The processing unit performs the step of reconstructing the video by replacing the first three-dimensional model data with the second three-dimensional model data in a video template data that has been pre-configured to include the second three-dimensional model data. The output unit outputs information indicating the source from which the reconstructed video image was acquired. Image processing methods, including those mentioned above.
[0070] (2) The plurality of images include markers for identifying the orientation and position of the model, The image processing method according to (1), wherein the first three-dimensional model data is generated by correcting its orientation and position based on the marker.
[0071] (3) Multiple template data sets are provided for the video template data, depending on the type of model. The image processing method according to (1) or (2), wherein the reconstruction step involves reconstructing the video image using template data selected by the user.
[0072] (4) The image processing method according to (3), wherein the types of models include a first type of model having an external shape that is a scaled-down version of a life-size figure, and a second type of model having an external shape that is a deformed version of a life-size figure.
[0073] (5) The image processing method according to (3) or (4), wherein in the generation step, the first three-dimensional model data is generated by adjusting its size to match the content of the moving image.
[0074] (6) The image processing method according to (5), wherein the size adjustment is performed on the first three-dimensional model data of the same type as the type of model in the template data of the moving image, but not on the first three-dimensional model data of a different type than the type of model in the template data of the moving image.
[0075] (7) The template data of the video is associated with a recommended pose for the model, The image processing method according to any one of (1) to (6), wherein the plurality of images are generated by photographing the appearance of the model in the recommended pose.
[0076] (8) The image processing method according to any one of (1) to (7), further comprising the step of the processing unit receiving the selection of template data for the moving image.
[0077] (9) The image processing method according to (8), further comprising the step of causing the processing unit to display a recommended pose for the video on a display device in response to receiving the selection of template data for the video.
[0078] (10) The image processing method according to any one of (1) to (9), wherein the multiple directions are determined to cover the entire circumference of the model.
[0079] (11) The image processing method according to any one of (1) to (10), wherein the plurality of images are generated by the imaging unit imaging the model illuminated by a first light source and at least one second light source that cancels out shadows caused by the light emitted from the first light source.
[0080] (12) The image processing method according to any one of (1) to (11), further comprising the step of the processing unit receiving the plurality of images from the imaging unit.
[0081] (13) The image processing method according to any one of (1) to (12), wherein the template data has information on at least one of the position, movement, camera angle, and lighting of the second three-dimensional model data in three-dimensional space set in accordance with the time axis of the video.
[0082] (14) A generation means for generating first three-dimensional model data of the model contained in the images from multiple images generated by taking images of the model's exterior from multiple directions, A reconstruction means for reconstructing a video by replacing the first three-dimensional model data with the second three-dimensional model data in a video template data that has been pre-configured to include second three-dimensional model data, A control means that controls the output unit to output information indicating the source of the reconstructed video image. An information processing device equipped with the following features.
[0083] (15) A computer program that causes a computer to perform any one of the image processing methods described in (1) through (13).
[0084] The invention is not limited to the embodiments described above, and various modifications and changes are possible within the scope of the gist of the invention.
Claims
1. An image processing method for reconstructing moving images, The process involves a processing unit capturing images of the model's exterior from multiple directions, generating multiple images, and then generating a first three-dimensional model data of the model included in those images. The processing unit performs the step of reconstructing the video by replacing the first three-dimensional model data with the second three-dimensional model data in a video template data that has been pre-configured to include the second three-dimensional model data. The processing unit includes the step of receiving the selection of template data for the video, In response to receiving the selection of template data for the video, the processing unit performs the step of displaying recommended poses for the video on the display device, Image processing methods, including those mentioned above.
2. The aforementioned multiple images include markers for identifying the orientation and position of the model, The image processing method according to claim 1, wherein the first three-dimensional model data is generated by correcting its orientation and position based on the marker.
3. The template data for the aforementioned video is provided in multiple templates depending on the type of model. The image processing method according to claim 1, wherein in the reconstruction step, the video is reconstructed using template data selected by the user.
4. The image processing method according to claim 3, wherein the types of models include a first type of model having an external shape that is a scaled-down version of a life-size figure, and a second type of model having an external shape that is a stylized version of a life-size figure.
5. The image processing method according to claim 3, wherein in the generation step, the first three-dimensional model data is generated by adjusting its size to a size corresponding to the content of the moving image.
6. The image processing method according to claim 5, wherein the size adjustment is performed on the first three-dimensional model data of the same type as the type of model in the template data of the moving image, but not on the first three-dimensional model data of a different type than the type of model in the template data of the moving image.
7. The template data for the aforementioned video is associated with a recommended pose for the aforementioned model. The image processing method according to claim 1, wherein the plurality of images are generated by photographing the appearance of the model in the recommended pose.
8. The image processing method according to claim 1, further comprising the step of receiving a selection of template data for the moving image.
9. The image processing method according to claim 8, further comprising the step of causing the processing unit to display a recommended pose for the video on a display device in response to receiving the selection of template data for the video.
10. The image processing method according to claim 1, wherein the multiple directions are determined to cover the entire circumference of the model.
11. The image processing method according to claim 1, wherein the plurality of images are generated by the imaging unit imaging the model illuminated by a first light source and at least one second light source that cancels out shadows caused by the light emitted from the first light source.
12. The image processing method according to claim 1, further comprising the step of the processing unit receiving the plurality of images from the imaging unit.
13. The image processing method according to claim 1, wherein the template data includes information on at least one of the position, movement, camera angle, and lighting of the second three-dimensional model data in three-dimensional space, set in accordance with the time axis of the video.
14. A generation means for generating first three-dimensional model data of the model contained in the images from multiple images generated by photographing the exterior of the model from multiple directions, A reconstruction means for reconstructing a video by replacing the first three-dimensional model data with the second three-dimensional model data in a video template data that has been pre-configured to include a second three-dimensional model data, A receiving means for receiving the selection of template data for the aforementioned video image, A display means that, upon receiving the selection of template data for the aforementioned video image, displays a recommended pose for the video image, An information processing device equipped with the following features.
15. A computer program for causing a computer to perform the image processing method described in any one of claims 1 to 13.
Citation Information
Patent Citations
Animation generation system, animation generation device, animation generation method, program, and storage medium
JP2006244306A
Computer program and portable terminal device
JP2019179481A
Image generation system and program
JP2020107251A
Image generation system and program
JP2020107252A
Synthetic dynamic image generation apparatus
JP2020137012A