Projection method, projection system, and program
By using positional and shape information to correct image data and employing AI for generating content images synchronized with music effects, the method enhances the artistic quality of projection mapping.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2026-04-02
AI Technical Summary
Existing projection technologies, such as those described in Patent Document 1, do not adequately meet the demand for enhanced artistic quality in projection mapping.
A method and system that utilize positional and shape information to correct image data, employing an AI model to generate content images optimized for projection targets, and apply effects synchronized with music, enhancing artistic quality.
Enables the projection of highly artistic content images tailored to the projection target's shape and position, improving the quality of projection mapping.
Smart Images

Figure 2026056918000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a projection method, a projection system, and a program.
Background Art
[0002] Patent Document 1 discloses an image generation method for generating a projection image projected from a projector to a projection target. In the image generation method disclosed in Patent Document 1, an information processing device that supplies projection image data representing the projection image to the projector causes the projector to project a pattern image for three-dimensional measurement onto the projection target. The information processing device generates projection image data representing the projection image by correcting a material image according to the three-dimensional shape measured based on an imaging image obtained by imaging the projection target onto which the pattern image for three-dimensional measurement is projected, using a camera from a determined position.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The material image in Patent Document 1 is prepared in advance. In recent years, in projection mapping, there is a demand to further enhance the artistic quality, but the technology disclosed in Patent Document 1 does not sufficiently meet this requirement.
Means for Solving the Problems
[0005] One aspect of the projection method of this disclosure includes generating correction data in which image data is corrected based on positional information indicating the positional relationship between a projector and a projection target and shape information indicating the shape of the projection target; generating content image data corresponding to the content image using a first image generation AI model into which specification information specifying the form of the content image and the correction data are input; and projecting the content image based on the content image data, which is the image data, from the projector onto the projection target.
[0006] One aspect of a projection system in this disclosure includes a projector and a processing device for controlling the projector, the processing device performing the following: generating correction data in which image data is corrected based on positional information indicating the positional relationship between the projector and the projection target and shape information indicating the shape of the projection target; generating content image data corresponding to the content image using a first image generation AI model into which specification information specifying the aspect of the content image and the correction data are input; and causing the projector to project the content image based on the content image data onto the projection target.
[0007] One aspect of the program of this disclosure involves a computer performing the following actions: generating correction data in which image data is corrected based on positional information indicating the positional relationship between a projector and the projection target and shape information indicating the shape of the projection target; generating content image data corresponding to the content image using a first image generation AI model into which specification information specifying the form of the content image and the correction data are input; and causing the projector to project the content image based on the content image data onto the projection target. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an example configuration of projection system 1 according to one embodiment of the present disclosure. [Figure 2] This figure shows an example of a projection target SC onto which content images are projected from projector 10. [Figure 3] This is a block diagram showing an example configuration of the information processing device 40 included in the projection system 1. [Figure 4] This figure shows an example of a base image. [Figure 5] This figure shows an example of a base image for generation. [Figure 6] This figure shows an example of a content image and individual objects extracted from that content image. [Figure 7] This figure shows an example of Effect B1. [Figure 8] This figure shows an example of Effect B2. [Figure 9] This figure shows an example of Effect B4. [Figure 10] This is a flowchart showing the processing flow in the projection method executed by the processing unit 410 of the information processing device 40. [Figure 11] This is a flowchart showing the processing flow in the content image generation process SA120. [Modes for carrying out the invention]
[0009] The embodiments described below are subject to various technically preferred limitations. However, the embodiments of this disclosure are not limited to those described below.
[0010] A. Embodiment Figure 1 shows an example configuration of a projection system 1 according to one embodiment of the present disclosure. The projection system 1 includes a projector 10 that projects an image onto a projection target SC, a camera 20, a speaker 30, and an information processing device 40. As shown in Figure 1, the projector 10, camera 20, and speaker 30 are each connected to the information processing device 40 by wire or wireless. The information processing device 40 is also connected by wire or wireless to a communication network NW, such as the Internet. An image generation server 50 and a segmentation server 60 are connected to the communication network NW. The information processing device 40 communicates with the image generation server 50 and the segmentation server 60, respectively, via the communication network NW.
[0011] Figure 2 shows an example of a projection target SC in this embodiment. In this embodiment, the projection target SC is, for example, the wall surface of a room and has irregularities. The projection target SC is formed in a circular shape and has a portion A1 that protrudes forward, and portions A2 to A5 located around portion A1, each formed in an arc shape and recessed toward the back. In Figure 2, the portion A1 that protrudes forward is given vertical hatching, and each of the recessed portions A2 to A5 is given horizontal hatching. The projector 10, camera 20, speaker 30, and information processing device 40 are installed in a room where the wall surface is the projection target SC. The projection target SC may be a flat surface or a three-dimensional object.
[0012] The projector 10 includes a generation unit that generates image light corresponding to externally supplied image data, and an optical system that guides the image light generated by the generation unit to the projection target SC. In Figure 1, the generation unit and optical system are not shown. Specific examples of the generation unit include a display panel that includes an optical modulation element such as an LCD (Liquid Crystal Display), LCOS (Liquid Crystal on Silicon), or DMD (Digital Mirror Device). The optical system includes, for example, a projection lens and a spectral prism.
[0013] The camera 20 includes a CMOS (Complementary Metal-Oxide-Semiconductor) or a CCD (Charge Coupled Device) image sensor. In this embodiment, the imaging area of the camera 20 is preset so as to include the entire projection target SC. The camera 20 images the imaging area under the control of the information processing device 40. The camera 20 outputs image data representing the captured image to the information processing device 40. The speaker 30 is supplied with an audio signal from the information processing device 40. The speaker 30 outputs a sound represented by a sound waveform according to the audio signal supplied from the information processing device 40.
[0014] The image generation server 50 has an image generation AI model. The image generation AI model in this embodiment receives an input of first data which is image data representing an original image and second data which represents a motif or the like of an image to be generated, and generates output image data representing a new image by changing the image represented by the first data based on the second data. The image generation server 50 receives an input of the first data and the second data via the communication network NW, and returns the generated output image data to the transmission sources of the first data and the second data. The image generation AI model possessed by the image generation server 50 is an example of the first image generation AI model in the present disclosure.
[0015] The segmentation server 60 is a device that executes segmentation processing. The segmentation processing means dividing an image into a plurality of regions and grouping a plurality of regions having the same meaning of the image. The segmentation server 60 performs segmentation processing on the image represented by the image data received from the information processing device 40 via the communication network NW, and returns image data representing the image of each region classified, that is, loop-divided, to the information processing device 40.
[0016] The information processing device 40 is, for example, a personal computer. FIG. 3 is a block diagram showing a configuration example of the information processing device 40. As shown in FIG. 3, the information processing device 40 includes a processing device 410, an external device interface (hereinafter referred to as IF) 420, a communication device 430, a storage device 440, and a bus 450 that mediates data transfer between these components. In addition to the processing device 410, the external device IF 420, the communication device 430, the storage device 440, and the bus 450, the information processing device 40 includes a display device that displays various images under the control of the processing device 410, and an input device that receives operations by the user and provides operation content data representing the content of the operation to the processing device 410. Since the display device and the input device have little relevance to the present disclosure, they are not shown in FIG. 3.
[0017] The processing device 410 includes one or more processors. The processor is, for example, a computer such as a CPU (Central Processing Unit). Although details will be described later, the processing device 410 functions as a control center of the information processing device 40 by operating according to a program PR1 stored in advance in the storage device 440.
[0018] The external device IF 420 includes an interface circuit for connecting other devices by wire or wirelessly. A specific example of the external device IF 420 is a USB interface. The projector 10, the camera 20, and the speaker 30 are connected to the external device IF 420 by wire or wirelessly, respectively. The external device IF 420 receives the image data transmitted from the camera 20 by communicating with the camera 20. The external device IF 420 delivers the image data received from the camera 20 to the processing device 410. Further, the external device IF 420 transmits the image data provided from the processing device 410 to the projector 10 by communicating with the projector 10. Further, the external device IF 420 converts the sound data (sample sequence obtained by sampling the sound waveform) provided from the processing device 410 into a sound signal, and outputs the sound signal obtained by this conversion to the speaker 30.
[0019] The communication device 430 is connected to the communication network NW by wire or wireless connection. The communication device 430 is a communication interface circuit that transmits and receives data with other devices connected to the communication network NW. Other devices that transmit and receive data with the communication device 430 via the communication network NW include the aforementioned image generation server 50 and segmentation server 60.
[0020] The storage device 440 includes non-volatile memory such as flash ROM (Read Only Memory) and volatile memory such as RAM (Random Access Memory). The non-volatile memory of the storage device 440 stores program PR1, which causes the processing unit 410 to function as the control center of the information processing unit 40, and music data D1, which is sampling data for music to be played in synchronization with the projection of projected content from the projector 10 to the projection target SC. The non-volatile memory of the storage device 440 is used by the processing unit 410 as a work area when executing program PR1.
[0021] When the power to the information processing device 40 (not shown in Figure 3) is turned on and the user instructs the execution of program PR1 through an operation on the input device, the processing device 410 reads program PR1 from non-volatile memory to volatile memory. Then, the processing device 410 starts executing program PR1 that has been read into volatile memory. The processing device 410, operating according to program PR1, functions as an acquisition unit 410a, a first generation unit 410b, a second generation unit 410c, and a projection control unit 410d. In other words, each of the acquisition unit 410a, first generation unit 410b, second generation unit 410c, and projection control unit 410d shown in Figure 2 is a software module realized by operating the computer according to program PR1. The roles of each of the acquisition unit 410a, first generation unit 410b, second generation unit 410c, and projection control unit 410d are as follows.
[0022] The acquisition unit 410a acquires positional information indicating the positional relationship between the projector 10 and the projection target SC, and shape information indicating the three-dimensional shape of the projection target SC. In this embodiment, the acquisition unit 410a controls the projector 10 to sequentially project multiple different measurement patterns (for example, grayscale periodic patterns) onto the projection target SC. The acquisition unit 410a also causes the camera 20 to capture an image of the projection target SC with the measurement pattern projected onto it for each measurement pattern, and acquires image data representing each captured image from the camera 20. The acquisition unit 410a then analyzes the multiple image data acquired from the camera 20 to calculate shape information indicating the three-dimensional shape of the projection target SC and positional information indicating the relative position of the projection target SC and the projector 10. The positional information includes, for example, at least one of the rotation matrix or position vector of the projection target SC with respect to the projector 10 (projection lens). The acquisition unit 410a may acquire internal parameters of the projector 10, internal parameters of the camera 20, and external parameters indicating the positional relationship between the camera 20 and the projector 10. The information indicating the positional relationship between the camera 20 and the projector 10 includes, for example, at least one of the rotation matrix or position vector of the projector 10 relative to the camera 20. For the calculation of shape information and position information based on the captured image of the projection target SC onto which multiple measurement patterns are sequentially projected, well-known 3D measurement techniques may be used as appropriate. The acquisition unit 410a may acquire shape information and position information not by 3D measurement using multiple measurement patterns, but by acquiring measurement data measured in advance by the user. For example, the shape information may be 3D model data of the projection target SC, and the position information may be data indicating the measured values of the inclination angle and distance between the projector 10 and the projection target SC in real space.
[0023] The first generation unit 410b generates base image data representing the base image based on the shape information acquired by the acquisition unit 410a and the position information also acquired by the acquisition unit 410a. The base image is a projected image (projected image data) in which the shape of the 2D image of the projection target SC has been corrected to match the 3D shape indicated by the shape information, assuming that the image is projected from the projector 10 at the position indicated by the position information with respect to the position of the projection target SC. The image data of the projection target SC before the shape of the 2D image of the projection target SC is an example of image data in this disclosure. The base image data is an example of the concept of correction data in this disclosure. The base image, i.e., the 2D image of the projection target SC, is an image obtained by converting the 3D shape of the projection target SC to the 2D shape of the viewpoint of the projector 10. In other words, the base image is a 2D image of the projection target SC obtained by converting the 2D image of the projection target SC as seen from, for example, the camera 20, to the viewpoint of the projector 10 based on the position information and shape information. The viewpoint of the projector 10 is, for example, the viewpoint viewed from the principal point of the projection lens. By converting to the viewpoint of the projector 10, when actually projecting an image from the projector 10, the image can be made to correspond to the projection target SC, which has a nearly three-dimensional shape. Note that the base image may not be a two-dimensional image of the projection target SC, but rather an image in which the shape of a separately prepared two-dimensional image has been corrected. If the base image is a separately prepared image, it is preferable that the image is designed to conform to the three-dimensional shape of the projection target SC. If the base image is an image in which the shape of a separately prepared two-dimensional image has been corrected, the separately prepared two-dimensional image is another example of the image data in this disclosure. Figure 4 is a diagram showing an example of the base image G1. As shown in Figure 4, the base image G1 has a region GA1 corresponding to part A1, and regions GA2 to GA5 corresponding one-to-one to parts A2 to A5. In Figure 4, as in Figure 2, regions that appear to protrude towards the foreground are given vertical hatching, and regions that appear to be recessed towards the background are given horizontal hatching. As is clear from comparing Figure 4 and Figure 2, the contours of the base image and the projected SC are reversed.
[0024] Next, the first generation unit 410b generates a plurality of generation base image data that correspond one-to-one with each of the plurality of generation base images based on the base image data. Each of the plurality of generation base images is a modulated image in which at least one parameter defining the aspect of the base image is different from the base image. The plurality of generation base images are another example of the concept of correction data in this disclosure, and further an example of the plurality of modulated image data in this disclosure. The first generation unit 410b generates a single generation base image by transmitting the base image data to the segmentation server 60, dividing the base image into a plurality of regions, and placing a grayscale pattern in each region. This grayscale pattern represents the gradation in the base image and represents the sense of depth in the base image. In this embodiment, at least one parameter is gradation. Note that at least one parameter may be a parameter that defines the aspect of the base image, such as the edge shape of the base image. The first generation unit 410b then generates a plurality of generation base images by appropriately adjusting the grayscale placed in each region for each region. Figure 5 is a diagram showing an example of a plurality of generation base images in this embodiment. Multiple base images for generation may be generated using the base image directly, or they may be generated using a shaped base image in which the edges representing the outline of the base image have been shaped using Bézier curves or the like. In Figure 5, grayscale is represented by black circles, and the larger the black circle, the lower the gradation and the more recessed the area is in the background. By adjusting the grayscale placed in each area, a base image for generation may be obtained in which the sense of depth does not necessarily match the shape of the projection target SC, but this discrepancy can be an interesting feature of the content image.
[0025] The second generation unit 410c generates multiple content image data that correspond one-to-one with each of the multiple content images that are sequentially projected onto the projection target SC, based on each of the multiple base images for generation. More specifically, the second generation unit 410c first displays a UI screen on the display device prompting the user to input text data that specifies the characteristics of the content image in words (natural language or prompts), and accepts the input of such text data through operation of an input device. The user inputs words that constitute the content image they imagine. These words may specify a concrete object such as a dog or a cat, or they may abstractly specify the composition of the content image such as color or style. If the words are the same, slightly different content images will be obtained within the range of expression of their meaning. The text data is an example of the designation information in this disclosure. Note that the designation information may be audio data or image data instead of text data. Also, the designation information may be a combination of at least two of the text data, audio data, and image data. Furthermore, the second generation unit 410c may accept input of multiple text data, not just one.
[0026] The second generation unit 410c sends the above text data as the first data to the image generation server 50, and also sends one of the multiple base image data for generation to the image generation server 50. This process is repeated N times (where N is a predetermined integer of 2 or more). The second generation unit 410c then generates multiple content image data by acquiring each of the N output image data returned from the image generation server 50 as content image data. In other words, the multiple content image data are generated by the image generation AI model of the image generation server 50.
[0027] Furthermore, the second generation unit 410c generates interpolated content image data corresponding to an intermediate image, which is an image between two content images arranged one after the other in a row of multiple content images, using an image generation AI similar to the content image data. Of the two content images arranged one after the other in a row of multiple content images, the content image data representing the content image projected first is an example of the first content image data in this disclosure, and the content image data representing the content image projected subsequently is an example of the second content image in this disclosure. The interpolated content image data is intended to reduce discontinuities when the multiple content images are played sequentially as a video.
[0028] Furthermore, the second generation unit 410c transmits image data representing each of the multiple content images and the interpolated content images to the segmentation server 60, recognizes semantically identical regions within each image, and writes the images of each region as files to the storage device 440. For example, suppose that text data containing phrases such as "clear skies," "park," "bicycle," "lunchbox," and "cat" is input as the aforementioned text data, and image G2 shown in Figure 6 is generated as the content image. As shown in Figure 6, image G2 is an image obtained by placing individual objects OB1 to B5 in each of the regions GA1 to GA5 in the base image for generation. Individual object OB1 is an image of "cat." Individual object OB2 is an image of "clear skies." Individual object OB3 is an image of "lunchbox." Individual object OB4 is an image of "bicycle." Individual object OB5 is an image of "park." The second generation unit 410c recognizes individual objects contained in image G2 using segmentation processing and writes individual files F1 to F5 (see Figure 6) corresponding to each of the individual objects OB1 to B5 to the storage device 440. The reason for decomposing and storing multiple content images and interpolated content images into individual objects is to reduce the processing load when adding effects, which will be described later.
[0029] When an operation to instruct the input device to start projecting the content image is performed, the projection control unit 410d reads the aforementioned individual files from the storage device 440, edits the image data based on these individual files, and supplies it to the projector 10, thereby projecting the content image onto the projection target SC.
[0030] Furthermore, when an operation to instruct the input device to start projecting the content image is performed, the projection control unit 410d reads the music data D1 from the storage device 440 and supplies an audio signal corresponding to the music data to the speaker 30 via the external device IF unit 420. At this time, the projection control unit 410d detects in real time the changes in speed and sound pressure of the music played according to the music data D1 and applies either effect A or effect B, or a combination thereof, to the content image projected onto the projection target SC. Effect A is an effect applied to the content image according to the speed of the music played according to the music data D1, and effect B is an effect applied according to changes in sound pressure in the music played according to the music data D1 (changes in sound pressure such as the sound of percussion or lead guitar, or the vocalization of a singer).
[0031] (A) Effect A A concrete example of effect A is the application of a fisheye lens effect. The fisheye lens effect is achieved by image processing that gives the entire projected image a three-dimensional curvature centered on the image's center. This curvature is expressed as a function of the elapsed time from the start of applying the fisheye lens effect. The period of the curvature change may match the speed (BPM: Beats Per Minute) of the music being played at the time the fisheye lens effect is applied, or it may be N times or 1 / N times the BPM. N is an integer greater than or equal to 2.
[0032] (B) Effect B A concrete example of Effect B is the enlargement of an individual object (hereinafter referred to as Effect B1). Enlarging an individual object means increasing the size of the individual object in accordance with the elapsed time from the start of applying the effect, while maintaining the center coordinates of the individual object relative to the content image, as shown in Figure 7. Figure 7 shows an example of applying Effect B1 to individual object OB2. Note that when an individual object is enlarged, it overlaps with other adjacent individual objects, so in order to avoid other individual objects becoming invisible, a process may be included to increase the transparency of the individual object as time passes in accordance with the enlargement of the individual object. In Figure 7, the increased transparency is shown by drawing individual object OB1 with a dotted line.
[0033] Another specific example of Effect B is the movement of individual objects (hereinafter referred to as Effect B2). As shown in Figure 8, the movement of individual objects refers to the movement of an object from the outer edge of the content image toward the drawing position of the individual object in the content image. Figure 8 shows an example of applying Effect B2 to individual object OB3. More specifically, the trajectory of the movement of the individual object is defined as a straight line passing through the center coordinates of the content image and the coordinates of the drawing position of the individual object, and the starting position of the movement is determined by the radius of a circle centered on the center coordinates of the content image. The movement speed is determined by the BPM of the music, and the movement speed should be determined so that the movement of the individual object is completed in one beat. Note that while an individual object is moving, it overlaps with other individual objects, so in order to avoid other individual objects becoming invisible, the individual object to be moved may be made transparent at the start of the movement, and the transparency of the individual object may be reduced as time passes.
[0034] Another specific example of Effect B is the hue change of individual objects (hereinafter referred to as Effect B3). Specifically, this involves sequentially rotating the hue of individual objects up to 360° according to the elapsed time from the moment the effect is applied. For example, the hue of an individual object can be rotated once (360°) in one beat.
[0035] Another specific example of Effect B is the rotation of individual objects (hereinafter referred to as Effect B4). Rotation of individual objects, as shown in Figure 9, means rotating the individual object within the screen plane within a specified angular range, with the center position of the individual object in the content image as the rotation center. Figure 9 shows an example of applying Effect B4 to individual object OB1. Note that rotation may occur in the positive direction (e.g., clockwise) or the negative direction (e.g., counterclockwise) within a single bead, and the number of rotations for each direction may also be indicated.
[0036] Another specific example of Effect B is sequentially switching between multiple created content images on an image-by-image basis (hereinafter referred to as Effect B5).
[0037] Effects B1 to B5 may be pre-associated with parts in the music played according to music data D1. For example, effect B1 may be associated with the percussion part, and effect B1 may be added when a change in the sound pressure of the percussion part is detected. Similarly, effect B2 may be associated with the vocal part, and effect B2 may be added when a change in the sound pressure of the vocal part is detected. The identification of which part of the music played according to music data D1 has experienced a change in sound pressure can be done, for example, based on the frequency distribution of the sound whose sound pressure has changed. In this embodiment, music data D1 was sampling data representing the waveform of a sound, but music data D1 may also be MIDI (Musical Instrument Digital Interface) data describing the sound on (note on) and off (note off) for each part. If music data D1 is MIDI data, it becomes easier to associate effects B1 to B5 with parts. Effects B1 to B5 may also be associated with the structure of the music. For example, effect B1 could be associated with the A section, effect B2 with the B section, and effect B3 with the chorus.
[0038] Furthermore, with respect to Effects B1 to B5, instead of always applying them when a change in sound pressure occurs in the music played according to the music data D1, the execution probability of each effect may be set, and a pseudo-random number greater than or equal to 0 and less than 1 may be used to determine whether or not to execute the effect. Specifically, the projection control unit 410d may generate the above-mentioned pseudo-random number each time it detects a change in sound pressure in the music played according to the music data D1, and then execute the addition of an effect whose execution probability is greater than or equal to the said pseudo-random number. In addition, the number of effects to be executed at once when a change in sound pressure occurs in the music played according to the music data D1 may be specified in advance by the user. In this case, the projection control unit 410d may use a pseudo-random number or the like to select the specified number of effects from Effects B1 to B5 each time it detects a change in sound pressure in the music played according to the music data D1. Furthermore, the number of individual objects to which the effects are applied may be specified in advance by the user.
[0039] Furthermore, the processing unit 410, operating according to program PR1, performs a content image projection method that prominently demonstrates the features of this disclosure. Figure 10 is a flowchart showing the processing flow in this projection method.
[0040] In acquisition process SA110, the processing unit 410, which operates according to program PR1, functions as an acquisition unit 410a. In acquisition process SA110, the processing unit 410 acquires positional information indicating the positional relationship between the projector 10 and the projection target SC, and shape information indicating the three-dimensional shape of the projection target SC. In acquisition process SA110, the processing unit 410 controls the projector 10 to sequentially project multiple different measurement patterns onto the projection target SC. The processing unit 410 also causes the camera 20 to capture an image including the projection target SC in the state where the measurement pattern is being projected, for each measurement pattern, and acquires image data representing each captured image from the camera 20. Then, by analyzing the multiple image data acquired from the camera 20, the processing unit 410 calculates shape information indicating the three-dimensional shape of the projection target SC and positional information indicating the relative position between the projection target SC and the projector 10.
[0041] The content image generation process SA120, which follows the acquisition process SA110, is the process of generating a content image to be projected from the projector 10 onto the projection target SC. Figure 11 is a flowchart showing the processing flow in the content image generation process SA120. As shown in Figure 11, the content image generation process SA120 includes a first generation process SA1210 and a second generation process SA1220.
[0042] In the first generation process SA1210, the processing unit 410, which operates according to program PR1, functions as the first generation unit 410b. In the first generation process SA1210, the processing unit 410 generates base image data representing the base image based on the shape information and position information acquired in the acquisition process SA110. Next, in the first generation process SA1210, the processing unit 410 generates multiple base image data for generation that correspond one-to-one for each of the multiple base images for generation, based on the generated base image data. In this embodiment, the processing unit 410 performs multiple segmentation on the base projection image, places a grayscale pattern in each segment, and generates multiple base images for generation by appropriately adjusting the grayscale placed in each segment for each segment.
[0043] In the second generation process SA1220, which follows the first generation process SA1210, the processing unit 410, operating according to program PR1, functions as the second generation unit 410c. In the second generation process SA1220, the processing unit 410 first generates multiple content image data that correspond one-to-one with each of the multiple content images to be sequentially projected onto the projection target SC, based on each of the multiple base images for generation generated in the first generation process SA1210. More specifically, the processing unit 410 displays a UI screen on the display device prompting the input of text data specifying the form of the content image in words, and accepts the input of this text data by operating the input device. Next, the processing unit 410 sends the above text data as the first data to the image generation server 50, and also sends one of the multiple base image data for generation to the image generation server 50. This process is repeated N times (where N is a predetermined integer of 2 or more). That is, the processing unit 410 sends N base image data for generation to the image generation server 50. The processing unit 410 then generates multiple content image data by acquiring each of the N output image data returned from the image generation server 50 as content image data. The processing unit 410 also generates interpolated content image data corresponding to intermediate images, which are images between two content images that are placed one after the other in a sequence of multiple content images, using the same image generation AI model as the content image data. The processing unit 410 then uses segmentation for each of the multiple content images and the interpolated content image data, and writes the images of each region as files to the storage device 440. The above is the processing flow in the content image generation process SA120.
[0044] Returning to Figure 10, the projection process SA130, which follows the content image generation process SA120, is executed when an operation to instruct the input device to start projecting the content image is performed. In the projection process SA130, the processing unit 410, which operates according to program PR1, functions as the projection control unit 410d. In the projection process SA130, the processing unit 410 reads the aforementioned individual files from the storage device 440, edits the image data based on these individual files, and supplies it to the projector 10 to project the content image onto the projection target SC, and simultaneously executes in parallel the process of reading the music data D1 from the storage device 440 and supplying an audio signal corresponding to the music data to the speaker 30 via the external device IF unit 420. The processing unit 410 detects changes in the speed and sound pressure of the music played according to the music data D1 in real time and applies effect A and one of effects B1 to B5, or a combination thereof, to the content image projected onto the projection target SC.
[0045] As described above, according to this embodiment, content images to be projected onto the projection target SC can be easily created using an image generation AI model. Furthermore, in this embodiment, the data input to the image generation AI model is not the image data itself, but corrected data obtained by correcting the image data based on positional information and shape information. Therefore, according to this embodiment, the content image projected onto the projection target SC is optimized according to the positional relationship between the projector 10 and the projection target SC, and the shape of the projection target SC. According to this embodiment, highly artistic content images suitable for projection mapping can be projected onto the projection target SC.
[0046] B. Transformation The above embodiment can be modified as follows. (1) In the above embodiment, the first data was data that defined the form of the image to be generated in natural language, but it may be any of the following: a photograph, sound data representing music or voice, biometric information representing the user's pulse, pupil, gaze, blood pressure, or heart rate, or any combination of these.
[0047] (2) In the above embodiment, music represented by music data D1 was played in synchronization with the projection of the content image onto the projection target SC, but the playback of the music may be omitted. In the embodiment in which music playback is omitted, various effects may be applied to the content image according to the elapsed time from the start of projection of the content image. Note that the application of various effects to the content image is not mandatory and may be omitted. Furthermore, the generation of the interpolated content image in the above embodiment is not mandatory and may be omitted. Furthermore, the image generation AI model that generates the interpolated content image data corresponding to the interpolated content image may be the same as or different from the image generation AI model that generates the content image data. That is, the image generation AI model that generates the interpolated content image data is an example of the second image generation AI model in this disclosure, and the image generation AI model that generates the content image data is an example of the first image generation AI model in this disclosure. In this case, the image generation server 50 may have the first image generation AI model and the second image generation AI model, or the image generation server 50 may have the first image generation AI model and an image generation server different from the image generation server 50 may have the second image generation AI model.
[0048] (3) Either the camera 20 or the speaker 30, or both, may be included in the information processing device 40. The camera 20 may also be included in the projector 10. The image generation AI model may also be stored in the storage device 440 of the information processing device 40. In the configuration in which the image generation AI model is stored in the storage device 440 of the information processing device 40, the information processing device 40 may also serve as the image generation server 50. The information processing device 40 may also serve as the segmentation server.
[0049] (4) In the above embodiment, the acquisition unit 410a, the first generation unit 410b, the second generation unit 410c, and the projection control unit 410d were software modules. However, at least one of the acquisition unit 410a, the first generation unit 410b, the second generation unit 410c, and the projection control unit 410d may be a hardware module such as an ASIC (Application Specific Integrated Circuit). Even if at least one of the acquisition unit 410a, the first generation unit 410b, the second generation unit 410c, and the projection control unit 410d is a hardware module, the same effects as in the above embodiment will be achieved.
[0050] (5) Program PR1 may be manufactured as a standalone product and may be provided for a fee or free of charge. Specific methods of providing Program PR1 include providing it by writing it to a computer-readable recording medium such as flash ROM, or providing it by downloading it via a telecommunications line such as the Internet. The information processing device 40 may also function as an ASP (Application Service Provider) server that provides a service of generating and returning content image data when it receives location information, shape information, and text data via a communication network NW.
[0051] C. Summary of this disclosure This disclosure is not limited to the embodiments and modifications described above, and can be implemented in various forms without departing from its spirit. For example, this disclosure can also be implemented in the following forms. The technical features in the embodiments described above that correspond to the technical features in each of the forms described below can be replaced or combined as appropriate in order to solve some or all of the problems of this disclosure or to achieve some or all of the effects of this disclosure. Furthermore, if such technical features are not described as essential in this specification, they can be deleted as appropriate. A summary of this disclosure is provided below.
[0052] (Note 1) One aspect of the projection method of this disclosure includes: generating correction data in which image data is corrected based on positional information indicating the positional relationship between a projector and a projection target and shape information indicating the shape of the projection target; generating content image data corresponding to the content image using a first image generation AI model into which specification information specifying the form of the content image and the correction data are input; and projecting the content image based on the content image data onto the projection target from the projector. According to this aspect, the data input to the first image generation AI model is not the image data itself, but correction data in which the image data is corrected based on the positional information and shape information, so the content image projected onto the projection target is optimized based on the positional information and shape information. Accordingly, according to this aspect, projection of highly artistic content images that are more suitable for projection mapping and are generated using the first image generation AI model is realized.
[0053] (Note 2) A more preferred embodiment of the projection method is the projection method described in (Note 1), wherein the correction data is a plurality of modulated image data in which the image data is corrected based on the position information and the shape information, and at least one parameter defining the form of the corrected image shown by the correction data is different from each other, the generation of the content image data includes generating the content image data for each of the plurality of modulated image data by the first image generation AI model into which the designation information and the plurality of modulated image data are input, and the projection of the content image from the projector to the projection target includes sequentially projecting a plurality of content images based on a plurality of content image data that correspond one-to-one with the plurality of modulated image data. According to this embodiment, by generating content image data for each of the plurality of modulated image data and sequentially projecting a plurality of content images based on the plurality of content images, a projection method of content images more suitable for projection mapping can be provided.
[0054] (Note 3) A more preferred embodiment of the projection method is the projection method described in (Appendix 2), wherein at least one parameter includes the gradation of the corrected image indicated by the correction data.
[0055] (Note 4) Another, more preferred embodiment of the projection method is the projection method described in (Note 2) or (Note 3), wherein the plurality of content image data includes a first content image data and a second content image data, and generating the content image data includes generating interpolated content image data to be interpolated between the first content image data and the second content image data by a first image generation AI model or a second image generation AI model different from the first image generation AI model, based on the first content image data and the second content image data, and projecting the content image from the projector onto the projection target includes sequentially projecting the content image based on the first content image data, the content image based on the interpolated content image data, and the content image based on the second content image data. According to this embodiment, by projecting the interpolated content image data, a method of projecting content images that is less unnatural for the user can be provided.
[0056] (Note 5) Another, more preferred embodiment of the projection method is the projection method according to (Note 1), (Note 2), (Note 3), or (Note 4), wherein generating the content image data includes classifying the content image into multiple regions by applying a segmentation process to the content image, and generating image data corresponding to each of the multiple regions. According to this embodiment, since image data corresponding to each of the multiple regions obtained by applying segmentation to the content image is generated, the processing load when applying effects to each region is reduced.
[0057] (Note 6) Another, more preferred embodiment of the projection method is the projection method described in (Note 1), (Note 2), (Note 3), (Note 4), or (Note 5), which includes playing music in synchronization with the projection of the content image, and projecting the content image includes adding effects to the content image corresponding to the music. According to this embodiment, music can be played in synchronization with the projection of the content image, and effects corresponding to the music can be added to the content image.
[0058] (Note 7) Furthermore, one aspect of the projection system in this disclosure includes a projector and a processing device for controlling the projector, wherein the processing device performs the following: generates correction data in which image data is corrected based on positional information indicating the positional relationship between the projector and the projection target and shape information indicating the shape of the projection target; generates content image data corresponding to the content image using a first image generation AI model into which specification information specifying the aspect of the content image and the correction data are input; and causes the projector to project the content image based on the content image data onto the projection target. According to this aspect, similar to the projection method in (Appendix 1), projection of highly artistic content images that are generated using the first image generation AI model and are more suitable for projection mapping is realized.
[0059] (Note 8) Furthermore, one aspect of the program of this disclosure involves a computer generating correction data in which image data is corrected based on positional information indicating the positional relationship between a projector that projects onto a projection target and the projection target, and shape information indicating the shape of the projection target; generating content image data corresponding to the content image using a first image generation AI model into which specification information specifying the form of the content image and the correction data are input; and causing the projector to project the content image based on the content image data onto the projection target. According to this aspect, similar to the projection method in (Appendix 1), projection of a highly artistic content image that is generated using the first image generation AI model and is more suitable for projection mapping is realized. [Explanation of Symbols]
[0060] 1...Projection system, 10...Projector, 20...Camera, 30...Speaker, 30...Information processing device, 50...Image generation server, 60...Segmentation server, NW...Communication network, 410...Processing device, 410a...Acquisition unit, 410b...First generation unit, 410c...Second generation unit, 410d...Projection control unit, 420...Communication device, 430...Storage device, 440...Bus, PR1...Program.
Claims
1. The process involves generating corrected data by correcting image data based on positional information indicating the positional relationship between the projector and the projection target, and shape information indicating the shape of the projection target. The first image generation AI model, which receives specification information specifying the aspect of the content image and the correction data as input, generates content image data corresponding to the content image. A projection method comprising projecting the content image, which is the content image data as the image data, from the projector onto the projection target.
2. The correction data is a plurality of modulated image data in which the image data is corrected based on the position information and the shape information, and in which at least one parameter defining the form of the corrected image shown by the correction data is different from each other. Generating the aforementioned content image data means The first image generation AI model, which receives the specified information and the plurality of modulated image data as input, generates the content image data for each of the plurality of modulated image data, Projecting the aforementioned content image from the projector onto the projection target is, The projection method according to claim 1, comprising sequentially projecting a plurality of content images based on a plurality of content image data corresponding one-to-one with the plurality of modulated image data.
3. The projection method according to claim 2, wherein the at least one parameter includes the gradation of the corrected image indicated by the correction data.
4. The aforementioned plurality of content image data includes a first content image data and a second content image data, Generating the aforementioned content image data means This includes generating interpolated content image data to be inserted between the first content image data and the second content image data using the first image generation AI model or a second image generation AI model different from the first image generation AI model, based on the first content image data and the second content image data. Projecting the aforementioned content image from the projector onto the projection target is, The projection method according to claim 2, comprising sequentially projecting the content image based on the first content image data, the content image based on the interpolated content image data, and the content image based on the second content image data.
5. Generating the aforementioned content image data means By applying segmentation processing to the aforementioned content image, the content image is classified into multiple regions, The projection method according to claim 1, comprising generating image data corresponding to each of the plurality of regions.
6. This includes playing music in sync with the projection of the aforementioned content image, The projection method according to claim 1, wherein projecting the content image includes adding an effect corresponding to the music to the content image.
7. A projector and The processing unit for controlling the projector includes, The aforementioned processing apparatus is The process involves generating corrected data by correcting image data based on positional information indicating the positional relationship between the projector and the projection target, and shape information indicating the shape of the projection target. The first image generation AI model, which receives specification information specifying the aspect of the content image and the correction data as input, generates content image data corresponding to the content image. A projection system that performs the following actions: projecting the content image, based on the content image data as the image data, onto the projection target using the projector.
8. On the computer, The process involves generating corrected data by correcting image data based on positional information indicating the positional relationship between the projector and the projection target, and shape information indicating the shape of the projection target. The first image generation AI model, which receives specification information specifying the aspect of the content image and the correction data as input, generates content image data corresponding to the content image. A program that performs the following actions: projecting the content image, based on the content image data as the image data, onto the projection target using the projector.
Citation Information
Patent Citations
Image generation method, image generation system, and program
JP2023118230A