Virtual space generating device
The virtual space generation device addresses the lack of interactive gameplay by enabling users to control player characters and perform actions within virtual spaces, using image classifiers and large language models to enhance user engagement.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2026-04-02
AI Technical Summary
Existing virtual space systems lack interactive gameplay elements, as users merely passively enjoy images displayed on screen objects without actively controlling player characters.
A virtual space generation device that includes a display data storage unit, virtual space generation unit, player character operation processing unit, image display processing unit, determination unit, and determination result output unit, allowing users to control player characters and perform actions within the virtual space, with features like image classifiers and large language models to identify target areas and determine gameplay interactions.
Enables users to engage in interactive games by controlling player characters and performing actions within virtual spaces, adding game-like elements and enhancing user engagement through dynamic gameplay.
Smart Images

Figure JP2025020970_02042026_PF_FP_ABST
Abstract
Description
Virtual Space Generation Device
[0001] The present invention relates to a virtual space generation device.
[0002] In recent years, technologies that allow users to perform events, games, etc. by operating player characters in a virtual space have been widely used. By using virtual reality (VR) technology that attaches a terminal such as a head-mounted display (HMD) to the user and displays an image of the virtual space from the perspective of the player character on its display unit, it is possible to give the user a sense of immersion and let them enjoy events, games, etc.
[0003] Patent Document 1 describes a system in which a screen object is placed on a stage provided in a virtual space and videos of live events, sports, performances, etc. captured in the real space are displayed there. In this system, the user operates a player character that operates in the virtual space, and a video of the virtual space captured from the perspective of the player character is displayed on the head-mounted display worn by the user.
[0004] Japanese Unexamined Patent Application Publication No. 2022 - 68643
[0005] "[Generative AI] A thorough explanation of the features and usage examples of 44 major free and paid tools for LLM, image generation, code generation, video generation, etc.!", [online], September 12, 2024, AI Market Editorial Department, [Accessed September 27, 2024], Internet<URL:https: / / ai-market.jp / services / generative-ai-tools / > "What is YOLO? Explanation of its differences from other methods, advantages, and disadvantages", [online], February 7, 2024, AIsmiley Editorial Department, [Retrieved September 27, 2024], Internet<URL:https: / / aismiley.co.jp / ai_news / yolo / > "Object Detection Explained Simply! Detailed Explanation of Differences, Applications, and Methods", [online], Future Standard Co., Ltd., [Accessed September 27, 2024], Internet<URL:https: / / www.scorer.jp / blog / object-detection> Miyazawa Riyuya, "StrongSORT: DeepSORT is back and stronger! Upgraded tracking model!", [online], December 31, 2022, wevnal Co., Ltd., [Retrieved September 27, 2024], Internet<URL:https: / / ai-scholar.tech / articles / object-tracking / strongsort> "Hakuhodo DY Media Partners' AaaS Tech Lab develops technology to detect arbitrary objects in images with high accuracy using natural language through generative AI - to be used in RoomClip's services and media business areas", [online], May 14, 2024, Hakuhodo DY Media Partners Inc., [Retrieved September 27, 2024], Internet<URL:https: / / www.hakuhodody-media.co.jp / newsrelease / service / 20240514_34939.html>
[0006] In the system described in Patent Document 1, the user merely passively enjoys the images displayed on the screen object and does not cause any changes by moving the player character, thus lacking in gameplay.
[0007] The problem that this invention aims to solve is to provide a virtual space generation device that has unprecedented game-like features.
[0008] To solve the above problems, the virtual space generation device according to the present invention comprises: a display data storage unit that stores display data for a virtual space and display data for a player character that operates in the virtual space; a virtual space generation unit that generates a virtual space using the display data for the virtual space stored in the display data storage unit; a player character operation processing unit that operates the player character in the virtual space in response to an external input operation using the display data for the player character stored in the display data storage unit; an image display processing unit that generates an image in the virtual space based on image data including a target area; a determination unit that determines whether or not a predetermined operation performed on the image by the player character was performed on the target area; and a determination result output unit that outputs the determination result from the determination unit.
[0009] In the virtual space generation device according to the present invention, an image display processing unit generates an image including a target area within the virtual space. When a user performs an input operation to move a player character to this virtual space generation device using an external input device, the player character movement processing unit moves the player character in the virtual space in response to that input operation. The external input device is, for example, a head-mounted display (HMD), and its display unit displays an image of the virtual space as seen from the player character's perspective.
[0010] The image data generated in the virtual space is input from an external source. Information identifying the location of the target region in the image may or may not be included in the image data input from the external source (i.e., the target region included in the image may be undefined at the time the image display processing unit generates the image). In the latter case, the determination unit may include a classifier that uses a pre-trained model that has been trained on training data consisting of pairs of game images and target regions for various types of games, in order to identify the target region. Furthermore, this image data may be prepared in advance, or it may be generated by a large language model (LLM) or the like based on keywords entered by the user. In the latter case, the determination unit may also include a classifier that uses an LLM.
[0011] When a player character controlled by the user performs a predetermined action on an image generated in the virtual space, the determination unit determines whether or not the action was performed on a target area, and the determination result output unit outputs the determination result. For example, if the generated image relates to a shooting game, the predetermined action is a shooting action by the player character, and the determination unit determines whether or not the bullet fired by the player character's action hit a target area (enemy character, airplane, animal, etc.) included in the image. Alternatively, for example, if the image generated in the virtual space relates to a dance game, the predetermined action is a dance action by the player character, and the determination unit determines the degree of agreement between the player character's body movements and the target area (dancer's body) included in the image. This allows the user to enjoy a novel game in which they control a player character in accordance with images generated in the virtual space.
[0012] By using the virtual space generation device according to the present invention, it is possible to add game-like elements to a virtual space generation device that places screen objects in a virtual space and displays images there.
[0013] A diagram illustrating the main components of a game system, including one embodiment of the virtual space generation device according to the present invention. A diagram illustrating the state in which a player character performs an action on an image displayed on a screen object in a shooting game using the virtual space generation device of this embodiment. An example of generating a virtual space based on an image of a screen object in the virtual space generation device of this embodiment. A diagram illustrating judgment using a judgment panel in a shooting game using the virtual space generation device of this embodiment. A diagram illustrating the state in which a player character performs an action on an image displayed on a screen object in a dance game using the virtual space generation device of this embodiment. A diagram illustrating judgment using a judgment panel in a dance game using the virtual space generation device of this embodiment.
[0014] One embodiment of the virtual space generation device according to the present invention will be described below with reference to the drawings.
[0015] Figure 1 is a diagram showing the main components of a game system 1, including a virtual space generation device 10, according to this embodiment. The game system 1 includes the virtual space generation device 10 and a user terminal 40. Although only one user terminal 40 is shown in Figure 1, there may be multiple user terminals 40. The game system 1 and the user terminal 40 are interconnected via a wireless communication network such as the Internet, or a wired communication network.
[0016] The virtual space generation device 10 includes a storage unit 11. The storage unit 11 includes a user information storage unit 111, a virtual space display data storage unit 112, a character data storage unit 113, an effect data storage unit 114, an image generator storage unit 115, an image discriminator storage unit 116, and an image data storage unit 117. In this specification, the concept of an image includes not only still images but also moving images (for example, moving images composed of still images at 60 frames per second).
[0017] The user information storage unit 111 stores information such as a user ID and password that identify the user using the user terminal 40, the type of player character operated by that user, and a terminal ID that identifies the user terminal 40. The virtual space display data storage unit 112 stores display data for the virtual space. The character data storage unit 113 stores display data for player characters and non-player characters that operate in the virtual space. The effect data storage unit 114 stores data for various effects (sound, images, etc.) that occur in the virtual space.
[0018] The image generator storage unit 115 stores an image generator that generates images based on words or sentences input from an external source. The image generator can suitably utilize a generative AI such as a Large Language Model (LLM; e.g., Non-Patent Document 1). The image classifier storage unit 116 stores an image classifier that identifies target and non-target regions in images input from an external source or images generated by the image generator, and identifies the background in an image according to predetermined criteria. The image classifier can utilize, for example, a classifier based on object detection algorithms such as YOLO (You Only Look Once; e.g., Non-Patent Documents 2, 3) and SSD (Single Shot Multi-Box Detector; e.g., Non-Patent Document 3), or a trajectory estimation algorithm such as DeepSORT or StrongSORT (e.g., Non-Patent Document 4) that estimates the trajectory of an object identified by the classifier if it is a moving object. Furthermore, if the image generator is a generative AI such as an LLM, the image classifier can similarly utilize a generative AI such as an LLM (e.g., Non-Patent Document 5). In addition, algorithms can be used to identify moving objects in a video (objects whose position changes between temporally consecutive images) as target regions.
[0019] The image data storage unit 117 stores image data that can be used when allowing the user to play the game. The image data stored in the image data storage unit 117 may include both images to which information identifying the location of the target area contained in the image is associated, and images to which such information is not associated. The target area refers to the area that the user is playing the game on, for example, the display area of the target object in a shooting game, or the display area of the model dancer in a dance game. Furthermore, the image data stored in the image data storage unit 117 may include both images to which data stored in the virtual space display data storage unit 112 is associated, and images to which it is not.
[0020] The virtual space generation device 10 includes, as functional blocks, a user authentication unit 21, a virtual space generation unit 22, a player character motion processing unit 23, an image information input receiving unit 24, an image generation unit 25, an image identification unit 26, an image display processing unit 27, a determination unit 28, and an effect generation unit 29. Details of these functional blocks will be described later.
[0021] The virtual space generation device 10 is composed of, for example, a cloud server or a general personal computer, and each of the above-mentioned functional blocks is realized by executing pre-installed dedicated software (virtual space generation program) on the processor.
[0022] The user terminal 40 is equipped with a storage unit 41. The storage unit 41 stores terminal ID information for identifying the user terminal 40, and information about the user who will be using the user terminal 40 (such as user ID information).
[0023] The user terminal 40 also includes, as hardware, a posture acquisition unit 42, an audio input unit 43, an audio output unit 44, an input unit 45, and a display unit 46. The user terminal 40 is, for example, a head-mounted display (HMD), and the posture acquisition unit 42 has various sensors and a calculation processing unit that determines the user's position and posture based on the output signals of the sensors. The posture acquisition unit 42, the audio input unit 43, and the audio output unit 44 transmit and receive signals to and from the virtual space generation device 10 at a predetermined frequency (for example, 60 times / second). As a result, the user's position and posture information acquired by the posture acquisition unit 42, and the audio information input to the audio input unit 43 are transmitted to the virtual space generation device 10 and reflected in the position, posture, and voice of the player character in the virtual space. In addition, audio information generated in the virtual space is output from the audio output unit 44. The display unit 46 is positioned in front of the user's eyes when the user is wearing the user terminal 40 on their head. Furthermore, the user terminal 40 includes a data processing unit 48 as a functional block.
[0024] Next, the operation of the game system 1 of this embodiment will be described.
[0025] When a user accesses the virtual space generation device 10 by performing a predetermined operation on the user terminal 40, the user authentication unit 21 generates a screen for entering a user ID and password and transmits this data to the user terminal 40. On the user terminal 40, the data processing unit 48 processes the received data and displays the screen on the display unit 46. When the user enters a user ID and password, the user authentication unit 21 identifies the user based on the entered user ID and password. Here, the configuration is set up so that the user ID and password are entered, but the user ID and password stored in the storage unit 41 of the user terminal 40, or the terminal ID, may be transmitted automatically. In the following description as well, the display data of the screen generated by each part of the virtual space generation device 10 is transmitted to the user terminal 40, and the display data is processed by the data processing unit 48 and displayed on the display unit 46, but these series of processes will not be described repeatedly and will simply be described as "display the screen on the display unit 46," etc.
[0026] When the user authentication unit 21 identifies the user, the virtual space generation unit 22 reads the virtual space display data predetermined as the initial state from the virtual space display data storage unit 112 and generates the virtual space. The player character operation processing unit 23 identifies the type of player character that the user will operate based on the data stored in the user information storage unit 111. If multiple types of player characters are associated with the user, the user is prompted to select one of them. Then, the player character is made to appear in the virtual space based on the data stored in the character data storage unit 113. In addition to this, a pre-configured non-player character may also be made to appear.
[0027] Here, the user logs in by entering a user ID and password, but the login process may be omitted. Also, although the type of player character the user will control is pre-stored in the user information storage unit 111, the user may be presented with information on multiple types of player characters stored in the character data storage unit 113 without associating them with user information, and the user may be allowed to select the player character to control each time. Alternatively, the type of player character may be stored in the storage unit 41 of the user terminal 40, and the player character operation processing unit 23 may read that information.
[0028] Next, the image information input receiving unit 24 displays a screen on the display unit 46 that allows the user to select whether to use an image already stored in the image data storage unit 117 or a newly generated image as the image to be used in the game.
[0029] If the user chooses to use an existing image, the image information input receiving unit 24 displays a list of images (game footage) stored in the image data storage unit 117 on the display unit 46 and prompts the user to select one of them.
[0030] When a user selects an image, the image recognition unit 26 checks whether information identifying the location of a target region contained in the selected image is associated with the data of that image. If information identifying the location of the target region is associated, it is read along with the image. On the other hand, if information identifying the location of the target region is not associated, the image recognition unit reads and operates the image recognition unit stored in the image recognition unit storage unit 116.
[0031] Furthermore, the image recognition unit 26 checks whether the data of the selected image is associated with the display data of the virtual space stored in the virtual space display data storage unit 112. If the data of the selected image is associated with the data stored in the virtual space display data storage unit 112, the virtual space generation unit 22 reads that data. If the data of the selected image is not associated with the data stored in the virtual space display data storage unit 112, the image recognizer stored in the image recognizer storage unit 116 is read and operated. The display data of the virtual space associated with the image data is the display data of the virtual space that is displayed in synchronization with the image (game video) displayed on the screen object 5, which will be described later.
[0032] If the user chooses to use a newly generated image on the selection screen described above, the image generation unit 25 reads the image generator from the image generator storage unit 115 and asks the user what kind of image (game footage) to display. In response, the user inputs a word (for example, shooting game, forest, hunting) or a sentence (for example, a shooting game where you hunt in a forest) via voice or text from the user terminal 40, and the image generator generates image data based on the input voice or text.
[0033] For example, if the image created here is footage from a dance game, the system can generate a dance game in which the user can watch a demonstration and then imitate the movements. In this case, the user should input information from the user terminal 40, such as, "a dance game consisting of two phases: a demonstration phase (a phase in which the demonstration movements are displayed for the user to watch) and a play phase (a phase in which the user imitates the demonstration movements that the user has watched. In this phase, the demonstration movements may or may not be displayed)." In addition, in a dance game, if the physique of the model performing the demonstration movements does not match that of the player character, the result of the movement matching judgment may be poor. Therefore, it is advisable to input information about the gender and physique of the player character along with the above information. By inputting this information, the image generation unit 25 can generate an image of a dance game in which a model with a physique similar to that of the player character performs the demonstration movements. The information input from the user terminal 40 is stored in the storage unit 11 (for example, the image data storage unit 117).
[0034] Once the selection or generation of images to be used in the game is complete, the image display processing unit 27 virtually displays a transparent screen object 5 at a predetermined position in the virtual space 8 (for example, a position directly facing the player character 9 at a predetermined distance L; see Figure 2), and displays an image (starts video playback) based on the selected or generated data. Because the screen object 5 is transparent, when the player character 9 views the screen object 5 from their perspective, they can simultaneously see the virtual space located behind the screen object 5. As described above, the screen object 5 is virtually displayed in the virtual space, and for example, player characters and non-player characters operated by other users can freely pass through the screen object 5. Furthermore, if the image is associated with data stored in the virtual space display data storage unit 112, the virtual space generation unit 22 generates a virtual space based on the virtual space display data associated with the image, in parallel with the display of the image (video playback; the same applies hereinafter).
[0035] If the selected or generated image is not associated with the virtual space display data stored in the virtual space display data storage unit 112, the image identification unit 26 identifies the elements contained in the image using an image classifier in parallel with the display of the image, and generates a virtual space 8 that extends the image displayed on the screen object 5. For example, as shown in Figure 3, if the image has a sky 6 and ground 7 as a background, with the sky 6 containing the sun 61 and clouds 62, and the ground 7 containing houses 71 and roads 72, then a virtual space 8 is generated in the virtual space, with the sky 60 in the space above the boundary between the sky 6 and ground 7, and the ground 70 in the space below the boundary. In addition, clouds 621 are placed in the sky 60, houses 711 are placed in the ground 70, and roads 721 extended from the screen object 5 are displayed, similar to what is included in the image.
[0036] Furthermore, if the data of the selected or generated image does not contain information that identifies the location of the target region, the image recognition unit 26 identifies the target region 52 included in the image using an image classifier in parallel with the display of the image (see Figure 2).
[0037] When displaying video on screen object 5, the process of identifying the target area 52 may be performed at the same frequency as the video's frame rate (e.g., 60 frames / second) (e.g., 60 times / second). However, depending on the performance of the devices constituting the virtual space generation device 10 and the type of image classifier, such high-speed processing may be difficult.
[0038] Therefore, as an image classifier, it is advisable to use object detection algorithms such as YOLO (You Only Look Once) and SSD (Single Shot Multibox Detector) in combination with trajectory estimation algorithms such as DeepSORT and StrongSORT, which estimate and track the trajectories of objects detected by these object detection algorithms. By adopting this configuration, it becomes possible to perform object detection at a frequency lower than the video frame rate (for example, once per second), and to identify the target region 52 using the position of the object detected every second and the estimated trajectory of that object, making it possible to configure the virtual space generation device 10 using a general-purpose personal computer.
[0039] To avoid VR sickness caused by images displayed in a virtual space, a minimum frame rate of 60 frames per second is required. By using this technology, it is possible to realize games using high frame rate images. In addition, the variety of algorithms that can be selected as image classifiers increases. Alternatively, if all moving objects in the image are targets, a simpler algorithm can be used that detects the moving objects and identifies them as target regions 52.
[0040] When an image is displayed (video is played) on the screen object 5, and the user operates the player character 9 through the user terminal 40, the player character motion processing unit 23 moves the player character 9 in the virtual space 8 in response to the input operation. For example, the player character 9 moves in the virtual space in accordance with the user's movement, and when the user changes their posture, the player character 9 also changes to the same posture.
[0041] When the player character 9 (or user) performs a first predetermined action, the determination unit 28 is activated. Then, when the player character 9 (or user) performs a second predetermined action, the determination unit 28 recognizes that the action was performed on the image displayed on the screen object 5. Since the specific content of the predetermined action differs depending on the type and content of the game, the determination unit 28 determines the content of the first and second predetermined actions according to the content of the target area 52. Alternatively, if only a specific type of game is to be played, the content of the first and second predetermined actions may be predetermined.
[0042] In the case of a shooting game, the first predetermined action is, for example, the action of the player character 9 aiming a gun, and the second predetermined action is, for example, the action of operating the trigger (the action of firing a bullet from the gun). Alternatively, in the case of a dance game, the first predetermined action is the action of the player character 9 facing the screen object 5, and the second predetermined action is any action that changes posture from the state of facing the screen object 5. In the case of a dance game consisting of two phases, a demonstration phase and a play phase, the determination unit 28 detects the first predetermined action and the second predetermined action only during the play phase. Alternatively, the user may notify the virtual space generation device 10 of the timing when the player character 9 will perform a predetermined action on the image (the timing when the action will start) by performing a predetermined input operation on the user terminal 40, and the determination unit 28 may then operate for a predetermined time (for example, 10 seconds).
[0043] When the determination unit 28 recognizes that the operation of the player character 9 is performed on the image displayed on the screen object 5, it determines whether or not the operation is performed on the target area 52. For example, in a shooting game, it determines whether or not the bullet fired from the gun held by the player character 9 passes through the target area 52. In a game where some object is emitted from the player character 9 to the screen object 5, such as a shooting game, the effect generation unit 29 emits a light ray indicating the trajectory 91 of the bullet in the direction in which the gun barrel extends from the gun barrel of the gun held by the player character 9.
[0044] At this time, for example, the determination unit 28 virtually arranges a transparent determination panel 51 on the back surface of the screen object 5 (the surface opposite to the player character 9), and based on whether or not there is an overlap between the incident position 512 of the light ray passing through the screen object 5 on the determination panel 51 and the target area 511 on the determination panel 51 corresponding to the target area 52 on the screen object 5, it determines whether or not the bullet hits the target area 52 (see FIG. 4). This is an example of the hit determination by the determination unit 28, and the method of the hit determination can be appropriately changed.
[0045] Furthermore, for example, in the case of a dance game, as shown in Figure 5, the judgment unit 28 can be configured to virtually place a light source 54 behind the player character 9 (opposite side from the screen object 5), and virtually place a transparent judgment panel 51 behind the screen object 5, and score the dance performed by the player character 9 based on the degree of overlap between the silhouette area 513 of the player character 9 projected onto the judgment panel 51 and the target area 511 on the judgment panel 51 that corresponds to the target area 52 on the screen object 5, by irradiating the player character 9 with light from the light source 54 (see Figure 6). Using a light source 54 is not mandatory; any configuration can be adopted as long as it can project the silhouette area 513 of the player character 9 onto the judgment panel 51. For example, a camera can be virtually placed to capture the player character 9 from the side of the screen object 5, and the silhouette captured by the camera can be projected onto the judgment panel 51. As described above, in a dance game, if the physique of the model performing the exemplary movements does not match that of the player character, the judgment result of the degree of matching of movements may be poor. Therefore, if the image data for the dance game is not generated by inputting information about the gender and physique of the player character, it is advisable to perform processing such as matching the physiques (height and body width) of both when displaying the silhouette of the player character 9 on the judgment panel 51.
[0046] Here, the dance performed by the player character 9 is scored based on the degree of overlap between the target area 511 on the judgment panel 51 corresponding to the target area 52 on the screen object 5. However, various other methods can be employed. For example, the posture acquisition unit 42 of the user terminal 42 could be one capable of extracting the user's bone data from the detected user posture, and the player character motion processing unit 23 could be configured to move the player character 9 based on the user's bone data. Furthermore, the judgment unit 28 could extract the bone data of a model performing exemplary movements from the target area 52 and score the dance based on the degree of agreement with the user's bone data. However, if the judgment unit 28 extracts the model's bone data from the target area 52 in real time during gameplay and determines the degree of agreement with the user's bone data, the processing load on the judgment unit 28 may become excessive. Therefore, when the judgment unit 28 makes a judgment based on the degree of agreement between the user's and the model's bone data, it is advisable to perform the process of extracting the model's bone data from the image data in advance, after the image data has been selected or generated and before gameplay using that image data begins. Alternatively, in the case of a dance game consisting of two phases, a demonstration phase and a play phase, the bone data of the model can be extracted while the image data of the demonstration phase is being played back.
[0047] Alternatively, the posture acquisition unit 42 may extract information on changes in the position of a controller held in the user's hand or a specific object attached to the user's foot, and the player character motion processing unit 23 may be configured to reflect this position information in the movements (changes in the position of the player character 9's limbs), and further, the judgment unit 28 may extract changes in the position of the model's limbs from the target area 52 and score the dance based on the degree of agreement with the user's changes in the position of their limbs (that is, the dance may be scored based on the degree of agreement between the position of a specific part of the user and the change in the position of that specific part of the model included in the target area 52).
[0048] The determination result by the determination unit 28 is displayed as an image (or video) on the screen object 5 by the effect generation unit 29. For example, in the case of a shooting game, as shown in FIG. 2, the effect generation unit 29 can be configured to display an effect 53 representing a rupture or an explosion centered on the landing position within the target area, or display an effect that emits blood splashes from the landing position. Further, an effect sound indicating that a bullet has hit the target area 52 may be generated as an effect. Alternatively, it is also possible to indicate a hit on the target area by displaying the character "HIT!" near the landing position within the target area 52. Also, in the case of a dance game, as shown in FIG. 5, the effect generation unit 29 can be configured to display the determination result by the determination unit 28 as a score on the screen object 5, or display characters such as "Excellent!", "Good!", "Bad", etc. according to the determination result by the determination unit 28 at each time point. The data of these effects are stored in advance in the effect data storage unit 114 according to the assumed game type.
[0049] When the image (game video) used in the above game is newly generated by an image generator, after the game ends, the data of the image (game video) is stored in the image data storage unit 117 together with the information of the target area 52 identified by the image identifier from the game video. Also, the display data of the virtual space 8 generated by the image identifier is stored in the virtual space display data storage unit 112 in association with the data of the game video. Further, even when the game video used in the game is stored in the image data storage unit 117 in advance, after the game ends, the information of the target area 52 identified by the image identifier is stored in the image data storage unit 117, and the display data of the virtual space 8 generated by the image identifier is stored in the virtual space display data storage unit 112.
[0050] In the game system 1 of this embodiment, images input from an external source to the virtual space generation device 10 are displayed on the screen object 5, and the game can be enjoyed using these images. While the use of screen objects has been proposed before, it was limited to passively enjoying the images displayed there. In contrast, by using the virtual space generation device 10 and the game system 1 including the virtual space generation device 10 of this embodiment, it is possible to enjoy a novel game using images displayed on the screen object.
[0051] Furthermore, in the above embodiment, users can input words or sentences via voice or text to generate images (game footage) that suit their preferences and enjoy them.
[0052] The above embodiments are examples and can be modified as appropriate in accordance with the spirit of the present invention.
[0053] Although described as a game system 1 in the above embodiment, the virtual space generation device according to the present invention can be used for purposes other than games. For example, a system with the same configuration as in the above embodiment can be suitably used for training such as yoga or exercise. Furthermore, although the game system 1 was configured with the virtual space generation device 10 and the user terminal 40 in the above embodiment, some or all of the storage unit and functional blocks of the virtual space generation device 10 may be replaced by the storage unit and functional blocks of the user terminal 40.
[0054] In the above embodiment, an image including the target region was displayed on a screen object, but other methods can also be used. For example, a hologram of the image corresponding to the target region may be generated in virtual space without using a screen object. Also, in the above embodiment, a rectangular screen object 5 was placed and an image including the target region and non-target region was displayed, but the target region alone may be displayed on a screen object having the same outline as the target region. Since the screen object itself is placed virtually, its shape can be changed instantly and arbitrarily. In the above embodiment, an example in which the determination unit 28 uses a determination panel 51 was described, but determination may be performed without using a determination panel 51.
[0055] In the above embodiment, an example was described in which an image generator creates game footage and an image classifier detects the target region included in the game footage. However, it is also possible for an image former to generate an initial image containing the target region, and for the image generator to generate game footage by assigning a predetermined movement to that target region. For example, in the case of an animal, a trajectory of charging towards the player character 9 can be predetermined, and in the case of an airplane, a trajectory of turning while heading towards the player character 9 can be predetermined. In this case, the image classifier can easily track the target region included in the game footage using the movement data provided by the image generator. Furthermore, it is also possible to configure the system to launch a predetermined attack from the target region towards the player character according to the characteristics of the target region (airplane, animal, etc.).
[0056] In the above embodiment, an example of one user enjoying the game was described, but multiple users can also use their respective user terminals 40 to enjoy the game simultaneously in a single virtual space.
[0057] [Embodiments] It will be apparent to those skilled in the art that the exemplary embodiments described above are specific examples of the following embodiments.
[0058] (Section 1) A virtual space generation device according to one aspect of the present invention includes: a display data storage unit that stores display data for a virtual space and display data for a player character that operates in the virtual space; a virtual space generation unit that generates a virtual space using the display data for the virtual space stored in the display data storage unit; a player character operation processing unit that operates the player character in the virtual space in response to an external input operation using the display data for the player character stored in the display data storage unit; an image display processing unit that generates an image in the virtual space based on image data including a target area; a determination unit that determines whether or not a predetermined operation performed on the image by the player character was performed on the target area; and a determination result output unit that outputs the determination result from the determination unit.
[0059] In the virtual space generation device described in paragraph 1, the image display processing unit generates an image including the target area within the virtual space. When a user performs an input operation to move a player character to this virtual space generation device using an external input device, the player character movement processing unit moves the player character in the virtual space in accordance with that input operation. The external input device is, for example, a head-mounted display (HMD), and its display unit displays an image of the virtual space as seen from the player character's perspective.
[0060] The image data generated in the virtual space is input from an external source. Information identifying the location of the target region in the image may or may not be included in the image data input from the external source (i.e., the target region included in the image may be undefined at the time the image display processing unit generates the image). In the latter case, the determination unit may include a classifier that uses a pre-trained model that has been machine-learned to identify the target region included in the game images of various types of games. Furthermore, this image data may be pre-prepared, or it may be generated by a large language model (LLM) or the like based on keywords entered by the user. In the latter case, the determination unit may also include a classifier that uses an LLM.
[0061] When a player character controlled by the user performs a predetermined action on an image generated in the virtual space, the determination unit determines whether or not the action was performed on a target area, and the determination result output unit outputs the determination result. For example, if the generated image relates to a shooting game, the predetermined action is a shooting action by the player character, and the determination unit determines whether or not the bullet fired by the player character's action hit a target area (enemy character, airplane, animal, etc.) included in the image. Alternatively, for example, if the image generated in the virtual space relates to a dance game, the predetermined action is a dance action by the player character, and the determination unit determines the degree of agreement between the player character's body movements and the target area (dancer's body) included in the image. This allows the user to enjoy a novel game in which they control a player character on an image generated in the virtual space.
[0062] (Paragraph 2) The virtual space generation device according to Paragraph 2 is a virtual space generation device according to Paragraph 1, wherein the image display processing unit places screen objects in the virtual space and displays the image on the screen objects.
[0063] (Clause 3) The virtual space generation device according to paragraph 3 is the virtual space generation device according to paragraph 2, wherein the screen object is a transparent object, and the determination unit places a determination panel on the back of the screen object and makes the determination based on the range of actions performed on the image by the player character on the determination panel and the overlap of the target area.
[0064] In the virtual space generation device described in paragraph 1, one method of generating images in the virtual space is to display the images on screen objects placed in the virtual space, as described in paragraph 2. Furthermore, in the virtual space generation device described in paragraph 2, as described in paragraph 3, it is possible to determine whether or not the player character's actions were performed on a target area using a judgment panel placed behind the screen object.
[0065] (Article 4) The virtual space generation device according to Article 4 is a virtual space generation device according to any of Articles 1 to 3, further comprising: an image generator storage unit that stores an image generator that generates an image based on a word or sentence; and an image information input receiving unit that receives input of a word or sentence related to the image, wherein the image display processing unit generates the image data by inputting the word or sentence input to the image information input receiving unit into the image generator.
[0066] In the virtual space generation device described in paragraph 4, users can generate images of their choice according to their own preferences and enjoy games using those images.
[0067] (Paragraph 5) The virtual space generation device relating to Paragraph 5 is the virtual space generation device relating to Paragraph 4, wherein the image generator includes a large-scale language model.
[0068] As the image generator in the virtual space generation device according to paragraph 4, the large language models (LLMs) described in paragraph 5 can be suitably used.
[0069] (Paragraph 6) The virtual space generation device according to Paragraph 6 is a virtual space generation device according to any of Paragraphs 1 to 5, further comprising an image classifier storage unit that stores an image classifier that identifies objects contained in an image, which is composed of a trained model created by machine learning using training data consisting of a game image and a target region contained in the game image, and the determination unit identifies the target region contained in the image using the image classifier.
[0070] The virtual space generation device described in paragraph 6 allows users to enjoy games using any image, even if the target area is not specified in advance.
[0071] (Clause 7) The virtual space generation device according to paragraph 7 is a virtual space generation device according to paragraph 6, wherein the image classifier includes at least one of an object detection algorithm, an algorithm for estimating the trajectory of a moving object detected by the object detection algorithm, and a large-scale language model.
[0072] As the image classifier in the virtual space generation device according to paragraph 6, the object detection algorithm described in paragraph 7 (e.g., YOLO (You Only Look Once), SSD (Single Shot Multibox Detector)), an algorithm for estimating the trajectory of a moving object detected by the object detection algorithm (e.g., DeepSORT, StrongSORT), and / or large language models (LLM) can be suitably used.
[0073] 1...Game System 10...Virtual Space Generation Device 11...Memory Unit 111...User Information Memory Unit 112...Virtual Space Display Data Memory Unit 113...Character Data Memory Unit 114...Effect Data Memory Unit 115...Image Generator Memory Unit 116...Image Identifier Memory Unit 117...Image Data Memory Unit 21...User Authentication Unit 22...Virtual Space Generation Unit 23...Player Character Motion Processing Unit 24...Image Information Input Reception Unit 25...Image Generation Unit 26...Image Identification Unit 27...Image Display Processing Unit 28...Judgment Unit 29...Effect Generation Unit 40...User Terminal 41...Memory Unit 42...Posture Acquisition Unit 43...Audio Input Unit 44...Audio Output Unit 45...Input Unit 46...Display Unit 48...Data Processing Unit 5...Screen Object 51...Judgment Panel 511...Target Area on Judgment Panel 512...Incident Position 513...Player Character Silhouette Area 52...Target Area 53...Effects 54...Light source 6, 60...Sky 61...Sun 62, 621...Clouds 7, 70...Ground 71, 711...House 72, 721...Road 8...Virtual space 9...Player character 91...Trajectory
Claims
1. A virtual space generation device comprising: a display data storage unit that stores display data for a virtual space and display data for a player character that operates in the virtual space; a virtual space generation unit that generates a virtual space using the display data for the virtual space stored in the display data storage unit; a player character operation processing unit that operates the player character in the virtual space in response to an external input operation using the display data for the player character stored in the display data storage unit; an image display processing unit that generates an image in the virtual space based on image data including a target area; a determination unit that determines whether or not a predetermined operation performed on the image by the player character was performed on the target area; and a determination result output unit that outputs the determination result from the determination unit, wherein the image display processing unit places a screen object in the virtual space and displays the image on the screen object.
2. The virtual space generation device according to claim 1, wherein the screen object is a transparent object, and the determination unit places a determination panel on the back of the screen object and performs the determination based on the range of actions performed by the player character on the image on the determination panel and the overlap of the target area.
3. The virtual space generation device according to claim 1, further comprising: an image generator storage unit storing an image generator that generates an image based on a word or sentence; and an image information input receiving unit that receives input of a word or sentence related to the image, wherein the image display processing unit generates image data by inputting the word or sentence input to the image information input receiving unit to the image generator.
4. The virtual space generation apparatus according to claim 3, characterized in that the image generator includes a large-scale language model.
5. The virtual space generation device according to claim 1, further comprising an image classifier storage unit that stores an image classifier that identifies objects contained in an image, the image classifier being composed of a trained model created by machine learning using training data consisting of a game image and a target region contained in the game image, and the determination unit identifies a target region contained in the image using the image classifier.
6. The virtual space generation device according to claim 5, characterized in that the image classifier includes at least one of an object detection algorithm, an algorithm for estimating the trajectory of a moving object detected by the object detection algorithm, and a large-scale language model.