Display image generation device and display image generation method
The display image generation device and method address the challenge of accurately displaying virtual objects interacting with users by employing a spring model to adjust node positions, achieving low latency and high accuracy in virtual object representation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies face challenges in accurately displaying virtual objects that interact with a user's body in a virtual world due to temporal constraints, leading to impaired realism and accuracy.
A display image generation device and method that utilize a spring model to adjust the positions of nodes in a skeletal model based on real-world object state information, ensuring low latency and high accuracy in displaying virtual objects.
Enables the display of virtual objects with low latency and high accuracy, maintaining natural motion and interaction with the user's movements.
Smart Images

Figure 2026046724000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a display image generation device and a display image generation method that generate a display image including a virtual object. [Background technology]
[0002] Technologies that provide a sense of immersion in virtual spaces using head-mounted displays and other devices are becoming commonplace across various fields. For example, the sense of presence in a virtual space can be enhanced by moving virtual objects on the display in a way that interacts with the user's movements, or by providing haptic feedback. In content such as electronic games, using user movements as a means of control allows for more intuitive operation compared to using input devices such as controllers. For example, by reflecting the user's hand movements on a virtual hand in the displayed world, it becomes possible to treat objects in the displayed world in the same way as in the real world. [Overview of the project] [Problems that the invention aims to solve]
[0003] When displaying virtual objects that interact with the user's body in a virtual world, even slight display errors can impair the sense of realism. In particular, when the user's movements are instantly reflected in the displayed virtual objects, it becomes difficult to display the virtual objects accurately due to temporal constraints.
[0004] This invention was made in view of these problems, and its purpose is to provide a technology that can display virtual objects that interact with the user with low latency and high accuracy. [Means for solving the problem]
[0005] One aspect of the present invention relates to a display image generation device. This display image generation device is characterized by comprising: a state information acquisition unit that acquires state information of an object in the real world in three-dimensional space; a skeletal model control unit that adjusts the position of nodes by applying a spring model whose natural length is the ideal distance based on the skeletal model of a virtual object corresponding to the object to the positions corresponding to the skeleton between nodes represented by the state information; and a display image generation unit that generates a display image including a virtual object that reflects the skeletal model composed of the adjusted nodes.
[0006] Another aspect of the present invention relates to a method for generating a display image. This method for generating a display image is characterized by comprising the steps of: acquiring state information of an object in the real world in three-dimensional space; adjusting the positions of nodes by applying a spring model whose natural length is the ideal distance based on the skeletal model of a virtual object corresponding to the object to positions corresponding to the skeleton between nodes represented by the state information; and generating a display image that includes the virtual object reflecting the skeletal model composed of the adjusted nodes.
[0007] Furthermore, any combination of the above components, as well as conversions of the expression of the present invention between methods, apparatus, systems, computer programs, recording media containing computer programs, etc., are also valid embodiments of the present invention. [Effects of the Invention]
[0008] According to the present invention, virtual objects that interact with the user can be displayed with low latency and high accuracy. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of the appearance of a head-mounted display to which this embodiment can be applied. [Figure 2] This figure shows an example configuration of a content processing system to which this embodiment can be applied. [Figure 3]This is a diagram showing an example of a display image generated by the content processing device in this embodiment. [Figure 4] This is a diagram schematically showing the basic procedure for reflecting the actual hand state information in the hand object in this embodiment. [Figure 5] This is a diagram illustrating problems caused by displacement in the display of the hand object. [Figure 6] This is a diagram schematically showing the state of introducing a spring model into the setting of the hand skeleton model in this embodiment. [Figure 7] This is a diagram showing the internal circuit configuration of the content processing device in this embodiment. [Figure 8] This is a diagram showing the functional blocks of the content processing device in this embodiment. [Figure 9] This is a diagram for explaining a specific example of a method in which the skeleton model control unit fits the state information to the skeleton model in this embodiment. [Figure 10] This is a diagram for explaining a specific example of a method in which the skeleton model control unit reflects the contact operation in the skeleton model in this embodiment. [Figure 11] This is a flowchart showing the processing procedure for the content processing device in this embodiment to generate and output a display image including a hand object that reflects the movement of the user's hand.
Mode for Carrying Out the Invention
[0010] This embodiment relates to a technique for representing at least a part of a user's body as a virtual object and making it联动 with actual movements. As long as it is within this scope, the means for detecting actual movements and the means for displaying images are not particularly limited. However, hereinafter, the description will mainly focus on the aspect of tracking the movement of the user's hand based on a captured image by a camera mounted on a head-mounted display and reflecting it in the movement of the hand in the display world.
[0011] Figure 1 shows an example of the appearance of a head-mounted display 100 to which this embodiment can be applied. In this example, the head-mounted display 100 consists of an output mechanism 102 and a mounting mechanism 104. The mounting mechanism 104 includes a mounting band 106 that wraps around the head when worn by the user to secure the device. The output mechanism 102 includes a housing 108 shaped to cover the left and right eyes when the user wears the head-mounted display 100, and has a display panel inside that faces the eyes when worn.
[0012] The housing 108 further includes an eyepiece positioned between the display panel and the user's eyes when the head-mounted display 100 is worn, which magnifies the image. The head-mounted display 100 may also be equipped with speakers or earphones positioned to correspond to the user's ears when worn. Furthermore, the head-mounted display 100 may incorporate motion sensors such as an accelerometer, gyroscope, and geomagnetic sensor to detect the translational and rotational movements of the user's head, as well as its position and orientation at each moment in time.
[0013] The head-mounted display 100 is equipped with cameras 110a, 110b, 110c, and 110d on the front of the housing 108 to capture video of the user and the surrounding real space. The number and arrangement of cameras 110a, 110b, 110c, and 110d are not particularly limited, but in the illustrated example, they are provided at the four corners of the front of the housing 108. Hereafter, cameras 110a, 110b, 110c, and 110d may be collectively referred to as camera 110. By sequentially analyzing each frame of the video image captured by camera 110, the movements of the user's hands within the field of view of camera 110 can be tracked in three-dimensional space. However, the part or unit to be tracked is not limited and may include the feet, upper body, lower body, or the entire body.
[0014] Furthermore, the images captured by camera 110 can also be used to obtain the position and orientation of the head-mounted display 100, and consequently the position and orientation of the user's head, using V-SLAM (Visual Simultaneous Localization and Mapping). V-SLAM is a technology that obtains the camera's position and orientation while creating an environmental map by repeatedly performing processes that estimate the 3D position of a real object from the positional relationship of images of the same real object captured from multiple viewpoints, and then estimating the camera's position and orientation based on the position of the real object's image in the captured image.
[0015] By changing the field of view of the images displayed on the head-mounted display 100 to correspond to the user's head position and orientation acquired by V-SLAM, the user can experience a sense of immersion in the displayed world. Furthermore, by immediately displaying images captured by a portion of the camera 110 on the head-mounted display 100, a see-through mode can be provided that shows the real space in the direction the user is facing.
[0016] Figure 2 shows an example configuration of a content processing system to which this embodiment can be applied. The head-mounted display 100 is connected to the content processing device 200 via wireless communication or an interface for connecting peripheral devices such as USB Type-C. The content processing device 200 may also be connected to a server via a network. In that case, the server may provide the content processing device 200 with online applications such as games that multiple users can participate in via the network.
[0017] The content processing device 200 basically processes the content program, generates display image and audio data, and transmits it to the head-mounted display 100. The head-mounted display 100 receives the display image and audio data and outputs it as the content image and audio. Here, the content processing device 200 sequentially acquires frame data of the moving image captured by the camera 110 of the head-mounted display 100, and based on that, immediately acquires information about the state of the user's hands.
[0018] The content processing device 200 displays a virtual hand object on the display image and sequentially reflects the user's hand state information. This allows for the display of an image of a hand that moves in the same way as the user's hand. As mentioned above, the object whose state is tracked using the captured image is not limited to the hand, so the virtual object that is linked may vary depending on the object being tracked. The processing that the content processing device 200 should perform using this mechanism is also not particularly limited. For example, the content processing device 200 may generate a display image in which a virtual object is lifted or moved in response to hand movements. Alternatively, the content processing device 200 may recognize the user's hand gestures as command inputs and perform corresponding information processing.
[0019] The content processing device 200 may also sequentially acquire information on the position and orientation of the user's head using technologies such as V-SLAM as described above, and generate a display image in the corresponding field of view. In this case, the content processing device 200 may acquire measurement values from the motion sensor built into the head-mounted display 100 to acquire the position and orientation of the user's head with greater accuracy.
[0020] Figure 3 shows examples of display images generated by the content processing device 200 in this embodiment. Both display images (a) and (b) assume that the user is in a virtual space 20 outdoors, and represent hand objects 22a and 22b. Based on the captured images transmitted from the head-mounted display 100, the content processing device 200 acquires information about the state of the hands in the real world and sequentially reflects this information in the state of the hand objects 22a and 22b.
[0021] The displayed image in (a) shows a scene in which a hand object 22a writes characters 24 in the virtual space 20. In this example, the content processing device 200 detects a gesture in which the tips of the middle and ring fingers are placed on the tip of the thumb, and the index and little fingers are extended, as the writing mode. In this mode, when the user moves their hand, the hand object 22a moves in conjunction, and the trajectories of the fingertips, such as the middle finger, are represented as characters 24.
[0022] At this time, the content processing device 200 sets a 3D model of the hand object 22a in a virtual 3D space in a state corresponding to the state information, and displays it on the display image along with other objects. Then, in accordance with the movement of the hand object 22a, linear objects representing the trajectory of the fingertips appear. As a result, the displayed characters 24 are defined as lines in 3D, so if the user wearing the head-mounted display 100 changes their viewpoint, the characters 24 can also be displayed as seen from an angle or from behind.
[0023] The displayed image in (b) shows a scene in which a keyboard 26 represented in virtual space 20 is operated with a hand object 22b. When the user moves their hand to press a desired key on the keyboard 26 while looking at the displayed image, the hand object 22b moves in the same way, and the key operation is performed. In this case, the content processing device 200 identifies the key to be operated by collision detection between the keyboard 26 and the fingertip in the virtual 3D space, based on the hand state information.
[0024] In parallel with this, the content processing device 200 sets a 3D model of the hand object 22b in a virtual 3D space in a state corresponding to the state information, and displays it on the display image together with the keyboard 26, etc. The content processing device 200 may also displace or change the color of the key on the keyboard 26 that is being operated, so as if it were being pressed by the hand object 22b. This makes it possible to represent the movement of the hand object 22b and the keyboard 26 in conjunction with the user's hand. Note that the illustrated display image is merely an example, and it will be understood by those skilled in the art that various representations can be realized using the hand object.
[0025] Figure 4 schematically shows the basic procedure for reflecting actual hand state information in a hand object in this embodiment. The content processing device 200 first acquires images 40 captured by the camera 110 from the head-mounted display 100. In practice, the content processing device 200 may acquire as many captured images as there are cameras 110 on the head-mounted display 100 at each time step corresponding to the frame rate.
[0026] The content processing device 200 extracts the hand region from the captured image using well-known techniques such as pattern matching, and then acquires the three-dimensional positional information of its feature points as state information 41 (S10). In the example shown in the figure, the positional coordinates of the nodes that determine the shape of the hand, such as joints, fingertips, and wrists (e.g., nodes 42a, 42b), and the positional orientation of the skeleton connecting the nodes (e.g., skeletons 44a, 44b) are identified in real space (XYZ space).
[0027] The method by which the content processing device 200 acquires state information from the captured image 40 is not particularly limited, but one example is the use of a Deep Neural Network (DNN). In this case, a DNN model data that takes hand images as input and outputs state information is prepared in advance by performing deep learning using a large number of hand images as training data. It will be understood by those skilled in the art that various types of neural networks and learning algorithms can be constructed using deep learning.
[0028] However, the means by which the content processing device 200 acquires state information are not limited to deep learning. For example, the content processing device 200 may determine the 3D position coordinates of feature points using the principle of triangulation, based on the position coordinates of corresponding feature points in multiple captured images taken from different viewing directions. Alternatively, the content processing device 200 may acquire hand state information using means other than captured images, such as a motion sensor attached to the hand.
[0029] The content processing device 200 stores model data of a hand object in its internal storage device 30. In the diagram, the model data schematically shows superficial data 32 such as polygon data and texture data, and a skeletal model 34 for controlling the state of the hand, i.e., its shape, position, and posture. However, this is not intended to limit the object's model data to only these. The content processing device 200 fits the actual hand state information 41 obtained in S10 to the hand object model (S12, S14).
[0030] In other words, the content processing device 200 maps the nodes in the hand state information 41 (e.g., nodes 42a, 42b) to the corresponding nodes in the hand object's skeletal model 34 (e.g., nodes 47a, 47b). The content processing device 200 also maps the skeletons in the hand state information 41 (e.g., skeletons 44a, 44b) to the corresponding skeletons in the hand object's skeletal model 34 (e.g., skeletons 48a, 48b).
[0031] The 3D model of an object is typically defined within the content or provided via an API (Application Programming Interface). Therefore, differences may exist between the user's hand and the object's hand in terms of hand shape, such as finger length and thickness, palm size, palm-to-finger ratio, and finger length ratio. Furthermore, the state information 41 obtained from the captured image may contain detection errors. For this reason, the content processing device 200 needs to derive a skeletal model 46 that is as close as possible to the state information 41 and represents a natural state. In this embodiment, this process is referred to as "fitting."
[0032] The content processing device 200 renders a hand object 49 in a virtual three-dimensional space (X'Y'Z' space) by applying polygon data and texture data to the fitted skeletal model 46 (S16). By repeating the illustrated process at a predetermined rate, the content processing device 200 can display a hand object 49 that moves in the same way as a real hand. On the other hand, slight discrepancies may occur in the hand object 49 due to differences in shape from a real hand, fitting errors, errors in state information, etc. This problem is particularly likely to become apparent in situations requiring detailed representation, such as hand gestures and interactions with other objects, as shown in Figure 3.
[0033] Figure 5 illustrates a problem caused by a display misalignment of a hand object. (a) assumes a gesture where the middle and ring fingers touch the thumb, as shown in Figure 3(a). If such a gesture is calculated to have occurred based on state information obtained from the captured image, it is expected that the display should also reflect this state, as shown in object 50a. However, due to the factors mentioned above, if a gap 52 occurs between the fingertips, as in object 50b, the gesture may appear not to have been performed.
[0034] (b) represents the scenario in which a key 54 in a virtual space is pressed with a fingertip, as shown in Figure 3(b). Based on the state information obtained from the captured image, if contact with the key 54 by the index finger is calculated, it is expected that the display should show that state, as in object 56a. However, due to the factors mentioned above, if the index finger has not reached the key 54 or is misaligned, as in object 56b, it will not appear as if the key has been pressed.
[0035] While it is conceivable to correct misalignments by changing the hand object in response to contact detection between fingertips or with other objects, this could lead to other problems such as distortion of the hand's shape as defined in the object model or abrupt, unnatural movements. Therefore, in this embodiment, a spring model is introduced between nodes and contact objects during fitting to the skeletal model and during contact movements. This facilitates fitting while representing contact movements with natural motion.
[0036] Figure 6 schematically shows how a spring model is introduced into the setting of a hand skeleton model. (a) shows an example of setting the spring model when fingertip contact is not considered. As described above, the content processing device 200 acquires state information 60, which includes the position coordinates of the nodes indicated by black circles (e.g., nodes 62a, 62b, 62c) and the position and orientation of the skeleton between them. In order to fit the state information 60 to the object's skeleton model, the content processing device 200 sets a spring model (e.g., spring model 64) at the position corresponding to the skeleton between the nodes included in the state information.
[0037] Here, "applying the spring model" means adjusting the position coordinates of the nodes by applying an attractive force if the nodes are farther apart than the natural length of the spring, and a repulsive force if they are closer, with a force proportional to the magnitude of the difference. By applying the spring model between each node, even if the distance between some nodes is farther in the state information, for example, the excess distance can be distributed among the other nodes in an appropriate balance according to the shape of the object defined in the 3D model. Although the figure shows springs only between some nodes, the number and position are not limited, and preferably the spring model is applied between all nodes. The same applies to the figures (b) and (c) described later.
[0038] (b) shows an example of the spring model settings when a gesture of touching the index finger and thumb is predicted. The content processing device 200 adjusts the distance between the node 62a corresponding to the tip of the index finger and the node 62b corresponding to the tip of the thumb so that when the fingertips actually come into contact, the fingertip surfaces of the object having thickness make contact. To this end, the content processing device 200 predicts the contact between the fingertips based on state information obtained from the captured image.
[0039] The content processing device 200 then introduces a spring model 66 between nodes 62a and 62b corresponding to the fingertips where contact is predicted. By setting the natural length of the spring model 66 to the ideal distance between the nodes when the fingertips of the objects come into contact, the nodes 62a and 62b are controlled to pull against each other and eventually stop at the ideal distance. As a result, a gap 52 like the one shown for object 50b in Figure 5(a) does not occur. Furthermore, since spring models (e.g., spring model 64) are also applied between each other node, the force from the spring model 66 is distributed, and all nodes can be positioned in a suitable balance.
[0040] (c) shows an example of the spring model settings when contact of the index finger with key 54 is predicted. The content processing device 200 adjusts the distance between the node 62a corresponding to the tip of the index finger and the contact point on key 54 so that when the actual finger reaches the position corresponding to key 54, the fingertip surface of the object having thickness makes contact with key 54. To this end, the content processing device 200 predicts the contact of the index finger with key 54 based on state information obtained from the captured image.
[0041] The content processing device 200 then introduces a spring model 68 between the node 62a corresponding to the tip of the index finger and the contact point on the key 54. By setting the natural length of the spring model 68 to the ideal distance, which is the distance between the node 62a and the contact point on the key 54 when the fingertip of the object contacts the key 54, the node 62a can be controlled to be pulled towards the contact point and eventually stop at the ideal distance. As a result, no misalignment occurs with the key 54, as shown in Figure 5(b) for object 56b. Furthermore, since spring models (e.g., spring model 64) are also applied between each other node, the force from the spring model 68 is distributed, and all nodes can be positioned in a suitable balance.
[0042] Figure 7 shows the internal circuit configuration of the content processing unit 200. The content processing unit 200 includes a CPU (Central Processing Unit) 222, a GPU (Graphics Processing Unit) 224, and main memory 226. These components are interconnected via a bus 230. An input / output interface 228 is further connected to the bus 230. A communication unit 232, a storage unit 234, an output unit 236, an input unit 238, and a recording medium drive unit 240 are connected to the input / output interface 228.
[0043] The communication unit 232 includes peripheral device interfaces such as USB, and network interfaces such as wired LAN or wireless LAN. The storage unit 234 includes a hard disk drive, non-volatile memory, etc. The output unit 236 outputs data to the head-mounted display 100. The input unit 238 receives data input from the head-mounted display 100. The recording medium drive unit 240 drives removable recording media such as magnetic disks, optical disks, or semiconductor memory.
[0044] The CPU 222 controls the entire content processing device 200 by executing the operating system stored in the memory unit 234. The CPU 222 also executes various programs that are read from the memory unit 234 or removable recording medium and loaded into the main memory 226, or downloaded via the communication unit 232. The GPU 224 has the functions of both a geometry engine and a rendering processor, performs drawing processing according to drawing commands from the CPU 222, and outputs the drawing results to the output unit 236. The main memory 226 is composed of RAM (Random Access Memory) and stores the programs and data necessary for processing.
[0045] Figure 8 shows the functional blocks of the content processing device 200. Each device may perform general information processing such as application progress and communication with a server, but Figure 8 specifically shows functional blocks related to display image generation processing, including the rendering of virtual objects. From this perspective, the content processing device 200 can be realized as a display image generation device. At least some of the functions of the content processing device 200 shown in Figure 8 may be implemented on a server connected to the content processing device 200 via a network, or on the head-mounted display 100.
[0046] Furthermore, the multiple functional blocks shown in Figure 8 can be realized in hardware terms using the various circuits shown in Figure 7, and in software terms using a computer program that implements the functions of the multiple functional blocks. Therefore, it will be understood by those skilled in the art that these functional blocks can be realized in various ways using hardware alone, software alone, or a combination thereof, and are not limited to any one of these.
[0047] The content processing device 200 includes an image acquisition unit 70 that acquires data of captured images, an operation information acquisition unit 72 that acquires information related to the content of user operations, a state information acquisition unit 76 that acquires hand state information from captured images, a contact prediction unit 78 that predicts contact actions based on the state information, an object data storage unit 80 that stores data of objects to be displayed, and a three-dimensional space control unit 82 that controls the three-dimensional space of the display target. The content processing device 200 further includes an information processing unit 74 that performs information processing based on the content of user operations and hand state information, a display image generation unit 84 that generates a display image, and an output unit 86 that outputs data of the display image.
[0048] The captured image acquisition unit 70 immediately acquires frame data of images captured by the head-mounted display's camera 110 at a predetermined rate. The captured image acquisition unit 70 may further detect the area of a hand in the captured image using pattern matching or the like, and cut out that area. The operation information acquisition unit 72 acquires the content of user operations on the content being performed using the head-mounted display 100 or a controller (not shown). The operation information acquisition unit 72 also acquires information related to the position and orientation of the head-mounted display 100, and consequently the position and orientation of the user's head, based on the V-SLAM and various sensor data described above.
[0049] The state information acquisition unit 76 acquires hand state information at each time step based on the captured images acquired by the captured image acquisition unit 70. For example, the state information acquisition unit 76 extracts hand feature points such as contours and joints from multiple images simultaneously captured by multiple cameras 110, and then determines the 3D position coordinates of the feature points using the principle of triangulation based on the position coordinates of the corresponding feature points in each image. Alternatively, the state information acquisition unit 76 may acquire hand state information using the aforementioned DNN or motion sensors attached to the hand, or it may integrate state information acquired by multiple means.
[0050] The contact prediction unit 78 predicts, based on the hand state information acquired by the state information acquisition unit 76, whether a part of the hand, such as the fingertips, will come into contact with an object within a predetermined time, and identifies a candidate contact target if contact occurs. Here, the contact candidate can be any other part of the actual hand, the other hand of the actual left or right hand, or an object existing in the virtual space. In other words, the contact prediction unit 78 may predict either contact in the real space or contact in the virtual space, as long as the contact should be reflected in the hand object. Hereafter, objects that can be candidates for contact targets, regardless of whether they are in the real space or virtual space, will be collectively referred to as "other objects".
[0051] The method by which the contact prediction unit 78 predicts contact with other objects is not particularly limited. For example, the contact prediction unit 78 predicts contact with another object when that object enters a predetermined range in real space or virtual space, based on the position of the fingertip indicated by the hand state information. Alternatively, the contact prediction unit 78 may predict subsequent movements based on the history of fingertip movement in real space or virtual space, and consider other objects within a predetermined range from the destination point after a predetermined time as contact candidates.
[0052] In any case, the contact prediction unit 78 widens the detection range of contact candidates the faster the finger movements determined from the state information. Also, the contact prediction unit 78 widens the detection range of contact candidates the greater the processing time required inside the content processing device 200 and the delay time until image display. This makes it possible to prepare a spring model without missing any other objects that may come into contact, and reduces problems such as sudden changes in the fingertips due to unexpected contact.
[0053] On the other hand, if the fingers are moving slowly, expanding the detection range of contact candidates unnecessarily may create unnecessary constraints with other objects, potentially causing jitter where the fingertips repeatedly fluctuate with even slight finger movements. Therefore, the contact prediction unit 78 may temporarily suspend its prediction operation when the speed of the hand or fingertips is below a threshold. In this case, other objects that have been predicted to come into contact up to that point may be retained as contact candidates. Note that the contact prediction unit 78 is not limited to predicting contact with the fingertips.
[0054] When predicting fingertip contact, the contact prediction unit 78 may predict contact for all five fingertips, or it may limit contact prediction to fingers involved in the operation, such as the index finger. Alternatively, the contact prediction unit 78 may vary the range in which it detects contact candidates for each finger, depending on the likelihood of the finger being involved in the operation. The contact prediction unit 78 may switch the selection rules for fingers involved in the operation and the detection range of contact candidates set for each finger depending on the content and the scene being displayed. Furthermore, the contact prediction unit 78 may temporarily suspend its prediction function during periods when the fingertips are hidden, such as when the hand is in a fist.
[0055] The 3D space control unit 82 controls the virtual 3D space, including the hand object, based on the latest state information determined by the state information acquisition unit 76. When setting the hand object in the 3D space, the 3D space control unit 82 includes a skeleton model control unit 88 that controls its skeleton model. The skeleton model control unit 88 performs a process to optimize the position coordinates of the nodes in the latest state information using a spring model at a predetermined frequency. Specifically, the skeleton model control unit 88 applies a spring model between the nodes and fits it to the skeleton model of the hand object. The skeleton model control unit 88 also applies a spring model between contact candidates and nodes so that contact can be represented without misalignment.
[0056] In applying the spring model, calculation methods known in various fields, such as physical simulations, can be used. As described above, the skeletal model control unit 88 applies forces to the nodes in the state information such that the distance between nodes and the distance between the contact point and the node in the contact candidate approach ideal values, and derives the three-dimensional position coordinates of all nodes when a suitable balance is obtained using the spring model. Specific examples of the processing performed by the skeletal model control unit 88 will be described later. The object data storage unit 80 stores data of the three-dimensional model of an object that exists in the display world. This data includes hand model data, including the skeletal model 34 shown in Figure 4.
[0057] The information processing unit 74 processes information about the content, such as an electronic game, based on the content of user operations acquired by the operation information acquisition unit 72, the hand state information acquired by the state information acquisition unit 76, and the contact actions predicted by the contact prediction unit 78. For example, the information processing unit 74 determines command input by hand gestures based on the hand state information and performs the corresponding processing. Alternatively, the information processing unit 74 may realize interaction with the hand object by appropriately changing the state of other objects that have been confirmed to have contact with the hand. The content and purpose of the processing performed by the information processing unit 74 are not particularly limited.
[0058] The information processing unit 74 may request the 3D space control unit 82 to reflect the results of the information processing in the 3D space of the displayed world. This allows not only the hand movements in the real world to be reflected in the hand object, but also other objects to change based on the progress of the content and interaction with the hand object.
[0059] Furthermore, if the information processing unit 74 detects a situation in which contact between the hand and another object is predicted during the information processing process, it may notify the contact prediction unit 78 accordingly. For example, if the hand object is to be moved separately by a controller or the like, the information processing unit 74 will obtain the details of that operation from the operation information acquisition unit 72, predict contact between the hand object and another object based on that information, and then notify the contact prediction unit 78 accordingly. In this case, the contact prediction unit 78 should then notify the 3D space control unit 82 of the other object that has been notified as a contact candidate.
[0060] The display image generation unit 84 renders an image representing the virtual 3D space controlled by the 3D space control unit 82 at a predetermined frame rate. In this process, the display image generation unit 84 may change its field of view of the virtual 3D space in accordance with the user's head movements. The output unit 86 sequentially outputs the frame data of the generated display image to the head-mounted display 100.
[0061] Figure 9 illustrates a specific example of the method by which the skeletal model control unit 88 fits state information to the skeletal model. The skeletal model control unit 88, for example, displaces the nodes included in the state information by applying a spring model, and determines the position coordinates of the nodes when an appropriate balance is achieved, through the following calculation.
[0062]
number
[0063] Here x i , x j This refers to the three-dimensional position coordinates of two nodes 90a and 90b connected by a single skeleton (edge), ||bij || is the length of the corresponding edge 92 in the skeletal model of the object, that is, the distance between nodes. As shown in (a) of the figure, F spring is, ||b ij When taking || as a reference, the spring force acting in the length direction of the edge on the nodes 90a and 90b with position coordinates x i , x j . Also, r i , r j are the initial values of the three-dimensional position coordinates of the two nodes 90a and 90b. As shown in (b) of the figure, F direction is the stress (elastic force) in the rotational direction with respect to the direction of the initial edge 94. Due to F direction , the positional relationship of the nodes displaced by F spring deviates from the original positional relationship, so as to prevent the direction of the edge from becoming unnatural.
[0064] α spring , α direction are coefficients representing the weights applied to F spring and F direction respectively. The skeletal model control unit 88 repeats the above calculations for all node pairs a predetermined number of times (for example, 32 times) to converge the values of the position coordinates, and uses them as the final position coordinates. Qualitatively, the larger the coefficients α spring , α direction , the faster the convergence but the higher the risk of oscillation. Conversely, the smaller the coefficients, the slower the convergence but the lower the risk of oscillation. Therefore, by appropriately setting each coefficient in advance, the values can be made to converge within a predetermined number of calculations.
[0065] As a result of adjusting the position coordinates of the nodes by the above calculations, the skeletal model control unit 88 checks whether the obtained finger angle (the angle formed by consecutive body segments) is realistic. If it is not realistic, the position coordinates of the nodes may be further adjusted. That is, as shown in (c) of the figure, the skeletal model control unit 88 uses the position coordinates x i , x j , x kFind the angle θ between the two edges 96a and 96b between the three nodes 90a, 90b, and 90c, and if it deviates from the upper or lower limits that represent a realistic range, adjust the position coordinate x so that it falls below the upper limit or above the lower limit. i , x j , x k Adjust.
[0066] Note that the angle θ can actually be the azimuth angle and zenith angle of one edge with respect to the other edge. The skeletal model control unit 88 performs the same check for all edge pairs connected by nodes and adjusts the position coordinates of the nodes as necessary. However, the timing of the skeletal model control unit 88 adjusting the nodes based on the angle between edges is not limited; qualitatively, the node position adjustment can be performed using a spring model or the like under constraints that the angle falls within a predetermined range.
[0067] Figure 10 is a diagram illustrating a specific example of how the skeletal model control unit 88 reflects contact motion in the skeletal model. When the contact prediction unit 78 detects a contact candidate, the skeletal model control unit 88 performs the following calculations in addition to the fitting calculations mentioned above. As a result, the skeletal model control unit 88 applies a spring model between the contact point of the contact candidate and the node of the finger where contact is predicted to occur, displacing the node and determining the position coordinates of the node that can represent natural contact.
[0068]
number
[0069] Note that the above equation is based on the position coordinate x i The fingertip with node x, and position coordinate x j This scenario assumes that another fingertip makes contact with the same node. For example, the so-called pinch gesture, where the thumb and index finger touch, can be considered. As shown in the diagram, L ij This is the ideal distance between those nodes 98a and 98b, and the distance between nodes 152a and 152b when the surfaces of the object's fingers 150a and 150b come into contact. That is, L ijThis parameter depends on the thickness of the object's fingers 150a and 150b. pinch is L ij When using as a reference, the position coordinate x i , x j This is the spring force acting along the length of the edge on nodes 98a and 98b. α pinch is F pinch This is a coefficient that represents the weight applied to α. spring , α direction Similarly, determine the appropriate value.
[0070] S ij This variable represents the degree to which contact has been achieved, starting at 0.0 in the initial state of the nodes, 1.0 when the fingertips are in contact, and monotonically increasing with decreasing distance between fingertips in states in between. T is the upper limit of the distance between fingertips when force is applied by the spring for contact action. In the above calculation, the maximum value operator increases the spring constant as the distance between fingertips decreases, and when the distance exceeds T, the spring force F pinch This has the effect of neutralizing the following: The former effect prevents the unnatural movement of fingertips suddenly attracting each other like magnets at a certain distance when they come close together. The latter effect prevents the spring force from being generated until a certain distance is approached, even if a spring model is applied between all contact candidates predicted by the contact prediction unit 78.
[0071] Therefore, the skeletal model control unit 88 may perform the above calculation for all pairs of touchable fingertips. Also, if the contact candidate is an object other than a hand, such as a virtual keyboard, the position coordinate x of one of the nodes in the above calculation j If we fix the value, the contact between the object and the fingertip can be naturally represented using a similar calculation. In this case, the ideal distance L ij This is the distance between the point of contact on the surface of the object's finger and the node corresponding to the fingertip when the surface of the object's finger touches the object's contact candidate.
[0072] Next, the operation of the content processing device 200 that can be realized in this embodiment will be described. Figure 11 is a flowchart showing the processing procedure by which the content processing device 200 generates and outputs a display image including a hand object that reflects the user's hand movements. This flowchart starts with the content processing device 200 having established communication with the head-mounted display 100 worn by the user, and acquiring frame data of the captured image, the content of the user's operation, and data related to the position and orientation of the user's head from the head-mounted display 100.
[0073] First, the state information acquisition unit 76 of the content processing device 200 acquires state information of the user's hand based on the frame of the captured image (S20). This state information includes at least the 3D position coordinates of the hand's nodes. If the contact prediction unit 78 has not detected any contact candidates based on the state information up to that point (N in S22), the skeletal model control unit 88 of the 3D space control unit 82 applies a spring model between the hand's nodes (S26) to determine the position coordinates of the nodes fitted to the object's skeletal model (S28).
[0074] If the contact prediction unit 78 detects a contact candidate (Y in S22), the skeletal model control unit 88 applies a spring model between the contact point of the contact candidate and the node of the finger where contact is predicted (S24), and then applies the spring model between other nodes as well (S26) to determine the position coordinates of the nodes (S28). This allows the position coordinates of nodes close to the object's skeletal model to be determined, while keeping the distance to the contact candidate as a constraint.
[0075] The 3D space control unit 82 sets the hand object in the virtual 3D space by applying polygons to the skeletal model whose nodes are the position coordinates determined in S28 (S30). In parallel with this, the 3D space control unit 82 may reflect the results of information processing in each object in the virtual 3D space according to the request of the information processing unit 74. The display image generation unit 84 generates frame data of the display image by drawing the latest state of the object in the virtual 3D space and outputs it sequentially to the head-mounted display 100 via the output unit 86 (S32).
[0076] If there is no need to stop the display due to content termination or user operation (N in S34), the content processing device 200 repeats the processing in S20 to S32 at a predetermined rate. This allows for accurate rendering of the hand object with low latency and natural representation of the fingertips touching other objects. The frequency of processing in S20 to S26 may be the same as the display frame rate or lower. In the latter case, the position coordinates of the nodes for each frame rate may be estimated by extrapolation based on the position coordinates in previous frames. If it becomes necessary to stop the display, the content processing device 200 terminates all processing (Y in S34).
[0077] According to the embodiment described above, in an embodiment in which the movement of an object in the real world is reflected in a display object, a spring model is applied between the nodes when reflecting the state of the object in the object's skeletal model. This makes it possible to represent an object with low latency and high accuracy while reflecting the state of the object and following the balance of the object's shape defined in the 3D model.
[0078] Furthermore, it predicts contact between fingers and other virtual objects, and applies a spring model between them. This prevents gaps or misalignments between contact objects on the display caused by differences in shape or state information between the actual object and the virtual object, allowing for a natural representation of gesture formation and contact. As a result, it can improve the quality of content representing objects that are linked to actual movements in various situations.
[0079] The present invention has been described above based on embodiments. The above embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their respective components and processing processes, and that such modifications also fall within the scope of the present invention.
[0080] For example, this embodiment mainly describes how to reflect hand state information into the skeletal model of a hand object, but as mentioned above, the same effect can be obtained by similar calculations when reflecting body parts other than the hand or the entire body into the skeletal model of those objects. For example, when reflecting the movement of the entire body into a human object, the node settings may be set in a larger unit than in the case of the hand.
[0081] This disclosure may include the following aspects: [Item 1] A display image generation device comprising a circuit configured as follows: The aforementioned circuit is By acquiring state information of an object in the real world in 3D space, By applying a spring model whose natural length is the ideal distance based on the skeletal model of the virtual object corresponding to the object to the position corresponding to the skeleton between the nodes represented by the state information, the position of the node is adjusted. A display image is generated that includes the virtual object, which reflects the skeletal model composed of the adjusted nodes. Display image generation device. [Item 2] The aforementioned circuit is Based on the aforementioned state information, contact candidates that are predicted to come into contact with the object are detected. The display image generation device according to item 1, wherein when adjusting the position of the node, a spring model is further applied between the node and the contact candidate, with the ideal distance at contact being the natural length. [Item 3] The aforementioned circuit is The display image generating device according to item 2, which adjusts the positions of both nodes by applying the spring model between the node corresponding to the fingertip of the hand which is the object and the node corresponding to another fingertip of the hand which is the contact candidate. [Item 4] The aforementioned circuit is The display image generation device described in item 3, which determines the ideal distance based on the thickness of the virtual object's finger. [Item 5] The aforementioned circuit is The display image generating device according to item 2, which adjusts the position of the node corresponding to the fingertip by applying the spring model between the node corresponding to the fingertip of the object and another virtual object that is a contact candidate. [Item 6] The aforementioned circuit is The display image generating device according to item 1, wherein when adjusting the position of the aforementioned nodes, a rotational stress is applied to the change in the direction of the skeleton between the nodes from the initial position of the aforementioned nodes. [Item 7] The aforementioned circuit is The display image generating device according to item 1, wherein the position of the node is adjusted under the constraint that the angle between the two skeletons connected by the node falls within a predetermined range. [Item 8] The aforementioned circuit is The display image generation device according to item 2, wherein the spring constant of the spring model applied between the node and the contact candidate is increased as the distance between them decreases. [Item 9] The aforementioned circuit is The display image generating device according to item 2, wherein when the distance between the node and the contact candidate exceeds a predetermined value, the force of the spring model applied between them is disabled. [Item 10] By acquiring state information of an object in the real world in 3D space, By applying a spring model whose natural length is the ideal distance based on the skeletal model of the virtual object corresponding to the object to the position corresponding to the skeleton between the nodes represented by the state information, the position of the node is adjusted. A display image is generated that includes the virtual object, which reflects the skeletal model composed of the adjusted nodes. Display image generation method. [Item 11] A function to acquire state information of an object in the real world in 3D space, The function adjusts the position of the nodes by applying a spring model whose natural length is the ideal distance based on the skeletal model of the virtual object corresponding to the object to the position corresponding to the skeleton between the nodes represented by the state information, A function to generate a display image including the virtual object that reflects the skeletal model composed of the adjusted nodes, A recording medium that stores a program to implement a computer. [Explanation of Symbols]
[0082] 100 Head-mounted display, 110 Camera, 70 Image acquisition unit, 72 Operation information acquisition unit, 74 Information processing unit, 76 State information acquisition unit, 78 Contact prediction unit, 80 Object data storage unit, 82 3D space control unit, 84 Display image generation unit, 86 Output unit, 88 Skeleton model control unit, 200 Content processing unit, 222 CPU, 224 GPU, 226 Main memory.
Claims
1. A state information acquisition unit that acquires state information of an object in the real world in three-dimensional space, A skeletal model control unit adjusts the position of the nodes by applying a spring model whose natural length is the ideal distance based on the skeletal model of the virtual object corresponding to the object to the position corresponding to the skeleton between the nodes represented by the state information, A display image generation unit generates a display image that includes the virtual object reflecting the skeletal model composed of the adjusted nodes, A display image generation device characterized by having the following features.
2. The system further includes a contact prediction unit that detects contact candidates that are predicted to come into contact with the object based on the aforementioned state information. The display image generation apparatus according to claim 1, characterized in that when the skeletal model control unit adjusts the position of the node, it further applies a spring model between the node and the contact candidate, the ideal distance at contact being the natural length.
3. The display image generating apparatus according to claim 2, characterized in that the skeletal model control unit adjusts the positions of both nodes by applying the spring model between the node corresponding to the fingertip of the hand which is the object and the node corresponding to another fingertip of the hand which is the contact candidate.
4. The display image generation device according to claim 3, characterized in that the skeletal model control unit determines the ideal distance based on the thickness of the fingers of the virtual object.
5. The display image generation apparatus according to claim 2, characterized in that the skeletal model control unit adjusts the position of the node corresponding to the fingertip by applying the spring model between the node corresponding to the fingertip of the object and another virtual object which is a contact candidate.
6. The display image generating apparatus according to claim 1 or 2, characterized in that when the skeletal model control unit adjusts the position of the nodes, it applies rotational stress to the change in the direction of the skeleton between the nodes from the initial position of the nodes.
7. The display image generation apparatus according to claim 1 or 2, characterized in that the skeletal model control unit adjusts the position of the node under constraint conditions such that the angle between the two skeletons connected by the node falls within a predetermined range.
8. The display image generation apparatus according to claim 2, characterized in that the skeletal model control unit increases the spring constant of the spring model applied between the node and the contact candidate as the distance between them decreases.
9. The display image generation apparatus according to claim 2 or 8, characterized in that the skeletal model control unit disables the force of the spring model applied between the node and the contact candidate when the distance between them exceeds a predetermined value.
10. Steps to obtain state information of an object in the real world in three-dimensional space, The steps include adjusting the position of the nodes by applying a spring model whose natural length is the ideal distance based on the skeletal model of the virtual object corresponding to the object to the position corresponding to the skeleton between the nodes represented by the state information, A step of generating a display image that includes the virtual object reflecting the skeletal model composed of the adjusted nodes, A method for generating a display image, characterized by including the following:
11. A function to acquire state information of an object in the real world in three-dimensional space, The function adjusts the position of the nodes by applying a spring model whose natural length is the ideal distance based on the skeletal model of the virtual object corresponding to the object to the position corresponding to the skeleton between the nodes represented by the state information, A function to generate a display image including the virtual object that reflects the skeletal model composed of the adjusted nodes, A computer program characterized by enabling a computer to implement the following.