Display image generation device, content processing system, and display image generation method

The display image generation device and method address the challenge of balancing frame rate and accuracy by dynamically switching between real-time and predicted state information, ensuring high-quality object representation in display images.

JP2026050170APending Publication Date: 2026-03-19SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing technologies face challenges in balancing the frame rate and accuracy of object tracking and display image generation, leading to potential deterioration in object movement quality and unnatural appearances.

Method used

A display image generation device and method that includes a state information acquisition unit, a state information control unit, and a display image generation unit, which dynamically switch between using real-time and predicted state information based on hand movement speed and accuracy to maintain high-quality object representation.

Benefits of technology

The solution ensures high-quality display images that accurately reflect object movement, maintaining image stability and reducing blurring and jitter, even with lower state information acquisition frequencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026050170000001_ABST
    Figure 2026050170000001_ABST
Patent Text Reader

Abstract

It generates high-quality display images that include objects reflecting the movement of the target object. [Solution] The content processing device starts acquiring state information of the object based on the captured image (S10). If state information corresponding to the time step of the display image to be generated has been acquired, the device generates and outputs a display image in which this information is reflected in the object (S20, S22). If state information corresponding to the time step has not been acquired, the device switches between using the most recent state information (Y in S14, S16) or predicting the state information (N in S14, S18) depending on the situation, and generates and outputs a display image in which this information is reflected in the object (S20, S22).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a display image generation device, a content processing system, and a display image generation method for generating a display image that reflects the movement of a real object.

Background Art

[0002] Techniques for providing an immersive experience in a virtual space using a head-mounted display or the like have become common in various fields. For example, by moving a virtual object on a display so as to interact with the movement of a user, or by providing tactile feedback, the sense of presence in the virtual space can be enhanced. In content such as electronic games, by using the movement of a user as an operation means, a more intuitive operation can be performed compared to using an input device such as a controller.

Summary of the Invention

Problems to be Solved by the Invention

[0003] In order to immediately reflect the movement of an object such as a user on an object in a display image, it is necessary to perform a process of tracking the state of the object at high speed and with high accuracy. For example, if the display frame rate is increased, higher image quality can be qualitatively expected, but on the other hand, the time allowed for the tracking process becomes shorter, and as a result, the accuracy of the movement of the object may deteriorate or the object may appear unnatural. Thus, when simultaneously performing the tracking process of an object and the generation of a display image, it is always a major issue to balance their qualities.

[0004] The present invention has been made in view of such problems, and an object thereof is to provide a technique for generating a display image including an object that reflects the movement of an object with high quality. <000002 >

Means for Solving the Problems

[0005] One aspect of the present invention relates to a display image generation device. This display image generation device is characterized by comprising: a state information acquisition unit that acquires state information of an object based on an image of the object in a video captured by an imaging device; a state information control unit that determines the state information to be used by switching whether or not to manipulate the state information according to the passage of time, depending on the situation; and a display image generation unit that generates a display image including a virtual object that reflects the movement of the object using the determined state information.

[0006] Another aspect of the present invention relates to a content processing system. This content processing system is characterized by including the above-mentioned display image generation device and a head-mounted display that acquires and displays display image data from the display image generation device.

[0007] A further aspect of the present invention relates to a method for generating a display image. This method for generating a display image is characterized by comprising the steps of: acquiring state information of an object based on an image of the object in a video captured by an imaging device; determining the state information to be used by switching whether or not to manipulate the state information according to the passage of time, depending on the situation; and generating a display image including a virtual object that reflects the movement of the object using the determined state information.

[0008] A further aspect of the present invention relates to an object tracking method. This object tracking method is characterized by including the steps of: analyzing the image of an object in an image captured by an imaging device and acquiring state information of the object; and determining whether the brightness of an evaluation region set according to a predetermined rule for the image of the object is within a predetermined appropriate range, and if it is not within the appropriate range, causing the imaging device to adjust the target brightness in AEAGC (Auto Exposure Auto Gain Control) processing.

[0009] Furthermore, any combination of the above components, as well as conversions of the expression of the present invention between methods, apparatus, systems, computer programs, recording media containing computer programs, etc., are also valid embodiments of the present invention. [Effects of the Invention]

[0010] According to the present invention, it is possible to generate a display image that includes an object reflecting the movement of the object, with high quality. [Brief explanation of the drawing]

[0011] [Figure 1] This figure shows an example of the appearance of a head-mounted display to which this embodiment can be applied. [Figure 2] This figure shows an example configuration of a content processing system to which this embodiment can be applied. [Figure 3] This figure shows the internal circuit configuration of the content processing device of this embodiment. [Figure 4] This is a block diagram showing the functional blocks of the content processing device of this embodiment. [Figure 5] This figure shows an example of a display image generated by the content processing device in this embodiment. [Figure 6] This diagram schematically shows the changes in the displayed image when the state information control unit of this embodiment does not predict the state information. [Figure 7] This figure illustrates the impact on the displayed image when state information is not predicted in this embodiment. [Figure 8] This diagram schematically shows the changes in the displayed image when the state information control unit of this embodiment predicts state information. [Figure 9] This flowchart shows the processing procedure for the content processing device in this embodiment to generate and output a display image that includes a hand object reflecting the user's hand movements. [Figure 10] This figure illustrates how, in this embodiment, the evaluation results of the accuracy of state information are used for content processing. [Figure 11] This figure illustrates the typical timing of processing in response to changes in the accuracy of state information in this embodiment. [Modes for carrying out the invention]

[0012] This embodiment relates to a technology for sequentially acquiring state information of a real object and immediately reflecting it in the state of an object on a displayed image. In this respect, the means for acquiring state information, the means for displaying the image, the type of real object, and the type of object to which the state is reflected are not limited. As an example, this embodiment will primarily describe a method in which the state of a user's hand is acquired based on an image captured by a camera mounted on a head-mounted display, and an image including a virtual object of the hand in the same state is displayed on the head-mounted display.

[0013] Figure 1 shows an example of the appearance of a head-mounted display 100 to which this embodiment can be applied. In this example, the head-mounted display 100 consists of an output mechanism 102 and a mounting mechanism 104. The mounting mechanism 104 includes a mounting band 106 that wraps around the head when worn by the user to secure the device. The output mechanism 102 includes a housing 108 shaped to cover the left and right eyes when the user wears the head-mounted display 100, and has a display panel inside that faces the eyes when worn.

[0014] The housing 108 further includes an eyepiece positioned between the display panel and the user's eyes when the head-mounted display 100 is worn, which magnifies the image. The head-mounted display 100 may also be equipped with speakers or earphones positioned to correspond to the user's ears when worn. Furthermore, the head-mounted display 100 may incorporate motion sensors such as an accelerometer, gyroscope, and geomagnetic sensor to detect the translational and rotational movements of the user's head, as well as its position and orientation at each moment in time.

[0015] The head-mounted display 100 further includes cameras 110a, 110b, 110c, and 110d on the front surface of the housing 108, which capture a moving image of the real space around the user. The number and arrangement of the cameras 110a, 110b, 110c, and 110d are not particularly limited. In the illustrated example, they are provided at the four corners on the front surface of the housing 108. Hereinafter, the cameras 110a, 110b, 110c, and 110d may be collectively referred to as the camera 110. By sequentially analyzing each frame of the moving image captured by the camera 110, the movement of the user's hand in the three-dimensional space within the field of view of the camera 110 can be tracked.

[0016] According to the tracking result, if a hand object simulating the actual hand movement is represented on the image, virtual reality or augmented reality in which the user picks up or moves other virtual objects can be realized. Also, when the user takes a specific pose with their hand, corresponding information processing can be performed and the result can be reflected in the display image. Note that the tracking target may be, in addition to the user's hand, another part of the user's body, the entire user's body, a real object that the user holds or wears, etc. Also, depending on the tracking target, various virtual objects that are interlocked may be used. Hereinafter, the information in the three-dimensional space obtained from each frame of the captured image, such as the position, posture, and shape of the object, is collectively referred to as "state information".

[0017] Note that the captured image by the camera 110 can also be used to obtain the position and posture of the head-mounted display 100, and thus the position and posture of the user's head, by V-SLAM (Visual Simultaneous Localization and Mapping). V-SLAM is a technology that repeatedly performs a process of estimating the three-dimensional position of a real object from the positional relationship of images of the same real object reflected in captured images from multiple viewpoints, and a process of estimating the position and posture of the camera based on the position of the image of the real object whose position has been estimated on the captured image, thereby obtaining the position and posture of the camera while creating an environmental map.

[0018] If the viewing field of the image to be displayed on the head-mounted display 100 is changed according to the position and orientation of the user's head obtained by V-SLAM, the user can obtain a sense of immersion in the displayed world. Also, by immediately displaying the captured image by a part of the camera 110 on the head-mounted display 100, it is possible to provide a see-through mode that shows the state of the real space in the direction the user is facing as it is.

[0019] FIG. 2 shows a configuration example of a content processing system to which the present embodiment can be applied. The head-mounted display 100 is connected to the content processing device 200 by an interface for connecting a peripheral device such as wireless communication or USB Type-C. The content processing device 200 may be further connected to a server via a network. In that case, the server may provide an online application such as a game in which a plurality of users can participate via the network to the content processing device 200.

[0020] Basically, the content processing device 200 processes the content program, generates display image and audio data, and transmits it to the head-mounted display 100. The head-mounted display 100 receives the display image and audio data and outputs them as the content image and audio. Here, the content processing device 200 sequentially acquires the frame data of the moving image captured by the camera 110 of the head-mounted display 100, and based on it, immediately acquires the state information of the user's hand.

[0021] The content processing device 200 generates a display image that includes a virtual hand object that moves in the same way as the user's hand, based on the acquired state information. In this embodiment, by allowing the acquisition of state information at a rate lower than the display image generation rate, the accuracy of state information acquisition can be maintained regardless of the frame rate. This makes it possible to improve the robustness of the accuracy of the object's movement in relation to the surrounding environment, such as illumination. As described above, the content processing device 200 may detect that a specific hand pose (gesture) has been taken based on the state information and perform corresponding information processing as a command input.

[0022] The content processing device 200 may also sequentially acquire information on the position and orientation of the user's head using technologies such as V-SLAM as described above, and generate a display image in the corresponding field of view. In this case, the content processing device 200 may acquire measurement values ​​from the motion sensor built into the head-mounted display 100 to acquire the position and orientation of the user's head with greater accuracy. It will be understood by those skilled in the art that there are various possibilities for the processing performed by the content processing device 200 and the display images generated using state information of objects such as hands.

[0023] Figure 3 shows the internal circuit configuration of the content processing unit 200. The content processing unit 200 includes a CPU (Central Processing Unit) 222, a GPU (Graphics Processing Unit) 224, and main memory 226. These components are interconnected via a bus 230. An input / output interface 228 is further connected to the bus 230. A communication unit 232, a storage unit 234, an output unit 236, an input unit 238, and a recording medium drive unit 240 are connected to the input / output interface 228.

[0024] The communication unit 232 includes peripheral device interfaces such as USB, and network interfaces such as wired LAN or wireless LAN. The storage unit 234 includes a hard disk drive, non-volatile memory, etc. The output unit 236 outputs data to the head-mounted display 100. The input unit 238 receives data input from the head-mounted display 100. The recording medium drive unit 240 drives removable recording media such as magnetic disks, optical disks, or semiconductor memory.

[0025] The CPU 222 controls the entire content processing device 200 by executing the operating system stored in the memory unit 234. The CPU 222 also executes various programs that are read from the memory unit 234 or removable recording medium and loaded into the main memory 226, or downloaded via the communication unit 232. The GPU 224 has the functions of both a geometry engine and a rendering processor, performs drawing processing according to drawing commands from the CPU 222, and outputs the drawing results to the output unit 236. The main memory 226 is composed of RAM (Random Access Memory) and stores the programs and data necessary for processing.

[0026] Figure 4 is a block diagram showing the functional blocks of the content processing device 200. The content processing device 200 may perform general information processing such as application progress and communication with a server, but Figure 4 specifically shows functional blocks related to the generation of display images based on hand state information. From this perspective, the content processing device 200 can be realized as a display image generation device. At least some of the functions of the content processing device 200 shown in Figure 4 may be implemented on a server connected to the content processing device 200 via a network, or on the head-mounted display 100.

[0027] Furthermore, the multiple functional blocks shown in Figure 4 can be realized in hardware terms using the various circuits shown in Figure 3, and in software terms using a computer program that implements the functions of the multiple functional blocks. Therefore, it will be understood by those skilled in the art that these functional blocks can be realized in various ways using hardware alone, software alone, or a combination thereof, and are not limited to any one of these.

[0028] The content processing device 200 includes an image acquisition unit 70 that acquires data of captured images, an operation information acquisition unit 72 that acquires information related to the content of user operations, a state information acquisition unit 76 that acquires hand state information from captured images, a state information control unit 78 that controls the time change of state information, an object data storage unit 80 that stores data of objects to be displayed, and a three-dimensional space control unit 82 that controls the three-dimensional space of the display target. The content processing device 200 further includes an information processing unit 74 that performs information processing based on the content of user operations and hand state information, a display image generation unit 84 that generates a display image, and an output unit 86 that outputs data of the display image.

[0029] The captured image acquisition unit 70 sequentially acquires frame data of images captured by the head-mounted display's camera 110 at a predetermined rate. The operation information acquisition unit 72 acquires the content of user operations performed on the content being played using the head-mounted display 100 or a controller (not shown). The operation information acquisition unit 72 also acquires information related to the position and orientation of the head-mounted display 100, and consequently the position and orientation of the user's head, based on the V-SLAM and various sensor data described above.

[0030] The state information acquisition unit 76 acquires hand state information at each time step based on the captured images acquired by the captured image acquisition unit 70. The state information acquisition unit 76 estimates the hand state information using, for example, a DNN (Deep Neural Network). In this case, the state information acquisition unit 76 internally stores DNN model data for estimating hand state information, which has been acquired in advance by performing deep learning using a large number of hand images as training data.

[0031] Those skilled in the art will understand that there are various types of neural networks and learning algorithms that can be constructed using deep learning. However, the means by which the state information acquisition unit 76 acquires state information are not limited to DNNs. For example, the state information acquisition unit 76 may acquire state information by fitting the positional relationship of the hand feature points on the captured image with a 3D model of the hand.

[0032] The state information control unit 78 controls the time change of state information for each time step corresponding to the frame rate of the displayed image. Specifically, the state information control unit 78 switches between using the state information obtained from the captured image directly or manipulating it according to the passage of time when generating the displayed image, depending on the speed of the hand or the accuracy of acquiring the state information.

[0033] For example, the state information control unit 78 uses the state information obtained from the most recent captured image as is when the hand is moving at a speed below the threshold at which it is considered stationary. When the hand is moving at a speed above the threshold at which it is considered to be moving, it adds a change to the state information according to the elapsed time since the image was captured. Hereafter, the process of changing the state information in accordance with the passage of time is called "prediction" of the state information.

[0034] This switching process allows for stable quality of the displayed image, including the hand object, even when state information is acquired at a frequency lower than the frame rate of the displayed image. The state information control unit 78 may also evaluate the accuracy of state information acquisition under predetermined conditions, predict state information when it is estimated that a certain level of accuracy has been achieved, and switch off from predicting state information when it is estimated that accuracy has not been achieved. The specific processing of the state information control unit 78 will be described later.

[0035] The 3D space control unit 82 controls the 3D space of the display world, including the hand object, based on the latest state information determined by the state information control unit 78. Here, the 3D space control unit 82 reflects the state of the hand, determined in a time step corresponding to the frame rate of the display image, in the state of the hand object. The object data storage unit 80 stores the 3D model of the object that exists in the display world.

[0036] The information processing unit 74 processes information about the content, such as an electronic game, based on the user operation acquired by the operation information acquisition unit 72 and the latest state information determined by the state information control unit 78. For example, the information processing unit 74 determines a command input by hand gesture based on the state information determined by the state information control unit 78 and performs the corresponding processing. Alternatively, the information processing unit 74 may perform collision detection between the hand object and other objects based on the state information and change the state of the other objects as appropriate to realize interaction with the hand object. The content and purpose of the processing performed by the information processing unit 74 are not particularly limited.

[0037] The information processing unit 74 may request the 3D space control unit 82 to reflect the results of information processing in the 3D space of the display world. This allows not only the hand movements in the real world to be reflected in the hand object, but also other objects to change according to the progress of the content and interaction with the hand object. The display image generation unit 84 renders the 3D space of the display world at a predetermined frame rate. In this case, the display image generation unit 84 may change the field of view of the display world according to the movement of the user's head. The output unit 86 sequentially outputs the frame data of the generated display image to the head-mounted display 100.

[0038] Figure 5 shows examples of display images generated by the content processing device 200 in this embodiment. Both display images (a) and (b) assume that the user is in a virtual space 20 outdoors, and represent hand objects 22a and 22b. Based on the captured images transmitted from the head-mounted display 100, the content processing device 200 acquires information about the state of the hands in the real world and sequentially reflects this information in the state of the hand objects 22a and 22b.

[0039] The displayed image (a) shows a scene in which a keyboard 24 represented in virtual space 20 is operated with a hand object 22a. When the user moves their hand to press a desired key on the keyboard 24 while looking at the displayed image, the hand object 22a moves in the same way, and the key operation is performed. In this case, the information processing unit 74 identifies the key to be operated by collision detection between the keyboard 24 and the fingertip in three-dimensional space, based on the state information of the hand.

[0040] In parallel with this, the 3D space control unit 82 sets a 3D model of the hand object 22a in 3D space in a state corresponding to the state information, and the display image generation unit 84 displays it on the display image along with the keyboard 24. The 3D space control unit 82 may also displace or change the color of the key on the keyboard 24 that is being operated, so that it appears as if it is being pressed by the hand object 22a. By repeating these actions at a predetermined rate, the movements of the hand object 22a and the keyboard 24 can be represented in conjunction with the user's hand.

[0041] The displayed image in (b) shows a scene in which the hand object 22b writes characters 26 in the virtual space 20. In this example, the information processing unit 74 detects a gesture in which the tips of the middle and ring fingers are placed on the tip of the thumb, and the index and little fingers are extended, as the writing mode. In this mode, when the user moves their hand, the hand object 22b moves in conjunction, and the trajectories of the fingertips, such as the middle finger, are represented as characters 26.

[0042] At this time, the 3D spatial control unit 82 moves the hand object 22b according to the state information and makes a linear object representing the trajectory of the fingertip appear. As a result, the characters 26 displayed by the display image generation unit 84 are defined as lines in three dimensions, so if the user wearing the head-mounted display 100 changes their viewpoint, the characters 26 can also be displayed as seen from an angle or from behind. It should be noted that the illustrated display image is merely an example, and it will be understood by those skilled in the art that various shapes of objects that reflect the state information of the hand and various modes that can be realized by such objects are possible.

[0043] Figure 6 schematically shows the changes in the displayed image when the state information control unit 78 does not predict state information. The horizontal axis of the figure is the time axis, and the upper section schematically shows the hand states 30a, 30b, 30c, ... obtained from the captured image at each time step, while the lower section schematically shows the frame sequence of the displayed image containing the hand object. As an example, the figure assumes a case where the state information acquisition unit 76 acquires state information from the captured image at a frequency of half the frame rate of the displayed image. However, the frequency of state information acquisition is not limited to this.

[0044] First, the displayed image at time t1 is generated based on the state information (state 30a) obtained from the most recent captured image. At the next time, t2, since the state information based on the captured image is not updated, the displayed image is generated based on the same state information (state 30a). In other words, at times t1 and t2, the hand object is represented with the same position, orientation, and shape in three-dimensional space.

[0045] At the next time t3, the state information is updated, and the display image is generated based on that state information (state 30b). At the next time t4, the state information based on the captured image is not updated, so the display image is generated based on the same state information (state 30b). In other words, at times t3 and t4, the hand object is represented in the same position, orientation, and shape in 3D space. Similarly, at the next times t5 and t6, the hand object is represented in the display image in the same position, orientation, and shape in 3D space based on the same state information (state 30c).

[0046] Figure 7 illustrates the impact on the displayed image when state information is not predicted. The figure schematically shows viewpoints 42a, 42b, and 42c in a three-dimensional virtual space 40. When an image is displayed on the head-mounted display 100, the viewpoint and field of view may change according to the user's head movements. The figure shows that at times t1, t2, and t3, the viewpoint changes as viewpoints 42a, 42b, and 42c, and the respective fields of view change as fields of view 44a, 44b, and 44c. It is also assumed that the hand object is moving in the direction of arrow A in three-dimensional space.

[0047] As shown in Figure 6, at times t1 and t2, the hand object 46a is represented in the virtual space 40 with the same position, orientation, and shape, and at the following time t3, the hand object 46b is represented with the updated position, orientation, and shape. In response to the time change at times t1, t2, and t3, the surrounding image, such as the background in the virtual space 40, is updated at that rate. On the other hand, since the state of the hand object 46a does not change at times t1 and t2, it does not appear to be moving in the direction of arrow A, and in some cases, it may even appear to be moving in the opposite direction relative to the movement of the background.

[0048] Therefore, if the hand object 46b moves suddenly at time t3, a problem may occur where the hand objects 46a and 46b are seen as doubled due to the afterimage from time t2. The inventors thus gained the unique insight that in a head-mounted display where the field of view can change freely, a difference between the update frequency of state information and the frame rate of the displayed image can cause the image of an object to appear blurred. Furthermore, even without a head-mounted display, a difference in the update frequency of the display of only certain objects can cause discomfort to the user. Therefore, the state information control unit 78 predicts the state information and adjusts the update frequency of the object's state to match the frame rate of the displayed image.

[0049] Figure 8 schematically shows the changes in the display image when the state information control unit 78 predicts state information. The representation in the figure and the frequency at which the state information acquisition unit 76 acquires state information from the captured image are the same as in Figure 6. First, the display image at time t1 is generated based on the state information (state 30a) acquired from the most recent captured image. At the next time t2, since the state information based on the captured image is not updated, the state information control unit 78 predicts the state information (state 32a) according to the elapsed time of the display image generation cycle Δt = t2 - t1.

[0050] The state information control unit 78 extrapolates the state information (state 32a) at time Δt later from the most recent state information (state 30a) based on changes in state information up to that point, i.e., changes in position, orientation, and shape. The 3D space control unit 82 sets the 3D model of the object using this predicted state information (state 32a), and the display image generation unit 84 draws the image to generate the display image at time t2. In the figure, the state information obtained by prediction is shown with a dashed line.

[0051] At the next time t3, state information is obtained from the captured image, and a display image is generated based on this state information (state 30b). At the next time t4, since the state information based on the captured image is not updated, the state information control unit 78 predicts the state information (state 32b) according to the elapsed time of the display image generation cycle Δt. The 3D space control unit 82 sets the 3D model of the object with this predicted state information (state 32b), and the display image generation unit 84 draws the image, thereby generating the display image for time t4. At the next time t5, a display image is generated based on the state information (state 30c) obtained from the captured image, and at the next time t6, a display image is generated based on the predicted state information (state 32c).

[0052] This procedure allows the update frequency of the object's state to match the frame rate of the displayed image, thus avoiding problems such as the object appearing blurry. On the other hand, state information based on captured images can contain errors and noise due to various factors such as ambient light and how the hand actually appears. Since these errors and noise can occur during various image processing steps, they are more difficult to control compared to controllers that can acquire state information based on motion sensors. If state information containing such errors and noise is used to predict further state information, the errors and noise will be amplified, resulting in fluctuations (jitter) in the object's image.

[0053] Therefore, as described above, the state information control unit 78 in this embodiment switches whether or not to predict state information according to the actual speed of the hand. During the period when the hand is moving at a speed above a threshold, even if jitter occurs in the image due to errors in the state information, it is unlikely to be recognized as a perceptual characteristic. Accordingly, the state information control unit 78 activates the state information prediction function to prevent the blurring of the object as described in Figure 7 from being visually apparent.

[0054] During periods when the speed is below the threshold at which the hand can be considered stationary, no visual blurring of the object, as explained in Figure 7, occurs, and therefore the state information control unit 78 does not activate the state information prediction function. This reduces image jitter caused by errors and noise. By switching the prediction function in this way, the image of the hand object can be represented with high quality regardless of hand movement or changes in the field of view.

[0055] According to the above principle, the magnitude of jitter becomes more pronounced as the error and noise in the state information increase. Therefore, the state information control unit 78 may evaluate the accuracy of the state information acquired by the state information acquisition unit 76, and if it meets the conditions for determining that the accuracy is low, it may choose not to activate the state information prediction function regardless of the speed of the hand. The decision of whether or not to activate the prediction function when taking the accuracy of the state information into consideration is summarized as follows.

[0056] [Table 1]

[0057] "No hand movement" indicates a speed below a threshold, while "movement present" indicates a speed above a threshold. Here, the speed threshold for determining whether movement is present or absent, and the speed threshold for determining whether movement is present or absent, may be the same or different. By using different thresholds and controlling the switching between operation and non-operation with hysteresis, the occurrence of jitter caused by repeated switching in a short period of time can be suppressed. The state information control unit 78 acquires the hand speed based on the rate of change of the state information up to that point. The state information control unit 78 may determine the overall hand speed based on a threshold, or it may determine the speed of a part of the hand, such as a finger, based on a threshold.

[0058] The state information control unit 78 may detect that the hand is about to stop and stop the state information prediction function at that time. For example, when a user makes a gesture of touching fingertips together, the speed of the fingers, which had been moving up to that point, suddenly becomes 0 the moment the fingers touch. Therefore, the state information control unit 78 may detect that the speed will become 0 after a small amount of time based on the gesture occurrence prediction and stop the state information prediction function at that point. This prevents the prediction from being made that there will be movement even after the fingers touch, thus preventing the fingertips of the object from overshooting.

[0059] The state information control unit 78 may detect when the hand object is about to come into contact with another object, such as a wall, not just when the fingertips are touching, and may stop the state information prediction function at that point. In this case as well, it is possible to prevent the fingertips of the object from overshooting.

[0060] To determine if the accuracy of the state information is low, for example, at least one of the following a to g can be introduced. a. When the speed of the entire hand or part of the hand exceeds the acceptable range that can maintain the accuracy of the state information. b. When a hand other than the hand whose state information is being acquired (the user's other hand or someone else's hand) is visible in the captured image. c. When at least a portion of the hand from which state information is to be obtained is obscured by another object. d. When at least a portion of the hand whose state information is to be acquired is outside the field of view of camera 110. e. When the average brightness of the captured image falls below the acceptable range for maintaining accuracy. f. When the distance of the hand to camera 110 falls below the acceptable range for maintaining accuracy. g. When the state information acquisition unit 76 detects low state information acquisition accuracy during processing.

[0061] When conditions a, e, and f are adopted, the threshold for the "acceptable range" is determined in advance. When conditions b, c, d, and g are adopted, the extent to which the accuracy of the state information is considered low is determined in advance. Furthermore, when multiple conditions are adopted, the scoring rules for each situation are determined in advance, and the state information control unit 78 may determine that the accuracy of the state information is low by summing the score values ​​assigned to each and comparing them with the threshold.

[0062] If the status information acquisition unit 76 fails to acquire status information, the status information control unit 78 may determine the latest status information using the most recently acquired status information for a predetermined time after detecting the failure. If normal status information can be obtained within that time, the hand object can continue to be displayed with minimal error. In this case as well, the status information control unit 78 may decide whether or not to activate the status information prediction function based on whether or not there is hand movement. The failure to acquire status information may be notified to the status information control unit 78 by the status information acquisition unit 76, or the status information control unit 78 may detect it itself by detecting an abnormal value in the status information.

[0063] Next, the operation of the content processing device 200 that can be realized in this embodiment will be described. Figure 9 is a flowchart showing the processing procedure by which the content processing device 200 generates and outputs a display image including a hand object that reflects the user's hand movements. This flowchart starts with the content processing device 200 having established communication with the head-mounted display 100 worn by the user, and acquiring frame data of the captured image, the content of the user's operation, and data related to the position and orientation of the user's head from the head-mounted display 100.

[0064] First, the state information acquisition unit 76 of the content processing device 200 starts acquiring state information of the user's hand based on the frame of the captured image (S10). If state information corresponding to the time step of the display image to be generated at that time is acquired directly using the captured image (Y in S12), the state information control unit 78 adopts the state information, and the 3D space control unit 82 reflects the state information in the hand object (S20).

[0065] If the state information corresponding to the time step of the display image to be generated has not been directly acquired using the captured image (N in S12), the state information control unit 78 determines whether or not to activate the state information prediction function (S14). That is, the state information control unit 78 determines whether or not to activate the prediction function according to the conditions set in the table above, based on the hand speed and the state information acquisition accuracy obtained up to that point. The state information control unit 78 may also determine whether or not to activate the prediction function based on only one of the hand speed or the state information acquisition accuracy.

[0066] If it is decided not to activate the prediction function (Y in S14), the state information control unit 78 adopts the most recent state information obtained using the captured image (S16). If it is decided to activate the prediction function (N in S14), the state information control unit 78 uses the state information up to that point to predict the state information corresponding to the time step of the display image to be generated (S18). The state information used for prediction may be limited to state information obtained directly from the captured image, or it may include state information that has been predicted up to that point.

[0067] In any case, the 3D space control unit 82 reflects the state information determined in S16 or S18 onto the hand object (S20). In parallel with this, the 3D space control unit 82 may reflect the results of information processing onto each object in the 3D space according to the request of the information processing unit 74. The display image generation unit 84 generates frame data for the display image by drawing the latest state object in the 3D space and outputs it sequentially to the head-mounted display 100 via the output unit 86 (S22).

[0068] If there is no need to stop the display due to content termination or user operation (N in S24), the content processing device 200 repeats the processing in S12 to S22 at a predetermined rate. As a result, the head-mounted display 100 displays a moving image including an object that reflects the user's hand movements. The frequency of the determination processing in S12 and S14 may be the same as the display frame rate or lower than the display frame rate. If it becomes necessary to stop the display, the content processing device 200 terminates all processing (Y in S24).

[0069] In this embodiment, the data on the accuracy of the state information acquired can be used not only to determine whether the state information prediction function is operating, but also in part of the content processing. Figure 10 is a diagram illustrating how the evaluation results of the state information accuracy are used in content processing. The figure shows that the content processing device 200 is composed of a system unit 92 that functions commonly regardless of the application, and an application unit 94 that processes the application program.

[0070] The system unit 92 includes the image acquisition unit 70, operation information acquisition unit 72, state information acquisition unit 76, state information control unit 78, display image generation unit 84, and output unit 86 shown in Figure 4, although some parts are omitted in the figure. The application unit 94 includes the information processing unit 74, the 3D space control unit 82, and the object data storage unit 80. However, the divisions shown in the figure are only examples.

[0071] As described above, the state information control unit 78 determines the latest state information as shown in Figure 9, based on the state information of the user's hand directly acquired from the captured image by the state information acquisition unit 76. This state information is supplied to the application unit 94 (S30), and the 3D space control unit 82 reflects it in the hand object in 3D space. The information processing unit 74 of the application unit 94 processes application programs such as electronic games, and the 3D space control unit 82 reflects the results in 3D space.

[0072] If the accuracy of the status information acquired by the status information acquisition unit 76 deteriorates during this process, the status information control unit 78 notifies the application unit 94 of this (S32). Here, deterioration in the accuracy of the status information includes failure to acquire status information and detection of abnormal values ​​in the status information. The status information acquisition unit 76 may also detect deterioration in accuracy if any of the above conditions a to g are met. In any case, the status information control unit 78 may also notify the application unit 94 of the basis for the determination of deterioration in accuracy. Even if the accuracy of the status information deteriorates, the status information control unit 78 may continue to determine the latest status information using the most recently obtained status information for a predetermined time and notify the application unit 94 of this.

[0073] The information processing unit 74 of the application unit 94 may perform various actions in response to notification of a deterioration in the accuracy of the status information. For example, if the cause of the deterioration in accuracy is that the distance of the hand to the camera 110 has fallen below the acceptable range for maintaining accuracy, the information processing unit 74 may request the display image generation unit 84 of the system unit 92 to warn the user that the hand is too close (S34). Similarly, if part of the hand is obscured or another hand is visible, the information processing unit 74 may request the user to be warned. The change in the display image requested by the application unit 94 is not limited to a warning to the user; it may also involve obscuring an object, etc.

[0074] Furthermore, the information processing unit 74 may decide whether or not to reflect the state information determined by the state information control unit 78 onto the object, according to criteria defined by the application. By giving the application the option to either not use state information supplied with low accuracy and keep the hand object in a fixed state, or to reflect the state information on the hand object even if it is of low accuracy, the movement of the hand object can be adapted to the circumstances and worldview of each content.

[0075] The status information control unit 78 also notifies the application unit 94 when the accuracy of the status information improves (S32). In practice, notification of deterioration or improvement in the accuracy of the status information may be performed by updating a flag stored in memory accessible to both the system unit 92 and the application unit 94. Alternatively, the system unit 92 may sequentially notify the application unit 94 of the accuracy data itself. In this case, the application unit 94 may change its information processing or change its requests to the system unit 92 according to the stage of accuracy. This allows for a more detailed and appropriate response to changes in accuracy.

[0076] Figure 11 illustrates typical processing timings in response to changes in the accuracy of state information. The first row of the figure shows the time change in the accuracy of the state information, the second row shows the on / off status of normal control of the state information, and the third row shows the on / off status of warning displays to the user, with the horizontal axis representing time. It should be noted that the "accuracy" of the state information is strictly an evaluation value of accuracy, and the specific value varies depending on the basis for evaluating accuracy. Therefore, this evaluation value may not only show a continuous change as shown in the figure, but may also show a discontinuous change between two values.

[0077] First, as long as the accuracy of the state information is equal to or greater than the threshold Th, the state information control unit 78 controls the state information as usual using the branch shown in Figure 9 and supplies the result to the application unit 94 (S40). The threshold for the accuracy of the state information that determines whether or not to activate the prediction function at this branch may be the same as the threshold Th shown in the figure, or it may be higher than the threshold Th.

[0078] At time T1, if the accuracy of the status information falls below the threshold Th, the status information control unit 78 notifies the application unit 94 of the deterioration in the accuracy of the status information. If the application unit 94 requests a warning in response, the display image generation unit 84 starts displaying a warning (S42). On the other hand, even if the accuracy of the status information falls below the threshold Th, the status information control unit 78 continues normal control of the status information and continues to supply the results to the application unit 94 (S40). At time T2, when a predetermined time has elapsed while the accuracy of the status information remains below the threshold Th, the status information control unit 78 determines that the acquisition of status information has failed and temporarily suspends normal control of the status information (S44).

[0079] At time T3, if the accuracy of the status information becomes equal to or greater than the threshold Th, the status information control unit 78 notifies the application unit 94 of the improvement in the accuracy of the status information. The display image generation unit 84 then stops displaying the warning (S46). The threshold for displaying the warning and the threshold for stopping the display may be the same or different, as shown in the figure. The stopping of the warning display may be performed in response to a request from the application unit 94, or it may be decided by the display image generation unit 84 itself.

[0080] Furthermore, at time T3, the state information control unit 78 resumes normal control of the state information (S48). Through the temporal control shown in the figure, even if the accuracy of the state information deteriorates, the processing by the application unit 94 is minimized, and the possibility of recovering accuracy in a short time while maintaining the object's state appropriately as much as possible is increased.

[0081] According to the embodiment described above, in an embodiment that reflects the movement of an object in the real world onto an object on display, the latest state information is predicted based on the acquired state information. As a result, even if the frequency of acquiring state information from captured images is lower than the display frame rate, the update frequency of the object's state can be matched to that frame rate. Consequently, the update frequency of the image of the object that reflects movement and the images of other objects are unified, and problems such as blurred images can be avoided.

[0082] On the other hand, considering the drawbacks of prediction, such as the amplification of errors and noise, and the resulting jitter in the object's image, the prediction function is stopped depending on the situation, such as the object's speed and the accuracy of acquiring state information. As a result, in synergy with ensuring sufficient time for acquiring state information, images including objects that reflect the object's movement can be displayed with high quality regardless of the situation.

[0083] Furthermore, information regarding the accuracy of state information is provided to the entity processing the content application. This allows for optimal countermeasures to be taken depending on the content, even if the accuracy of the state information deteriorates. For example, by providing options such as deciding whether or not to use state information with deteriorated accuracy, or warning the user to remove the cause of the deterioration, different responses can be taken depending on the accuracy and worldview required for the content. This allows for the optimal maintenance of the quality of objects that reflect the movement of the target object.

[0084] The present invention has been described above based on embodiments. The above embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their respective components and processing processes, and that such modifications also fall within the scope of the present invention.

[0085] This disclosure may include the following aspects: [Item 1] A display image generation device comprising a circuit configured as follows: The aforementioned circuit is Based on the image of the object captured by the imaging device during video recording, state information of the object is acquired. The state information to be used is determined by switching whether or not to operate on the aforementioned state information according to the passage of time, depending on the situation. Using the determined state information, a display image is generated that includes a virtual object reflecting the movement of the object. Display image generation device. [Item 2] The aforementioned circuit is The display image generating device described in item 1, which manipulates the state information in accordance with the passage of time when the aforementioned object is deemed to be in motion, according to predetermined conditions. [Item 3] The aforementioned circuit is The display image generation device described in item 1, which does not manipulate the state information in accordance with the passage of time when certain conditions are met under which the state information acquisition unit can be considered to have low accuracy of the state information. [Item 4] The aforementioned circuit is The display image generating device according to item 3, which determines that at least one of the following conditions for being considered to have low accuracy is not within a predetermined tolerance range: the speed of the object, the average value of the brightness of the captured image, or the distance of the object to the imaging device. [Item 5] The aforementioned circuit is The display image generating device according to item 3, which determines whether the conditions for low accuracy are met by evaluating the degree to which objects of the same type as the object are reflected in the captured image, the object is obscured, or the object extends beyond the field of view of the imaging device. [Item 6] The aforementioned circuit is The display image generation device described in item 1, which, after detecting that the acquisition of the state information based on the image of the object has failed, continues to use the most recently obtained state information to determine the state information to be adopted for a predetermined period of time. [Item 7] The aforementioned circuit is Further processing of the content application that defines the aforementioned display image, Information relating to the accuracy of the state information based on the image of the object is supplied to the application. A display image generation device according to item 1, which modifies the display image in accordance with a request from the application corresponding to the accuracy information. [Item 8] The aforementioned circuit is When certain conditions are met that indicate a deterioration in the accuracy of the aforementioned status information, the application is notified accordingly. The display image generating device described in item 7, which displays a warning to the user in accordance with a request from the application corresponding to the deterioration in accuracy. [Item 9] The display image generation device described in item 1, A head-mounted display that acquires and displays the data of the display image from the display image generation device, A content processing system, including a content processing system. [Item 10] Based on the image of the object captured by the imaging device during video recording, state information of the object is acquired. The state information to be used is determined by switching whether or not to operate on the aforementioned state information according to the passage of time, depending on the situation. Using the determined state information, a display image is generated that includes a virtual object reflecting the movement of the object. Display image generation method. [Item 11] The imaging device has a function to acquire state information of an object based on the image of the object in the image captured by the imaging device, The function determines which state information to use by switching whether or not to operate on the aforementioned state information according to the passage of time, depending on the situation. A function to generate a display image including a virtual object that reflects the movement of the object using the determined state information, A recording medium that stores a program to implement a computer. [Explanation of Symbols]

[0086] 70 Image acquisition unit, 72 Operation information acquisition unit, 74 Information processing unit, 76 State information acquisition unit, 78 State information control unit, 82 3D space control unit, 84 Display image generation unit, 100 Head-mounted display, 110 Camera, 200 Content processing unit, 222 CPU, 224 GPU, 226 Main memory.

Claims

1. A state information acquisition unit acquires state information of an object based on the image of the object in the image captured by the imaging device, A state information control unit determines which state information to use by switching whether or not to operate the aforementioned state information according to the passage of time, A display image generation unit generates a display image including a virtual object that reflects the movement of the object using the determined state information, A display image generation device characterized by having the following features.

2. The display image generation device according to claim 1, characterized in that the state information control unit manipulates the state information in accordance with the passage of time when predetermined conditions are met under which the object can be considered to be moving.

3. The display image generating apparatus according to claim 1 or 2, characterized in that the state information control unit does not manipulate the state information in accordance with the passage of time when it satisfies predetermined conditions that the state information acquired by the state information acquisition unit is deemed to have low accuracy.

4. The display image generating apparatus according to claim 3, characterized in that the state information control unit determines that at least one of the following conditions for low accuracy is that the speed of the object, the average value of the brightness of the captured image, or the distance of the object to the imaging device is not within a predetermined allowable range.

5. The display image generating apparatus according to claim 3, characterized in that the state information control unit determines whether or not the conditions for low accuracy are met by evaluating the degree to which an object of the same type as the object is reflected in the captured image, the object is obscured, or the object extends beyond the field of view of the imaging device.

6. The display image generation apparatus according to claim 1 or 2, characterized in that the state information control unit continues to determine which state information to adopt for a predetermined time after detecting that the state information acquisition unit has failed to acquire the state information, using the most recently obtained state information.

7. The system further includes an information processing unit that processes an application of content defining the aforementioned display image, The state information control unit supplies the information processing unit with information relating to the accuracy of the state information acquired by the state information acquisition unit. The display image generation device according to claim 1 or 2, characterized in that the display image generation unit modifies the display image in accordance with a request from the information processing unit corresponding to the information relating to the accuracy.

8. When the state information control unit meets predetermined conditions that indicate the accuracy of the state information has deteriorated, it notifies the information processing unit accordingly. The display image generation device according to claim 7, characterized in that the display image generation unit displays a warning to the user in accordance with a request from the information processing unit corresponding to the deterioration of accuracy.

9. A display image generation device according to claim 1 or 2, A head-mounted display that acquires and displays the data of the display image from the display image generation device, A content processing system characterized by including the following.

10. A step of acquiring state information of an object based on the image of the object in the image captured by the imaging device, The step of determining which state information to use by switching whether or not to operate the aforementioned state information according to the passage of time, Using the determined state information, a step of generating a display image including a virtual object that reflects the movement of the object, A method for generating a display image, characterized by including the following:

11. The imaging device has a function to acquire state information of an object based on the image of the object in the image captured by the imaging device, The function determines which state information to use by switching whether or not to operate on the aforementioned state information according to the passage of time, depending on the situation. A function to generate a display image including a virtual object that reflects the movement of the object using the determined state information, A computer program characterized by enabling a computer to implement the following.