Image processing device

JP2026142870APending Publication Date: 2026-09-08TECHLICO INC +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025030116
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2026-09-08

AI Technical Summary

Benefits of technology

【0032】 本発明によれば、ADLやIADLの向上を目的としたトレーニングシステムを提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026142870000001_ABST
    Figure 2026142870000001_ABST
Patent Text Reader

Abstract

This product provides an image processing device that offers training aimed at improving the user's ADL or IADL, and enables motion learning using virtual objects superimposed on the real world. [Solution] The image processing device 100 comprises a control device 110, a storage device 120, a depth-recognition camera 160, an inertial measurement unit 170, and an output device 140. The control device 110 presents tasks aimed at improving ADL / IADL stored in the storage device 120 to the user via the output device 140, and causes the user to memorize the actions to be performed. Furthermore, using the depth-recognition camera 160 and / or the inertial measurement unit 170, the device detects at least some of the physical movements performed by the user based on the tasks, and the determination means determines whether the user's movements are appropriate based on pre-set criteria. The determination result is fed back to the user using the output device 140, allowing the user to efficiently correct and learn the movements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus that superimposes a virtual object on the real world surrounding a user and performs operations after visually recognizing the superimposed object, and particularly relates to an image processing apparatus that presents tasks aiming at improving Activities of Daily Living (ADL) and Instrumental Activities of Daily Living (IADL), and provides training while determining the user's movements.

Background Art

[0002] In recent years, user training systems utilizing virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies have been developed. In particular, research on rehabilitation systems and work training systems utilizing these technologies has been advanced, and technologies for improving motor function and cognitive function by having users perform actions in a virtual environment have been proposed.

[0003] Patent Document 1 relates to a system for implementing rehabilitation for higher brain dysfunction. The present invention provides rehabilitation using virtual reality (VR), augmented reality (AR), or mixed reality (MR), and aims to record and manage the rehabilitation history as rehabilitation history information when a patient solves rehabilitation problems.

[0004] It is considered effective to train both motor function and higher cognitive function for the prevention of dementia. Patent Document 2 describes a training system for reducing dementia risk using augmented reality (AR). Patent Document 2 discloses a technology that provides a training menu to a user, displaying virtual objects in an AR environment and having the user select a virtual object that meets specific conditions. This technology aims to simultaneously train motor functions and higher-order cognitive functions by capturing the user's hand movements and determining whether they touched a virtual object that meets the conditions. Specifically, Patent Document 2 discloses the following configuration. Instruction Unit: Specifies the conditions for the virtual object that the user should interact with. Display Unit: Displays virtual objects that meet certain conditions and virtual objects that do not meet certain conditions in a predetermined three-dimensional space around the user. At the start of the training menu, the virtual objects are placed both within and outside the user's visible range. Camera Unit: Captures user movements and detects whether the hand has come into contact with a virtual object. Determination unit: Determines whether the user has touched a virtual object that meets the specified conditions. Furthermore, Patent Document 2 describes that by distributing the placement of virtual objects vertically, horizontally, and in other directions, and by placing objects in areas that the user cannot reach simply by extending their arms, it is possible to train the user's motor functions over a wider range. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] International public access number WO2020 / 152779 [Patent Document 2] Japanese Patent Publication No. 2019-076302 [Overview of the project] [Problems that the invention aims to solve]

[0006] However, these conventional technologies did not adequately consider the presentation of tasks aimed at improving users' ADL and IADL, or systems that support their performance. In other words, while conventional technologies had systems to support the performance of specific tasks, they were not specifically designed to improve ADL or IADL, and no mechanism was proposed to more comprehensively support the acquisition and performance of daily living activities.

[0007] Therefore, the present invention aims to provide a training system for improving ADL and IADL.

[0008] Patent Document 1 describes a system aimed at supporting rehabilitation for higher brain dysfunction, and is primarily composed of technology for recording and sharing the history of rehabilitation task completion. However, it does not provide an integrated series of training methods (memory → behavior → judgment → feedback) aimed at improving ADL / IADL, as is the case with the present invention.

[0009] Furthermore, Patent Document 2 describes providing training aimed at preventing dementia by utilizing an AR environment and having users select specific virtual objects. However, it does not mention improvements in specific daily living activities related to ADL and IADL. [Means for solving the problem]

[0010] To solve the above problems, the present invention has the following features. The present invention relates to an image processing apparatus that overlays virtual objects onto the real world surrounding the user, enabling the user to perform operations based on visual recognition. A task presentation means presents the user with tasks aimed at improving ADL (Activities of Daily Living) or IADL (Instrumental Activities of Daily Living), has the user memorize the actions to be performed according to the task, and encourages actions based on that memorization. Action recognition means for detecting at least some of the physical movements performed by the user based on the task, A determination means that determines whether the user's actions are appropriate actions for the task based on pre-set criteria, The system also includes a feedback means for providing feedback to the user regarding the results of the determination means.

[0011] Preferably, the task presentation means presents tasks that stimulate two or more levels among the basic, intermediate, and higher levels of the neuropsychological pyramid.

[0012] Preferably, the task that stimulates the basic level is to allow the user to view the virtual object superimposed on the real world and move at least a part of the body in order to improve the processing of visual information, spatial awareness, and motor function. The aforementioned task that stimulates the intermediate level is a task in which the user observes an example, memorizes the correct action, and then performs it, in order to improve learning and adaptive abilities. The tasks that stimulate the aforementioned higher levels of cognitive function are preferably tasks that enable the user to memorize procedures, perform planned actions, and learn through feedback in order to improve executive function.

[0013] Preferably, the motion recognition means is Using a depth-sensing camera, the position of the user's hand joints is tracked, and it is determined whether the user performed an action to grasp or release a virtual object. It is preferable to use an inertial measurement unit to detect the movement of the user's head or body and update the position of the virtual object according to the user's movements.

[0014] Preferably, the problem-presenting means presents the problem to the user using an avatar, text, voice, or visual effects.

[0015] Preferably, the feedback means preferably presents the determination result to the user by using at least one of visual feedback, audio feedback, tactile feedback, or physical feedback.

[0016] Preferably, the image processing apparatus is preferably capable of communicating with an external computer apparatus, and includes a communication means that transmits the user's field-of-view video to the external computer apparatus.

[0017] Preferably, the external computer apparatus is preferably capable of selecting a task to be presented to the user.

[0018] Preferably, the external computer apparatus preferably causes the image processing apparatus to display a hint.

[0019] Preferably, the task presenting means preferably reproduces a virtual object simulating the user's real-world environment in the virtual environment, and presents a task using the virtual object to the user.

[0020] Preferably, the task presenting means preferably causes the user to memorize an operation related to movement of a specified target object in the virtual environment, and causes the user to perform the task while actually moving.

[0021] Preferably, the determining means preferably determines whether a position of the target object moved and placed by the user is within a predetermined correct placement range.

[0022] Preferably, the determining means preferably analyzes an operation order of the user, and determines whether the target object has been moved in a correct order presented in the task.

[0023] Preferably, the motion recognition means recognizes a motion of the user moving while grasping the virtual object, the determining means preferably determines whether the user has transported the virtual object to a predetermined destination.

[0024] Preferably, the problem presentation means causes a virtual object of the target to fly into the user's field of view. It would be beneficial to present the user with a task that requires them to appropriately select the object.

[0025] Preferably, the determination means analyzes the relative distance between the world coordinates of the user's hand and the coordinates of the incoming object, and determines that the object has been selected when it comes within a certain distance.

[0026] Preferably, the task presentation means sets different types of incoming virtual objects and presents a task that requires the user to select only virtual objects of the specified type.

[0027] Preferably, the task presentation means instructs the user to select an object with either their right or left hand, and the determination means determines whether the user selected the object using the instructed hand.

[0028] Preferably, the motion recognition means analyzes hand joint data acquired by the depth recognition camera, determines whether the back of the hand or the palm is facing upward based on the wrist joint data, and determines whether the user's hand is the right hand or the left hand based on the thumb joint position according to the orientation of the hand.

[0029] Preferably, the task presentation means may, within the virtual environment, use the virtual objects to present tasks related to preparing meals, cleaning up, getting dressed, preparing for bathing, returning keys to their proper storage places, organizing shoes, correctly selecting specified items from a shopping list, cooking procedures, or using public transportation.

[0030] Furthermore, the present invention relates to an image processing apparatus that overlays virtual objects onto the real world surrounding the user, enabling the user to perform operations based on visual recognition. It comprises a storage device, a depth-sensing camera, an inertial measurement unit, a control device, and an output device. The control device is The process involves presenting tasks aimed at improving ADL (Activities of Daily Living) or IADL (Instrumental Activities of Daily Living), which are stored in the memory device, via the output device, and causing the user to remember the actions to be performed. A process to detect at least some of the body movements performed by the user based on the task using the depth recognition camera and / or the inertial measurement unit, A process to determine whether the user's actions are appropriate actions for the task based on pre-set criteria, The process of providing feedback of the determination result to the user using the output device is executed.

[0031] Furthermore, the present invention relates to a program to be executed by an image processing device that overlays virtual objects onto the real world surrounding the user, enabling the user to perform operations based on visual recognition. The image processing device comprises a storage device, a depth recognition camera, an inertial measurement unit, a control device, and an output device. The aforementioned image processing device, A task presentation step involves presenting tasks aimed at improving ADL (Activities of Daily Living) or IADL (Instrumental Activities of Daily Living), having the participant memorize the actions to be performed in accordance with the task, and encouraging actions based on that memorization. A motion recognition step that uses a depth-sensing camera and / or an inertial measurement unit to detect at least some of the body movements performed by the user based on the task, A determination step in which the user's actions are determined to be appropriate actions for the task based on pre-set criteria, Based on the result of the determination step, a feedback step is performed to provide feedback to the user. [Effects of the Invention]

[0032] According to the present invention, a training system aimed at improving ADL and IADL can be provided.

[0033] By presenting users with tasks to improve their ADL / IADL and having them perform actions corresponding to those tasks, the system can support the acquisition and improvement of daily living activities. The motion recognition system can detect the user's body movements and determine whether they are appropriate, providing accurate feedback and enhancing learning effectiveness. The feedback system allows users to immediately understand the correctness of their movements, improving the effectiveness of continuous training.

[0034] By presenting tasks that stimulate the foundational, intermediate, and higher levels of the neuropsychological pyramid, it is possible to comprehensively train memory, planning, and executive functions, rather than simply imitating actions. By simultaneously activating different areas of cognitive function, the effectiveness of improving ADL / IADL is maximized.

[0035] Strengthening foundational skills (visual, spatial awareness, and motor functions) improves the precision of daily activities. Strengthening the intermediate level (learning and adaptability) improves the ability to remember and properly execute procedures. Strengthening higher-level functions (executive functions) promotes planned behavior and learning through appropriate feedback, making improvements in ADL / IADL more effective.

[0036] By using a depth-sensing camera and inertial measurement unit, the system can detect the user's hand movements and actions with high precision, enabling more natural and intuitive interaction. This enables the adjustment of virtual object positions in response to user movements, providing a realistic experience in a mixed reality (MR) environment.

[0037] By using a variety of methods such as avatars, text, audio, and visual effects, it becomes possible to present appropriate tasks according to the user's level of understanding and learning style. This contributes to improving motivation and maintaining concentration.

[0038] By combining visual, auditory, tactile, and physical feedback, users can more intuitively understand their own actions. This encourages quick correction of incorrect actions and improves user learning.

[0039] By transmitting the user's visual field to an external computer, doctors and therapists can monitor and support the user in real time. This enables remote rehabilitation and allows for effective user learning through expert feedback.

[0040] Because experts can select the most suitable tasks based on the user's situation, personalized training becomes possible.

[0041] External computer devices can provide hints and support for task completion, which is expected to reduce user errors and improve learning effectiveness.

[0042] By recreating the user's real-world environment in a virtual space, training becomes possible that closely resembles real-life situations, thereby improving their ability to apply their knowledge in the real world.

[0043] By having users perform tasks while actually moving around, it becomes possible to train ADL / IADL movements that involve movement in the real world.

[0044] The system determines whether the user has placed an object in the appropriate location, encouraging the learning of tidying up and proper placement actions in daily life.

[0045] This system assesses whether tasks are being performed in the specified order and improves the ability to correctly remember and execute procedures.

[0046] By having users move objects while holding them and determining whether they placed them in the correct location, this system enables training to develop habits of carrying and organizing.

[0047] By training users to appropriately select objects flying into their field of view, we can improve their visual attention and reaction speed.

[0048] By analyzing the coordinates of the user's hand and the object, and determining whether the object was selected accurately, this system helps improve the precision of fine hand movements.

[0049] By selecting only objects of a specified type, classification and cognitive selection skills are improved.

[0050] Training to select an object with a designated hand helps improve executive function by making students aware of the difference between using their left and right hands.

[0051] By using a depth-sensing camera to analyze the shape of the hand and accurately distinguishing between the right and left hands, it becomes possible to train by using both hands independently.

[0052] By providing tasks that are tailored to diverse real-life situations, such as preparing and cleaning up meals, getting dressed, preparing for bathing, managing keys, organizing shoes, shopping, cooking, and using public transportation, we aim to improve a wide range of ADL / IADL.

[0053] The purpose, features, structure, operation, and effects of the present invention and its embodiments will become even clearer from the following detailed description in reference to the accompanying drawings. [Brief explanation of the drawing]

[0054] [Figure 1] This is a block diagram of the image processing device 100 according to the first embodiment. [Figure 2]This is an operation flow of the image processing device 100 according to the first embodiment. [Figure 3A] This is an example of a screen in an embodiment of the first embodiment. [Figure 3B] This is an example of a screen in an embodiment of the first embodiment. [Figure 3C] This is an example of a screen in an embodiment of the first embodiment. [Figure 3D] This is an example of a screen in an embodiment of the first embodiment. [Figure 3E] This is an example of a screen in an embodiment of the first embodiment. [Figure 3F] This is an example of a screen in an embodiment of the first embodiment. [Figure 3G] This is an example of a screen in an embodiment of the first embodiment. [Figure 3H] This is an example of a screen in an embodiment of the first embodiment. [Figure 3I] This is an example of a screen in an embodiment of the first embodiment. [Figure 3J] This is an example of a screen in an embodiment of the first embodiment. [Figure 3K] This is an example of a screen in an embodiment of the first embodiment. [Figure 4] This is the operation flow of the image processing device 100 according to the second embodiment. [Figure 5A] This is an example of a screen in an embodiment of the second embodiment. [Figure 5B] This is an example of a screen in an embodiment of the second embodiment. [Figure 6] This is a block diagram of the image processing apparatus 100 according to the third embodiment. [Figure 7] This is the operation flow of the image processing apparatus 100 according to the third embodiment. [Modes for carrying out the invention]

[0055] (First Embodiment) A first embodiment of the present invention relates to an image processing device 100 that allows a user to overlay virtual objects onto the real world surrounding them, enabling them to be visually recognized and manipulated. In particular, the image processing device 100 of the present invention aims to improve ADL (Activities of Daily Living) and IADL (Instrumental Activities of Daily Living), and facilitates the acquisition of daily living activities by allowing the user to manipulate virtual objects.

[0056] As shown in Figure 1, the image processing device 100 includes a computer with a control device 110, a storage device 120, an input device 130, an output device 140, and a communication device 150, as well as a depth-sensing camera 160 and an inertial measurement unit (IMU) 170. The image processing device 100 can be implemented as an HMD or smart glasses. In this embodiment, it will be described as an HMD, but it is not limited to that.

[0057] The control device 110 is a processor that controls the overall operation of the image processing device 100, and by executing programs stored in the storage device 120, it recognizes user movements, manages virtual objects, and provides feedback. While the content program 200 is running, the control device 110 appropriately executes the SLAM program 180 and the grip determination program 190 to determine the user's movement and hand movements.

[0058] The storage device 120 stores virtual environment data 210 and user operation history data 220. The storage device 120 also stores a content program 200, a SLAM program 180, and a grip determination program 190, and the control device 110 executes these to perform overall content analysis, user position estimation, and hand movement determination.

[0059] The input device 130 is a device that accepts user input. Since input can also be accepted through user input detected by the depth-sensing camera 160 and IMU 170, the depth-sensing camera 160 and IMU 170 may also be classified as input devices. These sensors detect the user's hand movements and walking motion, providing data for processing by the control device 110.

[0060] The output device 140 is a display and audio output device mounted on the HMD, providing the display of virtual objects and feedback to the user. The output device 140 displays the virtual environment in the user's field of view and presents an example of avatar movement. The output device 140 may be an optical see-through type or a video see-through type, as long as it displays the virtual environment in real space.

[0061] The communication device 150 is an interface for sending and receiving data with an external device via a network. In the first embodiment of the present invention, the communication device 150 is not essential as the HMD operates independently, but in other embodiments of the present invention, a configuration that connects to an external PC is also possible.

[0062] The depth-sensing camera 160 is used for 3D mapping of the environment and recognition of hand movements. In this invention, the depth camera (ToF) is used to acquire the three-dimensional coordinates of the environment and to appropriately position virtual objects. In addition, an IR camera is used to recognize the shape of the user's hand and to acquire the coordinates of each joint in real time.

[0063] The IMU170 detects the movement of the HMD and measures the user's walking movement in real time. Based on the data from the IMU170, the world coordinates of the HMD are updated and the coordinates of the user's hands are corrected to prevent positional drift caused by movement.

[0064] The SLAM program 180 is executed by the control unit 110, integrating data from the IMU 170 and the depth-sensing camera 160 to perform self-localization and environment mapping of the HMD. The SLAM program 180 calculates the amount of movement of the HMD and stably maintains the position of virtual objects.

[0065] The grip determination program 190 analyzes changes in the joint coordinates of the user's hand and determines whether the user is "grasp" or "release". This program analyzes the hand joint data acquired by the depth-sensing camera 160 and detects opening and closing movements of the fingers. It also records the movement trajectory of the user's hand and determines whether the object is correctly placed (OK) or incorrectly placed (NG).

[0066] The content program 200 is a program that controls the overall operation of the content executed by the user and is executed by the control device 110. This program displays a virtual environment in real space via the HMD's display 140, presents tasks to the user, and controls the presentation of examples by avatars, guidance on operating virtual objects, evaluation of the user's actions, and feedback.

[0067] The content program 200 has the function of managing scenarios within the virtual environment and presenting appropriate tasks to the user. When the user performs a task, a virtual avatar appears and demonstrates the correct actions. This program monitors the user's actions in real time and, in cooperation with the grip determination program 190, analyzes changes in the joint coordinates of the user's hand to determine whether the action is "grasping" or "releasing".

[0068] Furthermore, the content program 200 determines whether to proceed to the next step or prompt a retry based on the progress of the task. If the user correctly manipulates the virtual object, it provides visual and audio feedback and guides the user to the next task. On the other hand, if an error is detected, it provides appropriate instructions and supports the user in performing the correct action.

[0069] Content Program 200 allows users to experience actions similar to those in real life within a virtual environment placed in the real world, thereby facilitating the acquisition of ADL (Activities of Daily Living) and IADL (Instrumental Activities of Daily Living). This program improves the accuracy of users' action recognition and task performance, leading to more effective rehabilitation and training.

[0070] The virtual environment data 210 is data managed by the content program 200 and stores information such as the placement of objects in the virtual space, the layout of rooms, and task scenario information. This data includes placement information for virtual rooms (refrigerator, kitchen, trash can, washing machine, front door, etc.) and location information for virtual objects that the user interacts with (dish soap bottle, houseplant, trash, etc.). It also stores scenario data for each task, including information that defines a series of actions, such as "take milk from the refrigerator and place it in the kitchen" and "throw trash into the trash can."

[0071] The user operation history data 220 is data that records the actions the user takes when performing a task, and is managed by the content program 200. This data includes the user's task completion status, "grasp" and "release" action information recognized by the grip judgment program 190, and movement information detected by the IMU 170. In addition, the number of times the user performed the correct action and the history of NG judgments are also recorded, making it possible to evaluate the user's proficiency in actions and provide appropriate feedback.

[0072] By utilizing the virtual environment data 210 and user operation history data 220, the content program 200 can present appropriate tasks to the user and link actual actions with the operation of virtual objects. Furthermore, by analyzing the user operation history data 220, it becomes possible to provide customized content tailored to each user's proficiency level.

[0073] Based on the above, the detailed operation of the image processing device 100 while executing the content program 200 will be described below with reference to Figure 2. The content program 200 is executed by the control device 110 and manages the virtual environment data 210 and the user's operation history data 220, performing overall control such as displaying virtual objects, presenting user tasks, recognizing actions, and providing feedback.

[0074] First, the image processing device 100 uses the depth recognition camera 160 to scan the surrounding environment and acquire three-dimensional coordinates (S101). Based on this environment mapping, it refers to the virtual environment data 210 and places objects such as a refrigerator, kitchen, trash can, and washing machine in the virtual space (S102). It also executes the SLAM program 180 to determine the user's position and update the coordinates according to the movement of the HMD (S103).

[0075] The content program 200 refers to the virtual environment data 210 and loads a task scenario corresponding to the current user state (S104). Then, the content program 200 instructs the avatar to perform an exemplary action (S105), and the avatar demonstrates the steps the user should take (e.g., take milk from the refrigerator and place it in the kitchen). Note that the example may also show only the object moving, without the avatar actually moving.

[0076] Once the avatar's demonstration is complete, the content program 200 instructs the user to perform the task, and the user's task execution begins (S106).

[0077] When the user moves to perform a task, the IMU 170 detects the user's walking movement and records the amount of movement (S107). The SLAM program 180 updates the HMD's world coordinates based on the data from the IMU 170 and records the user's movement. In response to the user's movement, the program determines which area the user is in by referring to the virtual environment data 210.

[0078] Furthermore, the depth-sensing camera 160 tracks the hand's position, and the world coordinates of the hand are corrected by integrating the HMD's movement amount with the hand's position information (S108). This prevents the hand's coordinates from shifting due to movement and maintains an accurate positional relationship with virtual objects.

[0079] The image processing device 100 uses the depth recognition camera 160 to acquire the joint coordinates of the user's hand in real time and recognizes the shape of the hand (S109). The grip determination program 190 analyzes the changes in the joint coordinates of the hand and determines whether the action is grasping (grip) or releasing (release).

[0080] The system determines whether the joint coordinates of the hand are within the coordinate range of the virtual object, and if a closing motion of the fingers is recognized, it is determined that the object has "grabbed" (S110).

[0081] When the user moves the virtual object they are holding, the image processing unit 100 tracks the amount of movement of the HMD using the IMU 170 and the SLAM program 180 (S111). As the user moves, the world coordinates of the hand are also updated, and the virtual object follows the movement of the hand (S112).

[0082] It is also conceivable that the user will perform the task by moving only their hands without moving the HMD. In this case, the IMU170 does not correct for movement, and the world coordinates of the hands are calculated based on data from the depth-sensing camera 160, considering only the change in the relative coordinates of the hands relative to the HMD. Therefore, even when the HMD is not moving, the user's hand movements are correctly recognized and appropriate feedback is provided.

[0083] When the user reaches the designated placement location, the grip determination program 190 monitors changes in the joint coordinates of the hand, and if it detects that the fingers have been opened, it determines that the virtual object has been placed (S113).

[0084] The image processing device 100 determines whether the user has correctly placed the virtual object. The depth-sensing camera 160 obtains the world coordinates of the hand and determines whether the placed position is within a pre-set placement range (S114).

[0085] If the object is placed in the correct position, an "OK" judgment is made, and for example, the color of the virtual object changes as visual feedback, and an audio notification such as "Placed correctly" is given. On the other hand, if it is placed in the wrong position, an "NG" judgment is made, and an audio notification such as "Please place it in the designated location" is provided. Visual feedback is also acceptable.

[0086] This process is repeated for the next object until the placement of all objects has been determined (returning to S106). During this repetition, the image processing device 100 analyzes the user's sequence of actions and determines whether the objects were moved in the correct order presented in the task.

[0087] As an exception, if the user's hand goes outside the field of view of the depth-sensing camera 160, the image processing device 100 uses the data from the IMU 170 to update only the user's movement and correct the world coordinates of the hand (S115).

[0088] When the hand returns to the field of view, the world coordinates of the hand are corrected by adding the relative coordinates of the hand acquired by the depth-sensing camera 160 to the world coordinates of the HMD acquired by the IMU 170.

[0089] (Screen example of the first embodiment) The following describes a specific embodiment of the image processing device 100 with reference to Figures 3A to 3K. In this embodiment, the user is presented with the task of "throwing a bottle of dish soap into the trash can and opening the front door to leave the room," and the device shows the user performing the corresponding action in a virtual environment projected onto the real world.

[0090] Figure 3A shows the content program 200 explaining a task to the user. The output device 140 displays a virtual environment on top of a real-world room, and an avatar appears and instructs the user, "Throw the dish soap bottle in the trash can, open the front door, and leave the room." The instructions may be given by voice or displayed as text, as shown in Figure 3A.

[0091] Figure 3B shows how the avatar displays an arrow (hint arrow) to indicate the location of a dish soap bottle on a corresponding object. The user can use this hint to locate the object and initiate the appropriate action.

[0092] Figure 3C shows an avatar demonstrating the action of throwing a bottle of dish soap into a trash can. In this embodiment, the example is presented by having only the dish soap bottle move, while the avatar itself does not move. It is also possible to configure the avatar to move and throw the bottle into the trash can.

[0093] Figure 3D shows a scene where the avatar has completed its demonstration and the instruction "Please begin" is displayed or spoken aloud. Upon receiving this instruction, the user performs the task.

[0094] Figure 3E shows the user's hand approaching a bottle of dish soap. The image processing device 100 tracks the position of the user's hand using the depth recognition camera 160 and recognizes that the user's hand has approached the coordinate range of the virtual object.

[0095] Figure 3F shows a user grasping a bottle of dish soap and moving it. The grip detection program 190 detects that the joint coordinates of the hand are within the coordinate range of the bottle and that the fingers are closing, and determines that the user has "grabbed" the bottle.

[0096] As the user moves while holding the bottle, the HMD's movement is tracked using the IMU170 and SLAM program 180, and the hand's world coordinates are updated. This allows the virtual object to follow the user's hand movements as they move.

[0097] Figure 3G shows the user moving in front of the trash can and about to release the bottle. The image processing device 100 recognizes that the coordinates of the user's hand have reached the coordinate range of the trash can.

[0098] Figure 3H shows the bottle being placed in the trash can, which is considered an OK action. The grip detection program 190 detects the finger opening motion and determines that the bottle has been released. The depth recognition camera 160 is also used to confirm that the bottle is within the coordinate range of the trash can, and the program recognizes that the task has been completed correctly.

[0099] If the OK operation is completed, the content program 200 will, for example, change the color of the bottle as visual feedback and notify the user with the audio feedback "It was placed correctly."

[0100] Figure 3I shows a scenario where the user accidentally grabs a houseplant. The image processing device 100 recognizes that the joint coordinates of the hand are within the coordinate range of the houseplant, not the coordinates of the dish soap bottle.

[0101] Figure 3J shows a scenario where a user throws a houseplant into the trash can, which is an incorrect action. The depth recognition camera 160 is used to determine the type of object placed, and if an object different from the one specified in the task is placed in the trash can, an incorrect judgment is made.

[0102] If an "NG" (Not Good) judgment is made, the content program 200 will, for example, make the object glow red as visual feedback and notify the user with an audio message saying, "Please place the correct object."

[0103] Figure 3K shows a scene where a user is opening the front door and about to leave the room. This action is considered OK, and the image processing device 100 detects that the coordinates of the user's hand have reached the coordinate range of the door handle. The grip determination program 190 monitors the open / closed state of the hand, and if the door is opened properly, it determines that the task has been completed successfully.

[0104] Upon completion of an assignment, the content program 200 provides visual and audio feedback and guides the user to the next assignment.

[0105] As described above, Figures 3A to 3K sequentially show the process of task explanation by an avatar, instruction on the target object, presentation of exemplary actions, actual task execution by the user, detection of errors, and final completion judgment. The content program 200 provides appropriate instructions and supports the user in completing the task, while referring to the virtual environment data 210 and the user's operation history data 220.

[0106] By using the image processing device 100 of the present invention, users can learn actions necessary for daily life through a virtual environment and practice them repeatedly.

[0107] Although the image processing apparatus 100 of the present invention has been described assuming an MR (mixed reality) environment, it is not limited to this and can also be applied to system configurations that utilize other reality technologies such as VR (virtual reality), AR (augmented reality), and XR (cross-reality). The following describes the general operation of the present invention in each technological environment.

[0108] In an MR environment, virtual objects are placed on the coordinate axes of the real world, and users interact with these virtual objects while actually moving in the real space. In this invention, an HMD is used to overlay the real world and virtual objects, an avatar presents a task, and the user performs operations such as grasping and releasing. The user's movement is tracked by an IMU 170 and a SLAM program 180, and the position of the hands is corrected in real time by a depth-sensing camera 160.

[0109] In a VR environment, tasks are performed while fully immersed in a virtual space. Unlike an MR environment, there is no need to consider the coordinates of the real world, so the position of virtual objects is managed by the virtual environment data 210 within the HMD. In a VR environment, the following actions are performed:

[0110] First, the content program 200 loads the virtual environment, and the avatar demonstrates exemplary movements (Figures 3A-3D). When the user starts the task, the HMD's tracking function updates the user's movement and hand position in the virtual space (S107, S108). When the user performs an action to grasp an object using the controller, the grip determination program 190 analyzes the change in the controller and determines whether the user has "grabbed" or "released" the object (S109, S110). The program then determines whether the user has moved the object in the virtual space and placed it in the specified position, and performs an OK / NG judgment (S113, S114).

[0111] In an AR environment, virtual objects are superimposed onto the real world, allowing users to complete tasks within the real space. Unlike an MR environment, AR uses an external camera on the HMD to capture images of the real world and overlays virtual objects onto those images.

[0112] In the AR environment, similar to the MR environment, virtual objects are displayed on the HMD's display 140, and an avatar demonstrates (Figures 3A-3D). When the user's hand approaches an object, the depth recognition camera 160 detects the hand's coordinates, and the grip determination program 190 determines whether to "grasp" or "release" (S109, S110). The IMU 170 and SLAM program 180 are used to track the user's movement and correct the position of the virtual object (S107, S108).

[0113] In an AR environment, virtual objects are overlaid based on real-world images, so the position of the virtual objects depends on the HMD's viewpoint. Therefore, the relative position of the objects is corrected to move in conjunction with the HMD's movement.

[0114] XR is a concept that integrates MR, VR, and AR, and the image processing device 100 is applicable in an XR environment. In an XR environment, interaction is possible that combines elements of VR, MR, and AR according to the user's experience and purpose.

[0115] For example, after conducting basic training in a VR environment, users could perform tasks in an MR environment that interact with the real world, or receive feedback in an AR environment that combines it with the real world. This invention utilizes user operation history data 220 to record the progress of tasks performed in each environment and provides the user with optimal training.

[0116] The image processing device 100 according to the first embodiment is a system that enables the user to learn tasks by presenting actions using an avatar, and to improve ADL (Activities of Daily Living) and IADL (Instrumental Activities of Daily Living) by allowing the user to remember and reproduce those actions. In this invention, by adopting a training method that emphasizes the linkage between memory and action, the integrated strengthening of cognitive and motor functions is achieved, rather than merely imitation of actions.

[0117] Performing ADL / IADL requires the ability to integrate cognitive functions (memory, attention, and judgment) and motor functions (hand movements and mobility). The image processing device 100 of the first embodiment contributes to the improvement of ADL / IADL in the following ways.

[0118] First, the ability to plan and execute actions (motor planning) is enhanced. In the first embodiment, the avatar shows the correct sequence of actions, and the user improves their motor planning ability by following those steps. For example, by performing a series of actions in order, such as "taking milk from the refrigerator," "throwing away unwanted items in the trash can," and "leaving the room through the front door," the user can learn the appropriate sequence of actions.

[0119] Next, in this system, users are required to memorize the actions of their avatar and perform actions based on that memory. This stimulates the activation of working memory (short-term information retention and processing) and contributes to improved cognitive function. Furthermore, a function is incorporated that allows users to correct errors through trial and error, enabling them to develop the ability to select appropriate actions while receiving feedback.

[0120] Furthermore, error detection and feedback functions enable adaptive training to prevent users from learning incorrect actions. For example, if a user mistakenly manipulates the wrong object, the system immediately detects it as incorrect and provides feedback such as, "Please select the correct object." This improves the accuracy of user perception and actions, and enhances their ability to perform ADL / IADL.

[0121] The neuropsychological pyramid is a model that illustrates the hierarchical structure of brain function, where functions are integrated from the lower levels (basic levels) to the upper levels (higher-order levels). The training method of the first embodiment is characterized by simultaneously stimulating both these basic and higher-order levels.

[0122] The basic level includes visual-spatial perception and motor functions (hand movements, locomotion). In the first embodiment, the following effects are observed: (1) Processing visual information (viewing an example of an avatar) Observing the movements of avatars promotes the ability to process visual information and understand necessary actions. (2) Spatial perception and motor control When a user reaches for an object, the system accurately recognizes the position of the hand and the distance to the object, and performs the appropriate action. This is expected to improve motor function and enhance visual and spatial awareness.

[0123] Higher levels include working memory, executive functions (planning, judgment, and inhibition), and learning and adaptive abilities. In the first embodiment, the following effects are observed: (1) Activation of working memory (memorizing the avatar's actions) The user's working memory is activated because they need to short-term remember the avatar's actions and then execute them. (2) Strengthening of executive functions To successfully complete a task, you need to plan, follow procedures, and "grab" and "release" at the right time. This improves executive function (Planning & Execution) and makes it easier to learn daily living activities.

[0124] The correspondence with the neuropsychological pyramid can be summarized as follows: (1) Basic level: Sensory and motor functions: Processing of visual information, spatial awareness, motor functions (grasping, releasing, moving) (2) Intermediate level: Learning and adaptability: Observe the avatar's example, memorize and perform the correct actions. (3) Higher levels: Executive functions (planning, judgment, inhibition): memorization of procedures, execution of planned actions, learning through feedback Thus, in the first embodiment, by activating these functions simultaneously, the acquisition of ADL / IADL is efficiently supported.

[0125] The first embodiment features the simultaneous training of working memory and executive function by memorizing avatar movements and acting accordingly, stimulating intermediate and higher-level cognitive functions. It also stimulates basic-level cognitive functions by integrating visual information processing, spatial awareness, and motor control. Furthermore, error detection and feedback functions enable adaptive learning, promoting improvement in ADL / IADL.

[0126] The system of the present invention is characterized by simultaneously stimulating the basic, intermediate, and higher levels of the neuropsychological pyramid, and is expected to improve cognitive function more effectively than simple exercise training or memory training.

[0127] (Supplement to the first embodiment) The following provides supplementary information on possible modifications to the first embodiment. The image processing device 100 not only provides appropriate feedback when the user performs an incorrect action, but also has an adaptive learning function that re-executes the presentation of an example using an avatar if an NG judgment is made a certain number of times in a row.

[0128] For example, if a user mistakenly picks up a potted plant instead of a bottle of dish soap (Figure 3I), the image processing device 100 will determine it's an incorrect selection and provide voice feedback saying, "Please select the correct object." If the user makes a certain number of incorrect operations, the content program 200 will perform the avatar's exemplary actions again, allowing the user to reconfirm the correct actions. Through this adaptive learning, the user can learn to perform more accurate actions without repeating incorrect ones.

[0129] Furthermore, by utilizing user operation history data 220, if a particular malfunction occurs frequently, the content program 200 can provide supplementary explanations tailored to the type of malfunction. For example, it may add text such as "This object is not meant to be thrown in the trash," or take other appropriate action depending on the type of error.

[0130] The image processing device 100 has a function to optimize feedback according to the user's skill level. For example, if the user's success rate exceeds a certain level, it is possible to skip the presentation of exemplary actions by the avatar and switch to a mode in which the user directly performs the task.

[0131] Specifically, if a user correctly performs a specific task (e.g., throwing a bottle of dish soap into the trash can) a certain number of times, the content program 200 switches to a mode that simplifies or completely omits the avatar demonstration. In this case, the task's steps (Figure 3A) and initial instructions (Figure 3D) are still displayed, but the avatar demonstration, as shown in Figure 3C, is skipped, allowing the user to start the task directly.

[0132] On the other hand, if a certain number of incorrect attempts occur, the system can be configured to provide more detailed guidance. For example, if a user frequently makes mistakes with the "grab" action, a feature can be added to display the avatar's example action in slow motion to highlight the details of the movement. It is also possible to adjust the intensity of the feedback in stages to provide support tailored to the user's skill level.

[0133] The image processing device 100 is not limited to the task of throwing a dish soap bottle into a trash can (Figures 3A-3K), but can be applied to support various ADL / IADL actions. Examples of specific tasks are shown below. (Preparing the meal) Remove ingredients from the refrigerator and place them in the appropriate cooking area. • Hold a knife and perform the action of cutting the specified vegetables. Place the pot or frying pan in the appropriate place and add the ingredients. (Washing Instructions) Remove the laundry from the washing machine and transfer it to the basket. • Put the laundry into the dryer and start the cycle. Fold and store clothes. (Operation of electronic devices) • Properly operate the TV and air conditioner remotes and change settings. • Use a computer or tablet to open the specified app. • Simulate specific operations on a smartphone (e.g., making a phone call, sending a message).

[0134] This system supports the learning and acquisition of various daily living activities by loading task scenarios corresponding to these application examples while referring to virtual environment data 210. It can also utilize user operation history data 220 to provide appropriate tasks tailored to each user's proficiency level.

[0135] (Second embodiment) The image processing apparatus 100 according to the second embodiment of the present invention, like the first embodiment, overlays virtual objects onto the real world surrounding the user, allowing for visual recognition and subsequent manipulation. In this embodiment, the user does not move but remains in a fixed position, selecting virtual objects solely through hand movements.

[0136] As shown in Figure 1, the image processing device 100 includes a computer with a control device 110, a storage device 120, an input device 130, an output device 140, and a communication device 150, as well as a depth-sensing camera 160 and an inertial measurement unit (IMU) 170. This embodiment is also implemented with an HMD, but it may also be implemented with smart glasses or other devices, and is not limited to this embodiment.

[0137] The control device 110 is a processor that controls the overall operation of the image processing device 100. By executing the content program 200 stored in the storage device 120, it displays virtual objects, recognizes user actions, and provides feedback. It also executes the SLAM program 180 and the grip determination program 190 as needed.

[0138] The storage device 120 stores virtual environment data 210 and user operation history data 220. The content program 200 provides tasks according to the user's actions and controls the display of appropriate virtual objects based on the virtual environment data 210.

[0139] In this embodiment, the user primarily remains in a fixed position while performing actions, thus eliminating the need for significant movement compensation as in the first embodiment. However, the same method as in the first embodiment, including the IMU170, can be used to determine the world coordinates of the hand.

[0140] The IMU170 is a sensor that detects the movement of the HMD and corrects its world coordinates. If the user changes their sitting or standing position, or moves the HMD slightly, this information can be used to correct the world coordinates of the hands.

[0141] However, in this embodiment, it is also possible to determine the world coordinates of the hand using only the depth-sensing camera 160 without using the IMU 170. Specifically, by employing a method in which the depth-sensing camera 160 directly measures the relative coordinates of the hand and calculates the world coordinates based on the fixed coordinates of the virtual environment, the position of the hand can be recognized without using the IMU 170.

[0142] In this embodiment, the user does not need to perform a "grasp" action on the object; selection is completed simply by bringing the hand close to the object. Therefore, the "grasp" and "release" processes in the grip detection program 190 are unnecessary, and instead, logic that only detects hand proximity is applied. Specifically, the depth recognition camera 160 tracks the user's hand position in real time, and when the hand approaches a designated object within a certain distance, it is determined that the object has been selected.

[0143] The depth-sensing camera 160 tracks the user's hand position in real time and measures the relative distance to the object. It also detects when the hand approaches a designated object within a certain range. The virtual object's position information is maintained by the control device 110 using the SLAM program 180 and is appropriately adjusted according to the movement of the HMD and changes in the user's viewpoint.

[0144] The output device 140 uses the HMD's display to show the user the virtual environment and objects. It also provides audio feedback, giving appropriate responses when the user makes a correct or incorrect choice.

[0145] The content program 200 presents the user with a task and instructs them on the type of object that will fly towards them and how to select it (right hand, left hand, or either). It also records the user's operation history data 220 and manages the number of correct answers, incorrect answers, and combo count.

[0146] The detailed operation of the image processing apparatus 100 according to the second embodiment of the present invention will be described below with reference to Figure 4. In this embodiment, the user does not move and remains in a fixed position, selecting a virtual object that flies towards them using only hand movements.

[0147] First, the content program 200 references the virtual environment data 210 and loads the tasks to be presented to the user (S201).

[0148] Next, the content program 200 determines the type of object (e.g., fruit, animal, etc.) and its flight pattern, and sets the object's operation parameters (appearance position, flight speed, trajectory, etc.) (S202). Based on these settings, objects appear one after another within the set angular range and are controlled to fly in the specified direction. This angular range may extend beyond the user's view of the screen, or it may be limited to the area within the screen.

[0149] Next, the user is shown instructions on which object to select with which hand (S203). These instructions are given via text display or voice output. For example, they may be presented in the format of "Select the apple with your right hand" or "You can select the cat with either hand." In other words, there are different types of incoming virtual objects, and the user is presented with a task to select the virtual object of the specified type.

[0150] When the task begins, the object starts flying forward from a distance according to the flight pattern set in S202 (S204). The frequency of the object's appearance and its flight speed can be adjusted by the content program 200, and dynamic adjustments are made, for example, by shortening the appearance interval according to the user's skill level.

[0151] In parallel, the depth-sensing camera 160 tracks the user's hand position in real time and determines the hand's world coordinates (S205). When the IMU 170 is used, the HMD's world coordinates and the hand's relative coordinates are integrated to correct the hand's world coordinates. On the other hand, if the IMU 170 is not used, the depth-sensing camera 160 alone determines the hand's world coordinates.

[0152] To determine whether the user's hand is approaching an object in flight, the image processing device 100 calculates the relative distance between the hand's world coordinates and the object's coordinates, and determines that contact has occurred if the hand is within a certain distance (S206).

[0153] Furthermore, when the user touches an object, the content program 200 determines which hand selected the object (S207). The determination of the right hand and left hand will be explained in detail in another section of this specification.

[0154] Next, it is determined whether the correct object has been selected (S208). If the correct object is selected: The number of correct selections is counted, and audio feedback (such as "ping") is output. Alternatively, the object may be displayed in a way that causes it to disappear. If the wrong object is selected: The number of incorrect selections is counted, and audio feedback (such as "buzz") is output. Alternatively, the object may be displayed in a way that makes it bounce back. If the target object is missed: The number of misses is counted, and audio feedback (such as "B") is output. If you select the correct target consecutively: The combo count will be counted and a visual effect (e.g., an increase in the score multiplier) will be applied.

[0155] In this embodiment, even before the determination process in S208 is completed, the next object continues to appear according to the flight pattern set in S204 (S209). Therefore, multiple objects exist on the screen simultaneously, and their determinations are performed in parallel.

[0156] For example, while the user is selecting one object, new objects may appear in their field of view and continue flying towards the user. This requires the user to complete the task by repeatedly selecting objects at a consistent rhythm.

[0157] The content program 200 monitors the progress of the task and determines that the task is complete when a certain number of object selections are completed or when the set time limit is reached (S210).

[0158] If the game is deemed complete, the number of correct answers, incorrect answers, and combo counts for the user are tallied, and the results are output (S211). These results are recorded in the user's operation history data 220 by the content program 200 and used to adjust the difficulty level of the next challenge.

[0159] (Regarding the determination of right hand / left hand) The image processing device 100 has the function of recognizing the user's hand movements in real time and determining whether the hand that is in contact with the object is the right hand or the left hand. This determination method varies depending on the type of HMD used.

[0160] When the image processing device 100 of the present invention is used with Microsoft's HoloLens 2, the HMD comes standard with a hand tracking function. By utilizing this hand tracking function, the HMD's internal processing can determine whether the hand recognized by the depth-sensing camera 160 is the right hand or the left hand.

[0161] Therefore, in this invention, the hand tracking function of HoloLens 2 is utilized to distinguish between the right and left hands by comparing the data of the hand that has come into contact with an object with the judgment result of the HMD. For example, if the task is "select an apple with your right hand," only when the hand recognized by HoloLens 2 as the right hand touches the object will it be counted as a correct answer.

[0162] When the image processing device 100 uses an HMD that does not have a hand tracking function, the identification of the right and left hands is performed by a proprietary program. In the proprietary program of the present invention, the position of the thumb joint is analyzed based on the joint data of the hand acquired by the depth recognition camera 160, and the determination of whether it is the right or left hand is made.

[0163] Specifically, if the thumb joint is to the left of the index finger joint, the hand is determined to be a right hand; conversely, if the thumb joint is to the right of the index finger joint, the hand is determined to be a left hand. However, this determination assumes that the back of the hand is facing upwards.

[0164] However, if the palm is facing upwards, the direction of the thumb will be reversed, which can lead to misidentification. To correct this, the present invention analyzes the movement of the wrist joint to determine whether the back of the hand or the palm is facing upwards. If it is determined that the palm is facing upwards, the determination regarding the direction of the thumb joint is reversed, thereby enabling correct identification of the right and left hands.

[0165] (Screen example of the second embodiment) Hereinafter, with reference to Figures 5A and 5B, a specific example of the operation of the image processing device 100 in the second embodiment of the present invention will be described. In this embodiment, the user remains in a fixed position and selects incoming virtual objects using only hand movements.

[0166] Figure 4A shows how virtual objects appear in real space. The content program 200 refers to the virtual environment data 210 and makes the object appear within the user's field of view. In this embodiment, the object (e.g., a pear, a ball, an apple, a cat, a dog, etc.) is displayed as if it is flying from the back of the user's field of view toward the front.

[0167] Furthermore, as shown in Figure 4A, score information is displayed at the top of the screen. The score information includes the following elements: Number of correct answers: The number of times the user made the correct selection. Number of errors: The number of times the target object was missed. Number of incorrect answers: The number of times the wrong object was selected. Combo count: The number of times the correct target was selected consecutively.

[0168] Score information is updated in real time by content program 200, and the user's performance is reflected immediately.

[0169] Figure 4B shows the user selecting the correct apple with their right hand. The depth-sensing camera 160 acquires the world coordinates of the user's hand in real time and calculates the relative distance to the object. When the hand approaches within a certain distance, it is determined that the object has been selected.

[0170] In this embodiment, the user is instructed to "select the apple with your right hand," and in Figure 4B, the user touches the apple with their right hand as instructed. Therefore, the content program 200 increases the number of correct selections, applies visual effects, and outputs audio feedback (such as "ping").

[0171] The image processing apparatus 100 according to the second embodiment of the present invention allows the user to perform a task by selecting a virtual object that flies towards them while remaining in a fixed position. In this embodiment, the improvement of ADL (Activities of Daily Living) and IADL (Instrumental Activities of Daily Living) is promoted by coordinating visual recognition, judgment, and motor response through hand movements.

[0172] The key feature of this embodiment is that it allows for the simultaneous training of attention, reaction speed, spatial awareness, and executive function through the selection of objects. Furthermore, by recording the number of correct answers, incorrect answers, and combo counts, training tailored to the user's proficiency level becomes possible.

[0173] (Improvement of the ability to make quick decisions) In this embodiment, multiple objects fly in succession, and the user must select the appropriate object according to the instructions. This operation enhances the following capabilities: (1) Decision Making The ability to instantly identify which object is being instructed and select it using the appropriate hand will improve. Application examples: Quickly recognizing the location of a desired product while shopping, correctly selecting necessary cooking utensils in the kitchen. (2) Cognitive Inhibition This helps hone your ability to eliminate unnecessary options and avoid pursuing the wrong object. Application examples: Choosing the correct ticket on public transport, selecting the desired remote control from multiple options.

[0174] (Improvement of spatial awareness and hand motor control) Because the object flies from the back of the field of view towards the foreground, the user needs to correctly perceive the distance and movement of objects in space and adjust their hand movements accordingly. (1) Eye-Hand Coordination Through training to correctly move one's hands to the position of an object seen with the eyes, the precision of hand movements improves. Application examples: Safely handling tableware, placing medicines in the correct location. (2)Spatial Awareness Understanding the distance and speed of incoming objects and selecting them at the appropriate time improves one's ability to comprehend three-dimensional space. Application examples: Improving distance perception when parking, and making more precise movements when retrieving objects inside the house.

[0175] (Improved reaction speed and attention) Because the objects appear at unpredictable times, users must concentrate and react accordingly. (1) Reaction time It hones your ability to make quick decisions and act decisively. Application examples: Quickly catching falling objects, avoiding obstacles that suddenly appear. (2) Sustained Attention They are trained to concentrate for a certain period of time and react the moment an object flies towards them. Application examples: Managing timing during cooking, processing multiple pieces of information simultaneously when answering the phone.

[0176] The Neuropsychological Pyramid is a model that illustrates the hierarchical structure of brain function, where functions are integrated from the lower levels (basic levels) to the upper levels (higher levels). This embodiment is a training method that simultaneously stimulates both the basic and higher levels.

[0177] Neuropsychological pyramid levels: Function: Correspondence in the second embodiment Basic level: Sensory and motor functions: Vision-motor coordination, spatial awareness, hand movement accuracy Intermediate level: Learning and adaptability: Remembering instructions and selecting the correct object. Higher level: Executive function (planning, judgment, inhibition): Responding appropriately with the designated hand (right or left).

[0178] In a second embodiment of the present invention, the user performs a task by moving their hands at the appropriate timing based on visual information, thereby suppressing incorrect choices. This activates a wide range of functions, from basic levels (sensory and motor functions) to higher levels (executive functions), contributing to improvements in ADL / IADL.

[0179] The effect of training in the second embodiment of the present invention on the improvement of ADL / IADL is as follows: ADL / IADL capabilities: In accordance with the present invention Selection and judgment ability: The ability to identify the correct object and avoid making incorrect choices. Spatial awareness: The ability to correctly perceive the distance and direction of incoming objects. Reaction speed and attention: Respond quickly to instructions and move your hands correctly. Visual-motor coordination: Moving the hands correctly based on visual information.

[0180] Similar to the first embodiment, the image processing device 100 according to the second embodiment is applicable not only to MR (mixed reality) environments but also to other reality technologies such as VR (virtual reality), AR (augmented reality), and XR (cross-reality). However, in the second embodiment, the user operates from a fixed position without moving, so the method of application in each environment differs from that of the first embodiment. In a VR environment, the behavior of incoming objects is entirely dependent on the virtual environment data 210, and selection is completed by the user's hand movements. In an AR environment, virtual objects are overlaid on real-world images, and object selection is performed in combination with hand movements. In an XR environment, it is possible to construct scenarios that combine elements of MR / VR / AR to maximize the learning effect.

[0181] (Third embodiment) The image processing apparatus 100 according to the third embodiment of the present invention, like the first and second embodiments, overlays virtual objects onto the real world surrounding the user, allowing for visual recognition and subsequent operation. In this embodiment, the image processing apparatus 100 is connected to an external computer device 300, displaying the user's field of view externally and enabling the selection of content from the external computer device 300 to provide feedback.

[0182] As shown in Figure 6, the image processing device 100 includes a computer which comprises a control device 110, a storage device 120, an input device 130, an output device 140, and a communication device 150, as well as a depth-sensing camera 160, an inertial measurement unit (IMU) 170, and a field-of-view camera 180.

[0183] The communication device 150 is used to communicate with the external computer device 300, and is usually intended to use wireless communication, but wired connection is also possible. The communication device 150 streams the user's field of view in real time and transmits the video data to the external computer device 300.

[0184] The field-of-view camera 180 captures images of the real world as seen by the user. The control device 110 generates an image by superimposing the real-world image with virtual objects and transmits it in real time to the external computer device 300 via the communication device 150.

[0185] The external computer device 300 has the function of displaying the user's HMD field of view on a display. Furthermore, the external computer device 300 is equipped with a content selection interface, allowing therapists and doctors to select content to provide to the user. Selectable content includes tasks from the first embodiment (ADL / IADL training), tasks from the second embodiment (enhancing judgment through hand selection actions), or other tasks.

[0186] Furthermore, the external computer device 300 has a function to record the user's execution results, saving not only score information (number of correct answers, number of incorrect answers, number of combos) but also the user's eye movements.

[0187] Furthermore, the external computer device 300 can send instructions to the content program 200 to display hint arrows in order to assist the user in completing the task. Specifically, if the user does not arrive at the correct choice, the external computer device 300 sends instructions to the content program 200 via the communication device 150, and the content program 200 displays hint arrows on the HMD's display 140. This function allows the user to continue training with appropriate guidance even if they repeatedly make incorrect choices.

[0188] The detailed operation of the image processing apparatus 100 according to the third embodiment of the present invention will be described below with reference to Figure 7.

[0189] First, the image processing device 100 establishes a connection with the external computer device 300 via the communication device 150 (S301).

[0190] Next, the image processing device 100 refers to the virtual environment data 210 and constructs a virtual space. This determines the positions of the virtual objects displayed within the user's field of view (S302). After the virtual environment is established, the field of view camera 230 captures the user's field of view, generates an image with the virtual objects superimposed, and transmits it in real time to the external computer device 300 via the communication device 150.

[0191] The external computer device 300 displays the received video feed from the user's field of view on a display, allowing the user's operating status to be confirmed.

[0192] Next, the therapist or physician selects a task to be provided to the user using the interface of the external computer device 300 (S303). Selectable content includes tasks from the first embodiment (behavioral memory training), tasks from the second embodiment (selective touch training), or other tasks.

[0193] When content is selected, the external computer device 300 sends an instruction to start the task to the content program 200 via the communication device 150, and the content program 200 starts the task (S304).

[0194] When the user performs a task, the image processing device 100 detects the user's actions (selected object, hand movements, and success or failure of selection) and transmits them in real time to the external computer device 300 (S305).

[0195] The external computer device 300 records the user's score (number of correct answers, number of incorrect answers, time taken, etc. in the first embodiment, and number of correct answers, number of incorrect answers, number of mistakes, combo count, time taken, etc. in the second embodiment) and eye movements (S306). This allows therapists and doctors to analyze the user's eye trajectory and provide appropriate feedback.

[0196] The external computer device 300 monitors the user's task completion status and determines whether appropriate feedback is needed (S307).

[0197] If the user fails to arrive at the correct choice, the external computer device 300 sends an instruction to the content program 200 via the communication device 150 to display a hint arrow (S308). However, this action is not mandatory and is performed only when necessary.

[0198] Upon receiving this instruction, the content program 200 displays a hint arrow on the HMD's display 140 to assist the user in making a selection (S309).

[0199] The task ends when a certain number of tasks are completed or when the set time limit is reached (S310).

[0200] After completion, the user's final score and eye-tracking data are saved to the external computer device 300 (S311). This data can then be used for subsequent training analysis and evaluation of the user's learning progress.

[0201] The image processing device 100 according to the third embodiment of the present invention transmits the user's visual field to an external computer device 300 in real time, allowing a therapist or physician to monitor the user's progress on the task via the external computer device 300 and provide appropriate feedback. Therefore, in this embodiment, unlike standalone training, it is possible to improve ADL (Activities of Daily Living) and IADL (Instrumental Activities of Daily Living) while receiving adaptive support from an external source.

[0202] In this embodiment, the external computer device 300 can monitor the user's operation status in real time, select appropriate tasks, understand progress, and provide feedback. This configuration improves the adaptability of training and maximizes the user's learning effectiveness.

[0203] The external computer device 300 is equipped with a content selection interface, allowing therapists and doctors to select appropriate training content according to the user's proficiency level and task performance. This enables individualized optimization based on the user's characteristics and progress, resulting in more effective ADL / IADL training. Application examples: If the focus of ADL / IADL training is on acquiring household chore skills, select the first implementation method. Select the second embodiment if you want to enhance instantaneous judgment and reaction speed.

[0204] The external computer device 300 records the user's score (number of correct answers, number of incorrect answers, number of combos) and eye movements, and saves them as user behavior history data. This makes it possible to quantitatively analyze how the user directs their attention, the time taken to arrive at the correct answer, and the tendency of errors. Application examples: If a specific object is overlooked during task completion, the system analyzes eye-tracking data to provide feedback. The system measures the user's decision-making speed and evaluates their learning progress.

[0205] In this embodiment, if the user encounters difficulties in completing the task, the external computer device 300 sends instructions to the content program 200, which then displays hint arrows on the HMD's display 140 to provide appropriate guidance.

[0206] The external computer device 300 improves the learning effect by displaying hint arrows on the HMD's display 140 when the user repeatedly makes incorrect choices or is unable to reach the correct answer for a long time. Application examples: If the user is unable to correctly locate the target object, a hint arrow will indicate the object's location and guide them to the correct answer. If a particular task is too difficult, the difficulty level can be adjusted externally to help users learn at an appropriate level.

[0207] In this way, feedback from the external computer device 300 makes it possible to prevent incorrect learning and implement adaptive training.

[0208] In the third embodiment, real-time monitoring and feedback functions via an external computer device 300 enable more adaptive and individually optimized ADL / IADL training compared to training conducted alone. In particular, it contributes to improvements in attention, memory, judgment, visual perception, and motor planning abilities. (1) Improvement of visual perception and attention The field-of-view camera 230 analyzes the user's gaze data, allowing for training to prevent overlooking specific objects. Hint arrows can be used to assist users in properly recognizing objects. (2) Improvement of reaction speed and executive function It can record user scores and evaluate reaction speed and task performance. The difficulty level of the tasks can be dynamically adjusted to help improve the user's performance. (3) The link between memory and behavior By repeating tasks, specific actions can be made into habits. This can enhance the learning of correct actions and facilitate the correction of errors.

[0209] The following points should be added regarding the third embodiment. (How to utilize eye-tracking data) The external computer device 300 analyzes the gaze data and obtains the following information: (1) Analysis of user gaze concentration areas If a user's gaze is concentrated on a specific area, it's highly likely they are lost, and the task needs to be adjusted accordingly. For example, if a user frequently moves their gaze but cannot find an object, improving the object's placement and visibility can enhance the learning effect. (2) Analysis of the randomness of user gaze If the gaze does not follow a consistent pattern and moves randomly, it may be difficult to recognize objects. In this case, the external computer device 300 can adjust the difficulty level based on the gaze data and configure settings to display the target object more clearly. (3) Record the time taken to arrive at the correct answer. The system measures the time it takes users to arrive at the correct choice and adjusts the difficulty of the task accordingly. For example, if a participant takes an unusually long time to select the correct answer to a particular task, adaptive feedback can be provided in the next session, such as displaying hint arrows earlier.

[0210] In this way, by utilizing the analysis of eye-tracking data, it becomes possible to not only record scores but also to customize the training according to the user's learning progress, enabling more effective training.

[0211] (Specific relationship with ADL / IADL training) In a third embodiment of the present invention, real-time monitoring and feedback functions via an external computer device 300 are expected to improve the ability to perform tasks in daily life. (1) Training to perform tasks under the direction of a third party In this embodiment, since task selection and feedback are provided from an external computer device 300, the user can gain experience in performing tasks while receiving instructions. This contributes to improving the ability to act appropriately while receiving instructions from others in daily life, such as at work or at home. Application examples: This allows for training similar to the situation where a caregiver gives instructions to an elderly person for daily activities. It can also be used as vocational training, where trainees perform appropriate actions while receiving work instructions. (2) Sequential improvement of decision-making ability The user modifies their actions in real time while receiving feedback from the external computer device 300. This helps to train the executive function, which is responsible for recognizing and appropriately correcting mistakes. Application examples: This can enhance the process of selecting the necessary ingredients in the correct order when cooking. This allows you to practice selecting the correct items based on a shopping list. (3) Improvement of simultaneous processing capability (multitasking) The user needs to receive instructions from an external computer device 300 and respond to feedback in real time while completing the task. This enhances the ability to process multiple pieces of information simultaneously (Dual-Task Performance), contributing to smoother task execution in real life. Application examples: Properly follow navigation instructions while driving. While working, I process multiple tasks simultaneously and recognize the necessary information.

[0212] Thus, the collaboration with the external computer device 300 in the third embodiment is not merely score management, but also supports the improvement of cognitive and motor abilities in daily life by enhancing executive functions such as understanding instructions, making judgments, and modifying actions.

[0213] (Common technical framework of the first and second embodiments)

[0214] The image processing apparatus 100 of the present invention adopts a fundamentally common technical framework in both the first embodiment (with movement) and the second embodiment (without movement). That is, it supports the improvement of ADL (Activities of Daily Living) and IADL (Instrumental Activities of Daily Living) through a series of processes in which a task is presented to the user in a virtual environment (memory), the user performs an action corresponding to the task (action), it is determined whether the action is appropriate (determination), and the result is fed back (feedback).

[0215] This "memory → action → judgment → feedback" technical framework applies to both the first and second embodiments. However, the first embodiment differs in that the user performs the task while moving through real space (e.g., taking ingredients from the refrigerator and placing them in the kitchen), while the second embodiment differs in that the user remains in a fixed location and selects objects using only hand movements (e.g., selecting a flying object with a designated hand).

[0216] Specifically, when this technical framework is applied to the operational flows of the first and second embodiments, it can be summarized as follows.

[0217] (1) Memory: Presentation of tasks The image processing device 100 presents the user with a task to be performed and allows the user to memorize guidelines for action. Methods for presenting the task include presenting exemplary actions using an avatar (first embodiment) and providing visual and auditory instructions (second embodiment). The relevant step in the first embodiment: S104 (Load the assignment scenario) S105 (Avatar demonstrates the task) The relevant step in the second embodiment: S201 (Load assignment) S202 (Set the motion pattern of the target object) S203 (Present instructions for the task to the user)

[0218] (2) Action: User execution of an action The user manipulates an object according to the instructions of a memorized task. In this invention, the user's movements are detected using MR technology and analyzed as movement (first embodiment) or hand movements (second embodiment). The relevant step in the first embodiment: S107 (Detects user movement and updates world coordinates) S108 (Correct the world coordinates of the hand) S109 (Recognizes hand shape and acquires joint data) S110 (grabbing the object) S111 (Hand movement tracking) S112 (Moving while holding the object) S113 (Place the object after moving it) The relevant step in the second embodiment: S204 (Object begins flight) S205 (Determines the world coordinates of the hand) S206 (Measures the distance between the object and the hand) S207 (Recognizes a hand touching an object)

[0219] (3) Judgment: Evaluate the results of the task. The image processing apparatus 100 evaluates an action performed by a user and determines whether a task has been correctly completed. The determination result is based on the position of an object (first embodiment) and the type of the selected object (second embodiment). Corresponding steps in the first embodiment: S113 (determine whether the object has been released) S114 (determine whether the placement of the object is correct) Corresponding steps in the second embodiment: S208 (determine whether the selected object is correct)

[0220] (4) Feedback: Provision of determination results The image processing apparatus 100 provides appropriate feedback to the user based on the determination result. Feedback means include visual feedback (color change of an object, arrow display) and audio feedback ("Correct", "Incorrect"), etc. Corresponding steps in the first embodiment: S114 (provide visual and audio feedback) When a negative determination is made, the avatar presents the model action again Corresponding steps in the second embodiment: S208 (provide feedback: "piron" for correct, "bubu" for incorrect, reflection of the object) S208 (update of score count) S211 (display of results)

[0221] As described above, in the image processing apparatus 100 of the present invention, the flow of "storage → action → determination → feedback" constitutes a consistently applied technical framework. This framework supports the improvement of ADL / IADL by allowing a user to learn appropriate actions in an MR environment and perform repeated training, and is applicable to both the first embodiment (a task involving movement) and the second embodiment (a selection task in a fixed position). Furthermore, various types of ADL / IADL training can be implemented based on this technical framework

[0222] The following outlines examples of various training methods that can be implemented using the unified technical framework described above, further illustrating how ADL or IADL can be improved by using the image processing device 100 of the present invention.

[0223] (1) Basic ADL training By using the image processing device 100 of the present invention, it is possible to support users in acquiring activities of daily living (ADL). A specific example of ADL training is shown below.

[0224] (Meal preparation training) The image processing device 100 of the present invention can provide training related to meal preparation in order to support the improvement of the user's daily living activities. In this training, a refrigerator, a cooking counter, and food ingredients are placed in a virtual environment, and the user learns how to handle the ingredients in the correct order.

[0225] (Modified version of the first embodiment: with movement) In the first embodiment, the user practices taking ingredients from a refrigerator and placing them properly on a countertop while actually moving around. Specifically, this is performed in the following steps: (Memory: S104, S105) The content program 200 uses an avatar to present the user with the instruction, "Take the ingredients out of the refrigerator and place them on the counter." The avatar demonstrates the action, and the user memorizes that action. (Action: S107~S113) The user moves to the refrigerator in the virtual environment and opens the door. Reaching out to the virtual food objects, the user grasps the appropriate food item. The user carries the grasped food item to the cooking counter and places it there. (Judgment: S113, S114) The control device 110 determines whether the food items are placed in the correct positions based on the data from the depth recognition camera 160. An OK judgment is made if the items are placed in the correct positions, and an NG judgment is made if they are placed in the wrong positions. (Feedback: S114) If the result is OK, the ingredients are fixed in the correct position as visual feedback, and the user is notified with the audio feedback "Placed correctly". If the result is NG, the avatar demonstrates the correct action again, prompting the user to reconfirm the correct action.

[0226] (Modified version of the second embodiment: no movement, only hand movements) In the second embodiment, the user remains in a fixed position, and training is conducted by having the user correctly select ingredients that are displayed as virtual objects flying towards them and place them on the cooking surface (place them within reach from the fixed position; the same applies when applying the second embodiment hereafter). (Memory: S201, S202, S203) Content program 200 displays a refrigerator and a cooking counter on the screen, and the avatar instructs, "Catch the incoming ingredients correctly and place them on the cooking counter." (Actions: S204, S205, S206, S207) Virtual food objects fly into the user's field of view, are caught with the appropriate hand, and moved towards the cooking counter. (Judgment: S208) The control device 110 determines whether the ingredients are placed in the correct position on the cooking counter. If it is correct, an OK judgment is made; if it is placed in the wrong place, an NG judgment is made. (Feedback: S208, S211) If the result is OK, the ingredients are fixed in place as visual feedback, and the message "Correctly placed" is given as audio feedback. If the result is NG, the instruction "Place on the correct cooking surface" is given to prompt correction of the malfunction.

[0227] Thus, in the image processing device 100 of the present invention, meal preparation training can be performed in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to whether or not movement is involved, it is possible to learn the appropriate procedures for meal preparation in daily life.

[0228] (Tidying up training) The image processing device 100 of the present invention can provide training related to tidying up in order to support the improvement of the user's daily living activities. In this training, dishes and trash objects are placed in a virtual environment, and the user learns how to tidy up in the appropriate procedure.

[0229] (Modified version of the first embodiment: with movement) In the first embodiment, the user undergoes training in carrying dishes and trash to the appropriate locations while actually moving around. Specifically, this is carried out in the following steps. (Memory: S104, S105) The content program 200 uses an avatar to present instructions to the user, such as "Please take the dishes to the washing area" and "Please dispose of the trash properly." The avatar demonstrates the actions, and the user memorizes those actions. (Actions: S107~S113) The user reaches out and grabs a dish or trash object in the virtual environment. The user takes the grabbed dish or trash to the washing area or trash can. The user performs a "release" action at the designated location, placing it in its designated position. (Judgment: S114) The control device 110 determines whether the dishes and trash are placed in the correct location based on the data from the depth recognition camera 160. An OK judgment is made if they are placed in the correct location, and an NG judgment is made if they are placed in the wrong location. (Feedback: S114) If the result is OK, visual feedback will be provided by fixing the dishes and trash in the correct positions, and audio feedback will be provided saying "It has been cleaned up correctly." If the result is NG, the avatar will demonstrate again to prompt the user to reconfirm the correct action.

[0230] (Modified version of the second embodiment: no movement, only hand movements) In the second embodiment, the user remains in a fixed position and undergoes training by appropriately selecting and placing virtual objects of flying dishes and trash in the appropriate locations. (Memory: S201, S202, S203) The content program 200 displays a washing area and a trash bin on the screen, and an avatar instructs: "Please correctly catch the incoming tableware and place them in the washing area." (Action: S204, S205, S206, S207) Virtual objects of tableware and trash fly into the user's field of view, the user catches them with an appropriate hand, and moves the caught tableware and trash toward the washing area or the trash bin. (Judgment: S208) The control device 110 judges whether the tableware and trash are placed at the correct placement positions. An OK judgment is made when the placement is correct, and an NG judgment is made when the placement is at an incorrect position. (Feedback: S208, S211) In the case of an OK judgment, the tableware and trash are fixed at appropriate positions as visual feedback, and a notification "Placed correctly" is given as audio feedback. In the case of an NG judgment, an instruction "Please place in the correct washing area or trash bin" is presented to prompt correction of the misoperation.

[0231] As described above, in the image processing apparatus 100 of the present invention, tidying-up operation training can be implemented in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to the presence or absence of movement, it is possible to learn an appropriate procedure for tidying-up operations in daily life.

[0232] (Dressing Changing Training) Changing clothes is one of the basic activities of ADL (Activities of Daily Living), and includes appropriate clothing selection, the order of putting on and taking off, and storage. The image processing apparatus 100 of the present invention provides a virtual training environment for a user to learn a series of dressing changing operations.

[0233] (Modification of the first embodiment: with movement) (Memory: S104, S105) The content program 200 arranges a virtual closet and storage space, and an avatar presents the instruction: "Please put on the jacket first, and then put on the pants." The user observes the movement of the avatar and memorizes the correct order. (Actions: S107~S113) The user moves to the closet and picks up the clothing displayed as a virtual object. The user puts on (removes) the clothing in the specified order and location on the body. (Judgment: S114) The control device 110 determines whether the user put on and took off the clothes in the order presented to it, and if it is correct, it makes an OK judgment, and if it is incorrect, it makes an NG judgment. (Feedback: S114) If OK is detected, you will be notified with "You have changed clothes correctly." If NG is detected, the avatar will demonstrate again with the message "Please check the order."

[0234] (Modified version of the second embodiment: no movement, only hand movements) (Memory: S201, S202, S203) Content program 200 displays a closet or storage space on the screen, and the avatar instructs, "Choose the clothes that are flying in correctly and put them on in the correct order." (Action: S204~S207) A clothing object flies into the user's field of view, is caught with the appropriate hand, and moved to a designated location on the body. (Judgment: S208) The control device 110 determines whether the user has grasped and put on the clothes in the correct order. (Feedback: S208, S211) If the result is OK, you will be notified, "You have changed clothes correctly." If the result is NG, you will be instructed, "Please check the order again."

[0235] (Bath preparation training) Preparing for a bath involves properly preparing items such as towels, shampoo, and body soap to facilitate a smooth bathing experience. The image processing device 100 of the present invention provides a virtual training environment for users to learn how to prepare for a bath.

[0236] (Modified version of the first embodiment: with movement) (Memory: S104, S105) Content program 200 sets up a virtual bathroom and storage space, and the avatar gives the instruction, "Please prepare towels, shampoo, and body soap." (Action: S107~S113) The user moves to the storage space and picks up the designated bathing supplies. They take the necessary items and place them in the appropriate locations in the bathroom. (Decision: S114) The control device 110 determines whether the user has placed the correct item in the correct position. (Feedback: S114) If the result is OK, you will be notified, "Bath preparations are complete." If the result is NG, the avatar will instruct you, "Please check the placement of the items."

[0237] (Modified version of the second embodiment: no movement, only hand movements) (Memory: S201, S202, S203) Content program 200 displays a storage space and a bathroom on the screen, and the avatar instructs, "Catch the incoming items correctly and place them in the appropriate locations." (Action: S204~S207) Bathing product objects fly into the user's field of view, catch them with the appropriate hand, and place the items in the designated location. (Judgment: S208) The control device 110 determines whether the user has placed the correct item in the correct place. (Feedback: S208, S211) If the result is OK, you will be notified, "Bath preparation is complete." If the result is NG, you will be instructed, "Please check the placement of items."

[0238] (Key management training) The image processing device 100 of the present invention can provide training, including key management, to support the improvement of the user's daily living activities. In this training, key and key hook objects are placed in a virtual environment, and the user learns how to properly manage the keys.

[0239] (Modified version of the first embodiment: with movement) In the first embodiment, the user undergoes training to return the key to the correct location while actually moving around. Specifically, this is done in the following steps:

[0240] (Memory: S104, S105) The content program 200 uses an avatar to present the user with the instruction, "When you get home, put the keys back on the key hook." The avatar demonstrates the action, and the user memorizes that action. (Actions: S107~S113) The user reaches out and grasps the key object in the virtual environment. After grasping the key, the user moves towards the key hook. At the location of the key hook, the user performs the action of "releasing" the key and placing it in the designated position. (Judgment: S114) The control device 110 determines whether the key is placed in the correct key hook position based on the data from the depth recognition camera 160. An OK judgment is made if the key is placed in the correct position, and an NG judgment is made if it is placed in the wrong position. (Feedback: S114) If OK is determined, the key is fixed as visual feedback and the user is notified with the audio feedback "Placed correctly". If NG is determined, the avatar demonstrates the correct action again to prompt the user to reconfirm the correct action.

[0241] (Modified version of the second embodiment: no movement, only hand movements) In the second embodiment, the user remains in a fixed position, receives the key as a flying virtual object, and performs training by placing it in the appropriate location. (Memory: S201, S202, S203) Program 200 for content displays a key hook on the screen, and the avatar instructs, "When the key flies towards you, catch it and hang it on the key hook." (Action: S204~S207) A virtual key object flies into the user's field of view. The user catches it with the appropriate hand and moves the caught key towards the key hook. (Judgment: S208) The control device 110 determines whether the key is placed in the correct key hook. If it is correct, an OK judgment is made; if it is placed in the wrong location, an NG judgment is made. (Feedback: S208, S211) If OK is determined, the key is secured to the key hook as visual feedback, and the message "Properly placed" is given as audio feedback. If NG is determined, the instruction "Please place it on the correct key hook" is given, prompting the user to correct the malfunction.

[0242] Thus, in the image processing device 100 of the present invention, key management training can be performed in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to whether or not movement is involved, it is possible to improve the ability to properly manage keys in daily life.

[0243] (Shoe organization training) Organizing shoes is an important skill for developing good habits of tidiness in daily life. The image processing device 100 of the present invention enables users to learn the skills to properly store and organize their shoes.

[0244] (Modified version of the first embodiment: with movement) (Memory: S104, S105) Content program 200 virtually recreates the entrance area, and the avatar instructs, "Please arrange the shoes you are wearing and store them in the designated place." (Action: S107~S113) The user finds scattered shoes in the virtual entrance hall and picks them up. They then arrange the shoes in the appropriate place or put them in the shoe rack. (Decision: S114) The control device 110 determines whether the shoe is in the correct position. (Feedback: S114) If the result is OK, you will be notified, "The shoes have been properly organized." If the result is NG, the avatar will instruct you, "Please check the orientation of the shoes."

[0245] (Modified version of the second embodiment: no movement, only hand movements) (Memory: S201, S202, S203) Content program 200 displays shoes and storage space on the screen, and the avatar instructs, "Catch the flying shoes correctly and store them." (Action: S204~S207) A shoe object flies into the user's field of view, catch it with the appropriate hand, and place the shoe in the designated position. (Decision: S208) The control device 110 determines whether the user has stored the correct shoes in the correct position. (Feedback: S208, S211) If the result is OK, you will be notified that "The shoes have been properly organized." If the result is NG, you will be instructed to "Please check the storage location."

[0246] (2) IADL training The image processing device 100 of the present invention can also be applied to training that supports the performance of more advanced activities of daily living (IADL: Instrumental Activities of Daily Living).

[0247] (Shopping simulation) The image processing device 100 of the present invention provides training to improve the user's ability to perform shopping tasks necessary in daily life. In this training, a supermarket is constructed in a virtual environment, enabling the user to correctly select specified products and perform actions based on a shopping list.

[0248] (Modified version of the first embodiment: with movement) In the first embodiment, the user trains by physically moving around in a virtual supermarket, searching for items, and placing them in a shopping cart. Specifically, this is done in the following steps: (Memory: S104, S105) The content program 200 displays a shopping list and instructs the user to "Select the specified items and put them in the cart." An avatar demonstrates the actions, and the user memorizes the list. (Action: S107~S113) The user moves to the shelves of a supermarket in the virtual environment. The user reaches for the items listed on the shopping list and picks up the appropriate items. The user carries the picked-up items towards the basket and places them in the basket. (Judgment: S113, S114) The control device 110 determines, based on the data from the depth recognition camera 160, whether the product selected by the user matches one of those listed. If the selection is correct, an OK judgment is made; if an incorrect product is placed in the cart, an NG judgment is made. (Feedback: S115) If the selection is OK, the color of the items in the cart changes as visual feedback, and an audio message "Correctly selected" is given. If the selection is NG, the avatar demonstrates the correct selection again and provides feedback saying "Please select the correct item."

[0249] (Modified version of the second embodiment: no movement, only hand movements) In the second embodiment, the user remains in a fixed position and undergoes training by correctly selecting items listed on a shopping list from among virtual items flying towards them and placing them in a basket. (Memory: S201, S202, S203) Content program 200 displays a shopping list in the user's field of view, and an avatar instructs the user to "select the specified items from the incoming products." (Action: S204~S207) Multiple virtual product objects fly into the user's field of view. The user selects the correct product with the appropriate hand. If the selected product matches the shopping list, the user moves towards the shopping cart and places it there. (Judgment: S208) The control device 110 determines whether the selected product matches the list. If it matches, an OK judgment is made; if an incorrect product is selected, an NG judgment is made. (Feedback: S208, S211) If the result is OK, the selected product is fixed as visual feedback, and the message "Correctly selected" is given as audio feedback. If the result is NG, the instruction "Please select an item from the list" is presented to prompt correction of the malfunction.

[0250] Thus, in the image processing device 100 of the present invention, shopping simulation can be performed in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to whether or not movement is involved, it is possible to support the acquisition of planned behavior based on a shopping list and improve cognitive and executive functions.

[0251] (Cooking task) The image processing apparatus 100 of the present invention provides training to improve cooking skills necessary for daily life. In this training, a kitchen, ingredients, and cooking utensils are placed in a virtual environment, and the user learns how to perform cooking actions in the appropriate procedure.

[0252] (Modified version of the first embodiment: with movement) In the first embodiment, the user undergoes training in handling and properly cooking ingredients in a kitchen while actually moving around. Specifically, this is carried out in the following steps: (Memory: S104, S105) The content program 200 uses an avatar to present the user with the instruction, "Cut the vegetables and put them in the pot." The avatar demonstrates the actions, and the user memorizes the cooking procedure.

[0253] (Actions: S107~S113) The user moves to the refrigerator and food shelves in the kitchen within the virtual environment. The user reaches for the designated food item and picks up the appropriate item. The user places the picked-up item on the cutting board and uses a knife to cut it to the appropriate size. The user carries the cut food item to a pot and places it in the designated cooking area. (Judgment: S114) The control device 110 determines, based on the data from the depth recognition camera 160, whether the ingredients have been cut to the appropriate size and correctly placed in the pot. An OK judgment is made if the processing was done correctly, and an NG judgment is made if the cooking was done using an incorrect procedure. (Feedback: S114) If the result is OK, the cooked ingredients are placed appropriately as visual feedback, and the voice feedback is "Cooking complete." If the result is NG, the avatar demonstrates again and provides feedback saying, "Please cook using the correct procedure."

[0254] (Modified version of the second embodiment: no movement, only hand movements) In the second embodiment, the user remains in a fixed position and performs training by appropriately handling incoming virtual food items and placing them in the cooking area. (Memory: S201, S202, S203) Content program 200 displays ingredients and a cooking area in the user's field of view, and an avatar instructs, "Cut the flying ingredients correctly and put them in the pot." (Actions: S204~S207) A virtual food object flies into the user's field of view and is caught with the appropriate hand. The caught food is moved towards the cutting board and cut at the designated position. The cut food is moved towards the pot and placed in the appropriate position. (Judgment: S208, S211) The control device 110 determines whether the ingredients have been cut to the appropriate size and placed in the correct cooking area. If correct, it is judged as OK; if incorrectly placed or unprocessed ingredients are found, it is judged as NG. (Feedback: S208, S211) If the result is OK, visual feedback indicates that cooking was successful and audio feedback indicates that cooking is complete. If the result is NG, the instruction "Please follow the correct cooking procedure" is displayed to prompt correction of the malfunction.

[0255] Thus, in the image processing device 100 of the present invention, the cooking task can be performed in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to whether or not movement is involved, it is possible to support the improvement of cooking skills in daily life and to strengthen cognitive and executive functions.

[0256] (Use of public transportation) The image processing device 100 of the present invention provides training to help users improve their ability to use public transportation appropriately. In this training, stations and bus stops are displayed in a virtual environment, and the user is given the experience of performing the appropriate boarding procedures and traveling to their destination.

[0257] (Modified version of the first embodiment: with movement) In the first embodiment, the user trains on a series of actions, from purchasing a ticket to passing through the ticket gate and sitting down, while actually traveling. (Memory: S104, S105) Program 200 for content presents the instructions, "Purchase a train ticket, pass through the ticket gate, and sit in a suitable seat." An avatar demonstrates the actions, and the user memorizes the procedure.

[0258] (Actions: S107~S114) The user moves to a virtual train station and heads to the ticket counter. Selects and purchases a ticket for the designated destination. Passes through the automatic ticket gate and moves to the platform. Boards the correct train and sits in the appropriate seat (perceiving oneself as being in the designated position and the ticket as being placed) (S113). (Judgment: S114) The control device 110 determines whether the type of ticket purchased by the user, the status of passing through the ticket gate, and the seat selection are correct. If they are correct, an OK judgment is made; if an incorrect procedure was performed, an NG judgment is made. (Feedback: S115) If the result is OK, you will be notified that "You have boarded the train correctly." If the result is NG, the avatar will demonstrate again and provide feedback saying, "Please select the correct ticket and sit in your assigned seat."

[0259] Thus, in the image processing device 100 of the present invention, training on the use of public transportation can be carried out in the first embodiment (with movement). By involving movement, it is possible to correctly acquire the procedures necessary for actually using public transportation and to improve the ability to make appropriate decisions.

[0260] (3) Content that enhances cognitive function The image processing device 100 of the present invention can also be applied to training aimed at improving memory, judgment, and execution functions.

[0261] (Continuous operation memory) The image processing device 100 of the present invention provides training to improve a user's ability to remember a series of actions in order and reproduce them appropriately. This training aims to reproduce a living space in a virtual environment and enable the user to accurately perform the actions presented to them.

[0262] (Modified version of the first embodiment: with movement) In the first embodiment, the user moves while following the instructions of the avatar and performing the specified actions in order. (Memory: S104, S105) Content program 200 presents the instruction, "Take the milk from the refrigerator, pour it into a glass, and place it on the table." The avatar demonstrates the action, and the user memorizes the sequence of actions. (Actions: S107~S113) The user moves to the refrigerator in the virtual environment, opens the door, and takes out the milk. The user moves to the table with the milk and picks up a glass. The user pours the milk into the glass and places it appropriately on the table. (Judgment: S113, S114) The control device 110 determines whether the user performed the operations in the correct order. If the operations are completed in the correct order, an OK judgment is made; if the order is changed or some operations are omitted, an NG judgment is made. (Feedback: S114) If the result is OK, you will be notified that "You have completed the procedure correctly." If the result is NG, the avatar will demonstrate again and provide feedback that "Please perform the actions in the correct order."

[0263] (Modified version of the second embodiment: no movement, only hand movements) In the second embodiment, the user remains in a fixed position and performs training by appropriately handling incoming virtual objects and acting in a sequential manner. (Memory: S201, S202, S203) Content program 200 displays milk, a glass, and a table in the user's field of view, and the avatar instructs, "Catch the incoming milk, pour it into the glass, and place it on the table." (Actions: S204~S207) A virtual object of milk flies into the user's field of view and is caught with the appropriate hand. The caught milk is moved towards a glass and poured at the designated position. The glass containing the milk is moved towards the table and placed in the appropriate position. (Judgment: S208, S211) The control device 110 determines whether the user has completed the operations in the correct order. If the operations were performed in the correct order, an OK judgment is made; if the order is incorrect or any operations are missing, an NG judgment is made. (Feedback: S208, S211) If the result is OK, you will be notified that "You have completed the procedure correctly." If the result is NG, you will be instructed to "Please perform the operations in the correct order" and prompted to correct the malfunction.

[0264] Thus, in the image processing apparatus 100 of the present invention, the continuous motion memory task can be performed in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to the presence or absence of movement, it is possible to support the improvement of the ability to perform tasks in daily life and to strengthen cognitive and executive functions.

[0265] (Spot the difference task) The image processing device 100 of the present invention provides training to improve the user's visual recognition and memory abilities. This training aims to enhance observational and cognitive abilities by having the user identify changes in the arrangement of objects in a virtual space.

[0266] (Modified version of the first embodiment: with movement) In the first embodiment, the user undergoes training to find and point out changes within the virtual environment while actually moving around. (Memory: S104, S105) Content program 200 presents the instruction, "The placement of objects in the virtual space has been changed. Please point out the changes." The avatar shows the state before the change, and the user remembers the placement. (Action: S107~S113) The user moves around in the virtual space to find the changes (S107, S108). If a change is found, the user points to it with their finger or moves their hand to select the changed area (S109, S110). (Judgment: S113, S114) The control device 110 determines whether the location pointed out by the user matches the changes. If it is correct, it is judged as OK; otherwise, it is judged as NG. (Feedback: S114) If the result is OK, you will be notified with "That's correct" and proceed to the next task. If the result is NG, the avatar will provide a hint and give feedback such as "Please carefully observe the changes."

[0267] (Modified version of the second embodiment: no movement, only hand movements) In the second embodiment, the user remains in a fixed location and performs training by comparing objects in the virtual space and pointing out the changes. (Memory: S201, S202, S203) Content program 200 presents the user with the virtual space before and after the changes and instructs them to "find the changes and point them out." (Actions: S204~S207) After displaying the state before the change, the state after the change is displayed (S204). The user visually identifies the changes and selects them with a specified action (pointing or touching) (S206, S207). The system recognizes the selected changes (S208). (Judgment: S208) The control device 110 determines whether the changes specified by the user are correct. If correct, it is judged as OK; if incorrect, it is judged as NG. (Feedback: S208, S211) If the result is OK, you will be notified with "That's correct" and you will proceed to the next task. If the result is NG, you will be given feedback such as "Please look for the changes again."

[0268] Thus, in the image processing device 100 of the present invention, the spot-the-difference task can be performed in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to the presence or absence of movement, it is possible to support the improvement of observational skills and cognitive functions in daily life and strengthen adaptability.

[0269] (4) Improvement of spatial awareness and environmental adaptability The image processing device 100 of the present invention can also be applied to training aimed at improving spatial recognition ability.

[0270] (Perception of distance training) The image processing device 100 of the present invention provides training to improve the user's ability to perceive depth. In this training, virtual objects placed at different distances are displayed, and the user's ability to appropriately distinguish between distances is evaluated.

[0271] (Modified version of the first embodiment: with movement) In the first embodiment, the user undergoes training to identify the depth perception of objects within their field of view while actually moving. (Memory: S104, S105, S106) The content program 200 presents the instruction, "Look at the objects in your field of view, select the furthest object, and move it to the designated position." The avatar demonstrates the action, and the user memorizes the identification criteria. (Actions: S107~S113) The user observes multiple objects within their field of view and compares their distances. They perform an action to select a distant object in the specified manner. The user then moves the selected object to its designated position. (Judgment: S113, S114) The control device 110 determines whether the object selected by the user is actually at the furthest position and whether it has been placed in the predetermined position. If it is correctly selected and placed, an OK judgment is made; otherwise, an NG judgment is made. (Feedback: S114) If the result is OK, you will be notified with "You have selected and placed the objects correctly" and you will proceed to the next task. If the result is NG, the avatar will provide hints and give feedback such as "Please double-check the distance and placement of the objects."

[0272] (Modified version of the second embodiment: no movement, only viewpoint manipulation) In the second embodiment, the user remains in a fixed position and undergoes training by moving their viewpoint to identify depth and select the appropriate object. (Memory: S201, S202, S203) The content program 200 displays objects placed at different distances within the user's field of view and instructs the user to "select the furthest object and place it in the designated position."

[0273] (Action: S204~S207) Objects placed at different depths are displayed in the user's field of view. The user visually identifies the objects and selects the object that they determine to be the furthest away. The selected object is placed in the designated position (S207). (Judgment: S208, S210) The control device 110 determines whether the object selected by the user is truly the furthest away and whether it is placed in the designated position. If correct, an OK judgment is made; if incorrect, an NG judgment is made. (Feedback: S208, S211) If the result is OK, you will be notified with "You have selected and placed correctly" and will proceed to the next task. If the result is NG, you will be given feedback saying, "Please double-check the perspective and placement."

[0274] In the image processing apparatus 100 of the present invention, depth perception training can be performed in both the first embodiment (with movement) and the second embodiment (without movement). By providing appropriate training according to whether or not there is movement, it is possible to support the improvement of spatial recognition ability, depth judgment ability, and visual judgment ability, and to acquire more effective depth perception ability.

[0275] (Variations of feedback methods) The image processing apparatus 100 of the present invention is equipped with feedback means that provide appropriate feedback based on the results of determining the user's actions. The feedback is intended to allow the user to immediately recognize whether they have performed the task correctly, and is provided using visual, auditory, tactile, and physical means. The specific forms of feedback are described below.

[0276] (1) Visual feedback The image processing device 100 of the present invention can provide visual feedback using the HMD's display 140. The visual feedback intuitively communicates the judgment result to the user by changing objects and effects in the virtual environment. (Change in effect) When a user correctly manipulates an object, effects such as the object glowing, enlarging, or rotating are applied. For example, in the first embodiment, when a dish is placed in the correct position, the dish glows white and a completion sign is displayed. In the second embodiment, when the correct object is selected, the score increases and visual effects such as the background temporarily glowing are applied. If an incorrect operation occurs, visual feedback is provided, such as the object flashing red, an X mark appearing, or an arrow being displayed along with an error sound.

[0277] (2) Voice feedback The image processing device 100 of the present invention is equipped with a speaker and can provide voice feedback of the judgment result. Voice feedback can be applied in a wide range of ways, from simple sound effects to detailed guidance by AI. For example, if a correct action is performed, a success sound such as "ping" or "chime" will play. If an incorrect action is performed, an error sound such as "buzz" or "bang" will play. The AI ​​avatar provides situation-appropriate guidance to the user. For example, if the user repeatedly receives a "NG" (Not Good) rating, the avatar encourages them by saying, "Let's try again slowly." For users who have difficulty with certain actions, it provides adaptive feedback such as, "Would you like the avatar's movements to be shown in slow motion next time?"

[0278] (3) Tactile feedback The image processing device 100 of the present invention can provide haptic feedback to more intuitively support the user's hand movements. For example, a device worn on the user's hand (such as a glove or smartwatch) vibrates according to the detection result. For instance, a short vibration occurs when the object is correctly grasped, while a strong vibration occurs in the case of an incorrect action. Alternatively, resistance feedback could be used. For example, when grasping a specific virtual object, the hand device could generate slight resistance to provide a physical tactile sensation. For instance, when a user grasps a ball in a virtual space, a slight squeeze would occur, replicating the feeling of holding it in their hand.

[0279] (4) Physical feedback The image processing device 100 of the present invention is capable of providing feedback that generates a physical response to the user's actions. This enables a more immersive training experience. For example, there are pressure changes caused by actuators. When a user's hand touches a virtual object, the device generates pressure to reproduce the feeling of touch. Alternatively, pneumatic feedback could be used. In this case, wind pressure or vibration is used to emphasize the physical presence of the object. For example, if a user presses a virtual switch on a fan, wind is actually generated, allowing them to intuitively feel that they have pressed the switch.

[0280] (5) Adaptive feedback The image processing device 100 of the present invention can analyze the user's operation history and provide adaptive feedback. For example, feedback can be adjusted according to the user's proficiency level. If a user has a high success rate, the frequency of hints from the avatar can be reduced, and tasks of higher difficulty can be presented. If a user consistently fails, the avatar can demonstrate the correct actions in slow motion to provide a more conducive learning environment. Furthermore, guidance based on the number of times an "NG" (fail) is detected could also be considered. For example, if an "NG" is detected twice in a row, a voice guide would say, "Let's try a little slower next time." If it's detected three or more times in a row, the avatar would say, "Let's check your hand movements again," and demonstrate the correct movements again.

[0281] Thus, the image processing device 100 of the present invention provides a variety of feedback means for user actions. By combining visual, auditory, tactile, physical, and adaptive feedback, users can perform tasks more intuitively, contributing to improvements in ADL / IADL. In particular, by introducing adaptive feedback, it becomes possible to provide an optimal learning environment tailored to each user's level of proficiency.

[0282] (Overview of the demonstration experiment) To verify the effects of training using the image processing apparatus of the present invention on activities of daily living (ADL), instrumental activities of daily living (IADL), and cognitive function, a demonstration experiment was conducted in Gunma Prefecture (Maebashi City and Shibukawa City) under confidentiality.

[0283] (Implementation conditions) Implementation period: October 28, 2024 - February 7, 2025 Training content: ADL / IADL training using a Mixed Reality (MR) environment for 30 minutes once a week (the content used is the tasks shown in the first and second embodiments). Number of participants: 33 (36 at the start, 3 dropped out) Target audience attributes: Gender: 22 women, 11 men (all participants are elderly) Cognitive function level: MMSE-J (Mini-Mental State Examination): Average 26.90±2.27 points (18 / 33 people equivalent to MCI with 27 points or less = 54%) MoCA-J (Montreal Cognitive Assessment): Average 23.81±3.29 points (22 / 33 people equivalent to MCI with 25 points or less = 66%)

[0284] (Evaluation indicators and analysis methods) In this experiment, the changes before and after training were statistically evaluated using the Wilcoxon signed-rank test with the following indicators. 1. Indicators: MMSE-J (Dementia Screening Test) MoCA-J (Mild Cognitive Impairment Screening Test) TMT-A (Trail Making Test-A) TMT-B (Trail Making Test-B) SDMT(Symbol Digit Modalities Test) BI (Barthel Idexndex) FAI (French Activities Index) 2. Indicators related to Activities of Daily Living (ADL / IADL): Cleaning and tidying, heavy lifting, going out, gardening, home and car maintenance (muscle strength, balance, exercise endurance) Hobbies: Reading (concentration, information processing ability) Travel (activity level, self-efficacy)

[0285] (Experimental results) Changes in cognitive function (1) MMSE-J Healthy individuals (score 28 or higher): 15 people (45%) before training → 21 people (63%) after training MCI equivalent (27 points or less): 18 participants (54%) before training → 12 participants (36%) after training Number of patients who improved from MCI to a healthy level: 8 (44%) Average score: 26.90 points → 27.63 points (+0.73 points) [Table 1] (2) MoCA-J Healthy individuals (score of 26 or higher): 11 people (33%) before training → 18 people (54%) after training MCI equivalent (25 points or less): 22 participants (66%) before training → 15 participants (45%) after training Number of patients who improved from MCI to a healthy level: 7 (32%) Average score: 23.81 points → 25.63 points (+1.82 points) [Table 2]

[0286] (Items showing statistically significant improvement) Cognitive function related MMSE-J total (p=0.0135) MMSE-J understanding (p=0.0003) MoCA-J total (p=0.0001) MoCA-J: Visuospatial cognition (p=0.0007), language (p=0.0201), delayed recall (p=0.0074) TMT-B total (p=0.0167) TMT-B time (p=0.0167), number of errors (p=0.0297) SDMT (p=0.0002) Activities of daily living (ADL / IADL) FAI (Activities of Daily Living Score) (p=0.0001) Cleaning and tidying (p=0.0147), heavy lifting (p=0.0008), going out (p=0.0436), gardening (p=0.0020), house and car maintenance (p=0.0001) Hobbies (p=0.0004), Reading (p=0.0001), Traveling (p=0.0001) [Table 3]

[0287] (MoCA-J's MCID (Minimum Clinically Meaningful Change)) The average score for MoCA-J improved from 23.81 points before training to 25.63 points (+1.82 points), showing a greater improvement than MCID (1.73 points). This represents a clinically meaningful improvement in cognitive function, demonstrating the effectiveness of training.

[0288] (Consideration and Interpretation) 1. Reduction in the proportion of MCI: In this experiment, 54-66% of participants initially had MCI-like symptoms, but this decreased to 36-45% after training. This suggests that cognitive function improvement training utilizing Mixed Reality (MR) technology contributed to preventing cognitive decline in actual elderly individuals. 2. Improvement of Working Memory and Visuospatial Cognitive Function: In the system of the present invention, a complex task combining cognitive tasks using three-dimensional space and physical movements (reaching, walking) was performed. Particularly significant improvements were observed in visuospatial cognitive function (MoCA-J visuospatial) and working memory (SDMT, TMT-B), which is considered to be an effect that leverages the characteristics of Mixed Reality technology. 3. Improved Memory Function: A significant improvement in the MoCA-J delayed recall score (p=0.0074) suggests that training had a positive effect on memory function. Improvements in information processing speed and working memory may have contributed to improved performance on memory tasks. 4. Improvement in Quality of Daily Living (Improvement in FAI): The significant improvement in FAI scores (p=0.0001) indicates that the training contributed to an expansion of the range of activities in real life and an improvement in self-efficacy. This suggests that the improvement was not merely limited to paper-based test results, but was reflected in actual daily living activities.

[0289] (Conclusion) Training using the image processing apparatus of the present invention has been demonstrated to be effective in improving cognitive function (particularly visuospatial cognition, working memory, and memory), improving ADL / IADL, and expanding the range of activities. In particular, the reduction in the proportion of subjects with MCI-equivalent conditions and the improvement exceeding the MoCA-J MCID score are clinically significant results. In the future, it will be necessary to apply this system to a wider range of subjects and conduct long-term efficacy verification.

[0290] Although the present invention has been described in detail above, the above description is merely illustrative in all respects and is not intended to limit its scope. Needless to say, various improvements and modifications can be made without departing from the scope of the present invention. Each constituent element of the invention disclosed herein shall stand as an independent and standalone invention. Inventions that combine each constituent element in any way shall also be included in the present invention. The specific expressions in this specification are merely illustrative, and the present invention shall also include conceptualizations of such illustrative expressions. [Industrial applicability]

[0291] The present invention is an image processing apparatus that is industrially applicable. [Explanation of Symbols]

[0292] 100 Image Processing Devices 110 Control device 120 Storage device 130 Input device 140 Output device 150 Communication devices 160 Depth-sensing camera 170 Inertial Measurement Unit 180 SLAM Program 190 Grip Judgment Program 200 Content Programs 210 Virtual Environment Data 220 Operation History Data 230 Field of View Camera 300 External Computer Devices

Claims

1. An image processing device that overlays virtual objects onto the real world surrounding the user, enabling the user to perform operations based on visual recognition, A task presentation means that presents the user with tasks aimed at improving ADL (Activities of Daily Living) or IADL (Instrumental Activities of Daily Living), has the user memorize the actions to be performed according to the task, and prompts the user to take action based on that memorization, Action recognition means for detecting at least some of the physical movements performed by the user based on the task, A determination means that determines whether the user's actions are appropriate actions for the task based on pre-set criteria, An image processing apparatus comprising a feedback means for providing feedback to the user on the results of the determination means.

2. The image processing apparatus according to claim 1, characterized in that the task presentation means presents tasks that stimulate two or more levels among the basic level, intermediate level, and higher level of the neuropsychological pyramid.

3. The task that stimulates the aforementioned basic level involves having the user view the virtual object superimposed on the real world and moving at least a part of their body in order to improve the processing of visual information, spatial recognition, and motor function. The task that stimulates the intermediate level is a task in which the user observes an example, memorizes the correct action, and then performs it, in order to improve learning and adaptive abilities. The image processing apparatus according to claim 2, characterized in that the task stimulating the higher level is a task that causes the user to memorize procedures, perform planned actions, and learn through feedback in order to improve executive function.

4. The aforementioned motion recognition means is Using a depth-sensing camera, the position of the user's hand joints is tracked, and it is determined whether the user performed an action to grasp or release a virtual object. The image processing apparatus according to claim 1, characterized in that it detects the movement of the user's head or body using an inertial measurement unit and updates the position of a virtual object in accordance with the user's movements.

5. The image processing apparatus according to claim 1, characterized in that the problem presentation means presents the problem to the user using an avatar, text, voice, or visual effects.

6. The image processing apparatus according to claim 1, characterized in that the feedback means presents the determination result to the user using at least one of visual feedback, audio feedback, tactile feedback, or physical feedback.

7. The image processing apparatus according to claim 1, characterized in that the image processing apparatus is capable of communicating with an external computer device and includes communication means for transmitting the user's field of view image to the external computer device.

8. The image processing apparatus according to claim 7, characterized in that the external computer device is capable of selecting a task to present to the user.

9. The image processing apparatus according to claim 7, characterized in that the external computer device causes the image processing apparatus to display a hint.

10. The image processing apparatus according to claim 1, wherein the problem presentation means reproduces a virtual object that mimics the user's real-world environment within the virtual environment, and presents the user with a problem using the virtual object.

11. The image processing apparatus according to claim 10, characterized in that the task presentation means causes the user to store actions related to the movement of a specified object within the virtual environment, and allows the user to perform the task while actually moving.

12. The image processing apparatus according to claim 11, characterized in that the determination means determines whether the position of the object moved and placed by the user is within a predetermined correct placement range.

13. The image processing apparatus according to claim 11, characterized in that the determination means analyzes the user's sequence of actions and determines whether the object was moved in the correct sequence presented in the task.

14. The motion recognition means recognizes the user's movement while holding the virtual object, The image processing apparatus according to claim 11, wherein the determination means determines whether the user has transported the virtual object to a predetermined destination.

15. The aforementioned problem presentation means causes a virtual object of the target to fly into the user's field of view, The image processing apparatus according to claim 1, characterized in that it presents the user with a task requiring them to appropriately select the object.

16. The image processing apparatus according to claim 15, characterized in that the determination means analyzes the relative distance between the world coordinates of the user's hand and the coordinates of the incoming object, and determines that the object has been selected when it comes within a certain distance.

17. The image processing apparatus according to claim 15, characterized in that the task presentation means sets different types of incoming virtual objects and presents a task that requires the user to select only virtual objects of the specified type.

18. The image processing apparatus according to claim 15, characterized in that the task presentation means instructs the user to select an object with either their right or left hand, and the determination means determines whether the user selected an object using the instructed hand.

19. The image processing apparatus according to claim 18, characterized in that the motion recognition means analyzes hand joint data acquired by a depth recognition camera, determines whether the back of the hand is facing upward or the palm is facing upward based on the wrist joint data, and determines whether the user's hand is a right hand or a left hand based on the thumb joint position according to the orientation of the hand.

20. The image processing apparatus according to claim 1, wherein the task presentation means presents, within the virtual environment, a task related to preparing a meal, a task related to cleaning up, a task related to getting dressed, a task related to preparing for a bath, a task related to putting keys back in their proper place, a task related to organizing shoes, a task related to presenting a shopping list and correctly selecting specified items, a task related to cooking procedures, or a task related to using public transportation.

21. An image processing device that overlays virtual objects onto the real world surrounding the user, enabling the user to perform operations based on visual recognition, It comprises a storage device, a depth-sensing camera, an inertial measurement unit, a control device, and an output device. The control device is The process involves presenting tasks stored in the memory device, aimed at improving ADL (Activities of Daily Living) or IADL (Instrumental Activities of Daily Living), via the output device, and causing the user to remember the actions to be performed. A process for detecting at least some of the body movements performed by the user based on the task, using the depth recognition camera and / or the inertial measurement unit, A process to determine whether the user's actions are appropriate actions for the task based on pre-set criteria, An image processing device that performs a process of feeding back the determination result to the user using the output device.

22. A program to be executed by an image processing device that overlays virtual objects onto the real world surrounding the user, enabling the user to perform operations based on visual recognition, The image processing device comprises a storage device, a depth recognition camera, an inertial measurement unit, a control device, and an output device. The aforementioned image processing device, A task presentation step involves presenting tasks aimed at improving ADL (Activities of Daily Living) or IADL (Instrumental Activities of Daily Living), having the user memorize the actions to be performed according to the task, and encouraging actions based on that memorization. A motion recognition step that uses a depth-sensing camera and / or an inertial measurement unit to detect at least some of the body movements performed by the user based on the task, A determination step in which the user's actions are determined to be appropriate actions for the task based on pre-set criteria, A program for performing a feedback step that provides feedback to the user based on the result of the judgment step.

Citation Information

Patent Citations

  • Ar device, method, and program

    JP2019076302A

  • Rehabilitation system and image processing device for higher brain dysfunction

    WO2020152779A1