Information processing device, learning device, information processing method, learning data generation method, and recording medium

The information processing device addresses the limitations of arrow-based guidance by generating and outputting detailed textual instructions to guide users to a desired posture, effectively conveying complex movements and maintaining posture without requiring direct visual attention.

WO2026042618A1PCT designated stage Publication Date: 2026-02-26NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/028203
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-21
Filing Date
2025-08-07
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing guidance methods using arrows to indicate movements are insufficient for conveying complex posture adjustments, particularly three-dimensional movements or combinations of multiple movements, and require the user to gaze at a display, making it difficult to maintain a predetermined posture.

Method used

An information processing device that acquires reference and target posture information, generates difference information, and outputs instructional text to guide the user towards the desired posture, using a combination of hardware and software components to determine and convey detailed movement instructions.

Benefits of technology

Enables the conveyance of complex posture adjustments through textual instructions, allowing users to maintain the desired posture without direct gaze at a display, expanding the range of conveyable content beyond what arrows can achieve.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025028203_26022026_PF_FP_ABST
    Figure JP2025028203_26022026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to the present disclosure comprises a reference orientation information acquisition unit, a guidance target orientation information acquisition unit, a difference information generation unit, and an output unit. The reference orientation information acquisition unit acquires reference orientation information indicating a reference orientation. The guidance target orientation information acquisition unit acquires guidance target orientation information indicating the orientation of a guidance target. The difference information generation unit generates difference information indicating the difference between the reference orientation and the orientation of the guidance target. The output unit uses the difference information to output guidance text indicating, in text, guidance content for bringing the orientation of the guidance target closer to the reference orientation.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, learning device, information processing method, learning data generation method, and recording medium

[0001] The present disclosure relates to an information processing device, a learning device, a learning data generation device, an information processing method, a learning method, a learning data generation method, a recording medium, and a program.

[0002] There is a technique for providing guidance to guide the posture of a target object to a predetermined posture using posture estimation technology, etc. The technique disclosed in Patent Document 1 identifies a difference between the posture of a target object and a sample posture, and indicates, with arrows, a movement to eliminate the difference.

[0003] Japanese Patent Application Laid-Open No. 2018-5727

[0004] There are cases where guidance using arrows to indicate movements is insufficient. An example of a purpose of this disclosure is to realize a new guidance method in a technology for providing guidance to guide a target posture to a predetermined posture.

[0005] According to this disclosure, an information processing device is provided that has: a reference posture information acquisition means for acquiring reference posture information indicating a reference posture; a target posture information acquisition means for acquiring target posture information indicating the posture of a target; a difference information generation means for generating difference information indicating the difference between the reference posture and the posture of the target; and an output means for using the difference information to output a teaching text indicating, in text, teaching content for bringing the posture of the target closer to the reference posture.

[0006] This disclosure also provides an information processing method in which one or more computers acquire reference posture information indicating a reference posture, acquire training target posture information indicating the posture of a training target, generate difference information indicating the difference between the reference posture and the posture of the training target, and use the difference information to output a training text indicating, in text, training content for bringing the posture of the training target closer to the reference posture.

[0007] Furthermore, according to this disclosure, there is provided a program that causes a computer to function as: a reference posture information acquisition means that acquires reference posture information indicating a reference posture; a training target posture information acquisition means that acquires training target posture information that indicates the posture of a training target; a difference information generation means that generates difference information that indicates the difference between the reference posture and the posture of the training target; and an output means that uses the difference information to output a training text that indicates, in text form, training content to bring the posture of the training target closer to the reference posture.

[0008] Furthermore, according to this disclosure, a learning device is provided that has a learning means for learning a learning model using learning data that includes a pair of posture pair information indicating a pair of two postures and an instruction text indicating in text the instruction content for bringing one of the two postures closer to the other.

[0009] This disclosure also provides a learning method in which one or more computers learn a learning model using learning data including pairs of posture pair information indicating a pair of two postures and instruction text indicating in text the instruction content for bringing one of the two postures closer to the other.

[0010] Furthermore, according to this disclosure, a program is provided that causes a computer to function as a learning means for learning a learning model using learning data that includes a pair of posture pair information indicating a pair of two postures and an instruction text indicating in text the instruction content for bringing one of the two postures closer to the other.

[0011] Furthermore, according to this disclosure, a learning data generation device is provided, which includes: an instruction text selection means for selecting an instruction text indicating in text a combination of a part of the body to be taught, a direction in which to move that part, and an amount by which to move that part; and a posture pair information generation means for generating the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means determines a first posture, and then generates a second posture that is a posture reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture reached when the instruction indicated in the selected instruction text is applied to the first posture with the direction of movement reversed, thereby generating the posture pair information indicating a pair of the first posture and the second posture.

[0012] Furthermore, according to this disclosure, a learning data generation method is provided in which one or more computers select instruction texts indicating in text a combination of a part of the body to be taught, a direction in which to move that part, and an amount by which to move that part, and generate the posture pair information to be paired with the selected instruction text, and in generating the posture pair information, the method involves determining a first posture, and then generating a second posture that is a posture reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture reached when the instruction indicated in the selected instruction text is applied to the first posture with the direction of movement reversed, thereby generating the posture pair information indicating a pair of the first posture and the second posture.

[0013] Furthermore, according to this disclosure, a program is provided that causes a computer to function as: an instruction text selection means that selects an instruction text indicating in text a combination of a part of the body to be taught, a direction in which to move that part, and an amount by which to move that part; and a posture pair information generation means that generates the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means determines a first posture, and then generates a second posture that is a posture that is reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture that is reached when the instruction indicated in the selected instruction text is applied to the first posture with the direction of movement reversed, thereby generating the posture pair information that indicates a pair of the first posture and the second posture.

[0014] Furthermore, according to this disclosure, a learning data generation device is provided, which includes: an instruction text generation means for selecting an instruction text indicating, in text form, a combination of a body part of a training target and a reference state of the body part; and a posture pair information generation means for generating the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means generates posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture, by: determining a first posture in which the body part of the training target is not in the reference state, and then generating a second posture in which the body part of the training target is in the reference state in the first posture; or determining a third posture in which the body part of the training target is in the reference state, and then generating a fourth posture in which the body part of the training target is in a state different from the reference state in the third posture.

[0015] Furthermore, according to this disclosure, a learning data generation method is provided in which one or more computers select a training text indicating, in text form, a combination of a body part of a training target and a reference state of the body part, and generate the posture pair information to be paired with the selected training text, and in generating the posture pair information, the method involves determining a first posture in which the body part of the training target is not in the reference state, and then generating a second posture in which the body part of the training target is in the reference state in the first posture, or determining a third posture in which the body part of the training target is in the reference state, and then generating a fourth posture in which the body part of the training target is in a state different from the reference state in the third posture, thereby generating posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture.

[0016] Furthermore, according to this disclosure, a program is provided that causes a computer to function as: an instruction text generation means that selects an instruction text indicating, in text, a combination of a part of the body to be taught and a reference state of that part; and a posture pair information generation means that generates the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means determines a first posture in which the part of the body to be taught is not in the reference state, and then generates a second posture in which the part of the body to be taught is in the reference state in the first posture, or determines a third posture in which the part of the body to be taught is in the reference state, and then generates a fourth posture in which the part of the body to be taught is in a state different from the reference state in the third posture, thereby generating posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture.

[0017] Furthermore, according to this disclosure, there is provided an information processing device having: a reference posture information acquisition means for acquiring reference posture information indicating a plurality of time-series reference postures; a training target posture information acquisition means for acquiring training target posture information indicating a plurality of time-series postures of the training target; a difference information generation means for generating difference information indicating a difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training target; a selection means for selecting one from a plurality of pre-prepared instruction texts indicating instruction content; and an output means for outputting a pair of the posture of the training target and the reference posture reached when the instruction indicated in the selected instruction text is applied to the posture of the training target.

[0018] Furthermore, according to this disclosure, an information processing method is provided in which one or more computers acquire reference posture information indicating a plurality of time-series reference postures, acquire training target posture information indicating a plurality of time-series postures of a training target, generate difference information indicating a difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training target, select one from a plurality of pre-prepared training texts indicating training content, and output a pair of the posture of the training target and the reference posture that will be reached when the instruction indicated in the selected training text is applied to the posture of the training target.

[0019] Furthermore, according to this disclosure, there is provided a program that causes a computer to function as: a reference posture information acquisition means for acquiring reference posture information indicating a plurality of time-series reference postures; a training target posture information acquisition means for acquiring training target posture information indicating a plurality of time-series postures of the training target; a difference information generation means for generating difference information indicating the difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training target; a selection means for selecting one from a plurality of pre-prepared training texts indicating training content; and an output means for outputting a pair of the posture of the training target and the reference posture that is reached when the instruction indicated in the selected training text is applied to the posture of the training target.

[0020] According to one aspect of the present disclosure, a new teaching method is realized in a technique for teaching a subject to a predetermined posture.

[0021] 1 is a diagram showing an example of a functional block diagram of an information processing device; FIG. 2 is a flowchart showing an example of a processing flow of an information processing device; FIG. 3 is a diagram showing an example of a hardware configuration of an information processing device; FIG. 4 is an example of a data flow diagram of an information processing device; FIG. 5 is an example of a functional block diagram of a learning device; FIG. 6 is an example of a data flow diagram of a learning device; FIG. 7 is a diagram showing another example of a functional block diagram of a learning device; FIG. 8 is a diagram showing an example of a method for generating learning data; FIG. 9 is a diagram showing another example of a method for generating learning data; FIG. 10 is a diagram showing an example of a functional block diagram of a learning data generating device; FIG. 11 is another example of a functional block diagram of an information processing device; FIG. 12 is another example of a data flow diagram of an information processing device.

[0022] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In this disclosure, the drawings relate to one or more embodiments. In all drawings, similar components are designated by similar reference numerals, and descriptions thereof will be omitted as appropriate.

[0023] First Embodiment Fig. 1 is a functional block diagram showing an overview of an information processing device 10. Fig. 2 is a flowchart showing an example of the flow of processing executed by the information processing device 10.

[0024] 1, the information processing device 10 includes a reference posture information acquisition unit 11, a training target posture information acquisition unit 12, a difference information generation unit 13, and an output unit 14. These functional units execute the processing of the flowchart in FIG.

[0025] In S10, the reference posture information acquisition unit 11 acquires reference posture information indicating a reference posture. In S11, the training target posture information acquisition unit 12 acquires training target posture information indicating the posture of the training target. In S12, the difference information generation unit 13 generates difference information indicating the difference between the reference posture and the posture of the training target. In S13, the output unit 14 uses the difference information to output training text indicating, in text form, training content for bringing the posture of the training target closer to the reference posture.

[0026] The processing order of S10 and S11 is not limited to the example in Fig. 2. S10 may be performed after S11, or S10 and S11 may be performed in parallel.

[0027] In this way, the information processing device 10 outputs an instruction text that indicates, in "text," instruction content for bringing the posture of the instruction target closer to the reference posture. The information processing device 10 that outputs text indicating instruction content can realize instruction based on information similar to information conveyed orally by an instructor (person).

[0028] Furthermore, the information processing device 10 that outputs "text" indicating the instruction content can also output the text as "audio." Such an information processing device 10 can provide instruction similar to oral instruction by an instructor (person). In this case, the person receiving instruction does not need to direct their gaze toward a display or the like that displays the instruction content. If they need to direct their gaze toward the display, their face will inevitably turn toward the display, which can make it difficult to maintain a predetermined standard posture. The information processing device 10 that can convey the instruction content by audio can solve this problem.

[0029] Furthermore, when instruction content is conveyed using arrows as disclosed in Patent Document 1, the range of instruction content that can be conveyed is limited. When instruction content is conveyed using text as with the information processing device 10, the range of instruction content that can be conveyed is broadened. By using text, it is possible to convey movements that cannot be conveyed using arrows. For example, when instruction content is conveyed using arrows, it is difficult to convey complex instruction content such as three-dimensional movements or combinations of multiple movements. Specifically, it is difficult to convey instruction content such as "Twist your right upper arm 30 degrees to the right, and while keeping your left arm straight, raise your left arm 30 degrees using your shoulder as a fulcrum" using arrows. By conveying instruction content using text as with the information processing device 10, it is possible to convey such complex instruction content.

[0030] <<Second Embodiment>> <Overview> An information processing apparatus 10 according to a second embodiment is a specific implementation of the configuration of the information processing apparatus 10 according to the first embodiment. A detailed description will be given below.

[0031] <Hardware Configuration> First, an example of the hardware configuration of the information processing device 10 will be described. Each functional unit of the information processing device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various variations in the realization method and device. The software includes programs that are pre-stored in the device before shipping, and programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet.

[0032] FIG. 3 is a block diagram illustrating an example of the hardware configuration of an information processing device 10. As shown in FIG. 3, the information processing device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The information processing device 10 does not necessarily have to have the peripheral circuit 4A. Note that the information processing device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices may have the above hardware configuration.

[0033] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to mutually transmit and receive data. The processor 1A is, for example, a central processing unit (CPU) or a graphics processing unit (GPU). The memory 2A is, for example, a random access memory (RAM) or a read-only memory (ROM). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes an interface for connecting to a communication network such as the Internet. Examples of input devices include a keyboard, mouse, microphone, physical buttons, and touch panel. Examples of output devices include a display, projection device, speaker, printer, and mailer. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.

[0034] <Functional Configuration> Next, the functional configuration of the information processing device 10 will be described in detail with reference to FIGS. 1 and 4. FIG.

[0035] 1 is an example of a functional block diagram of an information processing device 10. As shown in the figure, the information processing device 10 includes a reference posture information acquisition unit 11, a training target posture information acquisition unit 12, a difference information generation unit 13, and an output unit 14.

[0036] FIG. 4 is a data flow diagram showing an example of the flow of data in the process executed by the information processing device 10. As shown in FIG.

[0037] The reference attitude information acquisition unit 11 acquires reference attitude information indicating the reference attitude.

[0038] The "reference posture" is a posture that serves as a reference. The information processing device 10 outputs information (instruction text) for bringing the posture of the training target, which will be described later, closer to the reference posture. The reference posture may be the posture of a person, a robot, a doll, an animal, or any other object.

[0039] "Reference posture information" is data indicating the reference posture as described above. The reference posture information can represent the reference posture based on the relative positional relationship of multiple key points in the body of the posture detection target taking the reference posture. The key points may be expressed in two dimensions (on a plane) or three dimensions (on a solid). Widely known techniques can be used to configure the reference posture information. The posture detection target may be a person, robot, doll, animal, other object, etc.

[0040] The reference attitude information acquisition unit 11 can execute, for example, one or more of the following methods 1 to 3 for acquiring reference attitude information.

[0041] "Reference posture information acquisition method 1" In this example, there is a reference posture detection object that serves as a reference (model). The reference posture detection object may be a person, a robot, a doll, an animal, or other object. In this example, the posture that the reference posture detection object actually takes becomes the reference posture. That is, the reference posture information acquisition unit 11 detects the posture that the reference posture detection object actually takes, and generates reference posture information that indicates the detected posture.

[0042] The reference posture information acquisition unit 11 can detect the posture of the reference posture detection target using a motion capture technique or the like, and generate reference posture information indicating the detected posture. Examples of the motion capture technique include, but are not limited to, the following methods: A method of tracking the positions of markers attached to the reference posture detection target using multiple cameras A method of using data from sensors (acceleration sensors, angular velocity sensors, geomagnetic sensors, etc.) attached to the reference posture detection target An image analysis method such as OpenPose or MMPose

[0043] The reference orientation detection target may take one reference orientation, and the reference orientation information acquisition unit 11 may detect the one orientation and generate reference orientation information indicating the one orientation.

[0044] Alternatively, the reference posture detection target may perform a "reference movement" consisting of a plurality of time-series reference postures. Then, the reference posture information acquisition unit 11 may detect each of the plurality of time-series reference postures and generate reference posture information indicating each of them.

[0045] Note that another device that is physically and / or logically separated from the information processing device 10 may generate reference posture information using a motion capture technique or the like and input the generated reference posture information to the information processing device 10. Then, the reference posture information acquisition unit 11 may acquire the reference posture information input from the other device.

[0046] "Method 2 for Acquiring Reference Attitude Information" In this example, reference attitude information indicating the reference attitude is generated in advance and stored in a predetermined storage device. The predetermined storage device may be provided within the information processing device 10, or may be provided within another device configured to be able to communicate with the information processing device 10. The same assumptions regarding the predetermined storage device apply hereinafter.

[0047] Then, the reference attitude information acquisition unit 11 acquires the reference attitude information stored in a predetermined storage device.

[0048] Note that the predetermined storage device may store reference attitude information indicating one reference attitude, and the reference attitude information acquisition unit 11 may then acquire that reference attitude information.

[0049] Alternatively, a predetermined storage device may store reference posture information indicating a plurality of reference postures. Then, the reference posture information acquisition unit 11 may acquire reference posture information indicating one of the plurality of reference postures. One of the plurality of reference postures may be designated by various methods. For example, a user may designate one of the plurality of reference postures. Alternatively, the reference posture information acquisition unit 11 may designate one of the plurality of reference postures according to a predetermined rule.

[0050] Alternatively, a predetermined storage device may store reference posture information indicating a plurality of time-series reference postures. The reference posture information acquisition unit 11 may then acquire the reference posture information indicating a plurality of time-series reference postures. The plurality of time-series reference postures indicates a reference movement.

[0051] "Reference posture information acquisition method 3" In this example, an image (still image or moving image) of a reference posture detection target to be used as a reference (model) is generated in advance and stored in a predetermined storage device.

[0052] The reference posture information acquisition unit 11 generates reference posture information by analyzing the image using an image analysis method such as OpenPose or MMPose to detect the posture of the reference posture detection target.

[0053] When the image is a still image, the reference orientation information acquisition unit 11 can detect one reference orientation shown in the still image and generate reference orientation information indicating that reference orientation.

[0054] When the image is a moving image, the reference posture information acquisition unit 11 can detect each of a plurality of time-series reference postures and generate reference posture information indicating each of the detected reference postures. The plurality of time-series reference postures indicate a reference movement.

[0055] Note that another device that is physically and / or logically separate from the information processing device 10 may analyze the image to generate reference posture information and input the generated reference posture information to the information processing device 10. Then, the reference posture information acquisition unit 11 may acquire the reference posture information input from the other device.

[0056] The training target posture information acquisition unit 12 acquires training target posture information indicating the posture of the training target.

[0057] The "training subject" is a subject to be trained to bring its posture closer to the reference posture. The training subject may be a person, a robot, a doll, an animal, or any other object.

[0058] "Training target posture information" is data indicating the posture of the training target as described above. The training target posture information can represent the posture of the training target using the same means as the reference posture information described above. For example, the training target posture information can represent the posture of the training target based on the relative positional relationship of multiple key points in the body or other part of the posture detection target assuming an arbitrary posture. The key points may be expressed in two dimensions (on a plane) or three dimensions (on a solid). Widely known techniques can be used to configure the training target posture information. The posture detection target may be a person, robot, doll, animal, other object, etc.

[0059] The training target posture information acquisition unit 12 can execute, for example, one or more of the following training target posture information acquisition methods 1 to 3.

[0060] "Method 1 for Acquiring Training Target Posture Information" In this example, the training target posture information acquisition unit 12 detects the posture that the training target actually assumes, and generates training target posture information indicating the detected posture. The training target posture information acquisition unit 12 can detect the posture of the training target using a motion capture technique or the like, and generate training target posture information indicating the detected posture.

[0061] The training target may take one posture, and the training target posture information acquisition unit 12 may detect the one posture and generate training target posture information indicating the one posture.

[0062] Alternatively, the training target may perform a movement consisting of a plurality of time-series postures, and the training target posture information acquisition unit 12 may detect each of the plurality of time-series postures and generate training target posture information indicating each of the detected postures.

[0063] Note that another device that is physically and / or logically separate from the information processing device 10 may generate training target posture information using a motion capture technique or the like and input the generated training target posture information to the information processing device 10. Then, the training target posture information acquisition unit 12 may acquire the training target posture information input from the other device.

[0064] "Method 2 for Acquiring Training Target Posture Information" In this example, training target posture information indicating the posture of a training target is generated in advance and stored in a predetermined storage device. Then, the training target posture information acquisition unit 12 acquires the training target posture information stored in the predetermined storage device.

[0065] Note that the predetermined storage device may store training target posture information indicating one posture of the training target, and the training target posture information acquisition unit 12 may acquire the training target posture information.

[0066] Alternatively, training target posture information indicating multiple postures of the training target may be stored in a predetermined storage device. Then, the training target posture information acquisition unit 12 may acquire training target posture information indicating one of the multiple postures of the training target. One of the multiple postures of the training target may be specified by various methods. For example, the user may specify one of the multiple postures of the training target. Alternatively, the training target posture information acquisition unit 12 may specify one of the multiple postures of the training target according to a predetermined rule.

[0067] Alternatively, training target posture information indicating a plurality of time-series postures of the training target may be stored in a predetermined storage device. Then, the training target posture information acquisition unit 12 may acquire the training target posture information indicating a plurality of time-series postures of the training target. The plurality of time-series postures of the training target indicate the movement of the training target.

[0068] "Method 3 for Acquiring Training Target Posture Information" In this example, an image (still image or moving image) of the training target is generated in advance and stored in a predetermined storage device.

[0069] The training target posture information acquisition unit 12 generates training target posture information by analyzing the image using an image analysis method such as OpenPose or MMPose to detect the posture of the training target.

[0070] When the image is a still image, the training target posture information acquisition unit 12 can detect one posture shown in the still image and generate training target posture information indicating that posture.

[0071] When the image is a moving image, the training target posture information acquisition unit 12 can detect each of a plurality of time-series postures and generate training target posture information indicating each of the detected postures. The plurality of time-series postures indicates the movement of the training target.

[0072] Note that another device that is physically and / or logically separate from the information processing device 10 may analyze the image to generate training target posture information and input the generated training target posture information to the information processing device 10. Then, the training target posture information acquiring unit 12 may acquire the training target posture information input from the other device.

[0073] The difference information generating unit 13 generates difference information indicating the difference between the reference posture and the posture of the training target. The difference information generating unit 13 includes a difference vector acquiring means shown in Fig. 4 and a posture encoder ("PoseEncoder" in Fig. 4).

[0074] The difference vector acquisition means generates a difference vector indicating the difference between the reference posture and the posture of the training subject.

[0075] The "difference vector" is calculated using information on key points that represent posture. For example, the difference vector can indicate at least one of a difference vector of the positions of key points and a difference vector of the angle of an angle formed by three key points. The angle of an angle formed by three key points represents, for example, the angle of a body joint.

[0076] The difference information generating unit 13 can calculate a difference vector between the position in the reference posture and the position in the posture of the training subject for each key point. The key points are identified according to the detected body parts, such as a head key point, a right shoulder key point, etc. The difference information generating unit 13 can calculate a difference vector of the above position for each of the various key points identified in this way.

[0077] Furthermore, the difference information generation unit 13 can calculate a difference vector between the angle in the reference posture and the angle in the posture of the training target for each angle (for each joint) formed by three key points. As described above, key points are distinguished from one another according to the detected body parts. Therefore, the above angles (joints) are also distinguished from one another according to the three key points that form each angle. The difference information generation unit 13 can calculate a difference vector of the above angle for each of the various angles thus distinguished.

[0078] After calculating the various difference vectors (the difference vectors of the positions of the various key points and the difference vectors of the angles of the various angles) as described above, the difference information generation unit 13 generates difference information indicating the difference between the reference posture and the posture of the training target based on the various calculated difference vectors. The difference information generation unit 13 calculates "feature values ​​of the difference vectors" as the difference information. The difference information generation unit 13 generates the difference information by inputting the difference vectors to a posture encoder ("PoseEncoder" in FIG. 4) generated in advance by machine learning and acquiring the feature values ​​of the difference vectors output from the posture encoder. The posture encoder is, for example, a DNN (Deep Neural Network).

[0079] The posture encoder is trained in advance to generate appropriate feature vectors for input difference vectors. The "appropriate feature vectors" generated for the input difference vectors have a relatively high similarity to the feature vectors of appropriate instruction texts, and a relatively low similarity to the feature vectors of inappropriate instruction texts.

[0080] The "feature amount of an appropriate instruction text" is a feature amount of an instruction text indicating an instruction content that eliminates (reduces) the difference between the reference posture indicated by the input difference vector and the posture of the instruction target.

[0081] The "feature amount of inappropriate instruction text" is a feature amount of an instruction text that does not indicate instruction content that eliminates (reduces) the difference between the reference posture indicated by the input difference vector and the posture of the instruction target.

[0082] There are various methods for training a posture encoder having such characteristics, but one example will be described in the following embodiment.

[0083] In addition, when the reference posture information and the training target posture information each indicate one posture, the difference information generation unit 13 can generate difference information for a pair of one reference posture indicated by the reference posture information and one posture of the training target indicated by the training target posture information.

[0084] Furthermore, when the reference posture information and the posture information of the target of training each indicate multiple time-series postures, the difference information generating unit 13 can time-synchronize the time-series data and generate difference information for pairs of reference postures and postures of the target of training that correspond to each other in timing.

[0085] There are various means for time synchronization. For example, the difference information generation unit 13 may time-synchronize the reference posture information and the training target posture information based on the elapsed time from the timing when the movement represented by the multiple time-series postures started. Alternatively, in a case where the reference posture information and the training target posture information are generated by detecting in real time the postures actually taken by the reference posture detection target and the training target, the difference information generation unit 13 may generate difference information for a pair of the latest reference posture and the latest posture of the training target. As described above, there are various time synchronization methods, and the difference information generation unit 13 can adopt a widely known time synchronization method. The time synchronization methods exemplified here are merely examples and are not limited to these.

[0086] The output unit 14 uses the difference information generated by the difference information generating unit 13 to output an instruction text showing, in text form, instruction content for bringing the posture of the instruction target closer to the reference posture.

[0087] The output unit 14 determines an instruction text to be output from among a plurality of instruction texts prepared in advance (N instruction texts in Figure 4) based on the similarity between the feature values ​​of each of the plurality of instruction texts prepared in advance and the feature values ​​of the difference vector indicated by the difference information generated by the difference information generation unit 13.

[0088] The "instruction text" indicates in text the instruction content for bringing the posture of the instruction target closer to the reference posture. One instruction text may indicate instruction for one body part. Alternatively, one instruction text may indicate instruction for multiple body parts. Furthermore, the instruction text may indicate instruction for one-dimensional movement, two-dimensional movement, or three-dimensional movement.

[0089] The instruction text may indicate, for example, "a combination of a part of the body to be taught, a direction to move that part, and an amount to move that part." Alternatively, the instruction text may indicate, in text form, "a combination of a part of the body to be taught and a state of that part." Note that the content of the instruction text is not limited to these.

[0090] The "combination of the part of the body to be taught, the direction of movement of that part, and the amount of movement of that part" indicates the content of the instruction as a change from the current state of the subject to be taught (a change based on the current state). An example of the instruction text is "Raise your right arm 30 degrees."

[0091] On the other hand, the "combination of the part of the training subject and the state of that part" indicates the training content in the final state that the training subject should be in. An example of the training text is "Bend your right elbow at a 90-degree angle."

[0092] A plurality of instruction text templates are prepared in advance (N instruction texts in FIG. 4). The templates may be generated by a user. Alternatively, a predetermined device may generate the templates. For example, the device may analyze audio data of a scene in which an instructor or the like is actually giving instruction to detect the content of the instruction that was actually given, and generate a template indicating the content of the instruction detected from the actual instruction.

[0093] The "features of instruction text" are generated by inputting the instruction text into a text encoder ("TextEncoder" in FIG. 4) that has been generated in advance by machine learning, and acquiring the features of the instruction text output from the text encoder. Each of the multiple instruction texts is input into the text encoder, and the features of each of the multiple instruction texts are generated. The text encoder is, for example, a DNN. The information processing device 10 or another device can generate the features of the instruction text in advance and store them in a predetermined storage device.

[0094] The text encoder is trained in advance to generate appropriate feature quantities for the input teaching text. The "appropriate feature quantities for the teaching text" generated for the input teaching text will have a relatively high similarity to the feature quantities of appropriate difference vectors and a relatively low similarity to the feature quantities of inappropriate difference vectors. The feature quantities of the difference vectors here are the feature quantities of the difference vectors generated by the difference information generator 13 described above.

[0095] The "appropriate difference vector feature" is a difference vector feature that eliminates (reduces) the difference between the reference posture indicated by the difference vector and the posture of the person being taught when the teaching content indicated in the input teaching text is applied.

[0096] The "features of an inappropriate difference vector" are the features of a difference vector in which the difference between the reference posture indicated by the difference vector and the posture of the person being taught is not eliminated (reduced) when the teaching content indicated in the input teaching text is applied.

[0097] There are various methods for training a text encoder having such characteristics, but one example will be described in the following embodiment.

[0098] The "similarity between the feature amounts of the instruction text and the feature amounts of the difference vector" can be calculated using widely known techniques. The output unit 14 can calculate, for example, the cosine similarity between the feature amounts of the instruction text and the feature amounts of the difference vector, or the difference between the feature amounts of the instruction text and the feature amounts of the difference vector.

[0099] The output unit 14 determines, as an instruction text to be output, an instruction text whose similarity to the feature quantity of the difference vector indicated by the difference information generated by the difference information generating unit 13 satisfies the output condition. The output unit 14 searches for an instruction text that satisfies the output condition from among a plurality of instruction text templates (N instruction texts in FIG. 4) prepared in advance, and determines the instruction text that satisfies the output condition as an instruction text to be output.

[0100] The output condition is one of the following: (Output condition 1) The similarity to the feature of the difference vector is the highest. (Output condition 2) The similarity to the feature of the difference vector is equal to or greater than a reference value. (Output condition 3) The similarity to the feature of the difference vector is included in a predetermined top number in the ranking sorted in descending order of similarity to the feature of the difference vector.

[0101] When output condition 1 is adopted, the number of instruction texts to be output is one, as shown in the example of Figure 4. When output condition 2 is adopted, the number of instruction texts to be output is not fixed, and may be 0, 1, or 2 or more. When output condition 3 is adopted, the number of instruction texts to be output is the "predetermined number" defined in output condition 3.

[0102] The output unit 14 may output a screen showing the instruction text via an output device such as a display or a projection device. Furthermore, the output unit 14 may output the instruction text aloud via a speaker. In this case, the output unit 14 may include a means for reading the text aloud. Alternatively, audio data of a plurality of prepared instruction texts read aloud may be generated in advance and stored in a predetermined storage device. The output unit 14 may then play back this audio data. Furthermore, the output unit 14 may be capable of both outputting the screen and outputting the audio as described above. The output device for outputting such instruction text may be included in the information processing device 10, may be external to the information processing device 10, or may be included in an external device configured to be able to communicate with the information processing device 10.

[0103] The output unit 14 may output the similarity between the feature of the instruction text and the feature of the difference vector as the reliability of the instruction text together with the instruction text that has been decided to be output. The output unit 14 may display the reliability on a screen or output it by voice.

[0104] Other configurations of the information processing apparatus 10 of the second embodiment are similar to those of the information processing apparatus 10 of the first embodiment.

[0105] <Effects> According to the information processing device 10 of the second embodiment, the same effects as those of the information processing device 10 of the first embodiment are achieved.

[0106] Furthermore, the information processing device 10 can output instruction text that indicates "a combination of a part of the body to be taught, a direction to move that part, and an amount to move that part" or "a combination of a part of the body to be taught and a state of that part" in text form. Such an information processing device 10 can provide instruction based on information similar to that conveyed orally by an instructor (person). This makes it possible to provide instruction that cannot be achieved by instruction using arrows as disclosed in Patent Document 1.

[0107] Furthermore, the information processing device 10 can determine the instruction text to be output by a characteristic method that uses the "difference vector acquisition means," "PoseEncoder," "TextEncoder," and "similarity calculation by the output unit 14" shown in Fig. 4. Such an information processing device 10 can generate and output, with high accuracy, an instruction text that can bring the posture of the instruction target closer to the reference posture.

[0108] Furthermore, the information processing device 10 can output instructional text that satisfies the above-described characteristic output conditions. According to such information processing device 10, it is possible to generate and output, with high accuracy, instructional text that can bring the posture of the instruction target closer to the reference posture.

[0109] Furthermore, the information processing device 10 can determine the instruction text to be output using a difference vector indicating at least one of the difference vector of the positions of the key points and the difference vector of the angle of the angle formed by the three key points. By determining the instruction text to be output based on such characteristic information, it is possible to generate and output with high accuracy an instruction text that can bring the posture of the target person closer to the reference posture.

[0110] <<Third Embodiment>> <Overview> A training device according to the third embodiment trains the pose encoder and text encoder described in the second embodiment. This will be described in detail below.

[0111] <Hardware Configuration> First, an example of the hardware configuration of a learning device will be described. Each functional unit of the learning device is realized by any combination of hardware and software. Those skilled in the art will understand that there are many variations in the implementation method and device. Software includes programs that are pre-loaded on the device before shipping, as well as programs downloaded from recording media such as CDs or servers on the Internet.

[0112] FIG. 3 is a block diagram illustrating the hardware configuration of a learning device. As shown in FIG. 3, the learning device has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The learning device does not need to have the peripheral circuit 4A. The learning device may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration. The processor 1A, memory 2A, input / output interface 3A, peripheral circuit 4A, and bus 5A are as described in the second embodiment.

[0113] <Functional Configuration> Next, the functional configuration of the learning device 20 will be described in detail with reference to FIG.

[0114] 5 is an example of a functional block diagram of the learning device 20. As shown in the figure, the learning device 20 includes a learning unit 21.

[0115] FIG. 6 is a data flow diagram showing an example of the flow of data in the processing executed by the learning device 20.

[0116] The learning unit 21 learns a learning model using learning data including pairs of posture pair information indicating a pair of two postures and instruction text indicating, in text form, instruction content for bringing one of the two postures closer to the other. The learning model learned by the learning unit 21 is the posture encoder and text encoder described in the second embodiment.

[0117] The training text included in the training data is correct training content for bringing one of the two postures indicated in the posture pair information paired with the training text closer to the other. The feature of such a training text can be called the "feature of an appropriate training text" in comparison with the feature of the difference vector between the two postures indicated in the posture pair information paired with the training text.

[0118] The training text included in the training data is inappropriate as training content for moving one of the two postures indicated by posture pair information that is not paired with the training text closer to the other. The features of such training text can be called "features of inappropriate training text" in comparison with the features of the difference vectors between the two postures indicated by posture pair information that is not paired with the training text.

[0119] In other words, the feature of the difference vector between two postures indicated by the posture pair information included in the training data can be said to be an "appropriate feature of the difference vector" with respect to the feature of the instruction text paired with the posture pair information.

[0120] The feature of the difference vector between the two postures indicated by the posture pair information included in the training data can be called an "inappropriate feature of the difference vector" in comparison with the feature of the instruction text that is not paired with the posture pair information.

[0121] The "feature amount of appropriate instruction text", "feature amount of inappropriate instruction text", "feature amount of appropriate difference vector", and "feature amount of inappropriate difference vector" are as described in the second embodiment.

[0122] The learning unit 21 learns the posture encoder and the text encoder so as to satisfy the following conditions: For posture pair information and instruction text that are paired in the learning data, the similarity between the feature of the difference vector and the feature of the instruction text is relatively high; For posture pair information and instruction text that are not paired in the learning data, the similarity between the feature of the difference vector and the feature of the instruction text is relatively low.

[0123] For example, the learning unit 21 dissolves pairs of posture pair information and instruction texts included in the learning data, and generates a set of multiple posture pair information and a set of multiple instruction texts. Then, the learning unit 21 generates multiple pairs of posture pair information and instruction texts based on the set of multiple posture pair information and the set of multiple instruction texts. Hereinafter, pairs of posture pair information and instruction texts generated by the learning unit 21 in this manner are referred to as "generated pairs." The generated pairs include pairs of posture pair information and instruction texts that were originally paired in the learning data, and pairs of posture pair information and instruction texts that were not paired in the learning data.

[0124] The learning unit 21 selects one of the multiple generated pairs thus generated, and generates the feature of the difference vector and the feature of the instruction text based on the posture pair information and the instruction text included in the selected generated pair.

[0125] Specifically, the learning unit 21 generates a difference vector indicating the difference between the two postures indicated by the posture pair information, and then inputs the generated difference vector to a posture encoder to generate a feature of the difference vector. This processing is similar to the processing by the difference information generator 13 described in the second embodiment.

[0126] Furthermore, the learning unit 21 generates feature quantities of the training text by inputting the training text into a text encoder.

[0127] The learning unit 21 then calculates the similarity between the feature amounts of the generated difference vector and the feature amounts of the instruction text. If the selected generation pair is "a pair of posture pair information and instruction text that was originally paired in the training data," the learning unit 21 adjusts the parameters of at least one of the posture encoder and the text encoder so that the calculated similarity becomes higher. On the other hand, if the selected generation pair is "a pair of posture pair information and instruction text that was not paired in the training data," the learning unit 21 adjusts the parameters of at least one of the posture encoder and the text encoder so that the calculated similarity becomes lower.

[0128] The learning unit 21 can learn the posture encoder and the text encoder by repeating this process using the plurality of generated pairs. As a variation, in the above learning method of the learning unit 21, one of the posture encoder and the text encoder may be learned in advance by a different method, and the parameters of the pre-trained encoder may not be adjusted, but only the encoder that has not been pre-trained may be adjusted.

[0129] The training data described above may be created by a person or may be generated by a computer. The following embodiment describes the generation of training data by a computer.

[0130] <Effects> The learning device 20 of the third embodiment can learn the pose encoder and text encoder described in the second embodiment. That is, the learning device 20 can learn the pose encoder and text encoder used by the information processing device 10 that realizes the effects described in the first and second embodiments.

[0131] <<Fourth Embodiment>> A learning device 20 according to a fourth embodiment has a function of generating learning data including pairs of posture pair information indicating a pair of two postures and instruction text indicating, in text form, instruction content for bringing one of the two postures closer to the other. This will be described in detail below.

[0132] 7 shows an example of a functional block diagram of the learning device 20. As shown in the figure, the learning device 20 includes a learning unit 21, an instruction text selection unit 22, and a posture pair information generation unit 23.

[0133] The instruction text selection unit 22 selects an instruction text that indicates, in text, a combination of the part of the body to be taught, the direction in which to move that part, and the amount by which to move that part. The "combination of the part of the body to be taught, the direction in which to move that part, and the amount by which to move that part" indicates the instruction content as a change from the current state of the teaching subject (a change based on the current state). An example of the instruction text is "Raise your right arm 30 degrees."

[0134] In one example, a plurality of instruction text templates are prepared in advance. The instruction text selection unit 22 can select one of the plurality of instruction text templates. This template may be generated by a user. Alternatively, a predetermined device may generate this template. For example, the device may analyze audio data or the like of a scene in which an instructor or the like is actually giving instruction to detect the content of the instruction that was actually given, and generate a template indicating the content of the instruction that was detected from the actual instruction.

[0135] The posture pair information generating unit 23 generates posture pair information to be paired with the instruction text selected by the instruction text selecting unit 22. As a result, learning data is generated in which the instruction text selected by the instruction text selecting unit 22 is paired with the posture pair information generated by the posture pair information generating unit 23.

[0136] The posture pair information generation unit 23 first determines a first posture. The content of the first posture is not particularly limited. The posture pair information generation unit 23 can determine the first posture by various means. For example, the posture pair information generation unit 23 may select one posture as the first posture from a plurality of postures generated in advance according to a predetermined rule (e.g., random). Alternatively, the posture pair information generation unit 23 may determine one posture specified by the user as the first posture. Alternatively, the posture pair information generation unit 23 may randomly determine the first posture.

[0137] After determining the first posture, the posture pair information generating unit 23 generates a second posture to be paired with the first posture. As a result, posture pair information indicating the pair of the first posture and the second posture is generated. The posture pair information generating unit 23 can generate at least one of the following postures as the second posture: A posture that is reached when the instruction indicated in the instruction text selected by the instruction text selecting unit 22 is applied to the first posture A posture that is reached when the instruction indicated in the instruction text selected by the instruction text selecting unit 22 is applied to the first posture with the "direction of moving the part to be taught" reversed

[0138] "Applying the instruction indicated in the instruction text to the first posture" means moving the part of the body to be taught in a predetermined direction by a predetermined amount as indicated in the instruction text.

[0139] "Applying the instruction indicated in the instruction text to the first posture by reversing the 'direction of moving the part of the body to be taught'" means moving the part of the body to be taught indicated in the instruction text in the direction opposite to the direction of movement indicated in the instruction text by the amount indicated in the instruction text.

[0140] A specific example will now be described with reference to FIG.

[0141] In the example of Fig. 8, the instruction text selection unit 22 selects the instruction text "Raise your right arm 30 degrees upward" as shown in Fig. 8(1). Then, the posture pair information generation unit 23 determines the first posture as shown in Fig. 8(2).

[0142] In this case, the posture pair information generating unit 23 can determine at least one of the two postures shown in FIG. 8(3) as the second posture.

[0143] The upper of the two postures shown in Fig. 8(3) is a posture that is reached when the instruction indicated in the instruction text selected by the instruction text selection unit 22 is applied to the first posture. Specifically, the upper of the two postures shown in Fig. 8(3) is a posture that is reached by raising the right arm 30 degrees upward in the first posture shown in Fig. 8(2).

[0144] On the other hand, the lower of the two postures shown in Fig. 8(3) is a posture that is reached when the instruction indicated in the instruction text selected by the instruction text selection unit 22 is applied to the first posture by reversing the "direction of moving the part to be taught." Specifically, the lower of the two postures shown in Fig. 8(3) is a posture that is reached by lowering the right arm 30 degrees downward in the first posture shown in Fig. 8(2).

[0145] As in this example, the posture pair information generating unit 23 can generate a second posture by moving only the body part of the target of instruction indicated in the instruction text, while leaving the body part other than the body part of the target of instruction indicated in the instruction text in the first posture.

[0146] In one example, a method for moving key points is defined in advance for each instruction text and stored in a predetermined storage device. The posture pair information generating unit 23 can generate a second posture from a first posture by moving the positions of predetermined key points in a first posture based on the definition.

[0147] The key point movement defined for each instruction text may be defined as a movement that moves the right arm exactly by the amount indicated by the instruction text. For example, if the instruction text is "Raise your right arm 30 degrees," the movement that moves the right arm exactly 30 degrees may be defined.

[0148] Alternatively, the key point movement method defined for each instruction text may be defined as a movement method that involves moving the key point by an amount indicated by the instruction text or any of the surrounding amounts. For example, if the instruction text is "Raise your right arm 30 degrees," the movement method may involve moving the right arm by any angle between 29 degrees and 31 degrees. In this case, there are various methods for selecting the actual movement amount from the defined numerical range. For example, the posture pair information generation unit 23 may select randomly. Alternatively, the posture pair information generation unit 23 may select based on a probability according to a pre-generated normal distribution.

[0149] Other configurations of the learning device 20 of the fourth embodiment are similar to those of the learning device 20 of the third embodiment.

[0150] The learning device 20 of the fourth embodiment achieves the same effects as the learning device 20 of the third embodiment. Furthermore, the learning device 20 allows a computer to generate learning data. Creating a large amount of learning data manually requires a great deal of effort. The learning device 20 of the fourth embodiment can alleviate this inconvenience.

[0151] <<Fifth Embodiment>> A learning device 20 according to a fifth embodiment has a function of generating learning data using a method different from that of the learning device 20 according to the fourth embodiment. This will be described in detail below.

[0152] 7 shows an example of a functional block diagram of the learning device 20. As shown in the figure, the learning device 20 includes a learning unit 21, an instruction text selection unit 22, and a posture pair information generation unit 23.

[0153] The instruction text selection unit 22 selects an instruction text that indicates, in text, a combination of the body part of the training target and the state of that body part. The "combination of the body part of the training target and the state of that body part" indicates the instruction content in the final state that the training target should be in. An example of such an instruction text is "Keep your right elbow at 90 degrees." Hereinafter, the state of the body part of the training target indicated in the instruction text will be referred to as the "reference state."

[0154] In one example, a plurality of instruction text templates are prepared. The instruction text selection unit 22 can select one from the plurality of instruction text templates. This template may be generated by a user. Alternatively, a predetermined device may generate this template. For example, the device may analyze audio data or the like of a scene in which an instructor or the like is actually giving instruction to detect the content of the instruction that was actually given, and generate a template indicating the content of the instruction detected from the actual instruction.

[0155] The posture pair information generating unit 23 generates posture pair information to be paired with the instruction text selected by the instruction text selecting unit 22. As a result, learning data is generated in which the instruction text selected by the instruction text selecting unit 22 is paired with the posture pair information generated by the posture pair information generating unit 23.

[0156] The posture pair information generation unit 23 first determines a first posture in which the training target body part is not in the reference state. The content of the first posture is not particularly limited except that the training target body part is not in the reference state. The posture pair information generation unit 23 can determine the first posture by various means. For example, the posture pair information generation unit 23 may select one posture from a plurality of postures generated in advance according to a predetermined rule (e.g., randomly). Then, the posture pair information generation unit 23 may determine, as the first posture, a posture in which the training target body part is in a state different from the reference state in the selected posture. Alternatively, the posture pair information generation unit 23 may determine, as the first posture, a posture specified by the user. The user specifies, as the first posture, a posture in which the training target body part is not in the reference state. Alternatively, the posture pair information generation unit 23 may randomly generate one posture. Then, the posture pair information generation unit 23 may determine, as the first posture, a posture in which the training target body part is in a state different from the reference state in one of the generated postures.

[0157] After determining the first posture, the posture pair information generating unit 23 generates a second posture to be paired with the first posture. As a result, posture pair information indicating the pair of the first posture and the second posture is generated. After determining the first posture in which the part of the training target is not in the reference state, the posture pair information generating unit 23 can generate the second posture by performing a process of generating a second posture in which the part of the training target is in the reference state in the first posture.

[0158] A specific example will now be described with reference to FIG.

[0159] In the example of Fig. 9, the instruction text selection unit 22 selects the instruction text "Bend the right elbow at a 90-degree angle" as shown in Fig. 9(1). Then, the posture pair information generation unit 23 determines a first posture in which the right elbow is not bent at a 90-degree angle as shown in Fig. 9(2).

[0160] In this case, the posture pair information generating unit 23 generates a second posture by tilting the right elbow 90 degrees in the first posture shown in FIG. 9(2), as shown in FIG. 9(3).

[0161] As in this example, the posture pair information generating unit 23 can generate a second posture by moving only the body part of the target of instruction indicated in the instruction text, while leaving the body part other than the body part of the target of instruction indicated in the instruction text in the first posture.

[0162] In one example, the states of key points are defined in advance for each instruction text and stored in a predetermined storage device. The posture pair information generating unit 23 can generate a second posture from a first posture by moving the positions of predetermined key points in a first posture based on the definition.

[0163] The key point state defined for each instruction text may define a movement method to accurately achieve the state indicated by the instruction text. For example, if the instruction text is "Bend your right elbow at 90 degrees," the key point state to accurately make the right elbow at 90 degrees may be defined.

[0164] Alternatively, the keypoint state defined for each instruction text may define a movement that results in one of the states indicated by the instruction text and its surrounding states. For example, if the instruction text is "Bend the right elbow to 90 degrees," a keypoint state may be defined in which the right elbow is bent to an angle between 89 degrees and 91 degrees. In this case, there are various methods for selecting the state to be actually adopted from the defined numerical range. For example, the posture pair information generation unit 23 may select randomly. Alternatively, the posture pair information generation unit 23 may select based on a probability according to a pre-generated normal distribution.

[0165] Here, a modified example will be described. In the above example, the posture pair information generating unit 23 determines a first posture in which the body part of the training target is not in the reference state, and then generates a second posture in which the body part of the training target is in the reference state in the first posture. The posture pair information generating unit 23 may generate posture pair information by performing the reverse process.

[0166] That is, the posture pair information generating unit 23 may generate posture pair information by determining a third posture in which the part of the training target is in a reference state, and then generating a fourth posture in which the part of the training target is in a state different from the reference state in the third posture. In this modified example, the posture pair information generating unit 23 can generate posture pair information indicating a pair of the third posture and the fourth posture.

[0167] Other configurations of the learning device 20 of the fifth embodiment are similar to those of the learning device 20 of the third embodiment. Note that the learning device 20 of the fifth embodiment may or may not include the configurations of the learning device 20 of the fourth embodiment.

[0168] The learning device 20 of the fifth embodiment achieves the same effects as the learning device 20 of the third embodiment. Furthermore, the learning device 20 allows a computer to generate learning data. Creating a large amount of learning data manually requires a great deal of effort. The learning device 20 of the fifth embodiment can alleviate this inconvenience.

[0169] <<Sixth Embodiment>> A training data generation device according to a sixth embodiment has the training data generation function described in the fourth and fifth embodiments. This will be described in detail below.

[0170] In the fourth and fifth embodiments, the learning device 20 has the learning data generation function (the teaching text selection unit 22 and the posture pair information generation unit 23). In the sixth embodiment, a learning data generation device 30, which is physically and / or logically separated from the learning device 20, has the learning data generation function (the teaching text selection unit 22 and the posture pair information generation unit 23).

[0171] 10 shows an example of a functional block diagram of the training data generation device 30. As shown in the figure, the training data generation device 30 has a training text selection unit 22 and a posture pair information generation unit 23. The configurations of the training text selection unit 22 and the posture pair information generation unit 23 are as described in the fourth and fifth embodiments. The training data generation device 30 may have one or both of the training data generation functions (the training text selection unit 22 and the posture pair information generation unit 23) described in the fourth and fifth embodiments.

[0172] Here, an example of the hardware configuration of the training data generation device will be described. Each functional unit of the training data generation device 30 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various variations in the implementation method and device. The software includes programs that are pre-loaded when the device is shipped, and programs downloaded from recording media such as CDs or servers on the Internet.

[0173] FIG. 3 is a block diagram illustrating an example of the hardware configuration of a training data generation device 30. As shown in FIG. 3, the training data generation device 30 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The training data generation device 30 does not necessarily have the peripheral circuit 4A. The training data generation device 30 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration. The processor 1A, memory 2A, input / output interface 3A, peripheral circuit 4A, and bus 5A are as described in the second embodiment.

[0174] The training data generation device 30 of the sixth embodiment allows a computer to generate training data. Manually creating a large amount of training data requires a great deal of effort. The training data generation device 30 of the sixth embodiment can alleviate this inconvenience.

[0175] <<Seventh Embodiment>> An information processing device 10 of the seventh embodiment uses the posture encoder and text encoder described in the previous embodiments to realize instruction different from that of the information processing device 10 of the first and second embodiments. This will be described in detail below.

[0176] 11 shows an example of a functional block diagram of the information processing device 10. As shown in the figure, the information processing device 10 includes a reference posture information acquisition unit 11, a training target posture information acquisition unit 12, a difference information generation unit 13, an output unit 14, and a selection unit 15.

[0177] FIG. 12 is a data flow diagram showing an example of the flow of data in the process executed by the information processing device 10. As shown in FIG.

[0178] The reference posture information acquisition unit 11 acquires reference posture information indicating a plurality of time-series reference postures. The training target posture information acquisition unit 12 acquires training target posture information indicating a plurality of time-series postures of the training target. The configurations of the reference posture information acquisition unit 11 and the training target posture information acquisition unit 12 are as described in the second embodiment.

[0179] The difference information generation unit 13 generates difference information indicating the difference between a reference posture and a posture of the training target. The difference information generation unit 13 generates a plurality of pairs, each including one reference posture and one posture of the training target, based on a plurality of time-series reference postures and a plurality of time-series postures of the training target. The difference information generation unit 13 then generates a difference vector indicating the difference between the reference posture and the posture of the training target for each generated pair, and can then generate a feature value of the difference vector for each generated pair. The difference information generation unit 13 includes a correspondence generation means, a difference vector acquisition means, and a posture encoder ("PoseEncoder" in FIG. 12) shown in FIG. 12.

[0180] The correspondence generating means generates a plurality of pairs, each including one reference posture and one posture of the target, based on a plurality of time-series reference postures indicated by the reference posture information and a plurality of time-series postures of the target indicated by the target posture information. Hereinafter, the pairs generated by the correspondence generating means will be referred to as "comparison target pairs."

[0181] For example, the correspondence generating means may be M 1The individual's reference posture and M 2 Any pairs generated from the postures of the training subjects may be generated as comparison pairs. In this case, the correspondence generating means 1 ×M 2 The correspondence generating means can generate M comparison pairs. 1 ×M 2 Instead of generating all of the pairs as comparison pairs, 1 ×M 2 For example, the correspondence generating means may generate a part of the M pairs as comparison pairs. 1 ×M 2 A randomly selected portion of the pairs may be generated as comparison pairs.

[0182] In addition, the correspondence generating means is M 1 A predetermined one of the reference postures and M 2 Any pairs generated from the postures of the training subjects may be generated as comparison pairs. In this case, the correspondence generating means 2 The correspondence generating means can generate M comparison pairs. 2 Instead of generating all of the pairs as comparison pairs, 2 For example, the correspondence generating means may generate a part of the M pairs as comparison pairs. 2 A randomly selected portion of the pairs may be generated as comparison pairs.

[0183] "M 1 There are various methods for selecting "a predetermined one of the reference postures." For example, when the reference posture information is generated by detecting in real time the posture that the reference posture detection target is actually performing, the correspondence generating means selects one of the M 1 The latest reference posture among the reference postures is defined as M 1 The reference posture may be selected as a predetermined one of the reference postures.

[0184] In addition, the correspondence generating means is M 2 One of the postures to be taught and M 1 In this case, the correspondence generating means may generate all pairs generated from the M reference postures as comparison pairs.1 The correspondence generating means can generate M comparison pairs. 1 Instead of generating all of the pairs as comparison pairs, 1 For example, the correspondence generating means may generate a part of the M pairs as comparison pairs. 1 A randomly selected portion of the pairs may be generated as comparison pairs.

[0185] "M 2 There are various methods for selecting "a predetermined one of the postures of the training target." For example, when the training target posture information is generated by detecting the posture that the training target actually takes in real time, the correspondence generating means 2 The latest posture of the individual training target is 2 The posture may be selected as a predetermined one of the postures to be taught.

[0186] The difference vector acquisition means generates a difference vector for each comparison pair generated by the correspondence generation means. Then, the difference information generation unit 13 inputs the generated difference vector to the posture encoder to generate difference information (feature amount of the difference vector) for each comparison pair generated by the correspondence generation means. This processing by the difference information generation unit 13 is as described in the second embodiment.

[0187] The selection unit 15 selects one of a plurality of prepared instruction texts indicating instruction content. The selection unit 15 may select one instruction text designated by a user input. Alternatively, the selection unit 15 may select one instruction text according to a predetermined rule.

[0188] The output unit 14 extracts and outputs a comparison pair that satisfies the following condition from among the plurality of comparison pairs generated by the correspondence generation unit 16: (Condition) A comparison pair that reaches a corresponding reference posture when the instruction indicated in the instruction text selected by the selection unit 15 is applied to the corresponding posture of the instruction target.

[0189] Here, a specific example of use of the information processing apparatus 10 of the seventh embodiment will be described.

[0190] The selection unit 15 selects an instruction text when there is no difference between the reference posture and the posture of the instruction target. The user may specify such an instruction text. Alternatively, the selection unit 15 may be configured to select such an instruction text in advance.

[0191] In this case, the output unit 14 extracts comparison pairs in which there is no difference between the postures of the training target and the reference posture. Then, based on the pairs of postures of the training target and the reference postures included in the extracted comparison pairs, the output unit 14 calculates the time deviations between the multiple time-series reference postures and the multiple time-series postures of the training target, and outputs the calculation results.

[0192] For example, the output unit 14 calculates the "deviation in elapsed time from the timing at which the movement shown in the plurality of time-series postures started" between the posture of the training target included in the extracted comparison pair and the reference posture. Hereinafter, the "elapsed time from the timing at which the movement shown in the plurality of time-series postures started" will be simply referred to as the "elapsed time from the start of the movement."

[0193] The output unit 14 outputs the elapsed time t from the start of the movement of the “posture of the training target” in the extracted comparison pair. 1 is the elapsed time t 2 If the time t is greater than the elapsed time from the start of the movement of the "posture of the training target" in the extracted comparison target pair, it can be determined that the "posture of the training target" is delayed from the "reference posture". Then, the output unit 14 can output information indicating this. 1 and the elapsed time t from the start of the movement of the "reference posture" 2 The difference between the two may be calculated and output as the delay time.

[0194] The output unit 14 also outputs the elapsed time t 1 is the elapsed time t 2If the time t 1 and the elapsed time t from the start of the movement of the "reference posture" 2 The difference between the two may be calculated and further output as the time advance.

[0195] The information processing device 10 of the seventh embodiment may or may not be equipped with the "configuration for outputting instructional text using difference information to indicate instruction content in text form for bringing the posture of the person being taught closer to the reference posture" described in the first and second embodiments.

[0196] "Effects" According to the information processing device 10 of the seventh embodiment, it is possible to extract pairs of "posture of a training target" and "reference posture that will be reached when the instruction indicated in the training text selected for the posture of the training target is applied" from among a plurality of comparison pairs. By using the information processing device 10 that performs such characteristic processing, it is possible to provide various types of instruction. For example, by selecting "training text when there is no difference between the reference posture and the posture of the training target" and extracting the comparison pair, it is possible to calculate and output the time shift between a plurality of time-series reference postures and a plurality of time-series postures of the training target.

[0197] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0198] In addition, in the flowcharts used in the above explanation, multiple steps (processes) are described in order. However, the order of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed as long as it does not cause any problems in terms of the content.

[0199] Some or all of the above embodiments can be described as, but are not limited to, the following notes: 1. An information processing device having: reference posture information acquisition means for acquiring reference posture information indicating a reference posture; training target posture information acquisition means for acquiring training target posture information indicating the posture of a target; difference information generation means for generating difference information indicating the difference between the reference posture and the posture of the target; and output means for outputting training text indicating, in text, training content for bringing the posture of the target closer to the reference posture, using the difference information. 2. The information processing device described in 1, wherein the training text indicates, in text, a combination of a part of the target, a direction in which to move the part, and an amount by which to move the part, or a combination of a part of the target and a state of the part. 3. The information processing device according to 1 or 2, wherein the difference information generation means generates a difference vector indicating a difference between the reference posture and the posture of the target person, and then generates a feature amount of the difference vector as the difference information, and the output means determines the teaching text to be output from a plurality of pre-prepared teaching texts based on a similarity between a feature amount of each of the plurality of pre-prepared teaching texts and a feature amount of the difference vector. 4. The information processing device according to 3, wherein the difference information generation means inputs the difference vector to a posture encoder generated in advance by machine learning and acquires the feature amount of the difference vector output from the posture encoder, and the output means determines the teaching text to be output based on the feature amount of each of the plurality of teaching texts generated by inputting each of the plurality of pre-prepared teaching texts to a text encoder generated in advance by machine learning. 5. The information processing device according to 3 or 4, wherein the difference vector indicates at least one of a difference vector of the position of a key point and a difference vector of the angle of an angle formed by three key points.6. The information processing device according to any one of 3 to 5, wherein the output means determines as the training text to be output the training text whose similarity to the feature amount of the difference vector satisfies an output condition, and the output condition is that the similarity to the feature amount of the difference vector is the highest, the similarity to the feature amount of the difference vector is equal to or greater than a reference value, or the training text is included in a predetermined top number in a ranking sorted in descending order of the similarity to the feature amount of the difference vector. 7. An information processing method wherein one or more computers acquire reference posture information indicating a reference posture, acquire training target posture information indicating the posture of a training target, generate difference information indicating the difference between the reference posture and the posture of the training target, and use the difference information to output a training text indicating, in text, training content to bring the posture of the training target closer to the reference posture. 9. A program that causes a computer to function as: reference posture information acquisition means for acquiring reference posture information indicating a reference posture; training target posture information acquisition means for acquiring training target posture information indicating the posture of a target; difference information generation means for generating difference information indicating the difference between the reference posture and the posture of the target; and output means for using the difference information to output a training text indicating, in text, training content to bring the posture of the target closer to the reference posture. 9. A learning device having learning means that trains a learning model using training data including pairs of posture pair information indicating a pair of two postures and training text indicating, in text, training content to bring one of the two postures closer to the other. 10. The learning device described in 9, wherein the learning means uses the training data to train a posture encoder that receives as input a difference vector indicating the difference between the two postures and outputs features of the difference vector, and a text encoder that receives as input the training text and outputs features of the training text.11. The learning device according to 10, wherein the learning means learns the posture encoder and the text encoder so that, for the posture pair information and the training text that are paired in the learning data, a degree of similarity between the feature amount of the difference vector and the feature amount of the training text is relatively high, and for the posture pair information and the training text that are not paired in the learning data, a degree of similarity between the feature amount of the difference vector and the feature amount of the training text is relatively low. 12. A learning device according to any one of items 9 to 11, comprising: an instruction text selection means for selecting the instruction text indicating in text a combination of a part of the body to be taught, a direction to move the part, and an amount to move the part; and a posture pair information generation means for generating the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means generates the posture pair information indicating a pair of the first posture and the second posture by determining a first posture and then generating a second posture that is a posture reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture reached when the instruction indicated in the selected instruction text is applied to the first posture with the movement direction reversed. 13. A learning device according to any one of claims 9 to 12, comprising: a training text selection means for selecting the training text which indicates, in text, a combination of a body part of a training target and a reference state of the body part; and a posture pair information generation means for generating the posture pair information to be paired with the selected training text, wherein the posture pair information generation means generates the posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture, by: determining a first posture in which the body part of the training target is not in the reference state, and then generating a second posture in which the body part of the training target is in the reference state in the first posture; or determining a third posture in which the body part of the training target is in the reference state, and then generating a fourth posture in which the body part of the training target is in a state different from the reference state in the third posture.14. A learning method in which one or more computers learn a learning model using learning data including pairs of posture pair information indicating a pair of two postures and instruction text indicating in text the instruction content for bringing one of the two postures closer to the other. 15. A program that causes a computer to function as learning means for learning a learning model using learning data including pairs of posture pair information indicating a pair of two postures and instruction text indicating in text the instruction content for bringing one of the two postures closer to the other. 16. 17. A learning data generation device comprising: an instruction text selection means for selecting an instruction text indicating in text a combination of a part of the body to be taught, a direction to move that part, and an amount to move that part; and a posture pair information generation means for generating the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means generates the posture pair information indicating a pair of the first posture and the second posture by determining a first posture and then generating a second posture that is a posture reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture reached when the instruction indicated in the selected instruction text is applied to the first posture with the direction of movement reversed. A learning data generation method in which one or more computers select instruction texts that indicate in text a combination of a part of the body to be taught, a direction in which to move that part, and an amount by which to move that part, and generate posture pair information to pair with the selected instruction text, and in generating the posture pair information, the method involves determining a first posture and then generating a second posture that is a posture that is reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture that is reached when the instruction indicated in the selected instruction text is applied to the first posture with the direction of movement reversed, thereby generating the posture pair information that indicates a pair of the first posture and the second posture.18. A program that causes a computer to function as: instruction text selection means that selects an instruction text indicating in text a combination of a part of the body to be taught, a direction to move that part, and an amount to move that part; and posture pair information generation means that generates the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means determines a first posture, and then generates a second posture that is a posture reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture reached when the instruction indicated in the selected instruction text is applied to the first posture with the direction of movement reversed, thereby generating the posture pair information that indicates a pair of the first posture and the second posture. a training text generation means for selecting a training text indicating, in text form, a combination of a body part of a training target and a reference state of the body part; and a posture pair information generation means for generating the posture pair information to be paired with the selected training text, wherein the posture pair information generation means generates posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture, by: determining a first posture in which the body part of the training target is not in the reference state, and then generating a second posture in which the body part of the training target is in the reference state in the first posture; or determining a third posture in which the body part of the training target is in the reference state, and then generating a fourth posture in which the body part of the training target is in a state different from the reference state in the third posture.20. A learning data generation method in which one or more computers select a training text indicating, in text form, a combination of a part of the body to be trained and a reference state of that part, and generate the posture pair information to be paired with the selected training text, and in generating the posture pair information, determine a first posture in which the part of the body to be trained is not in the reference state, and then generate a second posture in which the part of the body to be trained is in the reference state in the first posture, or determine a third posture in which the part of the body to be trained is in the reference state, and then generate a fourth posture in which the part of the body to be trained is in a state different from the reference state in the third posture, thereby generating posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture. 21. a program that causes a computer to function as: an instruction text generation means that selects an instruction text that indicates, in text, a combination of a part of a training target and a reference state of the part; and a posture pair information generation means that generates the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means generates posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture, by determining a first posture in which the part of the training target is not in the reference state, and then generating a second posture in which the part of the training target is in the reference state in the first posture, or by determining a third posture in which the part of the training target is in the reference state, and then generating a fourth posture in which the part of the training target is in a state different from the reference state in the third posture.22. An information processing device comprising: reference posture information acquisition means for acquiring reference posture information indicating a plurality of time-series reference postures; training target posture information acquisition means for acquiring training target posture information indicating a plurality of time-series postures of the training target; difference information generation means for generating difference information indicating a difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training target; selection means for selecting one from a plurality of pre-prepared training texts indicating training content; and output means for outputting a pair of the training target's posture and the reference posture reached when the training indicated in the selected training text is applied to the training target's posture. 23. The information processing device described in 22, wherein the selection means selects the training text when there is no difference between the reference posture and the training target's posture, and the output means extracts pairs of the training target's posture and the reference posture where there is no posture difference, and calculates a time shift between the plurality of time-series reference postures and the plurality of time-series postures of the training target based on the extracted pairs of the training target's posture and the reference posture, and outputs the calculation result. 24. 25. The information processing device according to any of 22 to 24, wherein the instruction text indicates, in text, a combination of a part to be moved, a direction in which the part is to be moved, and an amount by which the part is moved, or a combination of a part to be moved and a state of the part. 25. The information processing device according to any of 22 to 24, wherein the difference information generation means generates a plurality of pairs each including one reference posture and one posture of the target based on a plurality of time-series reference postures and a plurality of time-series postures of the target, generates a difference vector indicating the difference between the reference posture and the posture of the target for each generated pair, and then generates a feature amount of the difference vector for each generated pair, and the output means outputs the pair of the posture of the target and the reference posture based on the similarity between the feature amount of the instruction text selected by the selection means and each feature amount of the difference vector.26. The information processing device according to 25, wherein the difference information generation means inputs a plurality of difference vectors into a posture encoder generated in advance by machine learning and acquires feature quantities of the plurality of difference vectors output from the posture encoder, and the output means outputs a pair of the posture of the target and the reference posture based on feature quantities of the training text generated by inputting the text selected by the selection means into a text encoder generated in advance by machine learning. 27. The information processing device according to 25 or 26, wherein the difference vector indicates at least one of a difference vector of the position of a key point and a difference vector of the angle of an angle formed by three key points. 28. An information processing method in which one or more computers acquire reference posture information indicating a plurality of time-series reference postures, acquire training target posture information indicating a plurality of time-series postures of the target, generate difference information indicating the difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the target, select one from a plurality of training texts prepared in advance that indicate training content, and output a pair of the posture of the target and the reference posture reached when instruction indicated in the selected training text is applied to the posture of the target. 29. A program that causes a computer to function as: reference posture information acquisition means for acquiring reference posture information indicating a plurality of time-series reference postures; training target posture information acquisition means for acquiring training target posture information indicating a plurality of time-series postures of the training target; difference information generation means for generating difference information indicating a difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training target; selection means for selecting one from a plurality of pre-prepared training texts indicating training content; and output means for outputting a pair of the posture of the training target and the reference posture reached when the guidance indicated in the selected training text is applied to the posture of the training target.

[0200] Some or all of Appendices 2 to 6 that are dependent on the information processing device of Appendix 1 described above may also be dependent on the information processing method of Appendix 7 and the program of Appendix 8 in the same dependency relationship as Appendix 1 and Appendices 2 to 6. Furthermore, some or all of Appendices 10 to 13 that are dependent on the learning device of Appendix 9 may also be dependent on the learning method of Appendix 14 and the program of Appendix 15 in the same dependency relationship as Appendix 9 and Appendices 10 to 13. Furthermore, some or all of Appendices 23 to 27 that are dependent on the information processing method of Appendix 28 and the program of Appendix 29 in the same dependency relationship as Appendix 22 and Appendices 23 to 27. Furthermore, within the scope of each of the above-mentioned embodiments, some or all of the configurations described as appendices can be realized in various hardware, software, various recording means for recording software, or systems.

[0201] This application claims priority based on Japanese Patent Application No. 2024-139742, filed on August 21, 2024, the disclosure of which is incorporated herein by reference in its entirety.

[0202] REFERENCE SIGNS LIST 10 Information processing device 11 Reference posture information acquisition unit 12 Training target posture information acquisition unit 13 Difference information generation unit 14 Output unit 15 Selection unit 20 Learning device 21 Learning unit 22 Training text selection unit 23 Posture pair information generation unit 30 Training data generation device 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus

Claims

1. An information processing device having: a reference posture information acquisition means for acquiring reference posture information indicating a reference posture; a training target posture information acquisition means for acquiring training target posture information indicating the posture of a training target; a difference information generation means for generating difference information indicating the difference between the reference posture and the posture of the training target; and an output means for using the difference information to output a training text indicating, in text form, a training content for bringing the posture of the training target closer to the reference posture.

2. The information processing device according to claim 1, wherein the instruction text indicates, in text, a combination of the part to be taught, the direction in which the part is to be moved, and the amount by which the part is to be moved, or a combination of the part to be taught and the state of the part.

3. The information processing device described in claim 1 or 2, wherein the difference information generating means generates a difference vector indicating the difference between the reference posture and the posture of the subject of instruction, and then generates a feature of the difference vector as the difference information, and the output means determines the instruction text to be output from a plurality of instruction texts prepared in advance based on the similarity between the feature of each of the plurality of instruction texts prepared in advance and the feature of the difference vector.

4. The information processing device described in claim 3, wherein the difference information generation means inputs the difference vector into a posture encoder generated in advance by machine learning and acquires features of the difference vector output from the posture encoder, and the output means determines the instruction text to be output based on the features of each of the plurality of instruction texts generated by inputting each of the plurality of instruction texts prepared in advance into a text encoder generated in advance by machine learning.

5. The information processing device according to claim 3 or 4, wherein the difference vector indicates at least one of a difference vector of the position of a key point and a difference vector of the angle of an angle formed by three key points.

6. An information processing device according to any one of claims 3 to 5, wherein the output means determines as the training text to be output the training text whose similarity to the feature of the difference vector satisfies an output condition, and the output condition is that the similarity to the feature of the difference vector is the highest, the similarity to the feature of the difference vector is equal to or greater than a reference value, or the training text is included in a predetermined top number in a ranking sorted in descending order of the similarity to the feature of the difference vector.

7. An information processing method in which one or more computers acquire reference posture information indicating a reference posture, acquire training target posture information indicating the posture of a training target, generate difference information indicating the difference between the reference posture and the posture of the training target, and use the difference information to output a training text indicating in text the training content for bringing the posture of the training target closer to the reference posture.

8. A recording medium having recorded thereon a program that causes a computer to function as: a reference posture information acquisition means for acquiring reference posture information indicating a reference posture; a training target posture information acquisition means for acquiring training target posture information indicating the posture of a target; a difference information generation means for generating difference information indicating the difference between the reference posture and the posture of the target; and an output means for using the difference information to output a training text indicating in text the training content for bringing the posture of the target closer to the reference posture.

9. A learning device having a learning means for learning a learning model using learning data including pairs of posture pair information indicating a pair of two postures and instruction text indicating in text the instruction content for bringing one of the two postures closer to the other.

10. The learning device according to claim 9, wherein the learning means uses the learning data to train a posture encoder that receives a difference vector indicating the difference between the two postures as input and outputs the features of the difference vector, and a text encoder that receives the instruction text as input and outputs the features of the instruction text.

11. The learning device of claim 10, wherein the learning means trains the posture encoder and the text encoder so that, for the posture pair information and the instruction text that are paired in the learning data, the similarity between the feature of the difference vector and the feature of the instruction text is relatively high, and for the posture pair information and the instruction text that are not paired in the learning data, the similarity between the feature of the difference vector and the feature of the instruction text is relatively low.

12. A learning device as described in any one of claims 9 to 11, comprising: an instruction text selection means for selecting an instruction text indicating in text a combination of a part of the body to be taught, a direction in which to move that part, and an amount by which to move that part; and a posture pair information generation means for generating the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means generates the posture pair information indicating a pair of the first posture and the second posture by determining a first posture and then generating a second posture which is a posture reached when the instruction indicated in the selected instruction text is applied to the first posture, or a posture reached when the instruction indicated in the selected instruction text is applied to the first posture with the movement direction reversed.

13. A learning device as described in any one of claims 9 to 12, comprising: an instruction text selection means for selecting the instruction text which indicates in text a combination of a body part of the object to be taught and a reference state of that body part; and a posture pair information generation means for generating the posture pair information to be paired with the selected instruction text, wherein the posture pair information generation means generates the posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture, by: determining a first posture in which the body part of the object to be taught is not in the reference state, and then generating a second posture in which the body part of the object to be taught is in the reference state in the first posture; or determining a third posture in which the body part of the object to be taught is in the reference state, and then generating a fourth posture in which the body part of the object to be taught is in a state different from the reference state in the third posture.

14. A learning data generation method in which one or more computers select instruction texts indicating in text a combination of a part of the body to be taught, a direction in which to move that part, and an amount by which to move that part, and generate posture pair information to pair with the selected instruction text, and in generating the posture pair information, the method involves determining a first posture and then generating a second posture that is the posture reached when the instruction indicated in the selected instruction text is applied to the first posture, or the posture reached when the instruction indicated in the selected instruction text is applied to the first posture with the direction of movement reversed, thereby generating posture pair information indicating a pair of the first posture and the second posture.

15. A learning data generation method in which one or more computers select a training text indicating, in text form, a combination of a body part of a training target and a reference state of that body part, and generate posture pair information to pair with the selected training text, and in generating the posture pair information, determine a first posture in which the body part of the training target is not in the reference state, and then generate a second posture in which the body part of the training target is in the reference state in the first posture, or determine a third posture in which the body part of the training target is in the reference state, and then generate a fourth posture in which the body part of the training target is in a state different from the reference state in the third posture, thereby generating posture pair information indicating a pair of the first posture and the second posture, or a pair of the third posture and the fourth posture.

16. An information processing device having: a reference posture information acquisition means for acquiring reference posture information indicating a plurality of time-series reference postures; a training subject posture information acquisition means for acquiring training subject posture information indicating a plurality of time-series postures of a training subject; a difference information generation means for generating difference information indicating the difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training subject; a selection means for selecting one from a plurality of pre-prepared training texts indicating training content; and an output means for outputting a pair of the posture of the training subject and the reference posture reached when the instruction indicated in the selected training text is applied to the posture of the training subject.

17. The information processing device according to claim 16, wherein the difference information generating means generates a plurality of pairs each including one reference posture and one posture of the subject based on a plurality of time-series reference postures and a plurality of time-series postures of the subject, generates a difference vector indicating the difference between the reference posture and the posture of the subject for each generated pair, and then generates a feature of the difference vector for each generated pair, the output means outputs pairs of the posture of the subject and the reference posture based on the similarity between the feature of the instruction text selected by the selection means and each feature of the difference vector, the selection means selects the instruction text when there is no difference between the reference posture and the posture of the subject, and the output means extracts pairs of the posture of the subject and the reference posture that have no posture difference, and calculates the time shift between the plurality of time-series reference postures and the plurality of time-series postures of the subject based on the extracted pairs of the posture of the subject and the reference posture, and outputs the calculation result.

18. An information processing method in which one or more computers acquire reference posture information indicating a plurality of time-series reference postures, acquire training target posture information indicating a plurality of time-series postures of a training target, generate difference information indicating the difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training target, select one from a plurality of pre-prepared training texts indicating training content, and output a pair of the posture of the training target and the reference posture reached when the instruction indicated in the selected training text is applied to the posture of the training target.

19. A recording medium having recorded thereon a program that causes a computer to function as: reference posture information acquisition means for acquiring reference posture information indicating a plurality of time-series reference postures; training target posture information acquisition means for acquiring training target posture information indicating a plurality of time-series postures of the training target; difference information generation means for generating difference information indicating the difference between each of the plurality of time-series reference postures and each of the plurality of time-series postures of the training target; selection means for selecting one of a plurality of pre-prepared training texts indicating training content; and output means for outputting a pair of the posture of the training target and the reference posture reached when the instruction indicated in the selected training text is applied to the posture of the training target.

20. A recording medium as described in claim 19, having recorded thereon the program, wherein the difference information generating means generates a plurality of pairs each including one reference posture and one posture of the subject based on a plurality of time-series reference postures and a plurality of time-series postures of the subject, generates a difference vector indicating the difference between the reference posture and the posture of the subject for each generated pair, and then generates a feature of the difference vector for each generated pair; the output means outputs pairs of the posture of the subject and the reference posture based on the similarity between the feature of the teaching text selected by the selection means and each feature of the difference vector; the selection means selects the teaching text when there is no difference between the reference posture and the posture of the subject; and the output means extracts pairs of the posture of the subject and the reference posture that have no posture difference, and calculates the time deviation between the plurality of time-series reference postures and the plurality of time-series postures of the subject based on the extracted pairs of the posture of the subject and the reference posture, and outputs the calculation result.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    JP2022155037A