Learning device, inference device, program, learning method, and inference method
The learning device addresses the challenge of interpolating joint information from low-frame-rate video by using a skeleton information interpolation model that accounts for environmental context and attributes, improving the accuracy of high-frame-rate skeleton inference.
Patent Information
- Application Number
- PCT/JP2024/000651
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-17
AI Technical Summary
Conventional techniques struggle to accurately interpolate joint information from low-frame-rate video due to restrictions imposed by the surrounding environment, leading to inaccuracies in generating images reflecting the movement of articulated objects.
A learning device and method that utilizes a video acquisition unit to extract skeleton information at different frame rates, combined with surrounding environment analysis, to generate a skeleton information interpolation model for inferring high-frame-rate skeleton information from low-frame-rate information, incorporating attributes and environmental context.
Enables accurate inference of high-frame-rate skeleton information by considering the surrounding environment and attributes, enhancing precision in generating images of articulated objects.
Smart Images

Figure JP2024000651_17072025_PF_FP_ABST
Abstract
Description
Learning device, inference device, program, learning method, and inference method
[0001] The present disclosure relates to a learning device, an inference device, a program, a learning method, and an inference method.
[0002] A technique has been known for some time now in which information relating to the shapes of parts, clothing, movements, etc. of articulated objects present in an image is extracted from the image of the articulated object present in the image, and the extracted information is used to generate a new image that reflects the characteristics of the articulated object present in the image (see, for example, Patent Document 1).
[0003] Japanese Patent Application Laid-Open No. 2007-004732
[0004] Conventional technology extracts joint information from low-frame-rate video and generates images based on that information.
[0005] However, since a person's movements are limited by the surrounding environment, it is difficult to interpolate joint information with high accuracy when using only low frame rate video information, as in the past.
[0006] Therefore, one or more aspects of the present disclosure aim to make it possible to infer joint information of a person based on the surrounding environment in which the person is placed, etc.
[0007] A learning device according to one aspect of the present disclosure is characterized by comprising: a video acquisition unit that acquires video at a first frame rate; a skeletal information extraction unit that uses the video to extract, as first skeletal information, skeletal information indicating the skeletal structure of a learning subject contained in frames at the first frame rate, and that uses the video to extract, as second skeletal information, skeletal information indicating the skeletal structure of the learning subject contained in frames at a second frame rate that is lower than the first frame rate; a surrounding environment analysis unit that analyzes the environment around the learning subject from the video and generates surrounding environment information indicating the results of the analysis; and a learning unit that uses the second skeletal information and the surrounding environment information as input data and performs learning using the first skeletal information as correct answer data, thereby generating a skeletal information interpolation model for inferring skeletal information at the first frame rate from the skeletal information at the second frame rate.
[0008] An inference device according to one aspect of the present disclosure includes a video acquisition unit that acquires an inference target video, which is video at a second frame rate lower than a first frame rate; a skeletal information extraction unit that uses the inference target video to extract, as inference skeletal information, skeletal information that indicates the skeleton of a person to be inferred that is included in frames at the second frame rate; a surrounding environment analysis unit that analyzes an environment around the person to be inferred from the inference target video and generates inference surrounding environment information that indicates the results of the analysis; and a learning target video that is video at the first frame rate that extracts, as first skeletal information, skeletal information that indicates the skeleton of the person to be inferred that is included in frames at the first frame rate, and and an inference unit that uses a virtual video to extract, as second skeletal information, skeletal information indicating the skeleton of the learning subject contained in frames at the second frame rate, generates learning surrounding environment information indicating the results of analyzing the environment surrounding the learning subject from the learning subject video, and performs learning using the second skeletal information and the learning surrounding environment information as input data and the first skeletal information as correct answer data, thereby inferring skeletal information at the first frame rate by inputting the inferred skeletal information and the inferred surrounding environment information into a skeletal information interpolation model that has been generated to infer skeletal information at the first frame rate from the skeletal information at the second frame rate.
[0009] A program according to a first aspect of the present disclosure causes a computer to function as: a video acquisition unit that acquires video at a first frame rate; a skeletal information extraction unit that uses the video to extract, as first skeletal information, skeletal information indicating the skeletal structure of a learner that is contained in frames at the first frame rate, and that uses the video to extract, as second skeletal information, skeletal information indicating the skeletal structure of the learner that is contained in frames at a second frame rate that is lower than the first frame rate; a surrounding environment analysis unit that analyzes the environment around the learner from the video and generates surrounding environment information indicating the results of the analysis; and a learning unit that uses the second skeletal information and the surrounding environment information as input data and performs learning using the first skeletal information as correct answer data, thereby generating a skeletal information interpolation model for inferring skeletal information at the first frame rate from skeletal information at the second frame rate.
[0010] A program according to a second aspect of the present disclosure includes a computer including a video acquisition unit that acquires an inference target video, which is video at a second frame rate lower than a first frame rate; a skeletal information extraction unit that uses the inference target video to extract, as inference skeletal information, skeletal information indicating the skeleton of the person to be inferred that is included in frames at the second frame rate; a surrounding environment analysis unit that analyzes the environment surrounding the person to be inferred from the inference target video and generates inference surrounding environment information indicating the results of the analysis; and a learning target video that is video at the first frame rate that extracts, as first skeletal information, skeletal information indicating the skeleton of the person to be inferred that is included in frames at the first frame rate, and Using the learning target video, skeletal information indicating the skeleton of the learning subject contained in frames at the second frame rate is extracted as second skeletal information, learning surrounding environment information indicating the results of analyzing the environment around the learning subject from the learning target video is generated, and learning is performed using the second skeletal information and the learning surrounding environment information as input data and the first skeletal information as correct answer data, thereby inputting the inferred skeletal information and the inferred surrounding environment information into a skeletal information interpolation model generated to infer skeletal information at the first frame rate from skeletal information at the second frame rate, thereby functioning as an inference unit that infers skeletal information at the first frame rate.
[0011] A learning method according to one aspect of the present disclosure includes acquiring video at a first frame rate, using the video to extract, as first skeletal information, skeletal information indicating the skeletal structure of a learning subject that is included in frames at the first frame rate, using the video to extract, as second skeletal information, skeletal information indicating the skeletal structure of the learning subject that is included in frames at a second frame rate that is lower than the first frame rate, analyzing the environment surrounding the learning subject from the video, generating surrounding environment information indicating the results of the analysis, and performing learning using the second skeletal information and the surrounding environment information as input data and the first skeletal information as correct answer data, thereby generating a skeletal information interpolation model for inferring skeletal information at the first frame rate from the skeletal information at the second frame rate.
[0012] An inference method according to one aspect of the present disclosure includes acquiring an inference target video, which is video at a second frame rate lower than a first frame rate, using the inference target video to extract skeletal information indicating the skeleton of a person to be inferred that is included in frames at the second frame rate as inference skeletal information, analyzing the environment surrounding the person to be inferred from the inference target video, generating inference surrounding environment information indicating the results of the analysis, using a training target video, which is video at the first frame rate, to extract skeletal information indicating the skeleton of the person to be inferred that is included in frames at the first frame rate as first skeletal information, and using the training target video: The method is characterized in that skeletal information indicating the skeleton of the learner contained in frames at the second frame rate is extracted as second skeletal information, learning surrounding environment information indicating the results of analyzing the environment surrounding the learner from the learning target video is generated, and learning is performed using the second skeletal information and the learning surrounding environment information as input data and the first skeletal information as correct answer data, thereby inferring skeletal information at the first frame rate by inputting the inferred skeletal information and the inferred surrounding environment information into a skeletal information interpolation model generated to infer skeletal information at the first frame rate from skeletal information at the second frame rate.
[0013] According to one or more aspects of the present disclosure, joint information of a person can be inferred based on the surrounding environment in which the person is placed, etc.
[0014] FIG. 1 is a block diagram schematically showing the configuration of a learning inference system according to embodiments 1 and 2. FIG. 2 is a block diagram schematically showing the configuration of a learning device according to embodiment 1. FIG. 3 is a block diagram schematically showing the configuration of a PC. FIG. 4 is a block diagram schematically showing the configuration of an inference device according to embodiment 1. (A) and (B) are schematic diagrams for explaining an example of interpolation. FIG. 5 is a flowchart showing the operation of the learning device according to embodiment 1. FIG. 6 is a flowchart showing the operation of the inference device according to embodiment 1. (A) and (B) are schematic diagrams for explaining the influence of the environment. FIG. 7 is a block diagram schematically showing the configuration of a learning device according to embodiment 2. FIG. 8 is a block diagram schematically showing the configuration of the inference device according to embodiment 2. FIG. 9 is a flowchart showing the operation of the learning device according to embodiment 2. FIG. 10 is a flowchart showing the operation of the inference device according to embodiment 2.
[0015] 1 is a block diagram showing a schematic configuration of a learning inference system 100 according to embodiment 1. The learning inference system 100 includes a learning device 110 and an inference device 130. The learning device 110 and the inference device 130 are connected to a network 101, such as the Internet or a local area network (LAN), for example.
[0016] In the learning inference system 100, the learning device 110 uses low-frame-rate skeletal information to learn a learning model for inferring high-frame-rate skeletal information, and the inference device 130 uses the learning model to infer high-frame-rate skeletal information using the low-frame-rate skeletal information. Note that in the learning inference system 100, a predetermined frame rate is assigned to each of the high and low frame rates. The high frame rate is also referred to as the first frame rate, and the low frame rate is also referred to as the second frame rate.
[0017] 2 is a block diagram showing a schematic configuration of learning device 110 according to embodiment 1. Learning device 110 includes video acquisition unit 111, skeletal information extraction unit 112, surrounding environment analysis unit 113, attribute analysis unit 114, learning unit 115, skeletal information interpolation model storage unit 116, and communication unit 117.
[0018] The video acquisition unit 111 acquires video including learners who are people to be learned and non-learners who are people who are not people to be learned. The acquired video is video with a high frame rate and is also referred to as learner video.
[0019] For example, the image acquisition unit 111 may acquire images from the network 101 via the communication unit 117, or may acquire images from a camera (not shown) connected to the learning device 110 via a connection I / F (Interface) (not shown), or may acquire images stored in a storage unit (not shown). The acquired images are provided to the skeletal information extraction unit 112, the surrounding environment analysis unit 113, and the attribute analysis unit 114.
[0020] The skeletal information extraction unit 112 extracts skeletal information of the learner contained in frames at the high frame rate and skeletal information of the learner contained in frames at the low frame rate, using the high frame rate video provided by the video acquisition unit 111. The high frame skeletal information, which is skeletal information at the high frame rate, and the low frame skeletal information, which is skeletal information at the low frame rate, are provided to the learning unit 115. The low frame skeletal information is also provided to the attribute analysis unit 114. The high frame skeletal information is also referred to as first skeletal information, and the low frame skeletal information is also referred to as second skeletal information.
[0021] Skeletal information is information that indicates, by coordinates, the positions of a person's joints, such as the shoulders, elbows, and waist, in a frame included in the video. Skeletal information extraction unit 112 may extract the skeletal information of the learner using, for example, a known technique such as OpenPose. The learner in the frame may also be identified using a known technique such as pattern matching, or the user of learning device 110 may specify the learner present in the frame using an input unit (not shown).
[0022] The surrounding environment analysis unit 113 analyzes the environment surrounding the learner from the video provided by the video acquisition unit 111 and generates surrounding environment information indicating the results of the analysis. The surrounding environment information generated here is also referred to as learning surrounding environment information. Here, the surrounding environment analysis unit 113 identifies low frame rate video from the high frame rate video provided by the video acquisition unit 111 and uses the low frame rate video to analyze the surrounding environment of the learner. The environment analyzed here is at least one of the location of the learner and the magnitude of movement of people other than the learner.
[0023] For example, the surrounding environment analysis unit 113 identifies the location of the learner from the low frame rate video and generates location information that indicates the location. Here, the location information indicates a number that is pre-assigned to each location. Specifically, the surrounding environment analysis unit 113 generates text from frames included in the low frame rate video using img2text, and identifies the location of the learner from the location name included in the text.
[0024] Furthermore, the surrounding environment analysis unit 113 generates motion information indicating the magnitude of the motion of the non-learners from the low frame rate video. The non-learners may be identified using known techniques such as subtraction from the background. Specifically, the surrounding environment analysis unit 113 identifies the optical flow of the non-learners from two consecutive frames in the low frame rate video, and identifies the speed of the non-learners from the identified optical flow. The surrounding environment analysis unit 113 then generates motion information indicating the identified speed. The location information and motion information generated as described above are provided to the learning unit 115 as surrounding environment information.
[0025] The attribute analysis unit 114 analyzes the attributes of the learner from the high-frame-rate video or low-frame-rate skeletal information provided by the video acquisition unit 111, and generates attribute information indicating the results of the attribute analysis. The attributes may be at least one of the learner's gender, age, and body type. The attribute information generated here is also referred to as learning attribute information. Here, the attribute analysis unit 114 identifies low-frame-rate video from the high-frame-rate video provided by the video acquisition unit 111, and identifies the learner's attributes using the low-frame-rate video and the low-frame-rate skeletal information provided by the skeletal information extraction unit 112. For example, the attribute analysis unit 114 identifies the learner's gender, age, and height as attributes.
[0026] Specifically, the attribute analysis unit 114 may estimate gender and age from low frame rate video using a known technology such as OpenVINO. The attribute analysis unit 114 may also identify height from low frame rate skeletal information. The attribute information indicating the attributes identified in this manner is provided to the learning unit 115.
[0027] The learning unit 115 uses the low-frame skeletal information, surrounding environment information, and attribute information as input data, and performs learning using the high-frame skeletal information as correct answer data, thereby generating a skeletal information interpolation model, which is a learning model for inferring high-frame-rate skeletal information from low-frame-rate skeletal information. Here, the correct answer data is skeletal information of frames of the high-frame skeletal information that are not included in the low-frame skeletal information.
[0028] The skeleton information interpolation model storage unit 116 stores the skeleton information interpolation model generated by the learning unit 115 .
[0029] The communication unit 117 performs communication via the network 101. For example, the communication unit 117 transmits the skeletal information interpolation model stored in the skeletal information interpolation model storage unit 116 to the inference device 130.
[0030] The learning device 110 described above can be realized by, for example, a computer such as the PC 10 shown in Fig. 3. The PC 10 includes a storage 11 such as a hard disk drive (HDD) and a solid state drive (SSD), a memory 12, a processor 13 such as a central processing unit (CPU), a communication interface (I / F) 14 such as a network interface card (NIC), an input interface 15 such as a keyboard and a mouse, and a display 16.
[0031] For example, the skeletal information interpolation model storage unit 116 can be realized by the storage 11. The image acquisition unit 111, the skeletal information extraction unit 112, the surrounding environment analysis unit 113, the attribute analysis unit 114, and the learning unit 115 can be realized by loading a program stored in the storage 11 into the memory 12 and having the processor 13 execute the program. The communication unit 117 can be realized by the communication I / F 14.
[0032] The above programs may be downloaded to the storage 11 from a recording medium (not shown) via a reader / writer (not shown) or from the network 101 via the communication I / F 14, and then loaded onto the memory 12 and executed by the processor 13. Alternatively, the programs may be directly loaded onto the memory 12 from a recording medium via the reader / writer or from the network 101 via the communication I / F 14, and then executed by the processor 13. In other words, the programs may be provided by a program product such as a recording medium.
[0033] 4 is a block diagram showing a schematic configuration of the inference device 130 according to embodiment 1. The inference device 130 includes a communication unit 131, a skeletal information interpolation model storage unit 132, an image acquisition unit 133, a skeletal information extraction unit 134, a surrounding environment analysis unit 135, an attribute analysis unit 136, an inference unit 137, and a skeletal information output unit 138.
[0034] The communication unit 131 performs communication via the network 101. For example, the communication unit 131 receives a skeleton information interpolation model from the learning device 110.
[0035] The skeleton information interpolation model storage unit 132 stores the skeleton information interpolation model received by the communication unit 131 .
[0036] The video acquisition unit 133 acquires video including inference targets who are people to be inferred and non-inference targets who are people not to be inferred. The video acquired here is video with a low frame rate and is also referred to as inference target video.
[0037] For example, the image acquisition unit 133 may acquire images from the network 101 via the communication unit 131, may acquire images from a camera (not shown) connected to the inference device 130 via a connection I / F (not shown), or may acquire images stored in a storage unit (not shown). The acquired images are provided to the skeletal information extraction unit 134, the surrounding environment analysis unit 135, and the attribute analysis unit 136.
[0038] The skeletal information extraction unit 134 extracts skeletal information of the person to be inferred that is included in frames at the low frame rate, using the low frame rate video provided from the video acquisition unit 133. The extracted skeletal information is provided to the inference unit 137 and the attribute analysis unit 136. The skeletal information extracted here is also referred to as inference skeletal information.
[0039] The specific method for extracting skeletal information may be the same as the extraction method used by skeletal information extraction unit 112 of learning device 110. Note that the person to be inferred in the image may be identified using a known technique such as pattern matching, or the user of inference device 130 may specify the person to be inferred present in the image using an input unit (not shown).
[0040] The surrounding environment analysis unit 135 analyzes the surrounding environment of the person to be inferred using the low frame rate video provided from the video acquisition unit 133, and generates surrounding environment information indicating the analysis results. The surrounding environment information generated here is also referred to as inferred surrounding environment information.
[0041] Here, the specific method of analyzing the surrounding environment in the surrounding environment analysis unit 135 may be the same as the analysis method in the surrounding environment analysis unit 113 of the learning device 110. Note that, similar to the learning device 110, the surrounding environment information includes location information and movement information.
[0042] The attribute analysis unit 136 analyzes the attributes of the inference target person based on the low-frame-rate video provided by the video acquisition unit 133 or the skeletal information provided by the skeletal information extraction unit 134, and generates attribute information indicating the results of the analysis. The attribute information generated here is also referred to as inferred attribute information. Here, the attribute analysis unit 136 identifies the attributes of the inference target person using the low-frame-rate video provided by the video acquisition unit 133 and the skeletal information provided by the skeletal information extraction unit 134. For example, the attribute analysis unit 136 identifies the gender, age, and height of the inference target person as attributes. Note that the specific method for identifying attributes by the attribute analysis unit 136 may be the same as the identification method used by the attribute analysis unit 114 of the learning device 110. The attribute information indicating the attributes identified in this manner is provided to the inference unit 137.
[0043] The inference unit 137 infers high frame rate skeletal information by interpolating the low frame rate skeletal information using the low frame rate skeletal information, surrounding environment information, and attribute information.
[0044] For example, the inference unit 137 inputs skeletal information, surrounding environment information, and attribute information as input data into a skeletal information interpolation model stored in the skeletal information interpolation model storage unit 132, and then infers skeletal information of frames not included in the skeletal information from the skeletal information interpolation model.
[0045] The inference unit 137 may perform interpolation by a single interpolation, or may perform extrapolation, as shown in Fig. 5(A) . Furthermore, for example, as shown in Fig. 5(B) , when interpolating three frames of skeleton information 103a, 103b, and 103c between skeleton information 102a of a first frame and skeleton information 102b of a second frame that is the frame following the first frame using skeleton information at a low frame rate, the inference unit 137 may interpolate skeleton information 103b of the center frame in the first interpolation, interpolate skeleton information 103a using the skeleton information 102a and the skeleton information 103b in the second interpolation, and interpolate skeleton information 103c using the skeleton information 103b and the skeleton information 102b in the third interpolation. In other words, interpolation may be performed by performing interpolation multiple times.
[0046] Returning to Figure 4, the skeleton information output unit 138 generates high frame rate skeleton information by adding the skeleton information inferred by the inference unit 137 to the skeleton information provided by the skeleton information extraction unit 134, and outputs the high frame rate skeleton information.
[0047] The skeleton information output unit 138 may, for example, display the skeleton information on a display unit (not shown), or may transmit the high frame rate skeleton information to another device via the communication unit 131 .
[0048] The inference device 130 described above can also be realized by a computer such as the PC 10 shown in Fig. 3. For example, the skeletal information interpolation model storage unit 132 can be realized by the storage 11. The image acquisition unit 133, the skeletal information extraction unit 134, the surrounding environment analysis unit 135, the attribute analysis unit 136, the inference unit 137, and the skeletal information output unit 138 can be realized by loading a program stored in the storage 11 into the memory 12 and having the processor 13 execute the program. The communication unit 131 can be realized by the communication I / F 14.
[0049] 6 is a flowchart showing the operation of the learning device 110 in embodiment 1. The video acquisition unit 111 acquires high frame rate video (S10). The acquired video is provided to the skeletal information extraction unit 112, the surrounding environment analysis unit 113, and the attribute analysis unit 114.
[0050] The skeletal information extraction unit 112 extracts skeletal information of the learner for each frame corresponding to the high frame rate and skeletal information of the learner for each frame corresponding to the low frame rate using the high frame rate video provided from the video acquisition unit 111 (S11). The high frame skeletal information, which is skeletal information at the high frame rate, and the low frame skeletal information, which is skeletal information at the low frame rate, are provided to the learning unit 115. The low frame skeletal information is also provided to the attribute analysis unit 114.
[0051] The surrounding environment analysis unit 113 analyzes the surrounding environment of the subject using the high frame rate video provided from the video acquisition unit 111, and generates surrounding environment information indicating the analysis results (S12). Here, the surrounding environment information includes location information and movement information, and is provided to the learning unit 115.
[0052] The attribute analysis unit 114 uses the high frame rate video provided from the video acquisition unit 111 to identify the gender, age, and height of the subject as attributes, and generates attribute information indicating the identified attributes (S13). The generated attribute information is provided to the learning unit 115.
[0053] The learning unit 115 uses the low-frame skeletal information, surrounding environment information, and attribute information as input data, and performs learning using the skeletal information of frames of the high-frame skeletal information that are not included in the low-frame skeletal information as correct answer data, thereby generating a skeletal information interpolation model (S14). The skeletal information interpolation model generated in this manner is stored in the skeletal information interpolation model storage unit 116 and transmitted to the inference device 130 via the communication unit 117.
[0054] 7 is a flowchart showing the operation of the inference device 130 in embodiment 1. It is assumed here that the communication unit 131 receives the skeletal information interpolation model from the learning device 110, and that the skeletal information interpolation model storage unit 132 has already stored the skeletal information interpolation model received by the communication unit 131.
[0055] The image acquisition unit 133 acquires low frame rate images (S20). The acquired images are provided to the skeleton information extraction unit 134, the surrounding environment analysis unit 135, and the attribute analysis unit 136.
[0056] The skeletal information extraction unit 134 extracts skeletal information of the person to be inferred for each frame included in the low frame rate video provided from the video acquisition unit 133 (S21). The extracted skeletal information is provided to the inference unit 137 and the attribute analysis unit 136.
[0057] The surrounding environment analysis unit 135 analyzes the surrounding environment of the person to be inferred using the low frame rate video provided from the video acquisition unit 133, and generates surrounding environment information indicating the analysis result (S22). The surrounding environment information is provided to the inference unit 137.
[0058] The attribute analysis unit 136 identifies the attributes of the person to be inferred using the low-frame-rate video provided by the video acquisition unit 133 and the skeletal information provided by the skeletal information extraction unit 134, and generates attribute information indicating the identified attributes (S23). The attribute information is provided to the inference unit 137.
[0059] The inference unit 137 infers high frame rate skeletal information by interpolating the low frame rate skeletal information using the low frame rate skeletal information, surrounding environment information, and attribute information (S24).
[0060] The skeletal information output unit 138 generates high frame rate skeletal information by adding the skeletal information inferred by the inference unit 137 to the skeletal information provided by the skeletal information extraction unit 134, and outputs the high frame rate skeletal information (S25).
[0061] In general, when the movement of the person 104b that is not the subject of inference is large, as shown in Fig. 8(A), the movement of the person 104a that is the subject of inference is considered to be larger than when the movement of the person 104b that is not the subject of inference is small, as shown in Fig. 8(B). For example, the movement of a person inside a train is generally considered to be smaller than the movement of a person on a platform.
[0062] Furthermore, if the person to be inferred is elderly, the movements of that person are generally thought to be smaller.
[0063] As described above, according to the first embodiment, in addition to skeletal information at a low frame rate, skeletal information at a high frame rate is learned and inference is performed taking into consideration the attributes of the target person and the environment surrounding that person, thereby enabling highly accurate inference.
[0064] In the first embodiment described above, low-frame skeletal information, surrounding environment information, and attribute information are treated as input data when learning is performed, but the attribute information does not have to be used. In this case, the learning unit 115 performs learning using the low-frame skeletal information and surrounding environment information as input data and the high-frame skeletal information as correct answer data, thereby generating a skeletal information interpolation model for inferring high-frame-rate skeletal information from low-frame-rate skeletal information. The inference unit 137 infers skeletal information at the high frame rate by inputting the low-frame-rate skeletal information and surrounding environment information to the skeletal information interpolation model.
[0065] Second Embodiment As shown in FIG. 1, a learning and inference system 200 according to the second embodiment includes a learning device 210 and an inference device 230 .
[0066] 9 is a block diagram showing a schematic configuration of a learning device 210 according to Embodiment 2. The learning device 210 includes a video acquisition unit 111, a skeletal information extraction unit 112, a surrounding environment analysis unit 113, an attribute analysis unit 114, a learning unit 215, a skeletal information interpolation model storage unit 116, a communication unit 117, an audio acquisition unit 218, and an audio analysis unit 219.
[0067] The video acquisition unit 111, skeletal information extraction unit 112, surrounding environment analysis unit 113, attribute analysis unit 114, skeletal information interpolation model storage unit 116, and communication unit 117 of the learning device 210 in embodiment 2 are similar to the video acquisition unit 111, skeletal information extraction unit 112, surrounding environment analysis unit 113, attribute analysis unit 114, skeletal information interpolation model storage unit 116, and communication unit 117 of the learning device 110 in embodiment 1.
[0068] The audio acquisition unit 218 acquires audio collected in the space where the learner is located, in synchronization with the video acquired by the video acquisition unit 111. The audio acquired here is assumed to have been input to a directional microphone, for example, and the direction of arrival of the audio is identified.
[0069] For example, the voice acquisition unit 218 may acquire voice from the network 101 via the communication unit 117, or may acquire voice from a microphone (not shown) connected to the learning device 210 via a connection I / F (not shown), or may acquire voice stored in a storage unit (not shown). The acquired voice is provided to the voice analysis unit 219.
[0070] The voice analysis unit 219 analyzes the voice of the learning subject from the voice received from the voice acquisition unit 218, and generates voice analysis result information indicating the results of the analysis. At least one of the volume and content of the learning subject's voice is analyzed here.
[0071] Specifically, the voice analysis unit 219 identifies the voice of the learner based on the position of the microphone into which the voice was input and the direction from which the voice came, identifies attributes of the identified voice, and generates voice analysis result information indicating the identified attributes. The voice analysis result information generated here is also referred to as training voice analysis result information. The voice analysis result information is provided to the learning unit 215. In the second embodiment, the identified attributes are, for example, volume and voice content. A known voice recognition technology may be used for the voice content. Note that the position of the microphone may be predetermined, or the microphone may be included in the video acquired by the video acquisition unit 133, and the position of the microphone may be identified using that video.
[0072] The learning unit 215 receives the low frame skeletal information, surrounding environment information, attribute information, and audio analysis result information as input data, and performs learning using the skeletal information of frames of the high frame skeletal information that are not included in the low frame skeletal information as correct answer data, thereby generating a skeletal information interpolation model, which is a learning model for inferring skeletal information of frames that are not included in the low frame skeletal information. The generated skeletal information interpolation model is stored in the skeletal information interpolation model storage unit 116. Note that, as in the first embodiment, the attribute information does not need to be used.
[0073] The learning device 210 described above can also be realized by a computer such as the PC 10 shown in Fig. 3. For example, the speech acquisition unit 218 and speech analysis unit 219 can also be realized by loading a program stored in the storage 11 into the memory 12 and having the processor 13 execute the program.
[0074] 10 is a block diagram showing a schematic configuration of an inference device 230 according to embodiment 2. The inference device 230 includes a communication unit 131, a skeletal information interpolation model storage unit 132, a video acquisition unit 133, a skeletal information extraction unit 134, a surrounding environment analysis unit 135, an attribute analysis unit 136, an inference unit 237, a skeletal information output unit 138, an audio acquisition unit 239, and an audio analysis unit 240.
[0075] The communication unit 131, the skeletal information interpolation model storage unit 132, the video acquisition unit 133, the skeletal information extraction unit 134, the surrounding environment analysis unit 135, the attribute analysis unit 136, and the skeletal information output unit 138 of the inference device 230 in embodiment 2 are similar to the communication unit 131, the skeletal information interpolation model storage unit 132, the video acquisition unit 133, the skeletal information extraction unit 134, the surrounding environment analysis unit 135, the attribute analysis unit 136, and the skeletal information output unit 138 of the inference device 130 in embodiment 1.
[0076] The audio acquisition unit 239 acquires audio collected in the space where the person to be inferred is located, in synchronization with the video acquired by the video acquisition unit 133. The audio acquired here is assumed to have been input to a directional microphone, for example, and the direction of arrival of the audio is specified. The audio acquired here is also referred to as inference audio.
[0077] For example, the voice acquisition unit 239 may acquire voice from the network 101 via the communication unit 131, may acquire voice from a microphone (not shown) connected to the inference device 230 via a connection I / F (not shown), or may acquire voice stored in a storage unit (not shown). The acquired voice is provided to the voice analysis unit 240.
[0078] The voice analysis unit 240 analyzes the voice of the person to be inferred from the voice received from the voice acquisition unit 239, and generates voice analysis result information indicating the results of the analysis. The voice analysis result information generated here is also referred to as an inference voice analysis result. The attribute here may be at least one of the volume and content of the voice of the person to be inferred.
[0079] For example, the voice analysis unit 240 identifies the voice of the person to be inferred from the position of the microphone into which the voice was input and the direction from which the voice came, identifies attributes of the identified voice, and generates voice analysis result information indicating the identified attributes. The voice analysis result information generated here is also referred to as inferred voice analysis result information. The voice analysis result information is provided to the inference unit 237. In the second embodiment, the identified attributes are, for example, the volume and the content of the voice. A known voice recognition technology may be used for the content of the voice. Note that the position of the microphone may be predetermined, or the microphone may be included in the video acquired by the video acquisition unit 133, and the position of the microphone may be identified using that video.
[0080] The inference unit 237 infers high frame rate skeletal information by interpolating the low frame rate skeletal information using the low frame rate skeletal information, surrounding environment information, attribute information, and voice analysis result information. Note that, as in the first embodiment, the attribute information does not need to be used.
[0081] For example, the inference unit 237 inputs skeletal information, surrounding environment information, attribute information, and voice analysis result information as input data into a skeletal information interpolation model stored in the skeletal information interpolation model storage unit 132, and infers skeletal information of frames not included in the skeletal information from the skeletal information interpolation model.
[0082] The inference device 230 described above can also be realized by a computer such as the PC 10 shown in Fig. 3. For example, the voice acquisition unit 239 and the voice analysis unit 240 can also be realized by loading a program stored in the storage 11 into the memory 12 and having the processor 13 execute the program.
[0083] Fig. 11 is a flowchart showing the operation of the learning device 210 in embodiment 2. Among the steps included in the flowchart shown in Fig. 11, steps that perform the same processing as the steps in the flowchart shown in Fig. 6 are assigned the same reference numerals as in Fig. 6.
[0084] The processing in steps S10 to S13 in Fig. 11 is the same as the processing in steps S10 to S13 in Fig. 6. However, after step S13, the processing proceeds to step S34.
[0085] In step S34, the audio acquisition unit 218 acquires audio synchronized with the video acquired by the video acquisition unit 111. The acquired audio is provided to the audio analysis unit 219.
[0086] The voice analysis unit 219 analyzes the voice provided by the voice acquisition unit 218 and generates voice analysis result information indicating the analysis result (S35). The voice analysis result information includes the volume and content of the voice of the learning subject, and is provided to the learning unit 215.
[0087] The learning unit 215 uses the low-frame skeletal information, surrounding environment information, attribute information, and voice analysis result information as input data, and performs learning using skeletal information of frames of the high-frame skeletal information that are not included in the low-frame skeletal information as correct answer data, thereby generating a skeletal information interpolation model (S36). The skeletal information interpolation model generated in this manner is stored in the skeletal information interpolation model storage unit 116 and transmitted to the inference device 230 via the communication unit 117.
[0088] 12 is a flowchart showing the operation of the inference device 230 in embodiment 2. It is assumed here that the communication unit 131 receives the skeletal information interpolation model from the learning device 210, and that the skeletal information interpolation model storage unit 132 has already stored the skeletal information interpolation model received by the communication unit 131.
[0089] Furthermore, among the steps included in the flowchart shown in FIG. 12, steps that perform the same processing as the steps in the flowchart shown in FIG. 7 are assigned the same reference numerals as in FIG.
[0090] The processing in steps S20 to S23 in Fig. 12 is the same as the processing in steps S20 to S23 in Fig. 7. However, after step S23, the processing proceeds to step S44.
[0091] In step S44, the audio acquisition unit 239 acquires audio synchronized with the video acquired by the video acquisition unit 133. The acquired audio is provided to the audio analysis unit 240.
[0092] The voice analysis unit 240 analyzes the voice provided by the voice acquisition unit 239 and generates voice analysis result information indicating the analysis result (S45). The voice analysis result information includes the volume and content of the voice of the person to be inferred, and is provided to the inference unit 237.
[0093] The inference unit 237 infers high frame rate skeletal information by interpolating the low frame rate skeletal information using the low frame rate skeletal information, surrounding environment information, attribute information, and voice analysis result information (S46).
[0094] The skeletal information output unit 138 generates high frame rate skeletal information by adding the skeletal information inferred by the inference unit 237 to the skeletal information provided by the skeletal information extraction unit 134, and outputs the high frame rate skeletal information (S47).
[0095] Generally, it is thought that the movements of a person who speaks loudly or who is uttering abusive or other aggressive content will generally be large. For this reason, skeletal information is learned taking into account the volume and content of the speech, and inference is performed taking into account the volume and content, enabling highly accurate inference.
[0096] 100, 200 Learning inference system, 110, 210 Learning device, 111 Video acquisition unit, 112 Skeleton information extraction unit, 113 Surrounding environment analysis unit, 114 Attribute analysis unit, 115, 215 Learning unit, 116 Skeleton information interpolation model storage unit, 117 Communication unit, 218 Audio acquisition unit, 219 Audio analysis unit, 130, 230 Inference device, 131 Communication unit, 132 Skeleton information interpolation model storage unit, 133 Video acquisition unit, 134 Skeleton information extraction unit, 135 Surrounding environment analysis unit, 136 Attribute analysis unit, 137, 237 Inference unit, 138 Skeleton information output unit, 239 Audio acquisition unit, 240 Audio analysis unit.
Claims
1. A learning device comprising: a video acquisition unit that acquires video at a first frame rate; a skeleton information extraction unit that extracts, as first skeleton information, skeleton information indicating the skeleton of a person to be learned included in a frame at the first frame rate using the video, and also extracts, as second skeleton information, skeleton information indicating the skeleton of the person to be learned included in a frame at a second frame rate lower than the first frame rate using the video; a surrounding environment analysis unit that analyzes the environment around the person to be learned from the video and generates surrounding environment information indicating the result of the analysis; and a learning unit that generates a skeleton information interpolation model for inferring the skeleton information at the first frame rate from the skeleton information at the second frame rate by performing learning with the second skeleton information and the surrounding environment information as input data and the first skeleton information as correct data.
2. The learning device according to claim 1, wherein the environment is at least one of the location where the person to be learned is located and the magnitude of movement of a person other than the person to be learned.
3. The learning device according to claim 1 or 2, further comprising an attribute analysis unit that analyzes the attributes of the person to be learned from the video or the second skeleton information and generates attribute information indicating the result of the analysis of the attributes, wherein the learning unit generates the skeleton information interpolation model by adding the attribute information to the input data.
4. The learning device according to claim 3, wherein the attributes are at least one of the gender, age, and body type of the person to be learned.
5. The learning device according to any one of claims 1 to 4, further comprising: an audio acquisition unit that acquires audio collected in the space where the person to be learned is located in synchronization with the video; and an audio analysis unit that analyzes the audio of the person to be learned and generates audio analysis result information indicating the result of the analysis of the audio, wherein the learning unit generates the skeleton information interpolation model by adding the audio analysis result information to the input data.
6. The learning device according to claim 5, wherein the audio analysis result information indicates at least one of the volume and content of the audio of the person to be learned.
7. A video acquisition unit that acquires an inference target video which is a video with a second frame rate lower than the first frame rate; a skeleton information extraction unit that extracts, as inference skeleton information, skeleton information indicating the skeleton of an inference target person included in a frame at the second frame rate by using the inference target video; a peripheral environment analysis unit that analyzes the environment around the inference target person from the inference target video and generates inference peripheral environment information indicating the result of the analysis; a first skeleton information indicating the skeleton of a learning target person included in a frame at the first frame rate is extracted by using the learning target video which is a video with the first frame rate, and at the same time, skeleton information indicating the skeleton of the learning target person included in a frame at the second frame rate is extracted as second skeleton information by using the learning target video, learning peripheral environment information indicating the result of analyzing the environment around the learning target person is generated from the learning target video, the second skeleton information and the learning peripheral environment information are used as input data, and the first skeleton information is used as correct answer data to perform learning, and the inference skeleton information and the inference peripheral environment information are input to a skeleton information interpolation model generated for inferring the first frame rate skeleton information from the second frame rate skeleton information, and an inference unit that infers the skeleton information at the first frame rate. The inference device is characterized by comprising the above components.
8. The environment is at least one of the place where the learning target person and the inference target person are located and the magnitude of the movement of a person other than the learning target person and the inference target person. The inference device according to claim 7 is characterized by this.
9. The skeleton information interpolation model has also generated, as the input data, attribute information indicating the analysis result of the attributes of the learning target person analyzed from the learning target video. The inference device further comprises an attribute analysis unit that analyzes the attributes of the inference target person from the inference target video or the inference skeleton information and generates inference attribute information indicating the result of the analysis of the attributes of the inference target person. The inference unit inputs the inference attribute information to the skeleton information interpolation model. The inference device according to claim 7 or 8 is characterized by this.
10. The attribute is at least one of gender, age, and body type. The inference device according to claim 9 is characterized by this.
11. The skeleton information interpolation model also generates, in synchronization with the learning target video, voice analysis result information indicating the analysis result of the voice collected in the space where the learning target person is located as the input data. The voice acquisition unit acquires the inference voice, which is the voice collected in the space where the inference target person is located, in synchronization with the inference target video. The voice analysis unit further includes: performing analysis on the voice of the inference target person, and generating inference voice analysis result information indicating the result of the analysis of the voice of the inference target person. The inference unit inputs the inference voice analysis result information to the skeleton information interpolation model. The inference device according to any one of claims 7 to 10, characterized in that.
12. The inference voice analysis result information indicates at least one of the volume and content of the voice of the inference target person. The inference device according to claim 11, characterized in that.
13. A computer, a video acquisition unit that acquires a video at a first frame rate, using the video, the skeleton information indicating the skeleton of the learning target person included in the frame at the first frame rate is extracted as the first skeleton information, and using the video, the skeleton information indicating the skeleton of the learning target person included in the frame at a second frame rate lower than the first frame rate is extracted as the second skeleton information. A skeleton information extraction unit, a peripheral environment analysis unit that analyzes the environment around the learning target person from the video and generates peripheral environment information indicating the result of the analysis, and using the second skeleton information and the peripheral environment information as input data and the first skeleton information as correct data for learning, a learning unit that generates a skeleton information interpolation model for inferring the first frame rate skeleton information from the second frame rate skeleton information. A program characterized by functioning as.
14. A program that causes a computer to function as: a video acquisition unit that acquires an inference target video which is video at a second frame rate lower than a first frame rate; a skeleton information extraction unit that extracts, as inference skeleton information, skeleton information indicating the skeleton of an inference target person included in a frame at the second frame rate, using the inference target video; a surrounding environment analysis unit that analyzes the environment around the inference target person from the inference target video and generates inference surrounding environment information indicating the result of the analysis; and an inference unit that extracts, as first skeleton information, skeleton information indicating the skeleton of a learning target person included in a frame at the first frame rate, using a learning target video which is video at the first frame rate, and also extracts, as second skeleton information, skeleton information indicating the skeleton of the learning target person included in a frame at the second frame rate, using the learning target video, generates learning surrounding environment information indicating the result of analyzing the environment around the learning target person from the learning target video, inputs the second skeleton information and the learning surrounding environment information as input data, and inputs the first skeleton information as correct answer data to perform learning, and then inputs the inference skeleton information and the inference surrounding environment information to a skeleton information interpolation model generated to infer the first frame rate skeleton information from the second frame rate skeleton information, thereby inferring the skeleton information at the first frame rate.
15. A learning method characterized by: acquiring a video at a first frame rate; extracting, as first skeleton information, skeleton information indicating the skeleton of a learning target person included in a frame at the first frame rate, using the video; extracting, as second skeleton information, skeleton information indicating the skeleton of the learning target person included in a frame at a second frame rate lower than the first frame rate, using the video; analyzing the environment around the learning target person from the video and generating surrounding environment information indicating the result of the analysis; and generating a skeleton information interpolation model for inferring the first frame rate skeleton information from the second frame rate skeleton information by performing learning with the second skeleton information and the surrounding environment information as input data and the first skeleton information as correct answer data.
16. Obtain an inference target video that is a video with a second frame rate lower than the first frame rate, and using the inference target video, extract, as inference skeleton information, skeleton information indicating the skeleton of an inference target person included in a frame at the second frame rate. Analyze the environment around the inference target person from the inference target video, and generate inference surrounding environment information indicating the result of the analysis. Using a learning target video that is a video with the first frame rate, extract, as first skeleton information, skeleton information indicating the skeleton of a learning target person included in a frame at the first frame rate, and using the learning target video, extract, as second skeleton information, skeleton information indicating the skeleton of the learning target person included in a frame at the second frame rate. Generate learning surrounding environment information indicating the result of analyzing the environment around the learning target person from the learning target video. Use the second skeleton information and the learning surrounding environment information as input data, and use the first skeleton information as correct answer data to perform learning. Then, input the inference skeleton information and the inference surrounding environment information into a skeleton information interpolation model generated to infer the skeleton information at the first frame rate from the skeleton information at the second frame rate, thereby inferring the skeleton information at the first frame rate. This is a feature of the inference method.
Citation Information
Patent Citations
Image generation device and method
JP2007004732A
Information processing device, information processing method, and imaging system
WO2021205843A1