Program, information processing apparatus, and information processing method
A non-contact system using video and depth data analysis estimates ventilation indices, addressing the limitations of CPX tests by improving accuracy and reducing subject discomfort and costs.
Patent Information
- Application Number
- JP2025200886
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-26
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-10
AI Technical Summary
Existing cardiopulmonary exercise tests (CPX) for determining anaerobic threshold are burdensome for subjects, costly, and limited by equipment availability, with the exhalation mask being uncomfortable.
A non-contact method using a computer-based system that analyzes user video and depth data to estimate ventilation indices through an estimation model, applying skeletal data and health condition data for improved accuracy.
Enables non-contact estimation of ventilation indices, reducing subject burden and costs, while providing accurate assessment of respiratory effort and exercise tolerance.
Smart Images

Figure 2026021634000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a program, an information processing device, and an information processing method. [Background technology]
[0002] Cardiac rehabilitation aims to help heart disease patients regain their strength and self-confidence, return to comfortable home and social life, and prevent recurrence of heart disease or re-hospitalization through a comprehensive activity program including exercise therapy. Exercise therapy focuses on aerobic exercise such as walking, jogging, cycling, and aerobics. To perform aerobic exercise safely and effectively, patients should exercise at an intensity close to their anaerobic threshold (AT).
[0003] The anaerobic threshold is an example of a ventilation index and corresponds to a change point in cardiopulmonary function, i.e., an exercise intensity near the boundary between aerobic exercise and anaerobic exercise. The anaerobic threshold is generally determined by a cardiopulmonary exercise test (CPX test), in which a test subject is subjected to a progressively increasing exercise load while exhaled gas is collected and analyzed. In a CPX test, the anaerobic threshold is determined based on the results measured by exhaled gas analysis (e.g., oxygen intake, carbon dioxide output, tidal volume, respiratory rate, minute ventilation, or a combination thereof). In addition to the anaerobic threshold, a CPX test can also determine the maximum oxygen intake, which corresponds to an exercise intensity near maximum exercise tolerance.
[0004] However, CPX testing has some problems, such as placing a significant physical burden on test subjects, the cost of testing equipment, and limited facilities that can perform the test. Additionally, wearing an exhalation mask, which is required for exhaled gas analysis, is uncomfortable for test subjects.
[0005] Patent Document 1 describes capturing a three-dimensional image of a target area including the chest and abdomen of a subject, acquiring posture information from the three-dimensional image, determining a main respiratory area using the posture information, generating a movement waveform of the main respiratory area, generating a respiratory waveform from the movement waveform, and calculating the ventilation volume per unit time from the respiratory waveform. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Publication No. 2017-217298 Summary of the Invention [Problem to be solved by the invention]
[0007] The non-contact respiration measuring device described in Patent Document 1 may be useful for non-contact measurement of the ventilation volume per unit time of a subject.
[0008] An object of the present disclosure is to provide a novel technique for non-contact estimation of ventilation indices associated with a user's respiratory movements. [Means for solving the problem]
[0009] A program according to one aspect of the present disclosure causes a computer to function as a means for acquiring a user video showing the user's appearance, and a means for making an estimation regarding a ventilation index associated with the user's respiratory movement by applying an estimation model to input data based on the user video. [Effects of the Invention]
[0010] According to the present disclosure, a novel technique can be provided for non-contact estimation of ventilation indices associated with a user's respiratory movements. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a block diagram showing a configuration of an information processing system according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing the configuration of a client device according to the present embodiment. [Figure 3] FIG. 2 is a block diagram showing the configuration of a server according to the present embodiment. [Figure 4] FIG. 1 is an explanatory diagram of an overview of the present embodiment. [Figure 5] FIG. 2 is a diagram illustrating a data structure of a training data set according to the present embodiment. [Figure 6] 4 is a flowchart of information processing according to the present embodiment. [Figure 7] 10A and 10B are diagrams illustrating examples of screens displayed in the information processing of the present embodiment. [Figure 8] FIG. 10 is a diagram showing the data structure of a training data set according to the first modification. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. In the drawings for explaining the embodiment, the same components are generally designated by the same reference numerals, and repeated description thereof will be omitted.
[0013] (1) Information processing system configuration The configuration of the information processing system will now be described with reference to Fig. 1, which is a block diagram showing the configuration of the information processing system according to this embodiment.
[0014] As shown in FIG. 1, the information processing system 1 includes a client device 10 and a server 30. The client device 10 and the server 30 are connected via a network (for example, the Internet or an intranet) NW.
[0015] The client device 10 is an example of an information processing device that transmits a request to the server 30. The client device 10 is, for example, a smartphone, a tablet terminal, or a personal computer.
[0016] The server 30 is an example of an information processing device that provides the client device 10 with a response in response to a request transmitted from the client device 10. The server 30 is, for example, a web server.
[0017] (1-1) Client device configuration The configuration of the client device will now be described with reference to Fig. 2, which is a block diagram showing the configuration of the client device of this embodiment.
[0018] 2, the client device 10 includes a storage device 11, a processor 12, an input / output interface 13, and a communication interface 14. The client device 10 is connected to a display 15, a camera 16, and a depth sensor 17.
[0019] The storage device 11 is configured to store programs and data, and is, for example, a combination of a read-only memory (ROM), a random access memory (RAM), and a storage (for example, a flash memory or a hard disk).
[0020] The programs include, for example, the following programs: OS (Operating System) programs · Programs for applications that process information (e.g., web browsers, rehabilitation apps, or fitness apps)
[0021] The data includes, for example, the following data: Databases referenced in information processing Data obtained by performing information processing (i.e., the results of performing information processing)
[0022] The processor 12 is a computer that implements the functions of the client device 10 by running a program stored in the storage device 11. The processor 12 is, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Gate Array)
[0023] The input / output interface 13 is configured to acquire information (e.g., user instructions, images, sounds) from an input device connected to the client device 10, and to output information (e.g., images, commands) to an output device connected to the client device 10. The input device is, for example, a camera 16, a depth sensor 17, a microphone, a keyboard, a pointing device, a touch panel, a sensor, or a combination thereof. The output device is, for example, a display 15, a speaker, or a combination thereof.
[0024] The communication interface 14 is configured to control communications between the client device 10 and an external device (eg, the server 30). Specifically, the communication interface 14 may include a module for communication with the server 30 (eg, a WiFi module, a mobile communication module, or a combination thereof).
[0025] The display 15 is configured to display an image (a still image or a moving image). The display 15 is, for example, a liquid crystal display or an organic EL display.
[0026] The camera 16 is configured to take pictures and generate image signals.
[0027] The depth sensor 17 is, for example, a light detection and ranging (LIDAR) sensor. The depth sensor 17 is configured to measure the distance (depth) from the depth sensor 17 (i.e., a reference point) to a surrounding object (e.g., a user).
[0028] (1-2) Server configuration The configuration of the server will now be described with reference to Fig. 3, which is a block diagram showing the configuration of the server according to this embodiment.
[0029] As shown in FIG. 3, the server 30 includes a storage device 31, a processor 32, an input / output interface 33, and a communication interface .
[0030] The storage device 31 is configured to store programs and data, and is, for example, a combination of ROM, RAM, and storage.
[0031] The programs include, for example, the following programs: OS programs Application programs that perform information processing
[0032] The data includes, for example, the following data: Databases referenced in information processing - Results of information processing
[0033] The processor 32 is a computer that implements the functions of the server 30 by running a program stored in the storage device 31. The processor 32 is, for example, at least one of the following: ·CPU GPU ASIC FPGA
[0034] The input / output interface 33 is configured to obtain information (for example, a user's instruction) from an input device connected to the server 30 and to output information to an output device connected to the server 30. The input device is, for example, a keyboard, a pointing device, a touch panel, or a combination thereof. The output device is, for example, a display.
[0035] The communication interface 34 is configured to control communications between the server 30 and an external device (eg, the client device 10).
[0036] (2) Overview of the embodiment An outline of this embodiment will be described below with reference to Fig. 4.
[0037] As shown in Fig. 4, the camera 16 of the client device 10 captures an image of the appearance (e.g., the entire body) of the user US1. In the example of Fig. 4, the user US1 is shown performing a bicycle exercise, but the user US1 may perform any exercise (aerobic exercise or anaerobic exercise). Alternatively, the camera 16 may capture an image of the user US1 before or after exercise (including at rest).
[0038] As an example, the camera 16 captures an image of the appearance of the user US1 from the front or obliquely in front. The depth sensor 17 measures the distance (depth) from the depth sensor 17 to each part of the user US1. It is also possible to generate three-dimensional video data by combining, for example, video data (two-dimensional) generated by the camera 16 with, for example, depth data generated by the depth sensor 17.
[0039] The client device 10 analyzes the user's skeleton by referring to at least the video data acquired from the camera 16. The client device 10 may further refer to the depth data acquired from the depth sensor 17 to more appropriately analyze the user's skeleton. The client device 10 transmits data on the skeleton of the user US1 (hereinafter referred to as "user skeleton data") based on the analysis results of the video data (or the video data and the depth data) to the server 30.
[0040] The server 30 estimates the ventilation index associated with the respiratory movement of the user US1 by applying the learned model LM1 (an example of an "estimation model") to the acquired user skeletal data. The server 30 transmits the estimation result (e.g., a numerical value indicating the real-time ventilation index of the user US1) to the client device 10.
[0041] In this way, the information processing system 1 estimates the ventilation index of the user US1 by applying the learned model LM1 to input data based on a video (or video and depth) showing the appearance of the user US1. Therefore, this information processing system 1 can estimate the ventilation index associated with the respiratory movement of the user US1 without contacting the user.
[0042] (3) Training dataset The teacher dataset of this embodiment will now be described with reference to Fig. 5, which is a diagram showing the data structure of the teacher dataset of this embodiment.
[0043] As shown in Figure 5, the training data set includes multiple training data. The training data is used for training or evaluating a target model. The training data includes a sample ID, input data, and correct answer data.
[0044] The sample ID is information that identifies the training data.
[0045] The input data is data that is input to a target model during training or evaluation. The input data corresponds to example problems used during training or evaluation of the target model. As an example, the input data includes skeletal data of a subject. The skeletal data of a subject is data (e.g., feature values) related to the subject's skeleton at the time the subject's video was filmed.
[0046] The subject video data is data related to a subject video that shows the subject's appearance. The subject video is typically a video of the subject that includes at least the upper body of the subject (specifically, at least one of the subject's shoulders, chest, and abdomen) in the captured range. The subject video data can be obtained, for example, by capturing the subject's appearance (e.g., the entire body) from the front or from an angle (e.g., 45 degrees forward) with a camera (e.g., a camera mounted on a smartphone).
[0047] The camera can capture the subject's appearance during exercise, or before or after exercise (including at rest) to obtain subject video data. From the perspective of accurately correlating the correct answer data with the input data, subject video data may be obtained by capturing images of the subject during an exhaled gas test (e.g., a CPX test).
[0048] The subject depth data is data regarding the distance (depth) from the depth sensor to each part of the subject (typically at least one of the shoulders, chest, and abdomen). The subject depth data can be acquired by operating the depth sensor when shooting the subject video.
[0049] The subject may be the same person as the user whose respiratory movement-related ventilation index is estimated during operation of the information processing system 1, or may be a different person. By assuming that the subject and the user are the same person, the target model may learn the user's personality and improve estimation accuracy. On the other hand, allowing the subject to be a different person from the user has the advantage of making it easier to enrich the training data set. Furthermore, the subjects may be composed of multiple people, including the user, or multiple people excluding the user.
[0050] The skeletal data specifically includes data on the speed or acceleration of each part of the subject (which may include data on changes in the parts of the muscles used by the subject or data on the fluctuations felt by the subject).
[0051] At least a portion of the skeletal data can be obtained by analyzing the subject's skeletal structure during the subject video recording by referencing the subject video data (or the subject video data and subject depth data). For example, Vision, an SDK for iOS 14, or other skeletal detection algorithms can be used for skeletal analysis. Alternatively, skeletal data for a training dataset can be obtained by, for example, having the subject exercise while wearing motion sensors on each body part of the subject.
[0052] The skeletal data may also include data obtained by analyzing at least one of the following items with respect to the data relating to the speed or acceleration of each part of the subject described above. Movement (expansion) of the shoulders, chest (which may include the lateral chest), abdomen, or a combination thereof Inhalation time Exhalation time - Use of accessory respiratory muscles
[0053] The correct answer data is data corresponding to the correct answer for the corresponding input data (example question). The target model is trained (supervised learning) to output the input data in a way that is closer to the correct answer data. As an example, the correct answer data includes at least one of a ventilation index or an index that is used to determine the ventilation index. As an example, the ventilation index may include at least one of the following: Ventilation rate Ventilation volume Ventilation rate (i.e., the amount of ventilation per unit time, or the number of ventilations) Ventilation acceleration (i.e., the time derivative of the ventilation rate) However, the ventilation index may be any index for quantitatively grasping respiratory movement, and is not limited to the indexes exemplified here.
[0054] The correct answer data can be obtained, for example, from the results of a breath gas test administered to the subject when the subject's video was being filmed. A first example of a breath gas test is a test (typically a CPX test) conducted while the subject, wearing a breath gas analyzer, is performing exercise with increasing load (e.g., on an ergometer). A second example of a breath gas test is a test conducted while the subject, wearing a breath gas analyzer, is performing exercise with a constant or variable load (e.g., bodyweight exercise, gymnastics, strength training). A third example of a breath gas test is a test conducted while the subject, wearing a breath gas analyzer, is performing any activity. A fourth example of a breath gas test is a test conducted while the subject, wearing a breath gas analyzer, is at rest.
[0055] Alternatively, the correct answer data can be obtained from the results of a respiratory function test (e.g., a pulmonary function test or a spirometry test) performed on the subject when the subject video was shot. In this case, the respiratory function test may be performed using not only medical equipment but also commercially available testing equipment.
[0056] (4) Estimation model The estimation model used by the server 30 corresponds to a trained model created by supervised learning using a training dataset (FIG. 5), or a derived model or distilled model of the trained model.
[0057] (5) Information processing The information processing of this embodiment will be described below. Fig. 6 is a flowchart of the information processing of this embodiment. Fig. 7 is a diagram showing an example of a screen displayed in the information processing of this embodiment.
[0058] The information processing starts when, for example, any of the following start conditions is met. - Information processing is invoked by another process. The user performed an operation to call up information processing. The client device 10 enters a predetermined state (for example, a predetermined application is started). The appointed date and time has arrived. A certain amount of time has passed since a certain event.
[0059] As shown in FIG. 6, the client device 10 performs sensing (S110). Specifically, the client device 10 starts capturing a video of the user's appearance (hereinafter referred to as "user video") by enabling the operation of the camera 16. The user video is typically a video of the user captured so that at least the upper half of the user's body (specifically, at least one of the user's shoulders, chest, and abdomen) is included in the captured range.
[0060] Furthermore, by enabling the operation of the depth sensor 17, the client device 10 starts measuring the distance from the depth sensor 17 to each part of the user (hereinafter referred to as "user depth") when capturing the user video.
[0061] After step S110, the client device 10 executes data acquisition (S111). Specifically, the client device 10 acquires sensing results generated by the various sensors enabled in step S110. For example, the client device 10 acquires user video data from the camera 16 and user depth data from the depth sensor 17.
[0062] After step S111, the client device 10 executes the request (S112). Specifically, the client device 10 references the data acquired in step S111 and generates a request. The client device 10 transmits the generated request to the server 30. The request may include, for example, at least one of the following: Data acquired in step S111 (e.g., user video data or user depth data) Data obtained by processing the data acquired in step S111 User skeleton data acquired by analyzing the user moving image data (or the user moving image data and the user depth data) acquired in step S111
[0063] After step S112, the server 30 performs an estimation (S130) on the ventilation index. Specifically, the server 30 acquires input data for the estimation model based on a request received from the client device 10. The input data includes user skeletal data as well as training data. The server 30 estimates ventilation indices associated with the user's respiratory movement by applying the estimation model to the input data. As an example, the server 30 estimates at least one ventilation indices.
[0064] After step S130, the server 30 executes a response (S131). Specifically, the server 30 generates a response based on the result of the estimation in step S130. The server 30 transmits the generated response to the client device 10. As an example, the response may include at least one of the following: Data corresponding to the results of estimations on ventilation indices Data processed from the estimation results regarding ventilation indices (for example, data on a screen to be displayed on the display 15 of the client device 10, or data referenced to generate the screen)
[0065] After step S131, the client device 10 executes information presentation (S113). Specifically, the client device 10 causes the display 15 to display information based on the response acquired from the server 30 (that is, the result of estimation regarding the ventilation index of the user). However, the information may be presented to the user's instructor (e.g., a medical professional or trainer) on a terminal used by the instructor instead of or in addition to the user. Alternatively, the information may be provided to a computer that can use an algorithm or estimation model to evaluate the user's exercise tolerance based on the ventilatory index. This computer may be located inside or outside the information processing system 1.
[0066] As an example, the client device 10 displays a screen P10 (FIG. 7) on the display 15. The screen P10 includes a display object A10 and an operation object B10. The operation object B10 accepts an operation to specify a ventilation index to be displayed on the display object A 10. In the example of Fig. 7, the operation object B10 corresponds to a check box. The display object A10 displays a graph showing the change over time in the result of estimating the ventilation index of the user. In the example of Fig. 7, the display object A10 displays a graph showing the change over time in the result of estimating the ventilation rate, which is the ventilation index designated by the operation object B10, every minute (i.e., the estimated result of the minute ventilation). When multiple ventilation indices are specified in the operation object B10, the display object A10 may display a graph showing the time-dependent changes in the results of estimating the multiple ventilation indices superimposed thereon, or these graphs may be displayed individually.
[0067] After step S113, the client device 10 ends the information processing (FIG. 6). However, if estimation of the user's ventilation index is performed in real time while the user's video is being captured, the client device 10 may return to data acquisition (S111) after step S113.
[0068] (6) Summary As described above, the information processing system 1 according to the embodiment estimates the ventilation index associated with the user's respiratory movement by applying input data based on a user video showing the user's appearance to the estimation model. This allows the ventilation index associated with the user's respiratory movement to be estimated in a non-contact manner.
[0069] The estimation model may correspond to a trained model created by supervised learning using the training dataset (Figure 5) described above, or a derived model or distilled model of the trained model. This allows for efficient construction of the estimation model. Furthermore, the subject may be the same person as the user. This allows for highly accurate estimation using a model that has learned the user's personality.
[0070] The input data to which the estimation model is applied may include data on the user's skeleton at the time the user video was captured (i.e., user skeleton data), which can improve the estimation accuracy of the ventilation index.
[0071] The input data to which the estimation model is applied may include data on the depth from the reference point (i.e., depth sensor 17) to each part of the user when the user video was captured (i.e., user depth data), which can improve the estimation accuracy of the ventilation index.
[0072] The ventilation index may include at least one of ventilation volume, ventilation rate, or ventilation acceleration, which allows for proper assessment of the user's respiratory effort.
[0073] The user video may be a video of the user captured so that at least the upper body of the user (preferably at least one of the user's shoulders, chest, or abdomen) is included in the captured range, thereby improving the estimation accuracy of the ventilation index.
[0074] The information processing system 1 may present information based on the results of estimating the user's ventilation index. This allows the user or their instructor to be informed of the ventilation index associated with the user's respiratory movement. A human (e.g., an instructor such as a doctor) may evaluate the user's exercise tolerance based on the presented ventilation index. Alternatively, the information processing system 1 may present a predetermined algorithm or estimation model to an available computer, which may then evaluate the user's exercise tolerance using the algorithm or estimation model based on the ventilation index. In other words, the information processing system 1 can be used to support the evaluation of exercise tolerance. As an example, the information processing system 1 may present information regarding changes in the user's ventilation index over time. This can support a human or computer in evaluating the user's exercise tolerance.
[0075] (7) Variation 1 A description will be given of Modification 1. Modification 1 is an example in which input data for an estimation model is modified.
[0076] (7-1) Overview of Modification 1
[0077] An overview of Modification 1 will be described. In this embodiment, an example has been shown in which an estimation model is applied to input data based on a user video. In Modification 1, by applying the estimation model to input data based on both the user video and the user's health condition, it is also possible to estimate ventilation indices associated with the user's respiratory movement.
[0078] The health condition includes at least one of the following: ·age ·sex ·height ·body weight ·Body fat percentage Muscle mass ·Bone density History of current illness ·Past history ·Medication history Surgical history Life history (e.g., smoking history, drinking history, activities of daily living (ADL), frailty score, etc.) Family history Respiratory function test results Test results other than respiratory function tests (e.g., blood tests, urine tests, electrocardiograms (including Holter ECGs), cardiac ultrasounds, X-rays, CT scans (including cardiac morphological CT and coronary artery CT), MRI scans, nuclear medicine scans, PET scans, etc.) Data obtained during cardiac rehabilitation (including Borg index)
[0079] (7-2) Training Dataset The following describes the teacher dataset of Modification 1. Fig. 8 is a diagram showing the data structure of the teacher dataset of Modification 1.
[0080] 8, the training data set of the first modification includes a plurality of training data. The training data is used for training or evaluating a target model. The training data includes a sample ID, input data, and correct answer data.
[0081] The sample ID and the correct answer data are as described in this embodiment.
[0082] The input data is data that is input to the target model during training or evaluation. The input data corresponds to example problems used during training or evaluation of the target model. As an example, the input data is skeletal data of the subject (i.e., relatively dynamic data) and data on the subject's health condition (i.e., relatively static data). The skeletal data of the subject is as described in this embodiment.
[0083] Data on the subject's health condition can be obtained in various ways. The data on the subject's health condition may be obtained before, during, or after the subject's exercise. The data on the subject's health condition may be obtained based on a report from the subject or their doctor, by extracting information linked to the subject in a medical information system, or via the subject's app (e.g., a healthcare app).
[0084] (7-3) Estimation model In the first modification, the estimation model used by the server 30 corresponds to a trained model created by supervised learning using the training data set (FIG. 8), or a derived model or distilled model of the trained model.
[0085] (7-4) Information Processing The information processing of the first modification will be described with reference to FIG.
[0086] In the first modification, the client device 10 performs sensing (S110) in the same manner as in FIG.
[0087] After step S110, the client device 10 executes data acquisition (S111). Specifically, the client device 10 acquires sensing results generated by the various sensors enabled in step S110. For example, the client device 10 acquires user video data from the camera 16 and user depth data from the depth sensor 17.
[0088] Furthermore, client device 10 acquires data related to the user's health condition (hereinafter referred to as "user health condition data"). For example, client device 10 may acquire the user health condition data based on an operation (declaration) by the user or the user's doctor, or may acquire the user health condition data by extracting information linked to the user in a medical information system, or may acquire the user health condition data via the user's app (e.g., a healthcare app). However, client device 10 may acquire the user health condition data at a timing different from step S111 (e.g., before step S110, at the same timing as step S110, or after step S111).
[0089] After step S111, the client device 10 executes the request (S112). Specifically, the client device 10 references the data acquired in step S111 and generates a request. The client device 10 transmits the generated request to the server 30. The request may include, for example, at least one of the following: Data acquired in step S111 (e.g., user video data, user depth data, or user health status data) Data obtained by processing the data acquired in step S111 User skeleton data acquired by analyzing the user moving image data (or the user moving image data and the user depth data) acquired in step S111
[0090] After step S112, the server 30 performs an estimation (S130) on the ventilation index. Specifically, the server 30 acquires input data for the estimation model based on a request received from the client device 10. The input data includes user skeletal data and user health condition data, as well as training data. The server 30 estimates ventilation indices associated with the user's respiratory movement by applying the estimation model to the input data. As an example, the server 30 estimates at least one ventilation indices.
[0091] After step S130, the server 30 executes a response (S131) in the same manner as in FIG. After step S131, the client device 10 performs information presentation (S113) in the same manner as in FIG.
[0092] (7-5) Summary As described above, the information processing system 1 of the first modification estimates the ventilation index associated with the user's breathing movement by applying an estimation model to input data based on both the user's video and the user's health condition. This allows for highly accurate estimation by further taking the user's health condition into consideration. For example, even if there is a difference between the user's health condition and the health condition of the subject from which the training data was derived, a valid estimation can be made.
[0093] (8) Other variations The storage device 11 may be connected to the client device 10 via a network NW. The display 15 may be built into the client device 10. The storage device 31 may be connected to the server 30 via the network NW.
[0094] The information processing system of the embodiment and the first modification has been described as being implemented as a client / server system. However, the information processing system of the embodiment and the first modification may also be implemented as a stand-alone computer. As an example, the client device 10 may independently perform estimation of the ventilation index using an estimation model.
[0095] Each step of the above information processing can be executed by either the client device 10 or the server 30. As an example, the server 30, instead of the client device 10, may acquire user skeletal data by analyzing the user video (or the user video and user depth).
[0096] In the above description, an example has been given in which a user video is captured using the camera 16 of the client device 10. However, the user video may be captured using a camera other than the camera 16. In the above description, an example has been given in which the user depth is measured using the depth sensor 17 of the client device 10. However, the user depth may be measured using a depth sensor other than the depth sensor 17.
[0097] The information processing system 1 of this embodiment and Modification 1 can also be applied to a video game in which the game progress is controlled according to the player's physical movements. As an example, the information processing system 1 may estimate the user's ventilation index during game play and determine one of the following depending on the result of the estimation. This can enhance the effect that the video game has on improving the user's health. The quality (e.g., difficulty) or quantity of video game challenges (e.g., stages, missions, quests) provided to users The quality (e.g., type) or quantity of video game benefits (e.g., in-game currency, items, bonuses) provided to users
[0098] A microphone mounted on or connected to the client device 10 may receive sound waves emitted by the user (e.g., sounds generated by breathing or speaking) when capturing the user video, and generate sound data. The sound data, together with the user skeletal data, may constitute input data for the estimation model.
[0099] In the above description, a CPX test is exemplified as a test related to exhaled gas. In a CPX test, a gradually increasing exercise load is applied to the test subject. However, it is not necessary to gradually increase the exercise load applied to the user when recording the user video. Specifically, real-time ventilation indices can be estimated even when the user is subjected to a constant or variable exercise load, or even when the user is at rest. For example, the exercise performed by the user may be bodyweight exercise, calisthenics, or strength training.
[0100] In the first modification, an example was shown in which an estimation model was applied to input data based on a health condition. However, it is also possible to construct multiple estimation models based on (at least a part of) the subject's health condition. In this case, (at least a part of) the user's health condition may be referenced to select an estimation model. In this further modification, the input data for the estimation model may be data that is not based on the user's health condition, or may be data based on the user's health condition and a user video (e.g., user skeletal data).
[0101] Acceleration data can also be used as part of the input data for the estimation model. Alternatively, the user's skeletal structure can be analyzed by referring to the acceleration data. The acceleration data can be obtained, for example, by having the user carry or wear a client device 10 or a wearable device equipped with an acceleration sensor when capturing the user's video.
[0102] Oxygen saturation data can also be used as part of the input data for the estimation model. The oxygen saturation data can be obtained, for example, by having the user wear a wearable device equipped with a sensor (e.g., an optical sensor) capable of measuring blood oxygen levels or a pulse oximeter while recording the user's video. The oxygen saturation data can be estimated, for example, by performing rPPG (Remote Photoplethysmography) analysis on the user's video data.
[0103] Although the embodiments and modifications of the present invention have been described in detail above, the scope of the present invention is not limited to the above-described embodiments and modifications. Furthermore, the above-described embodiments and modifications can be improved or modified in various ways without departing from the spirit of the present invention. Furthermore, the above-described embodiments and modifications can be combined. [Explanation of symbols]
[0104] 1: Information processing system 10: Client device 11:Storage device 12: Processor 13: Input / output interface 14: Communication interface 15: Display 16: Camera 17: Depth sensor 30: Server 31:Storage device 32: Processor 33: Input / output interface 34: Communication interface
Claims
1. Computer, A means for acquiring a user video showing the user's appearance; a means for applying an estimation model to input data based on the user's video to estimate a ventilation index associated with the user's respiratory movement; A program that functions as a
2. The estimation model corresponds to a trained model created by supervised learning using a teacher dataset including input data including data on subject videos showing the appearance of the subject and correct answer data associated with each of the input data, or a derived model or distilled model of the trained model. The program according to claim 1.
3. The input data to which the estimation model is applied includes data related to the user's skeleton. The program according to claim 1 or 2.
4. The input data to which the estimation model is applied is further based on data regarding depths from a reference point to each part of the user. The program according to any one of claims 1 to 3.
5. The ventilation index includes at least one of ventilation volume, ventilation rate, or ventilation acceleration; 5. The program according to claim 1.
6. The user video is a video of the user captured so that at least the upper body of the user is included in the captured range.
6. The program according to claim 1.
7. the user's upper body includes at least one of the user's shoulders, chest, or abdomen; The program according to claim 6.
8. The subject is the same person as the user. The program according to claim 2.
9. and further causing the computer to function as a means for presenting information based on the result of the estimation of the ventilation index of the user.
9. The program according to claim 1.
10. The presenting means presents information regarding a change in the ventilation index of the user over time. The program according to claim 9.
11. A means for acquiring a user video showing the user's appearance; means for applying an estimation model to input data based on the user video to estimate a ventilation index associated with the user's respiratory movement; An information processing device comprising:
12. The computer Obtaining a user video showing the user's appearance; a means for applying an estimation model to input data based on the user's video to estimate a ventilation index associated with the user's respiratory movement; A program that functions as a
Citation Information
Patent Citations
Non-contact respiration measurement device and non-contact respiration measurement method
JP2017217298A