Posture analysis program, information processing device, and posture analysis method

JP2026144493APending Publication Date: 2026-09-09UNIVERSITY OF FUKUI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025031810
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2026-09-09

Smart Images

  • Figure 2026144493000001_ABST
    Figure 2026144493000001_ABST
Patent Text Reader

Abstract

This invention provides a posture analysis program, an information processing device, and a posture analysis method that do not require instructional learning and reflect the user's main symptom in the assessment results. [Solution] The posture analysis program 110 functions as a compression means 103 using a neural network that compresses structured data 112, which is a time series of one or more feature points extracted from video footage of multiple users, into compressed data 113 with fewer dimensions than n when the structured data is n dimensional; a classification means 104 that assigns labels to the compressed data 113 based on questionnaires given to multiple users and classifies the compressed data 113 based on the results of the label assignment; a restoration means 105 using a neural network that restores one coordinate of the classified compressed data 113 to structured data; and a playback means 106 that plays back the restored structured data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a posture analysis program, an information processing apparatus, and a posture analysis method.

Background Art

[0002] As a conventional technique, a posture analysis program that determines whether the posture of a VDT (Visual Display Terminal) user is good or needs improvement has been proposed (see, for example, Non-Patent Document 1).

[0003] The posture analysis program disclosed in Non-Patent Document 1 prepares images of a user during VDT work, performs supervised learning by attaching labels of "good" or "needs improvement" to the images, adopts the model having the highest matching rate between the determination results of researchers and the determination results of the learning model, and determines whether the posture of the VDT user is good or needs improvement using the model.

Prior Art Literature

Non-Patent Literature

[0004]

Non-Patent Literature 1

Summary of the Invention

Problem to be Solved by the Invention

[0005] Although the above posture analysis program determines whether the posture of a VDT user is good or needs improvement using a trained model, it has a problem that labeled images are required for training. Furthermore, since it adopts determination results from researchers, it has a problem that symptoms that are chief complaints of VDT users, such as shoulder stiffness and lower back pain felt by VDT users, cannot be reflected in the determination results.

[0006] Therefore, the object of the present invention is to provide a posture analysis program, an information processing device, and a posture analysis method that do not require taught learning and reflect the user's main complaint symptoms in the judgment results. [Means for solving the problem]

[0007] One aspect of the present invention provides the following posture analysis program, information processing device, and posture analysis method to achieve the above objective.

[0008] [1] Computers, A compression method using a neural network that compresses structured data, which is a time series of single or multiple feature points extracted from video footage of multiple users, into compressed data of a smaller dimension than n, when the structured data is n-dimensional, A classification means that assigns labels to the compressed data based on the surveys of the aforementioned multiple users, and classifies the compressed data based on the results of the label assignment, A restoration means using a neural network to restore one coordinate in the classified compressed data to structured data, A posture analysis program that functions as a playback means for reproducing the restored structured data. [2] The compression means is the posture analysis program described in [1], which determines the dimension of the compressed data according to the number of labels and / or the restoration accuracy of the restoration means. [3] The regeneration means is the posture analysis program described in [1] which regenerates the restored structured data by the difference between it and coordinates other than the first coordinate. [4] The posture analysis program according to [1], wherein the compression means compresses n-dimensional structured data, the restoration means restores the generated compressed data, and the program further functions as a learning means for training the compression means and the restoration means to reduce the difference between the restored n-dimensional structured data and the n-dimensional structured data before compression. [5] When structured data, which is a time series of one or more feature points extracted from video footage of multiple users, is n-dimensional, a compression method using a neural network is provided to compress the structured data into compressed data of a smaller dimension than n. A classification means that assigns labels to the compressed data based on the surveys of the aforementioned multiple users, and classifies the compressed data based on the results of the label assignment, A restoration means using a neural network to restore one coordinate in the classified compressed data to structured data, An information processing apparatus having a playback means for reproducing the restored structured data. [6] When structured data, which is a time series of one or more feature points extracted from video footage of multiple users, is n-dimensional, a compression step is performed using a neural network to compress the structured data into compressed data of less than n dimensions, A classification step in which labels are assigned to the compressed data based on the surveys of the aforementioned multiple users, and the compressed data is classified based on the results of the label assignment, A restoration step in which a coordinate in the classified compressed data is restored to structured data using a neural network, A posture analysis method comprising a playback step of reproducing the restored structured data. [Effects of the Invention]

[0009] According to the inventions of claims 1, 5, and 6, it is possible to reflect the user's main symptom in the judgment result without requiring taught learning. According to the invention of claim 2, the dimensions of the compressed data can be determined according to the number of labels and / or the restoration accuracy of the restoration means. According to the invention of claim 3, the restored structured data can be reconstructed using the difference between a coordinate other than one of the coordinates. According to the invention according to claim 4, a compression means compresses n-dimensional structured data, a restoration means restores the generated compressed data, and the compression means and the restoration means can be trained such that the difference between the restored n-dimensional structured data and the n-dimensional structured data before compression is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] [Figure 1] Figure 1 is a schematic diagram showing a configuration example of a system according to an embodiment. [Figure 2] Figure 2 is a block diagram showing a configuration example of an information processing apparatus according to an embodiment. [Figure 3] Figures 3(a) and 3(b) are schematic diagrams for explaining a feature point detection operation. [Figure 4] Figures 4(a) and 4(b) are schematic diagrams for explaining a structuring operation on detected feature points. [Figure 5] Figure 5 is a graph showing a display example of structured data. [Figure 6] Figure 6 is a schematic diagram showing an operation example of compressing structured data into compressed data. [Figure 7] Figure 7 is a graph showing a configuration example of compressed data. [Figure 8] Figure 8 is a schematic diagram showing an example of a classification operation on compressed data. [Figure 9] Figure 9 is a schematic diagram showing an example of data labeled with eye strain among the compressed data. [Figure 10] Figure 10 is a schematic diagram showing an example of data labeled with stiff shoulders among the compressed data. [Figure 11] Figure 11 is a schematic diagram showing an example of data labeled with back pain among the compressed data. [Figure 12] Figures 12(a) and 12(b) are graphs for explaining a restoration operation from compressed data to structured data. [Figure 13] Figure 13 is a schematic diagram for explaining a learning operation. [Modes for carrying out the invention]

[0011] [Embodiment] (System Configuration) Figure 1 is a schematic diagram showing an example of the system configuration according to the embodiment.

[0012] The system comprises an information processing device 1, such as a server or PC (Personal Computer), which processes information; a terminal 2, such as a PC, which acts as a VDT (Visual Display Terminal); and a network 3 that connects the information processing device 1 and the terminal 2 to enable communication. The terminal 2 is operated by a user 4 and has a camera 22 for capturing images of the user 4, a keyboard 23 for inputting characters and numbers, and a display 24 for displaying images. Alternatively, the information processing device 1 may be configured as an integrated system, with the functions of the information processing device 1 being executed by the terminal 2, rather than being configured separately.

[0013] In the system described above, as an outline of this embodiment, the information processing device 1 captures the posture of user 4, who is a VDT user, while he is working using camera 22, extracts feature points such as the eyes, nose, mouth, and shoulders of user 4 from the captured video data, performs a dimensionality reduction process on the time-series data of the feature points, and creates data for posture analysis by labeling the resulting compressed data based on user 4's chief complaint. It also provides a method for analyzing the parameter trends that appear when dimensionality is reduced by restoring the compressed data and replaying the restored time-series data of the feature points. This embodiment will be described in detail below.

[0014] (Configuration of information processing device) Figure 2 is a block diagram showing an example configuration of the information processing device 1 according to an embodiment.

[0015] The information processing device 1 comprises a control unit 10 which consists of a CPU (Central Processing Unit) and the like, controls each part and executes various programs; a storage unit 11 which consists of a storage medium such as flash memory and stores information; and a communication unit 12 which communicates with the outside via a network.

[0016] The control unit 10 functions as a video acquisition means 100, a feature point detection means 101, a structured data generation means 102, a compression means 103, a classification means 104, a restoration means 105, a playback means 106, a learning means 107, etc., by executing the posture analysis program 110 described later.

[0017] The video acquisition means 100 acquires the video captured by the camera 22 of the terminal 2 and stores it in the storage unit 11 as video data 111.

[0018] The feature point detection means 101 uses an artificial intelligence model to detect feature points (three dimensions: length, width, and depth), such as skeletal information, from image data extracted frame by frame from the video data 111. The artificial intelligence model and feature points will be described later.

[0019] The structured data generation means 102 structures the feature points detected by the feature point detection means 101 based on their coordinates and time, and stores them in the storage unit 11 as structured data 112. The structured data 112 will be described later.

[0020] The compression means 103 compresses the structured data 112 by reducing the number of parameters using a multilayer neural network model, and stores the compressed data 113 in the storage unit 11. Specifically, an Autoencoder is used as the multilayer neural network model.

[0021] The classification means 104 classifies the data by reflecting the label data 114, which represents the main symptoms of VDT users obtained through questionnaires, etc., into the compressed data 113.

[0022] The restoration means 105, like the compression means 103, uses a multilayer neural network model to restore the compressed data 113 as structured data.

[0023] When the combination of selected numerical values ​​that are parameters (latent variables, described later) of the compressed data 113 is restored by the restoration means 105, the restoration means 106 reproduces the structured data, which is the result of the restoration, as the movement of feature points.

[0024] The learning means 107 compresses certain structured data using the compression means 103, and then trains the compression means 103 and the restoration means 105 to minimize the difference between the result of restoring the compressed data using the restoration means 105 and the original structured data.

[0025] The memory unit 11 stores a posture analysis program 110, video data 111, structured data 112, compressed data 113, label data 114, etc., which cause the control unit 10 to operate as the means 100-107 described above.

[0026] (Operation of information processing device) Next, the operation of this embodiment will be explained.

[0027] First, user 4 operates terminal 2, which serves as a VDT (Visual Display Terminal). Specifically, user 4 operates the keyboard 23, a mouse (not shown), or a touchpad while checking the content displayed on the display 24.

[0028] Next, the camera 22 of terminal 2 takes a picture of user 4. In this embodiment, as an example, the camera takes a picture of user 4's upper body and outputs video data.

[0029] Next, the video acquisition means 100 of the information processing device 1 acquires the video data output from the results captured by the camera 22 of the terminal 2 and stores it in the storage unit 11 as video data 111.

[0030] Next, the feature point detection means 101 uses an artificial intelligence model to detect multiple feature points, such as skeletal information, in the image data extracted frame by frame from the video data 111. The feature points are detected as described in Figures 3(a) and (b) below. For example, OpenPose, PoseNet, etc., can be used as the artificial intelligence model.

[0031] Figures 3(a) and 3(b) are schematic diagrams illustrating the feature point detection operation.

[0032] As shown in Figures 3(a) and (b), 32 feature points 101a can be detected from the entire human body in the video data 111a. However, in this embodiment, it is sufficient to analyze the posture of the VDT user, so the feature points 101A to be detected are the upper body of the user, specifically the nose, left inner corner of the eye, left center of the eye, left outer corner of the eye, right inner corner of the eye, right center of the eye, right outer corner of the eye, left corner of the mouth, right corner of the mouth, left ear, right ear, left shoulder, and right shoulder, which are labeled with numbers 0 to 12 and detected as 13 feature points (in three dimensions: vertical, horizontal, and depth).

[0033] Next, as shown in Figures 4(a) and (b) below, the structured data generation means 102 structures the feature points detected by the feature point detection means 101 based on their coordinates and time, and stores them in the storage unit 11 as n=13×3(x,y,z)+1(t)=40-dimensional structured data 112.

[0034] Figures 4(a) and (b) are schematic diagrams illustrating the structuring process of the detected feature points.

[0035] As shown in Figure 4(a), each feature point 101b is detected from each frame 111b of the video data, and structured data 112b is obtained for each feature point, structured as time-series data of coordinates x, y, and z, as shown in Figure 4(b). This structured data 112b may be visualized as an image, as shown in Figure 5 below.

[0036] Figure 5 is a graph illustrating an example of how structured data is displayed.

[0037] Structured data 112b is represented as a graph where the vertical axis represents the type of feature point and the horizontal axis represents time, with the numerical values ​​of x, y, and z represented by varying shades of gray.

[0038] Next, the compression means 103 compresses the structured data 112 by reducing the number of parameters to less than n=40 dimensions using a multilayer neural network model, for example, by reducing it to 2 dimensions, and stores the compressed data 113 in the storage unit 11. Specifically, an Autoencoder is used as the multilayer neural network model for the reasons explained below, but other models may be used if similar results can be obtained.

[0039] Several methods have been proposed for data dimensionality reduction. Examples include principal component analysis (PCA) and singular value decomposition (SVD). However, since these methods are linear, they may be difficult to apply to data with a nonlinear structure. Furthermore, while it is possible to capture nonlinear relationships through preprocessing, determining the specific preprocessing to perform requires deep knowledge of the data. Therefore, in this embodiment, we have adopted an Autoencoder, which is a multilayer neural network model. We believe this will address the aforementioned challenges.

[0040] The model overview is as follows: An Autoencoder is a multilayer neural network model composed of two types of networks: an Encoder (compression model) and a Decoder (reconstruction model). For the most common Autoencoder model (AE), the number of input nodes in each of the Encoder and Decoder (Encoder input Decoder input ) and number of output nodes (Encoder output Decoder output The relationship between the two is as follows:

number

[0041] The input and output layers of the Encoder and Decoder are determined to satisfy equations (1)-(5) of the above equation 1. Furthermore, the configuration of the intermediate layers is chosen to suit the nature of the task and data being applied. For example, for image processing, a Convolutional Neural Network (CNN), which primarily exhibits spatial arrangement, may be used, while for time series processing, a Recurrent Neural Network (RNN), which primarily exhibits temporal progression, may be used. The Autoencoder in this embodiment is primarily composed of CNNs.

[0042] Next, we will explain parameter optimization (learning). Because the Autoencoder model uses lossy compression, an error occurs between the source data and the data compressed and restored by the Autoencoder. To minimize this error, we use backpropagation (BP method) to optimize the parameters of the intermediate layers of the Encoder and Decoder. In particular, since this proposal uses a CNN, we optimize the filters used during the convolution operation. This operation is performed by the learning means 107, which compresses a certain structured data using the compression means 103, and then trains the compression means 103 and the restoration means 105 so that the difference between the result of restoring the compressed data by the restoration means 105 and the original structured data becomes small.

[0043] Figure 13 is a schematic diagram illustrating the learning process.

[0044] The learning means 107 compresses the structured data 102b using the compression means 103, and then trains the compression means 103 and the restoration means 105 to reduce the difference between the resulting structured data 105b and the original structured data 102b.

[0045] Figure 6 is a schematic diagram illustrating an example of the process of compressing structured data into compressed data.

[0046] The structured data 112b is sequentially compressed using multiple layers 1030a, 1030b, and 1030c of the compression means 103 to reduce the number of parameters (dimensions) and obtain compressed data 113b. As an example, the case where the number of parameters is reduced to 2 is described below.

[0047] Figure 7 is a graph showing an example of the structure of compressed data.

[0048] The compressed data 113b is represented on a graph with the parameters obtained as a result of the compression operation (hereinafter referred to as "latent variable 1" and "latent variable 2") as the x-axis and y-axis, respectively. The meaning of latent variable 1 and latent variable 2 is currently unknown, and the trends will be understood through the classification operation described later.

[0049] Here, we will explain how to determine the dimensionality of the latent variables. Since the Autoencoder is a model aimed at information compression, the Encoder input / Encoder output A larger compression ratio, or a higher compression rate, indicates better performance. However, since compression using Autoencoder (AE) is irreversible, even with optimized parameters, a larger compression ratio tends to result in a larger error between the original input data and the data compressed and restored by the Autoencoder. Therefore, although it depends on the number of dimensions of the input data, it is desirable to use a few dimensions, taking into account the compression ratio and the error. Furthermore, when the space into which the input data is mapped by the Encoder is used as a latent variable space, a low number of dimensions in the latent variable space can potentially facilitate interpretation when interpreting the meaning of each dimension in this space.

[0050] Next, the classification means 104 reflects the label data 114 obtained from questionnaires, etc., into the compressed data 113 and classifies it. The questionnaire, for example, asks questions corresponding to the main complaints of VDT users, such as "I feel eye strain," "I feel stiff shoulders," and "I feel back pain," to which respondents answer "yes" or "no." However, it may also involve multiple-level answers, or the content of free-response answers may be classified, or symptoms diagnosed by a doctor or other professional may be assigned as labels instead of using a questionnaire.

[0051] Figure 8 is a schematic diagram illustrating an example of compressed data classification.

[0052] The graph in Figure 8 shows, for example, the result of the classification means 104 assigning labels "0", "1", and "3" to each data point. As a result, it can be seen that the compressed data 113b is divided into groups 113b1, 113b2, and 113b3. Here, for the three chief complaints of eye strain, stiff shoulders, and back pain, label "0" represents the group with no applicable chief complaint, "1" represents the group with one chief complaint, and "3" represents the group with all three chief complaints.

[0053] Figure 9 is a schematic diagram showing an example of compressed data to which the label for eye strain has been assigned.

[0054] The graph shown in Figure 9 shows, for example, the results of the classification means 104 assigning labels "0" to each data point to "experiencing eye strain" and "1" to data other than "experiencing eye strain." As a result of this assignment, it can be seen that group 113b1 of the compressed data 113b is the group that "experiencing eye strain."

[0055] Figure 10 is a schematic diagram showing an example of compressed data to which the label "shoulder stiffness" has been assigned.

[0056] Furthermore, the graph shown in Figure 10 shows, for example, the results of the classification means 104 assigning labels "0" to each data point to "feeling stiff shoulders" and "1" to data other than "feeling stiff shoulders." As a result of this assignment, it can be seen that groups 113b1 and 113b2 of the compressed data 113b are groups that "feel stiff shoulders."

[0057] Figure 11 is a schematic diagram showing an example of compressed data labeled with back pain.

[0058] Furthermore, the graph shown in Figure 11 shows, for example, the results of the classification means 104 assigning labels "0" to each data point, such as "experiencing back pain," and "1" to data other than "experiencing back pain." As a result of this assignment, it can be seen that groups 113b1 and 113b2 of the compressed data 113b are groups that "experiencing back pain."

[0059] As described above, by classifying the compressed data 113, the trends of latent variables 1 and 2 can be determined. The trends are that latent variable 1 tends to indicate eye strain, and latent variable 2 tends to indicate stiff shoulders and back pain. Another method for understanding the trends of latent variables 1 and 2 is to check what kind of structured data 112 the points (coordinates) on the graph correspond to. To perform this method, the restoration means 105 restores the compressed data 113 as structured data using a multilayer neural network model.

[0060] Furthermore, the playback means 106 reproduces the result of the restoration means 105 restoring the selected numerical combination of parameters from the compressed data 113 as the movement of feature points. This allows users to confirm how the points on the graph of the compressed data 113 correspond to the movement of feature points. Confirming this movement of feature points helps in understanding postures, etc., that correspond to symptoms. Alternatively, the playback means 106 may reproduce the difference between two points on the compressed data 113. This can be used, for example, when analyzing the difference between an ideal posture without symptoms and a posture with a certain symptom.

[0061] Figures 12(a) and (b) are graphs illustrating the process of restoring compressed data to structured data.

[0062] For example, as shown in Figure 12(a), if point 106b1 is specified, the restoration means 105 restores the parameters (latent variable 1, latent variable 2) = (0.5, 1.1) of point 106b1 to structured data 106B1 as shown in Figure 12(b). Note that the structured data 106B1 shown in Figure 12(b) only shows coordinates due to the representation constraints of the drawing, but in reality it also has a time dimension and involves movement.

[0063] (Effects of the embodiment) According to the embodiment described above, the posture of the VDT user is filmed, the filmed video is converted into structured data 112, the structured data 112 is compressed into compressed data 113 by the compression means 103 to reduce the number of parameters, and label data 114 such as symptoms is added to the compressed data 113 for classification. This makes it possible to grasp the trend of the parameters of the classification result. In other words, the symptoms that constitute the user's main complaint can be reflected in the judgment result.

[0064] Furthermore, since the system is designed to learn so that the data input to the compression means 103 matches the data output by the restoration means 105, labels for the VDT user's posture are unnecessary, meaning that taught learning is not required.

[0065] Furthermore, by specifying the coordinates of the compressed data 113, the restoration means 105 can restore the structured data 112 corresponding to those coordinates, and by playing back the structured data 112 with the playback means 106, the posture and movements of the VDT user corresponding to those coordinates can be confirmed.

[0066] [Other embodiments] It should be noted that the present invention is not limited to the embodiments described above, and various modifications are possible without departing from the spirit of the invention. For example, although the posture of VDT users was analyzed, the invention can be applied to any device that analyzes the posture or movements of users, not limited to VDTs. For example, it can be applied to detecting postures that lead to decreased vision during learning (labeled by visual acuity), detecting distractions, drowsiness, and intoxication in drivers (cars, trains, etc.) (labeled by each action), detecting fatigue in workers during desk work (labeled by degree of fatigue and type of fatigue), and evaluating the comfort of passenger seats in aircraft and trains (labeled by comfort and degree of comfort).

[0067] Furthermore, while we have described an example where the person reviewing the results examines the relationship between latent variables and labels, it is also possible to derive the relationship between latent variables and labels using an artificial intelligence model.

[0068] In the above embodiment, the functions of each means 100 to 107 of the control unit 10 were implemented by program, but all or part of each means may be implemented by hardware such as an ASIC. Furthermore, the program used in the above embodiment can be stored and provided on a recording medium such as a CD-ROM. Also, the steps described in the above embodiment can be rearranged, deleted, or added without altering the essence of the present invention. [Explanation of symbols]

[0069] 1: Information Processing Device 2: Terminal 3: Network 4: User 10: Control Unit 11: Storage section 12: Communications Department 22: Camera 23: Keyboard 24: Display 100: Video acquisition method 101: Feature point detection means 102: Structured Data Generation Method 103: Compression means 104: Classification means 105: Restoration means 106 :Reproduction means 107: Learning Methods 110: Posture Analysis Program 111: Video data 112: Structured data 113: Compressed data 114: Label data

Claims

1. Computers, A compression method using a neural network that compresses structured data, which is a time series of single or multiple feature points extracted from video footage of multiple users, into compressed data of a smaller dimension than n, when the structured data is n-dimensional, A classification means that assigns labels to the compressed data based on the surveys of the aforementioned multiple users, and classifies the compressed data based on the results of the label assignment, A restoration means using a neural network to restore one coordinate in the classified compressed data to structured data, A posture analysis program that functions as a playback means for reproducing the restored structured data.

2. The posture analysis program according to claim 1, wherein the compression means determines the dimension of the compressed data according to the number of labels and / or the restoration accuracy of the restoration means.

3. The posture analysis program according to claim 1, wherein the playback means reproduces the restored structured data using the difference between it and coordinates other than the first coordinate.

4. The posture analysis program according to claim 1, wherein the compression means compresses n-dimensional structured data, the restoration means restores the generated compressed data, and the restoration means further functions as a learning means for training the compression means and the restoration means to reduce the difference between the restored n-dimensional structured data and the n-dimensional structured data before compression.

5. A compression method using a neural network that compresses structured data, which is a time series of single or multiple feature points extracted from video footage of multiple users, into compressed data of a smaller dimension than n, when the structured data is n-dimensional, A classification means that assigns labels to the compressed data based on the surveys of the aforementioned multiple users, and classifies the compressed data based on the results of the label assignment, A restoration means using a neural network to restore one coordinate in the classified compressed data to structured data, An information processing apparatus having a playback means for reproducing the restored structured data.

6. When structured data, which is a time series of one or more feature points extracted from video footage of multiple users, is n-dimensional, a compression step is performed using a neural network to compress the structured data into compressed data of less than n dimensions. A classification step in which labels are assigned to the compressed data based on the surveys of the aforementioned multiple users, and the compressed data is classified based on the results of the label assignment, A restoration step in which a coordinate in the classified compressed data is restored to structured data using a neural network, A posture analysis method comprising a playback step of reproducing the restored structured data.