Work analysis device, program, and work analysis method

The work analysis device uses pattern classification and feature similarity to efficiently identify work processes, overcoming the inefficiencies of conventional methods by automating the process without manual criterion setting.

WO2026048085A1PCT designated stage Publication Date: 2026-03-05MITSUBISHI ELECTRIC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Conventional work analysis methods require time-consuming predefinition of criteria for identifying work processes, making them inefficient.

Method used

A work analysis device that classifies sensor values into patterns, calculates image and text features using pre-trained encoders, and identifies processes by similarity between image and text features, allowing for high-accuracy identification of work processes without manual criterion setting.

Benefits of technology

Enables rapid and accurate identification of work processes, reducing the time and effort required for work analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000964_05032026_PF_FP_ABST
    Figure JP2025000964_05032026_PF_FP_ABST
Patent Text Reader

Abstract

A work analysis device (120) comprises: a classification unit (123) that classifies a plurality of sensor values detected for work performed a plurality of times into a plurality of classes according to patterns so as to classify work time into a plurality of sections corresponding to the plurality of classes; an image feature calculation unit (127) that calculates a plurality of image feature amounts from a plurality of pieces of image data of interest obtained by imaging the work; a text feature calculation unit (129) that calculates a plurality of text feature amounts from a plurality of pieces of text data indicating the content of a plurality of steps constituting work; a comparison unit (130) that calculates the degree of similarity between the image feature amount and the text feature amount; a step identification unit (131) that identifies, for each image feature amount, a step in which the content is indicated by text data corresponding to the text feature amount having the highest degree of similarity; and a link unit (132) that identifies each step of a plurality of sections by the identified step.
Need to check novelty before this filing date? Find Prior Art

Description

Work analysis device, program, and work analysis method

[0001] The present disclosure relates to an activity analysis device, a program, and an activity analysis method.

[0002] In manufacturing and other industries, improving work processes to increase productivity is important. To improve work processes, the first step is to analyze the current situation and perform a work analysis that measures the time required for each process that makes up the work. Work analysis is often carried out visually based on images captured by a camera, but this takes a huge amount of time. In response to this, a work analysis device has been proposed that shortens the time required for work analysis.

[0003] For example, Patent Document 1 describes a work data management system in which criteria are set in advance, such as the positions through which a worker's body parts pass for each process, and the process performed at each time is identified when the criteria are met.

[0004] Japanese Patent Application Laid-Open No. 2019-016226

[0005] However, in the conventional technology, it is necessary to predetermine the criteria for each process, which is time-consuming.

[0006] Therefore, one or more aspects of the present disclosure aim to make it possible to easily identify work processes with high accuracy.

[0007] A work analysis device according to one aspect of the present disclosure includes a classification unit that classifies a plurality of sensor values, which are multiple values ​​obtained by detecting physical quantities related to a task performed multiple times in a time series, into a plurality of classes according to patterns, and classifies the time duration of each of the multiple tasks into a plurality of intervals corresponding to the multiple classes; an image feature calculation unit that calculates a plurality of image features, which are multiple feature values ​​corresponding to each of the multiple target image data, by inputting a plurality of target image data pieces, which are included in a video of the task performed multiple times, in time series for each of the multiple intervals, into an image encoder; a text feature calculation unit that calculates a plurality of text features, which are multiple feature values ​​corresponding to each of the multiple steps, by inputting a plurality of text data pieces, which respectively indicate the contents of a plurality of steps that make up the task, into a text encoder; The system comprises a comparison unit that calculates the similarity between each image feature and each of the plurality of text features; a process identification unit that identifies one or more processes in chronological order by identifying a process whose content is indicated by text data corresponding to the text feature with the highest similarity for each of the plurality of image features; and a link unit that identifies each process in the plurality of sections by assigning one of the one or more processes to each of the plurality of sections, wherein the image encoder and the text encoder are pre-trained models so that, for each process included in the plurality of processes, the image feature calculated by inputting image data showing an image of the one process into the image encoder is highly similar to the text feature calculated by inputting text data showing the content of the one process into the text encoder.

[0008] A program according to one aspect of the present disclosure includes a computer including: a classification unit that divides a plurality of sensor values, which are a plurality of values ​​obtained by detecting physical quantities related to a task performed a plurality of times in a time series, into a plurality of classes according to patterns, and classifies the time of each of the tasks performed a plurality of times into a plurality of intervals corresponding to the plurality of classes; an image feature calculation unit that calculates a plurality of image features, which are a plurality of feature values ​​corresponding to each of the plurality of target image data, by inputting a plurality of pieces of target image data, which show a plurality of target images included in a video of the task performed a plurality of times in a time series for each of the plurality of intervals, into an image encoder; a text feature calculation unit that calculates a plurality of text features, which are a plurality of feature values ​​corresponding to each of the plurality of steps, by inputting a plurality of pieces of text data, which show the contents of a plurality of steps constituting the task, into a text encoder; The image encoder and the text encoder function as a comparison unit that calculates the similarity between each of the image features and each of the plurality of text features, a process identification unit that identifies one or more processes in chronological order by identifying a process whose content is indicated by text data corresponding to the text feature with the highest similarity for each of the plurality of image features, and a link unit that identifies each process of the plurality of sections by assigning any of the one or more processes to each of the plurality of sections, and the image encoder and the text encoder are pre-trained models so that, for each process included in the plurality of processes, the image feature calculated by inputting image data showing an image of the one process into the image encoder and the text feature calculated by inputting text data showing the content of the one process into the text encoder are highly similar.

[0009] A work analysis method according to one aspect of the present disclosure includes: dividing a plurality of sensor values, which are a plurality of values ​​obtained by detecting physical quantities related to a work performed multiple times in a time series, into a plurality of classes according to patterns; classifying the time of each of the work performed multiple times into a plurality of intervals corresponding to the plurality of classes; inputting a plurality of pieces of target image data, which show a plurality of target images included in a video of the work performed multiple times in a time series for each of the plurality of intervals, into an image encoder; calculating a plurality of image features, which are a plurality of feature values ​​corresponding respectively to the plurality of target image data; inputting a plurality of pieces of text data, which show the contents of a plurality of steps that make up the work, into a text encoder; calculating a plurality of text features, which are a plurality of feature values ​​corresponding respectively to the plurality of steps; a process for each of the plurality of sections by calculating a similarity between each of the plurality of image features and each of the plurality of text features, and identifying a process whose content is indicated by text data corresponding to the text feature with the highest similarity for each of the plurality of image features, thereby identifying one or more processes in chronological order, and assigning one of the one or more processes to each of the plurality of sections; wherein the image encoder and the text encoder are pre-trained models so that, for each process included in the plurality of processes, the image feature calculated by inputting image data showing an image of the one process into the image encoder is highly similar to the text feature calculated by inputting text data showing the content of the one process into the text encoder.

[0010] According to one or more aspects of the present disclosure, it is possible to easily identify steps in a job with high accuracy.

[0011] 1 is a block diagram showing an outline of the configuration of a work analysis system according to embodiments 1 and 2. FIG. 2 is a schematic diagram showing an example of detection by a sensor. FIG. 3 is a graph showing an example of a sensor value output by a sensor. FIG. 4 is a block diagram showing an outline of the configuration of a work analysis device according to embodiment 1. FIG. 5 is a block diagram showing an outline of the configuration of a classification unit. FIG. 6 is a schematic diagram for explaining an example of the data structure of a class data sequence. (A) to (D) are a first example of a time series graph showing a class data sequence. FIG. 7 is a schematic diagram for explaining the data structure of a class data sequence generated by a second classification unit. (A) to (D) are a second example of a time series graph showing a class data sequence. FIG. 8 is a block diagram showing an outline of the configuration of a PC. FIG. 9 is a flowchart showing the operation of the work analysis device according to embodiment 1. FIG. 10 is a block diagram showing an outline of the configuration of a work analysis device according to embodiment 2. FIG. 11 is a flowchart showing the operation of the work analysis device according to embodiment 2.

[0012] 1 is a block diagram showing a schematic configuration of an activity analysis system 100 according to embodiment 1. The activity analysis system 100 includes a sensor 110 as a detection device, a camera 111 as an imaging device, and an activity analysis device 120.

[0013] The sensor 110, camera 111, and work analysis device 120 are connected to the network 101, but the first embodiment is not limited to this example. For example, at least one of the sensor 110 and the camera 111 may be connected to the work analysis device 120 via a connection interface compatible with a USB (Universal Serial Bus) or the like.

[0014] The sensor 110 detects physical quantities that indicate the state of work performed by the worker. FIG. 2 is a schematic diagram showing an example of detection by the sensor 110. Here, it is assumed that the worker 140 repeatedly performs work consisting of a series of multiple steps multiple times. For this reason, the sensor 110 detects physical quantities related to the work performed multiple times in a time series, and the detected multiple values ​​are used as multiple sensor values. The sensor 110 outputs the time series sensor values ​​detected by measuring the state of the work performed by the worker 140 as sensor data. For example, the sensor 110 is a depth sensor and is installed so as to be able to capture images of the state of the left hand 141 and right hand 142 of the worker 140.

[0015] The sensor 110 includes, for example, a light source (not shown) that emits infrared rays in a specific pattern and an imaging element (not shown) that receives infrared rays reflected from an object. The sensor 110 generates a depth image having pixel values ​​representing the depth to the object, and detects the height positions of the worker's 140 left hand 141 and right hand 142 from the depth image. The sensor 110 then outputs the detected height positions as sensor values, for example, every 200 milliseconds. A specific example of a depth sensor is an existing depth sensor such as Kinect (registered trademark). The process for detecting the hand positions from the depth image can be performed using existing processes used in depth sensors.

[0016] Fig. 3 is a graph showing an example of a sensor value output by the sensor 110. In Fig. 3, the horizontal axis indicates the time when the sensor value was acquired, and the vertical axis indicates the sensor value, i.e., the height positions of the left hand 141 and right hand 142. The sensor value is the height positions of the left hand 141 and right hand 142 of the worker 140, and is therefore a two-dimensional value.

[0017] In this embodiment, a depth sensor is used as the sensor 110, but the present embodiment is not limited to this example. For example, a video camera, a three-dimensional acceleration sensor, a three-dimensional angular velocity sensor, or the like may be used as the sensor 110. In addition, in this embodiment, the positions of the worker's right and left hands are detected, but the present embodiment is not limited to this example. For example, the detection target may be the position of the worker's head, the position or angle of one or more joints of the worker, or the worker's vital signs (for example, heart rate or respiratory rate).

[0018] The camera 111 is installed where the work is performed, captures video of the work, and generates video data representing the captured video. The camera 111 then transmits the video data to the work analysis device 120 via the network 101.

[0019] 4 is a block diagram showing a schematic configuration of the work analysis device 120 according to embodiment 1. The work analysis device 120 includes a communication unit 121, a sensor data acquisition unit 122, a classification unit 123, a section of interest identification unit 124, a video acquisition unit 125, an area of ​​interest extraction unit 126, an image feature calculation unit 127, a text acquisition unit 128, a text feature calculation unit 129, a comparison unit 130, a process identification unit 131, a link unit 132, and an output unit 133.

[0020] The communication unit 121 performs communication via the network 101. In this embodiment, the communication unit 121 receives sensor data from the sensor 110 via the network 101. The received sensor data is provided to the sensor data acquisition unit 122. The communication unit 121 also receives video data from the camera 111 via the network 101. The received video data is provided to the video acquisition unit 125.

[0021] The sensor data acquiring unit 122 acquires sensor data. The acquired sensor data is provided to the classifying unit 123. Here, the sensor data acquiring unit 122 acquires the sensor data from the sensor 110 via the communication unit 121, but the present embodiment is not limited to this example. For example, if the sensor data is stored in a storage unit (not shown), the sensor data acquiring unit 122 may acquire the sensor data by reading the sensor data from the storage unit. Furthermore, if the sensor data is stored in a server (not shown), the sensor data acquiring unit 122 may acquire the sensor data from the server via the communication unit 121.

[0022] The classification unit 123 classifies the intervals in which a certain process is performed based on the sensor values ​​indicated by the sensor data. For example, the classification unit 123 classifies the time periods of each task that is performed multiple times into multiple intervals corresponding to the multiple classes by dividing the multiple sensor values ​​detected in time series into multiple classes according to the patterns.

[0023] Specifically, the classification unit 123 performs the following process: The classification unit 123 classifies a plurality of pieces of sensor data included in one operation consisting of a plurality of steps into a sensor data sequence x d For example, by turning on the sensor 110 when a worker or the like starts a task and turning off the sensor 110 when the task ends, the sensor 110 can continuously output sensor data for each task. Furthermore, a user of the work analysis device 120 may use an input interface such as a keyboard or mouse (not shown) to input instructions that allow identification that the worker is performing a task.

[0024] Here, d is an identification number for identifying each of the multiple cycle tasks, and is an integer between 1 and D. D is the number of times the task is repeated. In this embodiment, the sensor data string x d = {x d (1), x d (2), ..., x d (N(d))}, where x d(n) is the sensor value indicated by the nth sensor data output in the task identified by identification number d. Furthermore, N(d) is the number of sensor data output in the task section identified by identification number d. For example, if the length of the task section identified by identification number d is 10 seconds, as described above, in this embodiment, the sensor 110 outputs sensor data every 200 milliseconds, so the number of sensor data N(d) is 50.

[0025] The classification unit 123 then classifies the plurality of sensor data sequences x acquired by the sensor data acquisition unit 122 in the work performed multiple times. d and a plurality of class data sequences s that satisfy a predetermined evaluation criterion. d Generate.

[0026] For example, the classification unit 123 classifies each sensor data sequence x d Based on the sensor values ​​contained in each sensor data sequence x d The classification unit 123 then determines a plurality of intervals into which the plurality of sensor data sequences x are divided in time, and a class indicating the type of temporal change (e.g., pattern) of the sensor values ​​included in each of the plurality of intervals. d In each of the above, a class data string s indicating the interval and the class of the interval is d Generate.

[0027] Here, the class data string s d = {s d,1 , s d,2 , ..., s d,m , ..., s d,M(d) M(d) is the sensor data sequence x in the task identified by the identification number d. d m is an identification number for identifying each of the divided sections, and is an integer between 1 and M(d). d,m is the sensor data sequence x d The class data sequence s in the m-th section d is an element of s d,m = {a d,m , b d,m , c d,m}.

[0028] a d,m is the sensor data sequence x d is an identification number for identifying the first sensor data included in the m-th section obtained by dividing b. d,m is the sensor data sequence x d is the number of sensor data included in the m-th section into which c is divided, and corresponds to the length of that section. d,m is the sensor data sequence x d is a class number for identifying the class into which the m-th section obtained by dividing

[0029] 5 is a block diagram showing a schematic configuration of the classification unit 123. The classification unit 123 includes a first classification unit 123a, a standard pattern generation unit 123b, a second classification unit 123c, and a class data evaluation unit 123d.

[0030] The first classification unit 123a classifies a plurality of sensor data sequences x d For each of the class data strings s d For example, the first classification unit 123a calculates the initial value of the plurality of sensor data sequences x d are divided into a plurality of intervals according to a predetermined rule so that the total number of classes, in other words, the total number of intervals, is J. Note that J is the total number of classes, and is desirably determined by the following formula (1): Here, l is the minimum length of a predetermined interval. min N(d) is the length of the shortest operation. floor(x) is a floor function as a function for finding the largest integer equal to or less than x. For example, when the length of the shortest cycle operation, min N(d) = 10, and the minimum length of the interval, l = 3, then J = floor(10 / 3) = 3. Note that J may be a predetermined value.

[0031] Here, the first classification unit 123a classifies the sensor data sequence x dHowever, even if the number of sensor data included in all sections cannot be uniform, the first classification unit 123a classifies the sensor data string x so that the difference between the maximum number of sensor data included in one section and the minimum number of sensor data included in one section does not exceed "1". d The sensor data included in the above can be divided into multiple sections.

[0032] 6 is a schematic diagram for explaining an example of the data structure of a class data string. FIG. 6 shows the initial values ​​of a class data string generated based on sensor data strings acquired in a task repeated D=4 times. In FIG. 6, the total number of classes is J=6, and each sensor data string x d is divided into six intervals of equal length as far as possible.

[0033] 7A to 7D are time series graphs showing the class data strings shown in Fig. 6. In Fig. 7A to 7D, the horizontal axis is the number n for identifying the sensor data output in time series in the work identified by the identification number d, and the rectangles containing the numbers "1" to "6" respectively represent the sensor data string x. d The numbers "1" to "6" represent the class numbers into which each section is classified.

[0034] 5, the standard pattern generator 123b generates a standard pattern. The standard pattern indicates, for each class, a standard change over time in the sensor values ​​indicated by the sensor data included in each section. For example, the standard pattern generator 123b generates a standard pattern for each class of the sensor data sequence x d and a plurality of class data strings s d Based on the above, a plurality of standard patterns g corresponding to each of a plurality of classes j are j Here, j is a number for identifying multiple classes and is an integer from 1 to J.

[0035] In this embodiment, the standard pattern generating unit 123b generates a standard pattern g as a set of Gaussian distributions of sensor values ​​indicated by sensor data at each time by using Gaussian process regression. jAt this time, the standard pattern g j can be obtained as a parameter of the Gaussian distribution of the sensor values ​​indicated by the sensor data in the section classified into class j.

[0036] Here, the standard pattern g j = {g j (1), g j (2), ..., g j (L)}. g j (i) is the parameter of the Gaussian distribution of the i-th sensor value in the section classified into class j, and g j (i) = {μ j (i), σ j 2 (i)}, where μ j (i) is the mean of the Gaussian distribution, and σ j 2 (i) is the variance of the Gaussian distribution, and L is the length of the standard pattern, in other words, the total number of sensor data included in each section obtained by dividing the sensor data string, and is an integer equal to or greater than 1.

[0037] As mentioned above, μ j (i) is the mean of the Gaussian distribution of the sensor value indicated by the i-th sensor data in the section classified into class j. j (i) is a two-dimensional value, similar to the sensor value. j 2 (i) is the variance of the Gaussian distribution of the sensor values ​​indicated by the i-th sensor data in the section classified into class j. In this embodiment, it is assumed that the variance of the Gaussian distribution of the sensor values ​​is the same in all dimensions, so σ j 2 (i) is a one-dimensional value.

[0038] Standard pattern g j is a set of sensor values ​​X indicated by the sensor data of the section classified into class j in the class data sequence. j and a set I of numbers at which sensor values ​​indicated by sensor data in a section classified into class j in the class data sequence are output. j It can be estimated using the following: j= {X j (1), X j (2), ..., X j (N2 j ), I j = {I j (1), I j (2), ..., I j (N2 j )}. For example, X j (1) is the interval classified into class j. j (1) is the sensor value outputted first. j is the set X j and I j In other words, N2 j is the sum of the number of sensor data included in the section classified into class j among the sections into which the D sensor data strings are divided.

[0039] In this embodiment, the standard pattern g is calculated by the following equations (2) and (3): j (i) = {μ j (i), σ j 2 (i)} is estimated.

[0040] Here, β is a predetermined parameter, and E represents a unit matrix. j is a matrix calculated by the following equation (4), and v j,i is a vector calculated by the following equation (5). Furthermore, k is a kernel function, and a Gaussian kernel shown in the following equation (6) can be used. 0 , θ 1 , θ 2 and θ 3 is a predetermined parameter.

[0041] The second classification unit 123c classifies the plurality of standard patterns g generated by the standard pattern generation unit 123b. j Using the above, multiple sensor data sequences x d For each of the class data strings s dHere, the second classification unit 123c divides the sensor data string into a plurality of intervals by using forward filtering-backward sampling (FF-BS), and classifies the time-series sensor data included in each of the divided intervals into one of a plurality of classes.

[0042] FF-BS consists of two steps: a probability calculation step for the FF step and a division and classification step for the BS step. In the FF step, the second classification unit 123c classifies the sensor data sequence x d The sensor value x indicated by the n-th sensor data in d (n) is the parameter g of the ith Gaussian distribution of the standard pattern corresponding to class j j (i) Probability P(x d (n)|X j , I j ) is calculated as a Gaussian distribution Normal using the following equation (7).

[0043] The second classification unit 123c classifies the sensor data sequence x into which the first to n-i sections have already been divided. d When further dividing the n-th section, the probability that the class of the n-th section is j is calculated using the following formula (8): d Calculate [n][i][j].

[0044] In this case, P(j|j') is the class transition probability, which is calculated by the following equation (9).

[0045] Also, N3 j’,j is the number of times that the mth section obtained by dividing the sensor data string is classified into class j' and the m+1th section is classified into class j in all sensor data strings. Also, N4j' is the number of times that the section obtained by dividing the sensor data string is classified into class j. γ is a predetermined parameter. Equation (8) is a recurrence formula, and the probability α d [n][i][j] can be calculated.

[0046] In the BS step, the second classification unit 123c classifies the sensor data sequence x d For the divided intervals, the class data string is sampled using the following formula (10).

[0047] In equation (10), b in the first row d,m’ and c d,m’ is a random variable obtained from the probability distribution on the right side, and the second line is the variable a d,m’ According to the formula (10), the class data sequence s d,m’ = {a d,m’ , b d,m’ , c d,m’} can be generated.

[0048] Here, M2(d) is the sensor data sequence x d is the number of intervals into which s is divided. d,m’ is the sensor data sequence x d In equation (10), the sensor data sequence x d The class data sequence in the divided section is called the sensor data sequence x d In other words, the sensor data sequence x d The class data sequence s in the m-th section d,m = {a d,m , b d,m , c d,m}={ad, M2(d)-m+1, bd, M2(d)-m+1, cd, M2(d)-m+1}.

[0049] 8 is a schematic diagram for explaining the data structure of the class data sequence generated by the second classification unit 123c. Fig. 8 shows a class data sequence generated based on multiple standard patterns from sensor data sequences acquired from tasks performed D=4 times. The total number of classes is J=6.

[0050] 9A to 9D are time series graphs showing the class data strings shown in Fig. 8. In Fig. 9A to 9D, the horizontal axis, like Fig. 6, is the number n at which time series sensor data was output in the task identified by the identification number d, and the rectangles containing the numbers "1" to "6" respectively represent the sections [a] to [c] into which the sensor data string xd is divided. d,m , a d,m +b d,m The numbers "1" to "6" written in each section represent the class number c into which each section is classified. d,m In the example of FIG. 9, the range of each section is updated from the initial value of the class data sequence generated by the first classifying unit 123a (see FIGS. 6 and 7). For example, the range of each section is updated from the initial value of the class data sequence generated by the first classifying unit 123a (see FIGS. 6 and 7). 4 is divided into M2(4) = 7 intervals. The class of each interval is also updated from the initial value.

[0051] The class data evaluation unit 123d evaluates each class data sequence generated by the second classification unit 123c based on a predetermined evaluation criterion. A class data sequence that is determined by the class data evaluation unit 123d to satisfy the evaluation criterion becomes the class data sequence classified by the classification unit 123c.

[0052] For example, the class data evaluation unit 123d compares the class data sequence previously generated by the second classification unit 123c with the class data sequence currently generated by the second classification unit 123c, and calculates a similarity indicating the percentage of matching class values ​​at each time between these class data sequences. If the similarity exceeds a predetermined threshold (e.g., 90%), the class data evaluation unit 123d determines that the class data sequence generated by the second classification unit 123c satisfies the evaluation criteria. Alternatively, the class data evaluation unit 123d may determine that the class data sequence generated by the second classification unit 123c satisfies the evaluation criteria when the number of evaluations by the class data evaluation unit 123d exceeds a predetermined threshold.

[0053] If the class data evaluation unit 123d determines that the evaluation criteria are not satisfied, the standard pattern generation unit 123b generates a standard pattern again based on the class data sequence newly generated by the second classification unit 123c, and a new class data sequence is generated by the second classification unit 123c.

[0054] As described above, the classifying unit 123 repeats the generation of a standard pattern by the standard pattern generating unit 123b, the generation of a class data sequence by the second classifying unit 123c, and the evaluation by the class data evaluating unit 123d until the class data sequence generated by the second classifying unit 123c satisfies the predetermined evaluation criteria. When the class data sequence generated by the second classifying unit 123c satisfies the evaluation criteria, the classifying unit 123 outputs the class data sequence to the noteworthy section identifying unit 124 and the linking unit 132, and then terminates its operation.

[0055] Returning to FIG. 4, the attention section identification unit 124 uses the multiple classes classified by the classification unit 123 to identify one attention section from one or more sections corresponding to the same class.

[0056] Here, the attention section identification unit 124 identifies, for each class, one section from each of a plurality of sections classified into the same class according to predetermined criteria, as an attention section, based on the class data sequence classified by the classification unit 123. The identified attention sections are notified to the video acquisition unit 125.

[0057] For example, the interest section identification unit 124 may identify, among multiple sections classified into the same class, one section whose temporal length is closest to the average of the multiple sections as the interest section.Furthermore, the interest section identification unit 124 may identify, among multiple sections classified into the same class, one section whose waveform of a motion feature vector formed by sensor values ​​indicated by sensor data in the section is closest to the average of the waveforms of the motion feature vectors of the multiple sections as the interest section.

[0058] The video acquisition unit 125 acquires video data. Here, the video acquisition unit 125 acquires the video data from the camera 111 via the communication unit 121, but the embodiment is not limited to this example. For example, if the video data is stored in a storage unit (not shown), the video acquisition unit 125 may acquire the video data by reading the video data from the storage unit. Furthermore, if the video data is stored in a server (not shown), the video acquisition unit 125 may acquire the video data from the server via the communication unit 121.

[0059] Furthermore, the video acquisition unit 125 identifies two or more frames corresponding to each of the multiple interest sections from the multiple frames constituting the video. Specifically, upon receiving notification of the identified interest sections from the interest section identification unit 124, the video acquisition unit 125 extracts, as interest section data, data indicating the video of the portion corresponding to the interest section from the video data acquired via the communication unit 121. The video acquisition unit 125 then provides the interest section data to the interest area extraction unit 126.

[0060] The attention area extraction unit 126 extracts one or more attention areas, which are one or more predetermined parts, from each of two or more frames that make up the attention section data, and composes multiple target image data using one or more attention area data, which are data that indicate the one or more attention areas.

[0061] Here, the attention area extraction unit 126 extracts one or more portions as one or more attention areas from each of two or more frames constituting the attention section data. For example, the attention area extraction unit 126 extracts at least one of the peripheral area of ​​a worker who is performing work and whose image is captured by the camera 111, the peripheral area of ​​the worker's hands, and a changed area from the previous frame as one or more attention areas.

[0062] Specifically, when a worker is working using his or her entire body, or when the worker's posture affects the content of the work, it is desirable to set the area around the worker as the region of interest. Furthermore, when a worker is working with his or her hands while maintaining a constant posture, such as while sitting, or when a worker is working with different tools in hand depending on the process, it is desirable to set the area around the hands as the region of interest. Furthermore, when work is performed using an inspection machine and the inspection results are displayed on the inspection machine monitor, it is desirable to set the changed area as the region of interest. Here, an example is described in which the area around the worker, the area around the hands, and the changed area are all extracted as multiple regions of interest, but it is sufficient if at least one of these is extracted. One or more pieces of region of interest data indicating the extracted one or more regions of interest (here, multiple regions of interest) are provided to the image feature calculation unit 127.

[0063] The image feature calculation unit 127 inputs a plurality of pieces of target image data, which are indicated by the attention section data and which represent a plurality of target images included in the video of the work in time series, into the image encoder, and calculates a plurality of image feature amounts, which are a plurality of feature amounts corresponding to the plurality of target images. In other words, the image feature calculation unit 127 uses a plurality of pieces of data representing a plurality of images in time series in each of the plurality of attention sections as the plurality of pieces of target image data to be input to the image encoder. Here, the target images are the attention area, but the frames constituting the video may also be the target images. In this case, the attention area extraction unit 126 is not necessary.

[0064] For example, the image feature calculation unit 127 calculates one or more image feature amounts, which are one or more feature amounts, by inputting each of one or more pieces of attention region data from the attention region extraction unit 126 to a known image encoder. The calculated one or more image feature amounts are provided to the comparison unit 130. Here, since multiple attention regions are provided from the attention region extraction unit 126, the image feature calculation unit 127 calculates the multiple image feature amounts by inputting data indicating each of the multiple attention regions to a known image encoder.

[0065] The text acquisition unit 128 acquires text data indicating text describing each of a plurality of steps included in the work imaged by the camera 111. The acquired text data is provided to the text feature calculation unit 129. For example, if the text data is stored in a storage unit (not shown), the text acquisition unit 128 may acquire the text data by reading the text data from the storage unit. Alternatively, if the text data is stored in a server (not shown), the text acquisition unit 128 may acquire the text data from the server via the communication unit 121.

[0066] The text feature calculation unit 129 inputs a plurality of pieces of text data indicating the contents of a plurality of steps that make up the work into a text encoder, and calculates a plurality of text features that are a plurality of features that respectively correspond to the plurality of steps. Here, the text feature calculation unit 129 inputs the text data provided by the text acquisition unit 128 into a known text encoder, and calculates the text features that are the features for each step.

[0067] Here, we will explain the image encoder used by the image feature calculation unit 127 and the text encoder used by the text feature calculation unit 129. The image encoder and text encoder used here are assumed to have been trained in advance to convert image data of a process included in a job and text data indicating the content of that process into similar feature quantities in the same space. The content of the process may be any content related to that process.

[0068] In other words, the image encoder and the text encoder are pre-trained models that, for each process included in the plurality of processes, increase the similarity between the image feature calculated by inputting image data representing an image of that process into the image encoder and the text feature calculated by inputting text data representing the content of that process into the text encoder. Such image encoders and text encoders are described, for example, in the following literature: Literature: Alec Radford et al., "Learning Transferable Visual Models From Natural Language Supervision," arXiv:2103.00020 [cs. CV], 26 February 2021

[0069] The comparison unit 130 calculates the similarity between each of the multiple image features and each of the multiple text features. For example, the comparison unit 130 calculates a similarity vector indicating the similarity between one or more image features and the text features for each process, for each frame, from one or more image features for each frame and the text features for each process. The similarity calculated here may be, for example, cosine similarity, but is not limited to this. The calculated similarity vector is provided to the process identification unit 131. Here, one similarity vector indicates the similarity between one image feature in a certain frame and a text feature for each process.

[0070] Here, the comparison unit 130 acquires a plurality of image feature amounts from the image feature calculation unit 127, and therefore calculates a similarity vector for each frame for each of the plurality of image feature amounts. Specifically, the comparison unit 130 calculates a similarity vector for each frame in which the peripheral portion of the worker is the region of interest, a similarity vector for each frame in which the peripheral portion of the hand is the region of interest, and a similarity vector for each frame in which the changed portion is the region of interest.

[0071] The process identification unit 131 identifies one or more processes in chronological order by identifying a process whose content is indicated by text corresponding to the text feature with the highest calculated similarity for each of a plurality of image features. Here, the process identification unit 131 identifies the process indicated by a frame from the similarity vector for that frame. For example, if only one region of interest is extracted, the process identification unit 131 may determine that the process described in the text data with the highest similarity indicated by the similarity vector for that frame is the process indicated by that frame.

[0072] On the other hand, when multiple regions of interest are extracted, the process identification unit 131 determines the process indicated in that frame from the multiple similarity vectors for each frame. For example, the process identification unit 131 may determine the process described in the text data with the highest similarity among the similarities indicated by the multiple similarity vectors for each frame as the process indicated in that frame. Alternatively, the process identification unit 131 may summarize the similarities indicated by the multiple similarity vectors for each frame for each text data, and, for example, determine the process described in the text data with the highest added or multiplied value as the process indicated in that frame.

[0073] Here, since the area of ​​interest is extracted from the video in the section of interest indicated by the section of interest data, the process identification unit 131 notifies the link unit 132 of the process identified for each frame in each section of interest.

[0074] The linking unit 132 assigns any one of the one or more processes identified by the process identifying unit 131 to each of the multiple sections classified by the classifying unit 123, thereby identifying each process in the multiple sections. The linking unit 132 then assigns the same process identified in one section of interest to other sections of the same class as the section of interest.

[0075] Specifically, the link unit 132 identifies the process in each section of interest based on the process identified for each frame in the section of interest. Here, the link unit 132 may determine the most frequently identified process among the processes identified in the frames included in the section of interest as the process of the section of interest.

[0076] The link unit 132 then identifies the processes in all sections in which video is captured in the video data by assigning the processes identified in the section of interest as processes in sections classified into the same class as the section of interest based on the class data sequence classified by the classification unit 123. The processes identified in this way for each section are notified to the output unit 133.

[0077] The output unit 133 outputs data indicating the process identified for each section. For example, the output unit 133 may display the data on a display (not shown) in a predetermined display format. The output unit 133 may also send such data to a predetermined destination via the communication unit 121.

[0078] The work analysis device 120 described above can be realized by, for example, a computer such as the PC 10 shown in Fig. 10. The PC 10 includes storage 11 such as a hard disk drive (HDD) and a solid state drive (SSD), memory 12, a processor 13 such as a central processing unit (CPU), a communication interface (I / F) 14 such as a network interface card (NIC), an input interface 15 such as a keyboard and a mouse, and a display 16.

[0079] For example, the communication unit 121 can be realized by the communication I / F 14. The sensor data acquisition unit 122, the classification unit 123, the attention section identification unit 124, the video acquisition unit 125, the attention area extraction unit 126, the image feature calculation unit 127, the text acquisition unit 128, the text feature calculation unit 129, the comparison unit 130, the process identification unit 131, the link unit 132, and the output unit 133 can be realized by loading a program stored in the storage 11 into the memory 12 and having the processor 13 execute the program.

[0080] The program may be downloaded to the storage 11 from a recording medium (not shown) via a reader / writer (not shown) or from the network 101 via the communication I / F 14, and then loaded onto the memory 12 and executed by the processor 13. Alternatively, the program may be directly loaded onto the memory 12 from a recording medium via the reader / writer or from the network 101 via the communication I / F 14, and then executed by the processor 13. In other words, the program may be provided by a computer program product such as a recording medium.

[0081] 11 is a flowchart showing the operation of the work analysis apparatus 120 in embodiment 1. First, the sensor data acquisition unit 122 acquires sensor data via the communication unit 121 (S10). The acquired sensor data is provided to the classification unit 123.

[0082] The classification unit 123 classifies a section in which a certain process is performed into a class based on the sensor value indicated by the sensor data (S11).

[0083] The attention section identification unit 124 identifies, for each class classified by the classification unit 123, one section from one or more sections classified into the same class according to a predetermined criterion as an attention section (S12). The identified attention sections are notified to the video acquisition unit 125.

[0084] The video acquisition unit 125 acquires video data (S13). The video acquisition unit 125 also selects one unselected section of interest from the sections of interest identified in step S12 (S14). Then, data representing the video in the section of interest from the video data acquired in step S13 is provided to the attention area extraction unit 126 as section of interest data.

[0085] The attention region extraction unit 126 selects one unprocessed frame in order from the multiple frames constituting the attention section data (S15). Then, the attention region extraction unit 126 extracts one or more attention regions from the selected frame (S16). Here, the description will be given assuming that multiple attention regions are extracted. Multiple attention region data indicating the extracted multiple attention regions is provided to the image feature calculation unit 127.

[0086] The image feature calculation unit 127 inputs each of the plurality of pieces of attention region data from the attention region extraction unit 126 into a known image encoder, thereby calculating a plurality of image feature amounts corresponding to each of the plurality of attention regions (S17). The calculated plurality of image feature amounts are provided to the comparison unit 130.

[0087] The text acquisition unit 128 acquires text data indicating text explaining each of the multiple steps included in the work imaged by the camera 111, and the text feature calculation unit 129 inputs the text data into a known text encoder to calculate text features, which are features for each step (S18). The calculated text features are provided to the comparison unit 130.

[0088] The comparison unit 130 calculates the similarity between each of the plurality of image features and the text feature for each process, thereby calculating a plurality of similarity vectors corresponding to the plurality of image features (S19). The plurality of similarity vectors are provided to the process identification unit 131.

[0089] The process identification unit 131 identifies the process indicated in the frame from the plurality of similarity vectors (S20). Here, the process identification unit 131 uses the similarity for each process indicated by each of the plurality of similarity vectors to identify the process that can be determined to have the highest similarity among the processes depicted in the frame, based on a predetermined determination criterion.

[0090] Then, the attention area extraction unit 126 determines whether or not all frames of the attention section data have been selected in step S15 (S21). If all frames have been selected (Yes in S21), the process proceeds to step S22. If there are frames that have not yet been selected (No in S21), the process returns to step S15.

[0091] In step S22, the video acquisition unit 125 determines whether or not all of the interest sections identified in step S12 have been selected. If all of the interest sections have been selected (Yes in S22), the process proceeds to step S23. If there are any interest sections that have not yet been selected (No in S22), the process returns to step S14.

[0092] In step S23, the link unit 132 identifies a process in each section of interest based on the process identified for each frame in that section of interest, and assigns the process identified in that section of interest as a process in a section classified into the same class as the section of interest (S23). In this way, the link unit 132 identifies processes in all sections in which video is captured in the video data.

[0093] Next, the output unit 133 outputs data indicating the process determined for each section (S24).

[0094] As described above, according to the first embodiment, the steps in a job can be easily identified with high accuracy.

[0095] 1, an activity analysis system 200 according to the second embodiment includes a sensor 110, a camera 111, and an activity analysis device 220. The sensor 110 and the camera 111 of the activity analysis system 200 according to the second embodiment are similar to the sensor 110 and the camera 111 of the activity analysis system 100 according to the first embodiment. However, in the second embodiment, the sensor 110 and the camera 111 transmit sensor data and video data to the activity analysis device 220 via the network 101.

[0096] 12 is a block diagram showing a schematic configuration of a work analysis device 220 according to embodiment 2. The work analysis device 220 includes a communication unit 121, a sensor data acquisition unit 122, a classification unit 123, a section of interest identification unit 124, a video acquisition unit 125, an area of ​​interest extraction unit 126, an image feature calculation unit 127, a text acquisition unit 128, a text feature calculation unit 129, a comparison unit 130, a process identification unit 131, a link unit 132, an output unit 133, and a process correction unit 234.

[0097] The communication unit 121, sensor data acquisition unit 122, classification unit 123, attention section identification unit 124, video acquisition unit 125, attention area extraction unit 126, image feature calculation unit 127, text acquisition unit 128, text feature calculation unit 129, comparison unit 130, process identification unit 131, link unit 132, and output unit 133 of the work analysis apparatus 220 in embodiment 2 are similar to the communication unit 121, sensor data acquisition unit 122, classification unit 123, attention section identification unit 124, video acquisition unit 125, attention area extraction unit 126, image feature calculation unit 127, text acquisition unit 128, text feature calculation unit 129, comparison unit 130, process identification unit 131, link unit 132, and output unit 133 of the work analysis apparatus 120 in embodiment 1. However, in the second embodiment, the process specifying unit 131 notifies the process correcting unit 234 of the process specified for each frame, and the linking unit 132 receives notification of the process corrected by the process correcting unit 234 .

[0098] The process correction unit 234 corrects errors contained in the processes for each frame identified by the process identification unit 131. For example, the process correction unit 234 corrects the predetermined number of processes by using the most frequently occurring process among a predetermined number of processes consecutively identified in time series. Specifically, the process correction unit 234 may determine the most frequently identified process for each predetermined number of frames as the process for the predetermined number of frames.

[0099] Furthermore, the process correction unit 234 may statistically correct errors in one or more identified processes. Specifically, the process correction unit 234 may statistically correct errors in the identified processes using a known algorithm such as a hidden Markov model. The process corrected by the process correction unit 234 is notified to the link unit 132.

[0100] The work analysis device 220 described above can also be realized by a computer such as the PC 10 shown in Fig. 10. In the second embodiment, the process correction unit 234 can also be realized by loading a program stored in the storage 11 into the memory 12 and having the processor 13 execute the program.

[0101] Fig. 13 is a flowchart showing the operation of the work analysis device 220 in embodiment 2. Note that, of the processes included in the flowchart shown in Fig. 13, the processes that are the same as the processes included in the flowchart shown in Fig. 11 are assigned the same reference numerals as those used in the flowchart shown in Fig. 11.

[0102] The processing of steps S10 to S22 in Fig. 13 is the same as the processing of steps S10 to S22 in Fig. 11. However, in Fig. 13, if it is determined in step S22 that all of the intervals of interest have been selected (Yes in S22), the processing proceeds to step S30.

[0103] In step S30, the process correction unit 234 corrects errors contained in the process for each frame identified by the process identification unit 131. The corrected process for each frame is provided to the link unit 132. Then, the process proceeds to step S23.

[0104] The processes in steps S23 and S24 in Fig. 13 are similar to the processes in steps S23 and S24 in the flowchart shown in Fig. 11. However, in Fig. 13, the linking unit 132 identifies each process in a plurality of sections by assigning any of the one or more corrected processes to each of the plurality of sections.

[0105] As described above, according to the second embodiment, the process is corrected for each frame, so that the process can be specified with higher accuracy.

[0106] In the above-described first and second embodiments, a process is identified for each frame, but the first and second embodiments are not limited to such examples. For example, one sample frame serving as one sample may be extracted for each of a predetermined number of frames included in a video, and a process may be identified for each sample frame by the above-described operation.

[0107] In the above-described first and second embodiments, the attention section is identified by the attention section identification unit 124. However, at least one of the first and second embodiments is not limited to such an example. For example, the attention section identification unit 124 may be omitted in at least one of the first and second embodiments. In this case, the video acquisition unit 125 provides the acquired video data to the attention area extraction unit 126. The attention area extraction unit 126 extracts one or more attention areas, which are one or more predetermined portions, from each of multiple frames constituting the video, and generates multiple sets of target image data using data indicating the one or more attention areas. Note that, if both the attention section identification unit 124 and the attention area extraction unit 126 are omitted, the video acquisition unit 125 provides the video data to the image feature calculation unit 127. The image feature calculation unit 127 inputs multiple sets of target image data representing multiple target images included in video captured of multiple tasks, in chronological order, for each of multiple classified sections, to an image encoder, thereby calculating multiple image feature amounts corresponding to the multiple sets of target image data.

[0108] 100, 200 Work analysis system, 110 Sensor, 111 Camera, 120, 220 Work analysis device, 121 Communication unit, 122 Sensor data acquisition unit, 123 Classification unit, 123a First classification unit, 123b Standard pattern generation unit, 123c Second classification unit, 123d Class data evaluation unit, 124 Noteworthy section identification unit, 125 Video acquisition unit, 126 Noteworthy area extraction unit, 127 Image feature calculation unit, 128 Text acquisition unit, 129 Text feature calculation unit, 130 Comparison unit, 131 Process identification unit, 132 Link unit, 133 Output unit, 234 Process correction unit.

Claims

1. A classification unit that classifies the time of each of the multiple operations into a plurality of intervals corresponding to the multiple classes by dividing a plurality of sensor values, which are a plurality of values ​​obtained by detecting physical quantities related to an operation performed multiple times in a time series, into a plurality of classes according to patterns; an image feature calculation unit that calculates a plurality of image features, which are a plurality of feature values ​​corresponding to each of the multiple operation image data, by inputting a plurality of object image data, which shows a plurality of object images included in a video of the operation performed multiple times in a time series for each of the multiple intervals, into an image encoder; a text feature calculation unit that calculates a plurality of text features, which are a plurality of feature values ​​corresponding to each of the multiple operation steps, by inputting a plurality of text data, which respectively show the contents of a plurality of operations constituting the operation, into a text encoder; a comparison unit that calculates the similarity between each of the multiple image features and each of the multiple text features; a process identification unit that identifies one or more operations in time series by identifying a process whose contents are shown by text data corresponding to the text feature with the highest similarity for each of the multiple image features; and a link unit that identifies each of the multiple operations in the multiple intervals by assigning one of the one or more operations to each of the multiple intervals, The work analysis device is characterized in that the image encoder and the text encoder are pre-trained models so that, for each process included in the plurality of processes, image features calculated by inputting image data showing an image of the one process into the image encoder and text features calculated by inputting text data showing the content of the one process into the text encoder have a high degree of similarity.

2. The work analysis device described in claim 1 further comprises a section of interest identification unit that identifies multiple sections of interest by identifying one section of interest from one or more sections corresponding to the same class for the multiple classes, wherein the image feature calculation unit uses multiple data representing the multiple target images in chronological order in each of the multiple sections of interest as the multiple target image data to be input to the image encoder, and the link unit identifies one process in each of the multiple sections of interest by assigning any of the one or more processes to each of the multiple sections of interest, and also assigns the one process identified in the one section of interest to other sections in the same class as the one section of interest.

3. The work analysis device described in claim 1, further comprising an attention area extraction unit that extracts one or more attention areas, which are one or more predetermined parts, from each of the multiple frames that make up the video, and constructs the multiple target image data using data indicating the one or more attention areas.

4. The work analysis device described in claim 1 further comprises: a section of interest identification unit that identifies multiple sections of interest by identifying one section of interest from one or more sections corresponding to the same class in the multiple classes; an image acquisition unit that identifies two or more frames corresponding to each of the multiple sections of interest from the multiple frames that make up the video; and a section of interest extraction unit that extracts one or more areas of interest that are one or more predetermined parts from each of the two or more frames and constructs the multiple target image data using data indicating the one or more areas of interest, wherein the link unit identifies one process in each of the multiple sections of interest by assigning any of the one or more processes to each of the multiple sections of interest, and assigns the one process identified in the one section of interest to other sections of the same class as the one section of interest.

5. A work analysis device according to claim 3 or 4, characterized in that the attention area extraction unit selects the area surrounding the worker performing the work as one of the one or more attention areas.

6. A work analysis device as described in claim 3 or 4, characterized in that the attention area extraction unit selects the area around the hand of the worker performing the work as one of the one or more attention areas.

7. A work analysis device as described in claim 3 or 4, characterized in that the attention area extraction unit selects a changed portion in the video that has changed from the previous frame as one of the one or more attention areas.

8. A work analysis device as described in any one of claims 1 to 7, further comprising a process correction unit that corrects errors in the one or more processes, and the link unit assigns the corrected one or more processes.

9. The work analysis device according to claim 8, characterized in that the process correction unit corrects the predetermined number of processes using the most frequently occurring process among a predetermined number of processes identified consecutively in time series.

10. The work analysis device according to claim 8, wherein the process correction unit statistically corrects errors in the one or more processes.

11. A computer is caused to function as: a classification unit that divides a plurality of sensor values, which are a plurality of values ​​obtained by detecting physical quantities related to an operation performed a plurality of times in a time series, into a plurality of classes according to patterns, and classifies the time of each of the operations performed a plurality of times into a plurality of intervals corresponding to the plurality of classes; an image feature calculation unit that calculates a plurality of image features, which are a plurality of feature amounts corresponding to each of the plurality of object image data, by inputting a plurality of object image data, which shows a plurality of object images included in a video of the operation performed a plurality of times in a time series in each of the plurality of intervals, into an image encoder; a text feature calculation unit that calculates a plurality of text features, which are a plurality of feature amounts corresponding to each of the plurality of steps, by inputting a plurality of text data, which respectively show the contents of a plurality of steps constituting the operation, into a text encoder; a comparison unit that calculates the similarity between each of the plurality of image features and each of the plurality of text features; a process identification unit that identifies one or more processes in time series by identifying a process whose content is shown by text data corresponding to the text feature with the highest similarity in each of the plurality of image features; and a link unit that identifies each of the plurality of steps by assigning any of the one or more processes to each of the plurality of intervals, The image encoder and the text encoder are pre-trained models that, for each process included in the plurality of processes, increase the similarity between image features calculated by inputting image data representing an image of the process into the image encoder and text features calculated by inputting text data representing the content of the process into the text encoder.

12. A work analysis method comprising: dividing a plurality of sensor values, which are a plurality of values ​​obtained by detecting physical quantities related to a task performed a plurality of times in a time series, into a plurality of classes according to patterns, thereby classifying the time of each of the multiple tasks performed a plurality of times into a plurality of intervals corresponding to the multiple classes; inputting a plurality of pieces of target image data, which show a plurality of target images included in video footage of the task performed a plurality of times in a time series for each of the multiple intervals, into an image encoder, thereby calculating a plurality of image features, which are a plurality of features corresponding respectively to the plurality of target image data; inputting a plurality of pieces of text data, which show the contents of a plurality of steps constituting the task, into a text encoder, thereby calculating a plurality of text features, which are a plurality of features corresponding respectively to the multiple steps; calculating the similarity between each of the multiple image features and each of the multiple text features; identifying one or more steps in time series by identifying the step whose content is shown by the text data corresponding to the text feature with the highest similarity for each of the multiple image features; and assigning any of the one or more steps to each of the multiple intervals, thereby identifying each of the multiple intervals, a work analysis method characterized in that the image encoder and the text encoder are pre-trained models so that, for each process included in the plurality of processes, image features calculated by inputting image data showing an image of the one process into the image encoder are highly similar to text features calculated by inputting text data showing the content of the one process into the text encoder.

Citation Information

Patent Citations

  • Image understanding method and device, equipment and medium

    CN114511043A

  • Motion detection model learning device and program thereof, and motion section detection device and program thereof

    JP2023007542A

  • Operation analysis device

    WO2019229943A1

  • Information processing method, information processing device, and program

    WO2022234692A1