Advice provision device, advice provision method, program, and evaluation device

The advice providing device addresses the limitation of pre-prepared advice by using video data evaluation and a language model to generate real-time advice, improving flexibility and reducing operational effort.

WO2026070024A1PCT designated stage Publication Date: 2026-04-02NEC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing systems require pre-prepared advice for every potential deviation operation, limiting their ability to provide advice for unexpected actions and increasing operational effort.

Method used

An advice providing device that acquires first and second video data, evaluates the appropriateness of operations using a language model, and generates advice on-the-fly by inputting request data into a language model, eliminating the need for pre-established advice correlations.

Benefits of technology

Enables real-time advice generation for unexpected actions, reducing operational effort by avoiding the need for pre-prepared advice and enhancing flexibility in providing guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025027334_02042026_PF_FP_ABST
    Figure JP2025027334_02042026_PF_FP_ABST
Patent Text Reader

Abstract

An advice provision device according to the present disclosure comprises: an acquisition means for acquiring first video data including a scene in which target work is performed by a target worker, and second video data including a scene in which the target work is carried out correctly; an evaluation means for using the first video data and the second video data to perform a motion evaluation regarding the suitability of a motion by the target worker; a generation means for generating request data which requests advice for the target worker regarding the target work on the basis of the motion evaluation results; and an output means for inputting the request data into a language model and using answer data obtained from the language model to output advice data representing the advice.
Need to check novelty before this filing date? Find Prior Art

Description

Advice Providing Device, Advice Providing Method, Program, and Evaluation Device

[0001] The present disclosure relates to an advice providing device, an advice providing method, a program, and an evaluation device.

[0002] A system has been developed that provides advice on work performed by an operator. For example, the system of Patent Document 1 analyzes the operator's actions by analyzing the operation in which the work by the operator is imaged, and detects a deviation operation that deviates from the reference operation. Then, the system extracts advice corresponding to the detected deviation operation from an advice table in which the deviation operation and the advice are previously associated, and provides it to the user.

[0003] Japanese Unexamined Patent Application Publication No. 2020-106954

[0004] In the system of Patent Document 1, it is necessary to prepare in advance advice for improving the deviation operation for each assumed deviation operation. The present disclosure has been made in view of this problem, and one of its purposes is to provide a new technology for providing advice on work.

[0005] The advice providing device according to the present disclosure includes an acquisition unit that acquires first video data including a scene in which a target operation is performed by a target operator and second video data including a scene in which the target operation is correctly performed, an evaluation unit that performs an operation evaluation regarding the appropriateness of the operation of the target operator using the first video data and the second video data, a generation unit that generates request data for obtaining advice for the target operator regarding the target operation based on the result of the operation evaluation, and an output unit that inputs the request data into a language model and outputs advice data representing the advice using the response data obtained from the language model.

[0006] The advice provision method relating to this disclosure is performed by a computer. The advice provision method includes: an acquisition step of acquiring first video data including a scene in which a target work is performed by a target worker and second video data including a scene in which the target work is performed correctly; an evaluation step of performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; a generation step of generating request data that seeks advice for the target worker regarding the target work based on the results of the action evaluation; and an output step of inputting the request data into a language model and outputting advice data representing the advice using response data obtained from the language model.

[0007] The program relating to this disclosure causes a computer to perform the following steps: an acquisition step of acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; an evaluation step of performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; a generation step of generating request data that seeks advice for the target worker regarding the target task based on the results of the action evaluation; and an output step of inputting the request data into a language model and outputting advice data representing the advice using response data obtained from the language model.

[0008] The evaluation device according to this disclosure includes: acquisition means for acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; evaluation means for calculating an action evaluation value representing the appropriateness of the target worker's actions using the first video data and the second video data; and output means for outputting the action evaluation value. The evaluation means calculates an action feature quantity representing the characteristics of the captured action for each first frame included in the first video frame and each second frame included in the second video frame; associates the first frame and the second frame based on the degree of similarity between the action feature quantity of the first frame and the action feature quantity of the second frame; calculates the action evaluation value using the result of the association; and the action feature quantity includes a gripping feature quantity representing whether or not an object is being gripped.

[0009] This disclosure provides a new technology for evaluating work using video.

[0010] This is a diagram illustrating the overview of the operation of the advice-providing device. This is a block diagram illustrating the functional configuration of the advice-providing device. This is a block diagram illustrating the hardware configuration of the computer that implements the advice-providing device. This is a flowchart illustrating the flow of processing performed by the advice-providing device. This is a diagram illustrating a frame pair. This is a graph illustrating the correspondence between the first frame and the second frame. This is a diagram illustrating the screen output by the advice-providing device. This is a diagram illustrating the functional configuration of the evaluation device. This is a flowchart illustrating the flow of processing performed by the evaluation device.

[0011] Embodiments of the present disclosure will be described in detail below with reference to the drawings. In each drawing, the same or corresponding elements are denoted by the same reference numerals, and redundant explanations are omitted as necessary for clarity. Unless otherwise specified, predetermined values ​​such as specified values ​​and thresholds are stored in advance in a storage device accessible from the device that uses those values. Furthermore, unless otherwise specified, the storage unit is composed of one or any number of storage devices.

[0012] <Overview> Figure 1 is a diagram illustrating the general operation of the advice-providing device 2000. Here, Figure 1 is a diagram intended to facilitate understanding of the general operation of the advice-providing device 2000, and the operation of the advice-providing device 2000 is not limited to what is shown in Figure 1.

[0013] The advice-providing device 2000 is used to provide advice to the worker. A worker refers to a person performing a task. A task, in this context, refers to a series of actions performed on an object. Examples of actions on an object include screwing, unscrewing, fitting, gluing, peeling, checking for scratches, or checking dimensions.

[0014] To provide advice to the worker, the advice-providing device 2000 acquires first video data 10 and second video data 20. First video data 10 is video data that includes scenes in which the work to be evaluated (target work) is being performed appropriately. First video data 10 is generated, for example, by capturing images with a camera of the target work being performed by a person who can perform the target work correctly (for example, a skilled worker). First video data 10 can also be called "example video data". Here, each frame (image) that makes up first video data 10 is also referred to as first frame 12.

[0015] The second video data 20 is video data that includes scenes in which the target work is being performed by the worker being evaluated (target worker). The second video data 20 is generated by capturing images with a camera of the target worker performing the target work. Here, each frame (image) that makes up the second video data 20 is also referred to as the second frame 22.

[0016] The advice-providing device 2000 evaluates the appropriateness of the actions performed by the target worker on the target task, and based on the results of the evaluation, generates advice data 90 representing advice for the target worker. This evaluation is called an action evaluation. The action evaluation is performed by comparing the first video data 10 and the second video data 20.

[0017] The language model 70 is used to generate the advice data 90. The language model 70 is configured to output answer data representing the answer to a question in response to the input data representing the question. For example, the language model 70 is a language model that can handle multimodal information, such as a language model classified as a Large Language Model (LLM) or a Large Vision and Language Model (VLM) that handles both video and language input.

[0018] The advice-providing device 2000 generates request data 60 based on the results of the operational evaluation. The request data 60 is data representing a question to the language model 70. The request data 60 includes text data representing the question and image data related to the question.

[0019] The advice-providing device 2000 inputs the request data 60 into the language model 70 and obtains response data 80 from the language model 70. Based on the response data 80, the advice-providing device 2000 outputs advice data 90. The advice data 90 may be the response data 80 itself, or it may be data generated based on the response data 80.

[0020] <Example of effect> According to the advice-providing device 2000, the actions of the target worker are evaluated by comparing the first video data 10, which contains an example of the target work, with the second video data 20, which contains the target work performed by the target worker. Furthermore, a request for advice is made to the language model 70 using request data 60 generated based on the results of the action evaluation. Then, advice data 90 based on the response data 80 obtained from the language model 70 is output.

[0021] This method eliminates the need to pre-establish a correspondence between actions that need improvement and advice for those actions. Therefore, since there is no need to prepare advice in advance, the effort required to operate the advice-providing device 2000 is reduced.

[0022] Furthermore, a method that pre-prepares advice on actions that need improvement cannot provide advice for unexpected actions. In contrast, the advice-providing device 2000 generates advice using the language model 70, so it can generate advice even if the target worker performs an unexpected action.

[0023] The advice-providing device 2000 of this embodiment will be described in more detail below.

[0024] <Example of Functional Configuration> Figure 2 is a block diagram illustrating the functional configuration of the advice provision device 2000. The advice provision device 2000 includes an acquisition unit 2020, an evaluation unit 2040, a generation unit 2060, and an output unit 2080. The acquisition unit 2020 acquires first video data 10 and second video data 20. The evaluation unit 2040 performs an action evaluation on the target worker using the first video data 10 and second video data 20. The generation unit 2060 generates request data 60 based on the results of the action evaluation. The output unit 2080 obtains response data 80 by inputting the request data 60 into the language model 70. The output unit 2080 then outputs advice data 90 based on the response data 80.

[0025] <Example of Hardware Configuration> Each functional component of the advice-providing device 2000 may be implemented by hardware that realizes each functional component (e.g., hardwired electronic circuits), or by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it). The following will further explain the case where each functional component of the advice-providing device 2000 is implemented by a combination of hardware and software.

[0026] Figure 3 is a block diagram illustrating the hardware configuration of the computer 1000 that implements the advice provision device 2000. The computer 1000 is any computer. For example, the computer 1000 is a stationary computer such as a PC (Personal Computer) or a server machine. Alternatively, the computer 1000 is a portable computer such as a smartphone or a tablet terminal. The computer 1000 may be a dedicated computer designed to implement the advice provision device 2000, or it may be a general-purpose computer.

[0027] For example, by installing a predetermined application on computer 1000, the various functions of the advice provision device 2000 are realized on computer 1000. The above application consists of programs for realizing each functional component of the advice provision device 2000. The method of obtaining the above program is arbitrary. For example, the program can be obtained from a storage medium on which it is stored. The storage medium on which the program is stored can be any storage medium such as a DVD (Digital Versatile Disk) or a USB (Universal Serial Bus) memory. Alternatively, for example, the program can be obtained by downloading it from a server device that manages the storage device on which the program is stored.

[0028] Computer 1000 includes a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, memory 1060, storage device 1080, input / output interface 1100, and network interface 1120 to send and receive data to and from each other. However, the method of connecting the processor 1040 and the other components is not limited to bus connection.

[0029] The processor 1040 is a variety of processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an FPGA (Field-Programmable Gate Array). The memory 1060 is a main memory device implemented using RAM (Random Access Memory), etc. The storage device 1080 is an auxiliary storage device implemented using a hard disk, SSD (Solid State Drive), memory card, or ROM (Read Only Memory), etc.

[0030] The input / output interface 1100 is an interface for connecting the computer 1000 with input / output devices. For example, input devices such as keyboards and output devices such as display devices are connected to the input / output interface 1100.

[0031] The network interface 1120 is an interface for connecting the computer 1000 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).

[0032] The storage device 1080 stores programs that implement each functional component of the advice provision device 2000 (programs that implement the aforementioned applications). The processor 1040 reads these programs into the memory 1060 and executes them to implement each functional component of the advice provision device 2000.

[0033] The advice-providing device 2000 may be implemented using one computer 1000 or multiple computers 1000. In the latter case, the configuration of each computer 1000 does not need to be the same and can be different.

[0034] Here, the language model 70 may operate on the advice-providing device 2000, or it may operate on a device other than the advice-providing device 2000. In the latter case, the hardware configuration of the device on which the language model 70 is executed is similar to the hardware configuration of the advice-providing device 2000, as shown in Figure 3, for example. However, the hardware configuration of the device on which the language model 70 is executed does not have to be the same as the hardware configuration of the advice-providing device 2000.

[0035] <Processing Flow> Figure 4 is a flowchart illustrating the processing flow performed by the advice provision device 2000. The acquisition unit 2020 acquires the first video data 10 and the second video data 20 (S102). The evaluation unit 2040 performs an action evaluation of the target work performed by the target worker (S104). The generation unit 2060 generates request data 60 based on the results of the action evaluation (S106). The output unit 2080 obtains response data 80 by inputting the request data 60 into the language model 70 (S108). The output unit 2080 outputs advice data 90 based on the response data 80 (S110).

[0036] <Acquisition of the first video data 10: S102> The acquisition unit 2020 acquires the first video data 10 (S102). Various methods can be used to acquire the video data to be processed. For example, the first video data 10 is stored in a storage device in advance in a manner that can be acquired from the advice providing device 2000. In this case, the acquisition unit 2020 acquires the first video data 10 by reading the first video data 10 from the storage device.

[0037] In addition, for example, the acquisition unit 2020 acquires the first video data 10 by receiving the first video data 10 transmitted from another device. The device that transmits the first video data 10 is, for example, the device that generated the first video data 10. The device that generated the first video data 10 is, for example, a camera.

[0038] Here, it is possible to select a target task from among several predetermined tasks. In this case, example video data is prepared in advance for each of the predetermined tasks. The user of the advice-providing device 2000 selects a task to be treated as the target task from among the multiple predetermined tasks. The advice-providing device 2000 acquires the example video data prepared in association with the selected task as the first video data 10.

[0039] <Acquisition of the second video data 20: S102> The acquisition unit 2020 acquires the second video data 20 (S102). The acquisition unit 2020 can use, for example, the various methods described above as the method by which the acquisition unit 2020 acquires the first video data 10 to acquire the second video data 20. If both the first video data 10 and the second video data 20 are stored in a storage device, the first video data 10 and the second video data 20 may be stored in the same storage device or in different storage devices. Also, if both the first video data 10 and the second video data 20 are transmitted from a device other than the advice providing device 2000, the device transmitting the first video data 10 and the device transmitting the second video data 20 may be the same device or in different devices.

[0040] <Execution of motion evaluation: S104> The evaluation unit 2040 uses the first video data 10 and the second video data 20 to perform a motion evaluation on the target work performed by the target worker (S104). The specific contents of the motion evaluation are exemplified below.

[0041] For example, the evaluation unit 2040 performs an operational evaluation by associating frames between the first video data 10 and the second video data 20. Specifically, the evaluation unit 2040 identifies the corresponding second frame 22 for each of the multiple first frames 12 contained in the first video data 10. Here, the combination of the first frame 12 and the second frame 22 that are associated with each other is also expressed as a frame pair.

[0042] FIG. 5 is a diagram illustrating a frame pair. In FIG. 5, the frame pair is associated with each other by the same arrow.

[0043] The association between frames (in other words, the generation of a frame pair) is performed based on the operation similarity. The operation similarity calculated for a certain frames F1 and F2 represents the degree of similarity between the operation captured in frame F1 and the operation captured in frame F2. A specific method for calculating the operation similarity will be described later.

[0044] For example, the evaluation unit 2040 calculates the operation similarity between each first frame 12 and each of all the second frames 22. For each first frame 12, the evaluation unit 2040 identifies the second frame 22 having the maximum operation similarity with the first frame 12 as the second frame 22 corresponding to the first frame 12.

[0045] Here, it is also conceivable that there is no second frame 22 corresponding to the first frame 12, such as when there is an error in the operation of the target operator. Therefore, for example, a lower threshold value may be provided for the above-described operation similarity. In this case, when the maximum operation similarity calculated for a certain first frame 12 is equal to or greater than the lower threshold value, the evaluation unit 2040 identifies the second frame 22 for which the maximum similarity is calculated as the second frame 22 corresponding to the first frame 12. On the other hand, when the maximum operation similarity calculated for the first frame 12 is less than the lower threshold value, the evaluation unit 2040 determines that there is no second frame 22 corresponding to the first frame 12. Therefore, no frame pair is generated for this first frame 12.

[0046] The second frame 22 for which the operation similarity is calculated with the first frame 12 may be a part of the second frames 22 included in the second video data 20. For example, assume that the target operation is composed of a plurality of steps. In this case, the evaluation unit 2040 generates a frame pair for each step.

[0047] Therefore, the evaluation unit 2040 divides the second video data 20 into partial video data for each process. For example, assume that the target operation is composed of processes P1, P2, and P3. In this case, the evaluation unit 2040 divides the second video data 20 into partial video data V21 including the state of process P1, partial video data V22 including the state of process P2, and partial video data V23 including the state of process P3.

[0048] Similarly, the evaluation unit 2040 can divide the first video data 10 into partial video data for each process. For example, when the target operation is composed of processes P1, P2, and P3 as described above, the first video data 10 is divided into partial video data V11 including the state of process P1, partial video data V12 including the state of process P2, and partial video data V13 including the state of process P3. However, since the first video data 10 is reference video data prepared in advance, it may be divided in advance for each process.

[0049] Assume that the first video data 10 and the second video data 20 are each divided into a plurality of partial video data in this way. In this case, the evaluation unit 2040 generates frame pairs from partial video data representing the same process for each other. For example, in the above example, the evaluation unit 2040 searches for corresponding second frames 22 from the second frames 22 included in the partial video data V21 for each first frame 12 included in the partial video data V11. By generating frame pairs from the same process for each other in this way, it is possible to prevent the operations of different processes from being erroneously associated with each other.

[0050] <<Regarding the operation similarity>> An example of a method for calculating the operation similarity will be described. For example, the evaluation unit 2040 calculates an operation feature amount representing the feature of the operation captured in each frame for each frame. Then, the evaluation unit 2040 calculates the similarity of the operation feature amounts of two frames as the operation similarity of those two frames.

[0051] Motion features are calculated using, for example, skeletal features, image features, or both. The evaluation unit 2040 calculates skeletal features representing the skeletal features of the worker's body included in the frame. Similarly, the evaluation unit 2040 calculates image features representing the features of the image region of the worker's body included in the frame.

[0052] The motion features may further include features related to grasping an object (hereinafter referred to as "grasping features"). Grasping features are features that indicate whether or not an object is being grasped by the operator. For example, if the operator's hand and the object overlap in a given frame, the evaluation unit 2040 calculates a grasping feature for that frame that indicates the object is being grasped by the operator. On the other hand, if the operator's hand and the object do not overlap in that frame, the evaluation unit 2040 calculates a grasping feature for that frame that indicates the object is not being grasped by the operator.

[0053] The method for determining whether an object is being held by a worker is not limited to determining whether the object and the worker's hand are overlapping. For example, whether an object is being held by a worker can be determined using a trained machine learning model (hereinafter referred to as a gripping determination model).

[0054] The grip determination model is configured to calculate a grip feature that indicates whether or not an object is being grasped in an image, in response to an input image. The evaluation unit 2040 inputs the frame to the grip determination model to obtain the grip feature for that frame.

[0055] The grasp determination model is pre-trained using multiple training samples. A training sample is a combination of a training image and a ground truth grasp feature that indicates whether or not an object is being grasped in that training image.

[0056] Here, if the motion features are calculated using multiple features (for example, two or more of skeletal features, image features, and grasping features), the evaluation unit 2040 calculates the motion features by combining these multiple features. For example, motion features can be represented by concatenation of multiple features. Alternatively, for example, motion features can be calculated by combining multiple features using a trained machine learning model.

[0057] <<Results of the motion evaluation>> The results of the motion evaluation are represented, for example, by the correspondence between the first frame 12 and the second frame 22. Specifically, the results of the motion evaluation can be represented by a numerical sequence {g[1], g[2], ..., g[N]} that represents a set of frame pairs. Here, N represents the total number of first frames 12 included in the first video data 10. g[i] represents the frame number of the second frame 22 that corresponds to the i-th first frame 12. Thus, the frame pair corresponding to the i-th first frame 12 is represented by a combination of frame numbers (i, g[i]). Hereafter, the frame pair corresponding to the i-th first frame 12 will also be expressed as the i-th frame pair.

[0058] In addition, the results of the motion evaluation are represented by the magnitude of the discrepancy between the ideal and the actual relationship in the correspondence between the first frame 12 and the second frame 22. The magnitude of the discrepancy between the ideal and the actual relationship in the frame correspondence will be explained below with reference to Figure 6.

[0059] Figure 6 illustrates a graph showing the correspondence between the first frame 12 and the second frame 22. The horizontal axis of the graph represents the frame number of the first frame 12. The vertical axis of the graph represents the frame number of the second frame 22. Each point on the graph (relationship point 100) represents a frame pair. For example, as mentioned above, suppose a frame pair is represented by a combination of frame numbers (i, g[i]). In this case, for each of i=1 to N, a relationship point 100 is plotted at the coordinate (i, g[i]).

[0060] Here, if the target work performed by the target worker is ideal (exactly as per the example), all relation points lie on the same straight line (ideal straight line 110). On the other hand, if there are actions in the target worker's actions that differ from the example, there will be relation points 100 that deviate from the original relation points 100. From this, it can be said that the degree to which each relation point 100 is close to the ideal straight line 110 represents the degree to which the target worker's action corresponding to that relation point 100 is close to the example.

[0061] For example, the evaluation unit 2040 calculates a motion evaluation score that represents the level of motion evaluation of the target worker based on the magnitude of the deviation between each related point 100 and the ideal straight line 110. The motion evaluation score is calculated, for example, by the following formula (1). In equation (1), Ea represents the motion evaluation score. i represents the frame number of the first frame 12. Si represents the difference between the relationship point 100 representing the i-th frame pair and the ideal line 110. A represents the adjustment coefficient that adjusts the minimum and maximum values ​​of the normal distribution exp to 0 and 100, respectively. σ represents the variance of the normal distribution exp.

[0062] The difference between point 100 and the ideal line 110 can be expressed, for example, by the shortest distance between point 100 and the ideal line 110 (the length of the perpendicular line drawn from point 100 to the ideal line 110). Alternatively, the difference between point 100 and the ideal line 110 can be expressed by the distance between point 100 and the ideal line 110 in the X-axis direction, or the distance between point 100 and the ideal line 110 in the Y-axis direction.

[0063] In calculating the motion evaluation score, weighting based on the motion similarity described above may be applied. In this case, the motion evaluation score is calculated using, for example, the following formula (2). wi represents the behavioral similarity of the i-th frame pair (the behavioral similarity calculated between the i-th first frame 12 and the second frame 22 associated with that first frame 12).

[0064] Here, the greater the operational similarity of a frame pair, the higher the reliability of the correspondence within that frame pair. Therefore, it is preferable that the operational evaluation score increases as the operational similarity of the frame pair increases. In equation (2), the difference between the ideal line 110 and the ideal line 110 is multiplied by the reciprocal of the operational similarity. As a result, the greater the operational similarity, the smaller the influence of the difference between the relationship point 100 and the ideal line 110 on the operational evaluation score, and thus the higher the operational evaluation score.

[0065] There are various ways to set the ideal straight line 110. For example, the evaluation unit 2040 may treat the straight line connecting the relationship point 100 corresponding to the first frame 12 and the relationship point 100 corresponding to the last frame 12 as the ideal straight line 110. Alternatively, the evaluation unit 2040 may also treat the straight line obtained by applying any straight line fitting method, such as the least squares method, to all relationship points 100 as the ideal straight line 110.

[0066] <Generation of Request Data 60> The generation unit 2060 generates request data 60 based on the results of the operation evaluation. The request data 60 is data that represents a request for advice based on a comparison with an example regarding the target work performed by the target worker. For example, the request data 60 includes text data of a sentence expressing the request for advice (hereinafter referred to as the request text) and reference data which is data useful for generating advice.

[0067] For example, the request text might include phrases like, "Generate advice to help the worker's actions more closely resemble the example." The request text may also specify further constraints, such as character limits, on the requested advice.

[0068] The request text may further include sentences that express the assumptions of the advice. These assumptions may include, for example, a description of the task in question. Including a description of the task in the request text enables the language model 70 to generate advice while taking the content of the task in question into consideration.

[0069] The description of the target task includes, for example, an overview of the task, a description of each of the multiple processes that make up the task, and the sequence of the processes. The following are specific examples of the description of the target task.

[0070] - We are assembling the product. - The work includes processes P1, P2, and P3 in that order. - Process P1 is the process of doing A. - Process P2 is the process of doing B. - Process P3 is the process of doing C.

[0071] The premise of the advice may include a description of the worker in question. This description may include, for example, information indicating the worker's skill level (e.g., years of experience).

[0072] The reference data includes, for example, one or more frame pairs generated during the motion evaluation. By including frame pairs in the reference data, the language model 70 can generate advice based on the differences in corresponding frames between the motion performed by the operator and the example motion.

[0073] Furthermore, the reference data may include descriptions for each frame pair. For example, for each of one or more frame pairs, the reference data may include text describing the first frame 12, image data of the first frame 12, text describing the second frame 22, and image data of the second frame 22. The descriptions for the first frame 12 and the second frame 22 may be "This is an image showing an example of the work being done" and "This is an image showing work being done by a worker," respectively.

[0074] The description of a frame may further indicate which process is being captured in that frame. For example, suppose process P1 is being performed in the second frame 22. In this case, the description of the second frame 22 could be a sentence such as, "Process P1 is being performed by an operator."

[0075] The description of a frame may further indicate whether or not an object is being held in that frame. For example, suppose an object is being held in the second frame 22. In this case, the description of the second frame 22 could be: "This is an image showing work being performed by a worker. An object is being held."

[0076] Here, the worker's actions differ significantly depending on whether or not they are grasping an object. Therefore, whether or not an object is being grasped is a very important characteristic of the action. By including information indicating whether or not an object is being grasped in the request data 60, the language model 70 can accurately grasp the characteristics of the action and generate advice.

[0077] The reference data may further include skeletal information extracted from the first frame 12 and skeletal information extracted from the second frame 22. The skeletal information, for example, shows the coordinates of each joint point extracted from the frame. Alternatively, the skeletal information may be image data representing the skeleton extracted from the frame. In such image data, the skeleton is represented by each joint point and lines connecting adjacent joint points. By including skeletal information in the reference data, the language model 70 can generate advice using more detailed information about the worker's movements and the example movements.

[0078] The reference data may include not only the first frame 12 and the second frame 22 included in the frame pair, but also the frames before and after them. For example, the reference data may include the first frame 12 and a predetermined number of first frames 12 before and after it, as well as the second frame 22 and a predetermined number of second frames 22 before and after it, as reference data in the request data 60. In this way, short clips of the first frame 12 and the second frame 22 that constitute the frame pair are included in the request data 60.

[0079] For example, suppose a short clip consists of five consecutive frames. In this case, the following sentences might be used as descriptions for the first frame 12 and the second frame 22: "These five images represent video of the example work," and "These five images represent video of the work performed by the worker," respectively.

[0080] Here, the reference data may include all of the frame pairs generated in the operational evaluation, or it may include only a portion of the frame pairs generated in the operational evaluation. In the latter case, for example, the generation unit 2060 includes a predetermined number of frame pairs in the reference data in ascending order of the magnitude of the difference between the corresponding relationship point 100 and the ideal straight line 110.

[0081] Here, the greater the difference between the frame pair and the ideal, the greater the difference between the worker's actions and the example actions, and therefore the more useful the advice is considered to be for the worker. Therefore, by using a predetermined number of frame pairs in ascending order of the magnitude of the difference between the corresponding relationship point 100 and the ideal straight line 110, the advice is generated by preferentially using frame pairs with higher usefulness.

[0082] <Output of advice data 90: S108> The output unit 2080 inputs the request data 60 to the language model 70 to obtain response data 80. The response data 80 is text data of a sentence representing advice.

[0083] The language model 70 may be a general-purpose language model or a dedicated language model created for use with the advice-providing device 2000. In the latter case, the language model 70 is pre-trained using multiple training samples consisting of training request data and ground truth response data to said request data. Alternatively, the language model 70 may be created by fine-tuning a general-purpose language model using multiple of the above training samples.

[0084] The output unit 2080 outputs advice data 90 based on the response data 80. The advice data 90 may be the response data 80 itself, or it may be data generated by processing the response data 80.

[0085] In the latter case, for example, the output unit 2080 generates advice data 90 by performing a text modification process (hereinafter referred to as text modification process) on the response data 80. The text modification process is, for example, a process that replaces inappropriate words or phrases (hereinafter referred to as words, etc.) with appropriate words, etc. Inappropriate words, etc. and appropriate words, etc. are, for example, informal words, etc. and formal words, etc. By performing the text modification process on the response data 80, the advice provision device 2000 can provide advice to its user in more appropriate sentences.

[0086] The method for implementing the document correction process is arbitrary. For example, the output unit 2080 has a machine learning model (hereinafter referred to as the document correction model) that implements the document correction process. The document correction model is pre-trained to output a second text data composed of more appropriate words, etc., in response to the input of a first text data. The output unit 2080 inputs the response data 80 into the document correction model and uses the text data output from the document correction model as advice data 90.

[0087] The output unit 2080 outputs advice data 90. The manner in which the advice data 90 is output is arbitrary. For example, the output unit 2080 stores the advice data 90 in any storage device. Alternatively, for example, the output unit 2080 transmits the advice data 90 to any device. Alternatively, for example, the output unit 2080 may display the advice data 90 on any display device.

[0088] The advice data 90 may be part of the data output by the output unit 2080. For example, the output unit 2080 outputs a screen that includes advice for the worker.

[0089] Figure 7 illustrates a screen output by the advice-providing device 2000. The screen 120 includes an area 122 that displays the correspondence between the first video data 10 and the second video data 20, an area 124 that displays advice to the worker, and an area 126 that displays video.

[0090] Area 122 shows the correspondence between each first frame 12 of the first video data 10 (example) and each second frame 22 of the second video data 20 (you). Specifically, it shows frame pairs with particularly high similarity in the "Same Action" column. The "Advice" column shows the frame from which the advice was generated. Frames marked with a circle represent the frame pair currently selected by the user.

[0091] Area 124 shows the advice data 90 generated for the selected frame pair. Area 126 shows the first frame 12 and the second frame 22, respectively, included in the selected frame pair. Here, skeletal information is superimposed on each frame. The user can play the first video data 10 and the second video data 20, respectively, from the current frame by pressing the play button.

[0092] <Regarding evaluations other than motion evaluation> The advice-providing device 2000 may further perform evaluations other than motion evaluation on the target work performed by the target worker. For example, the advice-providing device 2000 may use the first video data 10 and the second video data 20 to further perform an evaluation regarding the appropriateness of the speed of the target work performed by the target worker (hereinafter referred to as speed evaluation) and an evaluation regarding the appropriateness of the order of processes in the target work performed by the target worker (hereinafter referred to as order evaluation).

[0093] In the speed evaluation, a speed evaluation value is calculated that represents the appropriateness of the speed of the target work. The speed evaluation value decreases as the difference between the speed of the target work captured in the first video data 10 and the speed of the target work captured in the second video data 20 increases.

[0094] The speed comparison may be performed for each process. In this case, the advice-providing device 2000 calculates an evaluation value regarding the appropriateness of the speed for each process, and calculates a speed evaluation value from the statistical value of the multiple evaluation values ​​calculated. When a weighted average is calculated as the statistical value, the importance of each predetermined process is used as the weight.

[0095] In sequence evaluation, a sequence evaluation value is calculated that represents the appropriateness of the sequence of processes. For example, the advice-providing device 2000 calculates an index value that represents the difference between the sequence of processes in the example of the target work and the sequence of processes in the target work performed by the target worker. This index value is expressed, for example, as the edit distance between the permutation representing the sequence of processes in the target work performed by the target worker and the permutation representing the sequence of processes in the example of the target work.

[0096] The advice provider 2000 calculates an order evaluation value using the calculated edit distance. For example, the advice provider 2000 calculates an order evaluation value by normalizing the reciprocal of the calculated edit distance such that the minimum value is 0 and the maximum value is 100.

[0097] The advice-providing device 2000 generates screen data that, for example, displays operation evaluation values, speed evaluation values, sequence evaluation values, and advice data 90, and outputs the screen data in any manner. The screen data may also further display an overall evaluation value calculated based on multiple evaluation values. The overall evaluation value is, for example, a statistical value of the operation evaluation value, speed evaluation value, and sequence evaluation value.

[0098] [Embodiment 2] Instead of the advice data 90, a device for outputting the results of the operation evaluation may be provided. This device is called an evaluation device. The evaluation device acquires the first video data 10 and the second video data 20, and performs an operation evaluation using the first video data 10 and the second video data 20. The evaluation device then outputs evaluation data representing the results of the operation evaluation. The evaluation data includes, for example, an operation evaluation value. The output mode of the evaluation data is arbitrary, similar to the output mode of the advice data 90.

[0099] In this evaluation device, it is preferable to include gripping features in the motion features during motion evaluation. As mentioned above, a worker's movements differ significantly depending on whether or not they are gripping an object. Therefore, whether or not an object is being gripped is a very important characteristic of the movement. By including gripping features in the motion features, the movements of the target worker can be evaluated more accurately.

[0100] Figure 8 illustrates the functional configuration of the evaluation device. The evaluation device 3000 includes, for example, an acquisition unit 3020, an evaluation unit 3040, and an output unit 3060. The acquisition unit 3020 and the evaluation unit 3040 have the same functions as the acquisition unit 2020 and the evaluation unit 2040, respectively. The output unit 3060 outputs evaluation data representing the results of the operation evaluation.

[0101] Figure 9 is a flowchart illustrating the processing flow performed by the evaluation device 3000. The acquisition unit 3020 acquires the first video data 10 and the second video data 20 (S302). The method by which the acquisition unit 3020 acquires the first video data 10 and the second video data 20 is the same as the method by which the acquisition unit 2020 acquires the first video data 10 and the second video data 20. The evaluation unit 3040 performs an operation evaluation using the first video data 10 and the second video data 20 (S204). The method by which the evaluation unit 3040 performs the operation evaluation is the same as the method by which the evaluation unit 2040 performs the operation evaluation.

[0102] The output unit 3060 outputs evaluation data representing the results of the operation evaluation (S306). For example, the output unit 3060 outputs the operation evaluation value calculated by the evaluation unit 3040 as evaluation data.

[0103] The evaluation device 3000 may perform speed evaluation and sequence evaluation in addition to the operation evaluation. In this case, the evaluation data will include speed evaluation values ​​and sequence evaluation values ​​in addition to the operation evaluation values. The evaluation data may also include an overall evaluation value.

[0104] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0105] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments, rather than being associated with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps described in any of the drawings may be changed as appropriate.

[0106] In this disclosure, a program includes a set of instructions (or software code) that, when loaded into a computer, causes the computer to perform one or more of the functions described in the embodiments. A program may be stored on a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. A program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically or otherwise propagating signals.

[0107] Some or all of the above embodiments may also be described as follows, but are not limited to the following: (Note 1) An advice providing device comprising: an acquisition means for acquiring first video data including a scene in which a target work is performed by a target worker and second video data including a scene in which the target work is performed correctly; an evaluation means for performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; a generation means for generating request data for advice to the target worker regarding the target work based on the results of the action evaluation; and an output means for inputting the request data into a language model and outputting advice data representing the advice using response data obtained from the language model. (Note 2) The advice providing device according to Note 1, wherein the evaluation means calculates an action similarity score representing the degree of similarity of the captured actions between each first frame constituting the first video data and each second frame constituting the second video data; generates a plurality of pairs of the first frame and the second frame based on the action similarity score; and the output means generates the request data including one or more of the generated pairs. (Note 3) The advice-providing device according to Note 2, wherein the evaluation means calculates a value representing the difference from the ideal for each pair in which the second frame associated with the first frame differs from the ideal, and the output means generates the request data including a predetermined number of the pairs in ascending order of the calculated values. (Note 4) The advice-providing device according to Note 2 or 3, wherein the output means includes information on the skeleton of the person being imaged for each of the first frame and the second frame included in the request data. (Note 5) The advice-providing device according to any one of Notes 1 to 3, wherein the output means generates the request data including an explanation of the target work.(Note 6) The advice-providing device according to Note 2 or 3, wherein the evaluation means calculates motion feature quantities representing the characteristics of the motion being captured for each of the first frame and the second frame, calculates the degree of similarity between the motion feature quantities of the first frame and the motion feature quantities of the second frame as the motion similarity, and the motion feature quantities include a gripping feature quantity representing whether or not an object is being gripped. (Note 7) The advice-providing device according to any one of Notes 1 to 3, wherein the output means generates the advice data by performing a process on the response data to modify words, phrases, or both. (Note 8) A computer-based advice provision method comprising: an acquisition step of acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; an evaluation step of performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; a generation step of generating request data that seeks advice for the target worker regarding the target task based on the results of the action evaluation; and an output step of inputting the request data into a language model and outputting advice data representing the advice using response data obtained from the language model. (Note 9) A program that causes a computer to execute the following steps: an acquisition step of acquiring first video data including a scene in which the target work is performed by the target worker and second video data including a scene in which the target work is performed correctly; an evaluation step of performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; a generation step of generating request data that seeks advice for the target worker regarding the target work based on the results of the action evaluation; and an output step of inputting the request data into a language model and outputting advice data representing the advice using the response data obtained from the language model.(Note 10) An evaluation device comprising: an acquisition means for acquiring first video data including a scene in which a target work is performed by a target worker and second video data including a scene in which the target work is performed correctly; an evaluation means for calculating an action evaluation value representing the appropriateness of the target worker's actions using the first video data and the second video data; and an output means for outputting the action evaluation value, wherein the evaluation means calculates an action feature quantity representing the characteristics of the captured action for each first frame included in the first video frame and each second frame included in the second video frame; associates the first frame and the second frame based on the degree of similarity between the action feature quantity of the first frame and the action feature quantity of the second frame; calculates the action evaluation value using the result of the association; and the action feature quantity includes a gripping feature quantity representing whether or not an object is being gripped. (Note 11) An evaluation method performed by a computer, comprising: an acquisition step of acquiring first video data including a scene in which a target work is performed by a target worker and second video data including a scene in which the target work is performed correctly; an evaluation step of calculating an action evaluation value representing the appropriateness of the target worker's actions using the first video data and the second video data; and an output step of outputting the action evaluation value, wherein in the evaluation step, for each first frame included in the first video frame and each second frame included in the second video frame, an action feature quantity representing the characteristics of the captured action is calculated; an association is made between the first frame and the second frame based on the degree of similarity between the action feature quantity of the first frame and the action feature quantity of the second frame; the action evaluation value is calculated using the result of the association, and the action feature quantity includes a gripping feature quantity representing whether or not an object is being gripped.(Note 12) A program that causes a computer to perform the following steps: an acquisition step of acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; an evaluation step of calculating an action evaluation value representing the appropriateness of the target worker's actions using the first video data and the second video data; and an output step of outputting the action evaluation value; wherein in the evaluation step, for each first frame included in the first video frame and each second frame included in the second video frame, an action feature quantity representing the characteristics of the captured action is calculated; an association is made between the first frame and the second frame based on the degree of similarity between the action feature quantity of the first frame and the action feature quantity of the second frame; the action evaluation value is calculated using the result of the association; and the action feature quantity includes a gripping feature quantity representing whether or not an object is being gripped.

[0108] Some or all of the elements (e.g., configuration and function) described in Appendices 2 to 7 that are subordinate to Appendice 1 may also be subordinate to Appendices 8 and 9 in the same manner as those described in Appendices 2 to 7. Some or all of the elements described in any appendice may be applied to various hardware, software, recording means, systems, and methods for recording software.

[0109] This application claims priority based on Japanese Patent Application No. 2024-165880, filed on 25 September 2024, and incorporates all of its disclosures herein.

[0110] 10 First video data 12 First frame 20 Second video data 22 Second frame 60 Request data 70 Language model 80 Response data 90 Advice data 100 Relationship point 110 Ideal line 120 Screen 122 Area 124 Area 126 Area 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface 2000 Advice provider 2020 Acquisition unit 2040 Evaluation unit 2060 Generation unit 2080 Output unit 3000 Evaluation device 3020 Acquisition unit 3040 Evaluation unit 3060 Output unit

Claims

1. An advice providing device comprising: acquisition means for acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; evaluation means for performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; generation means for generating request data to request advice for the target worker regarding the target task based on the results of the action evaluation; and output means for inputting the request data into a language model and outputting advice data representing the advice using response data obtained from the language model.

2. The advice providing device according to claim 1, wherein the evaluation means calculates an action similarity score representing the degree of similarity of the captured action between each first frame constituting the first video data and each second frame constituting the second video data, generates a plurality of pairs of the first frame and the second frame based on the action similarity scores, and the output means generates the request data including one or more of the generated pairs.

3. The advice providing device according to claim 2, wherein the evaluation means calculates a value representing the difference from the ideal for each pair in which the second frame associated with the first frame differs from the ideal, and the output means generates the request data including a predetermined number of the pairs in ascending order of the calculated values.

4. The advice-providing device according to claim 2 or 3, wherein the output means includes information on the skeleton of the person being imaged in the request data for each of the first frame and the second frame to be included in the request data.

5. The advice-providing device according to any one of claims 1 to 3, wherein the output means generates the request data which includes a description of the target work.

6. The advice-providing device according to claim 2 or 3, wherein the evaluation means calculates motion feature quantities representing the characteristics of the motion being captured for each of the first frame and the second frame, calculates the degree of similarity between the motion feature quantities of the first frame and the motion feature quantities of the second frame as the motion similarity, and the motion feature quantities include a gripping feature quantity indicating whether or not an object is being gripped.

7. The advice providing device according to any one of claims 1 to 3, wherein the output means generates the advice data by performing a process on the response data to modify words, phrases, or both.

8. A computer-based advice provision method comprising: an acquisition step of acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; an evaluation step of performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; a generation step of generating request data that seeks advice for the target worker regarding the target task based on the results of the action evaluation; and an output step of inputting the request data into a language model and outputting advice data representing the advice using response data obtained from the language model.

9. A program that causes a computer to execute the following steps: an acquisition step of acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; an evaluation step of performing an action evaluation regarding the appropriateness of the target worker's actions using the first video data and the second video data; a generation step of generating request data that seeks advice for the target worker regarding the target task based on the results of the action evaluation; and an output step of inputting the request data into a language model and outputting advice data representing the advice using response data obtained from the language model.

10. An evaluation device comprising: an acquisition means for acquiring first video data including a scene in which a target worker performs a target task and second video data including a scene in which the target task is performed correctly; an evaluation means for calculating an action evaluation value representing the appropriateness of the target worker's actions using the first video data and the second video data; and an output means for outputting the action evaluation value, wherein the evaluation means calculates an action feature quantity representing the characteristics of the captured action for each first frame included in the first video frame and each second frame included in the second video frame; associates the first frame and the second frame based on the degree of similarity between the action feature quantity of the first frame and the action feature quantity of the second frame; calculates the action evaluation value using the result of the association; and the action feature quantity includes a gripping feature quantity representing whether or not an object is being gripped.

Citation Information

Patent Citations

  • Motion analysis device, motion analysis method and motion analysis program

    JP2020106954A

  • Learning assisting system, learning assisting device, and program

    JP2020144233A

  • Care support device, care support program, and care support method

    JP2023097545A

  • Medical Support Devices

    JP7454090B1

  • Methods for determining manufacturing waste to optimize productivity and devices thereof

    US20160239769A1