Work analysis system, work analysis method, and work analysis program

WO2026159976A1PCT designated stage Publication Date: 2026-07-30HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HITACHI LTD
Filing Date
2025-10-29
Publication Date
2026-07-30

Smart Images

  • Figure JP2025038060_30072026_PF_FP_ABST
    Figure JP2025038060_30072026_PF_FP_ABST
Patent Text Reader

Abstract

This work analysis system can be connected to a model that is generated by pre-training by using, as learning data, a plurality of modalities including images and text, and that receives, as input, data in which actions pertaining to a work by a worker are recorded and outputs the analysis result on the basis of the input. The work analysis system: identifies the type of the work on the basis of data pertaining to the work by the worker; inputs, to the model, a prompt instructing output of the analysis result based on the input data together with the input data when it is determined that the work is a target work; acquires the analysis result output by the model on the basis of the input data and the prompt; and outputs the analysis result via an output unit.
Need to check novelty before this filing date? Find Prior Art

Description

Work analysis system, work analysis method, and work analysis program

[0001] The present invention relates to a work analysis system, a work analysis method, and a work analysis program.

[0002] In recent years, due to the declining birthrate and aging population, a shortage of human resources has spread in workplaces of all industries, and the transfer of skills from skilled workers has become an issue.

[0003] Regarding the transfer of skills from skilled workers, for example, in Patent Document 1, tacit knowledge regarding the skills of caregivers / nurses in a caregiving site is accumulated in a form that can be used by others, and a technique for performing regression analysis on the importance of tacit knowledge based on whether the tacit knowledge can actually be used or not is disclosed. By the technique disclosed in Patent Document 1, caregivers / nurses at a caregiving site can efficiently transfer skills by using tacit knowledge with a high degree of importance.

[0004] Japanese Unexamined Patent Application Publication No. 2022-189528

[0005] However, in the above-mentioned conventional technology, there was still room for further improvement in terms of the transfer of skills from skilled workers to unskilled workers.

[0006] For example, in Patent Document 1, only pre-defined tacit knowledge is presented, and tacit knowledge that skilled workers cannot verbalize (formalize) or tacit knowledge that skilled workers are not aware of is not extracted. Therefore, there is tacit knowledge that is not transferred from skilled workers to unskilled workers, and there is a problem with the efficiency of skill transfer.

[0007] The present invention has been made in consideration of the above points, and an object thereof is to efficiently perform the transfer of skills from skilled workers to unskilled workers at a work site.

[0008] To achieve the above objective, the present invention, in one embodiment, provides a work analysis system that outputs analysis results for improving the work skills of an operator when performing a task, wherein the work analysis system is connected to a model that is generated by pre-training a plurality of modalities including images and text as learning data, and takes data recording the operator's actions related to the task as input and outputs the analysis results based on the input, wherein the processor of the work analysis system receives the operator's data as input data, identifies the type of task based on the input data, determines whether the task is a target task based on the type of task, and if it is determined that the task is a target task, inputs a prompt to the model along with the input data to instruct the output of the analysis results based on the input data, acquires the analysis results output by the model based on the input data and the prompt, and outputs the acquired analysis results via an output unit.

[0009] According to the present invention, for example, it is possible to efficiently transfer skills from skilled workers to unskilled workers at a work site.

[0010] A diagram showing the configuration of the work analysis system according to Embodiment 1. A diagram showing the prompts according to Embodiment 1. A flowchart showing the work analysis process according to Embodiment 1. A diagram showing the work analysis result display screen according to Embodiment 1. A diagram showing the configuration of the work analysis system according to Embodiment 2. A flowchart showing the work analysis process according to Embodiment 2. A diagram showing the configuration of the work analysis system according to Embodiment 3. A flowchart showing the work analysis process according to Embodiment 3. A diagram showing the know-how question screen according to Embodiment 3. A diagram showing the report output prompt according to Embodiment 3. A diagram showing the report output screen according to Embodiment 3. A diagram showing a modified version of Embodiment 3's report output prompt. A diagram showing a modified version of Embodiment 3's report output screen.

[0011] In the following explanation, "processor" refers to, for example, a CPU (Central Processing Unit), but may also refer to other types of processors such as a GPU (Graphics Processing Unit). Furthermore, "processor" may also refer to a broader definition of a processor, such as a hardware circuit that performs some or all of the processing (e.g., an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)).

[0012] Furthermore, in the following explanation, the processing will be described using "XXX section" as the subject. The XXX section performs defined processing as appropriate, using memory and / or interface devices, etc., when the program is executed by the processor. For this reason, the subject of the processing may also be the processor. The program may be installed from the program source into a device such as a computer. The program source may be, for example, a program distribution server or a computer-readable (e.g., non-temporary) recording medium. Also, two or more programs may be implemented as a single program, or one program may be implemented as two or more programs.

[0013] [Embodiment 1] (Configuration of the work analysis system 1 according to Embodiment 1) Figure 1 is a diagram showing the configuration of the work analysis system 1 according to Embodiment 1. The work analysis system 1 is a computer that outputs analysis results to improve the work skills of worker h when performing work. The work analysis system 1 has a processor 11, a storage unit 12, a memory 13, and a communication I / F (Interface) 14. A sensor 15 and an input / output unit 16 are connected to the work analysis system 1.

[0014] Furthermore, the work analysis system 1 is connected to the VLM (Vision and Language Model) server 2 via the network N. The VLM server 2 is equipped with VLM2M. VLM2M is a type of model that constitutes a generative AI (Artificial Intelligence) capable of recognizing multiple modalities, including images and text. It is a model generated by pre-training with multiple modalities, including images and text, as training data. Modalities include images and text, as well as sensor signals, etc. VLM2M is capable of describing input images and generates answers to questions based on multiple modalities, including images and text.

[0015] VLM2M is a model trained using various data related to the work of worker h with various work skills as training data. Alternatively, even if VLM2M has not been trained with various data related to the work of worker h, it can also use RAG (Retrieval-Augmented Generation), in which data related to the work of skilled workers with a certain level of work skill is retrieved by searching or other means during a query and input as additional information along with the input data and prompts.

[0016] Note that VLM2M may be provided in the work analysis system 1 instead of the VLM server 2.

[0017] The processor 11 includes a task identification unit 11a and a response acquisition unit 11b, which are realized by the execution of a program.

[0018] Sensor 15 is a device that acquires sensor data 15D representing the actions or postures of worker h related to their work, such as a camera, sensor, and microphone, at predetermined intervals. The camera acquires video or still image data of worker h's work. The sensor is, for example, attached to worker h and acquires physical quantities representing worker h's actions or posture. The microphone acquires audio data of worker h during their work.

[0019] The input / output unit 16 includes an input device such as a keyboard or mouse, and a display (image output device) and / or a speaker (audio output device). If the input / output unit 16 is a speaker, the information described later as being displayed on the display screen of the input / output unit 16 is output as audio from the speaker.

[0020] The work identification unit 11a identifies the type of work performed by worker h, based on the sensor data 15D acquired by the sensor 15. The sensor data 15D includes still and moving images taken from the perspective of worker h or a third party, data representing the posture and movements of worker h, and acceleration and pressure of equipment operated by worker h. The sensor data 15D can be image data, sensor data, or acoustic data.

[0021] The task identification unit 11a prepares, for example, a pattern of sensor data 15D that characterizes each task for each task, and when the sensor data 15D matches the pattern of the task to be fed back, it identifies the task type associated with this pattern. The task identification unit 11a stores the sensor data 15D corresponding to the period for which the task type has been identified in a predetermined storage area, associating it with the identified task type and the corresponding period.

[0022] When the type of work in the sensor data 15D is identified by the work identification unit 11a, the response acquisition unit 11b reads the prompt 12a from the storage unit 12 and modifies the prompt 12a or adds information to the prompt 12a as necessary. The response acquisition unit 11b then sends the modified or added prompt 12a, along with the sensor data 15D corresponding to the identified work type and period, to the VLM server 2. The VLM server 2 uses the prompt 12a and sensor data 15D received from the work analysis system 1 as input data to the VLM 2M and outputs analysis results based on the input data. The response acquisition unit 11b acquires the analysis results generated by the VLM 2M based on the prompt 12a and sensor data 15D.

[0023] The memory unit 12 is a non-volatile semiconductor memory device, etc., and stores the prompt 12a and the operation parameters 12b. The operation parameters 12b are parameters that define the operation when the work specification unit 11a operates.

[0024] Memory 13 is a volatile semiconductor memory device, etc. Memory 13 stores loaded program data and executes the program in cooperation with the processor 11. Communication I / F 14 is a communication device for the work analysis system 1 to communicate with the outside world.

[0025] (Prompt 12a according to Embodiment 1) Figure 2 shows a prompt 12a according to Embodiment 1. The prompt 12a has sections for instruction 12a1, question 12a2, and answer 12a3.

[0026] Instruction 12a1 includes the designation of a skills transfer task. The skills transfer task allows the VLM2M to identify areas for improvement in the work to enhance worker h's skills, based on sensor data 15D representing the work of worker h that has been entered.

[0027] Question 12a2 is a section that requests the VLM2M to provide an answer based on the sensor data 15D sent to the VLM server 2 along with prompt 12a, and the reason for this answer. Answer 12a3 is a section for inputting the answer output by the VLM2M based on prompt 12a and sensor data 15D. The VLM2M on the VLM server 2 takes prompt 12a and sensor data 15D as input and generates an answer based on sensor data 15D in the format specified by prompt 12a.

[0028] The data to be analyzed can be in any format. For example, the data to be analyzed is not limited to the waveform of the graph; it can also be data obtained by transforming the original data of the graph (for example, a spectrogram in the case of audio), or the numerical sequence of the original data of the graph itself.

[0029] (Work analysis process according to Embodiment 1) Figure 3 is a flowchart showing the work analysis process according to Embodiment 1. The work analysis process is executed at predetermined intervals by the work analysis system 1 when work analysis is instructed. The work analysis process may be executed in real time when the worker h performs work and sensor data 15D is acquired, or it may be executed while sequentially reading the accumulated sensor data 15D.

[0030] First, in step S11, the work identification unit 11a acquires sensor data 15D from the sensor 15. Next, in step S12, the work identification unit 11a identifies the type of work recorded in the sensor data 15D. Then, in step S13, the work identification unit 11a determines whether the type of work identified in step S12 corresponds to the target work. If it corresponds to the target work (step S13 YES), the work identification unit 11a moves the process to step S14. On the other hand, if it does not correspond to the target work (step S13 NO), the work identification unit 11a returns the process to step S11.

[0031] In step S14, the response acquisition unit 11b reads the prompt 12a from the storage unit 12. The response acquisition unit 11b then sends the sensor data 15D (or a graph or table based on the sensor data 15D) and the prompt 12a, which were determined to correspond to the target operation in step S13, to the VLM2M of the VLM server 2. The response acquisition unit 11b then receives the response generated by the VLM2M based on the sensor data 15D and the prompt 12a as feedback.

[0032] Next, in step S15, the response acquisition unit 11b outputs the feedback acquired in step S14 via the output screen of the input / output unit 16.

[0033] (Work analysis result display screen 16D1 according to Embodiment 1) Figure 4 is a diagram showing the work analysis result display screen 16D1 according to Embodiment 1. The work analysis result display screen 16D1 outputs the response generated by the VLM2M based on the input prompt 12a and sensor data 15D as feedback 16D11 to the work analysis result display screen 16D11. The feedback 16D11 may include a highlight display portion 16DH that visually highlights at least a part of the response. The highlight display portion 16DH allows the worker h to easily recognize the main points of the feedback 16D11.

[0034] This analysis (feedback) may include changes to worker h's work procedures, modifications to worker h's actions, changes to the tools worker h uses when performing the work, and adjustments to the work environment in which worker h performs the work.

[0035] (Effects of Embodiment 1) In Embodiment 1 described above, the work analysis system is generated by pre-training on multiple modalities, including images and text, as training data, and can be connected to a model that takes data recording the actions of workers related to their work as input and outputs analysis results based on the input. The work analysis system then inputs prompts to the model, along with the input data, instructing it to output analysis results that improve the work skills of workers when they perform their work based on the input data. The work analysis system then acquires the analysis results output by the model based on the input data and prompts, and outputs the acquired analysis results via an output unit. Thus, according to Embodiment 1, the skill transfer task can classify workers into skilled / unskilled, regress the workers' work proficiency, provide feedback to unskilled workers, and extract tacit knowledge.

[0036] Furthermore, in Embodiment 1 described above, the multiple types of data include at least one of image data, sensor data, and acoustic data. Therefore, according to Embodiment 1, the worker's work can be represented and analyzed using various types of data.

[0037] Furthermore, in Embodiment 1 described above, the model is trained using data from skilled workers with a certain level of work skill as training data. Therefore, according to Embodiment 1, few-shot prompting using in-context learning becomes possible based on a vast amount of prior training.

[0038] Furthermore, in Embodiment 1 described above, the work analysis system inputs data from skilled workers with a certain level of work skill as additional information, along with input data and prompts, into the model. Thus, according to Embodiment 1, in-context learning by RAG becomes possible based on extensive prior training.

[0039] Furthermore, in Embodiment 1 described above, the analysis results include at least one of the following: changes to the work procedure, modifications to the operation, changes to the tools used by the worker when performing the work, and adjustments to the work environment in which the work is performed. Therefore, according to Embodiment 1, measures necessary to improve the worker's work skills can be implemented from various perspectives.

[0040] Furthermore, in Embodiment 1 described above, the output unit of the work analysis system includes an image output device and / or an audio output device. Thus, according to Embodiment 1, workers can visually and / or audibly recognize the feedback necessary for improving their work skills while performing their work.

[0041] Furthermore, in the above-described embodiment 1, the work analysis system visually highlights a portion of the analysis results and outputs them via the output unit. Therefore, according to embodiment 1, the worker can easily identify important items in the feedback.

[0042] Furthermore, according to Embodiment 1 described above, the prompt includes specifying a skills transfer task to improve work skills and requesting the reason for the answer based on the input data. Thus, according to Embodiment 1, the worker who receives the feedback can efficiently acquire skills by knowing the reason "why the expert did it that way."

[0043] [Embodiment 2] (Configuration of Work Analysis System 1B According to Embodiment 2) FIG. 5 is a diagram showing the configuration of a work analysis system 1B according to Embodiment 2. The work analysis system 1B is different from the work analysis system 1 according to Embodiment 1 in that a plurality of types of sensors 15 are connected, the processor 11B includes a work identification unit 11Ba instead of the work identification unit 11a, and further includes an integrated analysis unit 11c.

[0044] The work identification unit 11Ba receives a plurality of types of sensor data 15D acquired by each of the plurality of types of sensors 15. The plurality of types of sensor data 15D are multi-dimensional signals or multi-modal data, and include at least one of image data, sensor data, and acoustic data. The work identification unit 11Ba identifies the work type of the work performed by the worker h captured by the plurality of types of sensors 15 based on the plurality of types of sensor data 15D.

[0045] The integrated analysis unit 11c integrates the plurality of types of sensor data 15D to generate a latent space. The latent space is obtained as an output of an integrated model generated by learning, with the plurality of spaces constructed by the plurality of types of sensor data 15D as input data. The latent space is a feature space composed of feature amount data with a lower dimension than the total dimension of the plurality of spaces constructed by the plurality of types of sensor data 15D. For example, the plurality of types of sensor data 15D are projected into a 10-dimensional latent space obtained by learning a 100-dimensional space of variable 1 and a 150-dimensional space of variable 2 among the plurality of types of sensor data 15D. Thereby, the dimension can be reduced from the dimension of the plurality of types of sensor data 15D to the dimension of the feature amount data (projection vector) of the latent space.

[0046] For example, when classifying the worker h into two values of skilled and unskilled, it is important to integrate the feature amounts into a latent space that can separate skilled and unskilled as clearly as possible. To discover an ideal latent space, it is necessary to consider non-linear data structures, multi-modal distributions, class dispersions, etc. in the feature space.

[0047] In the discovery of the latent space by binary classification, as the simplest method that does not consider any data structure, multimodal distribution, class dispersion, etc., there is a method of linearly combining all feature vectors and performing PCA: Principal Component Analysis.

[0048] Also, instead of binary classification of skilled and unskilled, multi-class classification / ordinal regression indicating "degree of skill" may be used. For the discovery of the latent space by multi-class classification / ordinal regression, there are autoencoders and generalized multivariate analysis. In an autoencoder, in the design of the loss function, non-linear data structures, multimodal distributions, class dispersions, etc. in the feature space are considered. In generalized multivariate analysis, in the design of the objective function of the optimization problem, non-linear data structures, multimodal distributions, class dispersions, etc. in the feature space are considered.

[0049] (Work analysis process according to Embodiment 2) FIG. 6 is a flowchart showing the work analysis process according to Embodiment 2. The work analysis process according to Embodiment 2 is different from the work analysis process according to Embodiment 1 in that step S13a is executed between step S13 and step S14, and the rest is the same as in Embodiment 1.

[0050] In step S13a, the integrated analysis unit 11c generates a latent space based on a plurality of types of sensor data 15D determined to correspond to the target work in step S13, and projects the plurality of types of sensor data 15D into the latent space to generate a projection vector. Note that the integration of the sensor data in step S13a may be performed immediately before step S12, and the work type in step S12 may be specified based on the integrated projection vector.

[0051] Next, in step S14, the response acquisition unit 11b reads the prompt 12a from the storage unit 12 and sends the projection vector (or graph or table based on the projection vector) and prompt 12a generated in step S13a to the VLM2M of the VLM server 2. The response acquisition unit 11b then receives the response generated by the VLM2M based on the projection vector and prompt 12a as feedback. The response acquisition unit 11b then outputs the response (analysis result) obtained from the VLM2M based on the projection vector and prompt 12a to the work analysis result display screen 16D1 of the input / output unit 16.

[0052] (Effects of Embodiment 2) In Embodiment 2 described above, the work analysis system receives multiple types of data recording the actions of a worker, integrates the multiple types of data to generate a latent space, and uses the feature data in the latent space corresponding to the multiple types of data as input data. Therefore, according to Embodiment 2, the work analysis system can reduce the dimensionality of the multimodal input data to a level where it can be visualized and used.

[0053] Furthermore, in Embodiment 2 described above, the latent space handled by the work analysis system is generated based on an autoencoder or generalized multivariate analysis. Therefore, according to Embodiment 2, the work analysis system can reduce the dimensionality of the input data of multimodal data using a known autoencoder or generalized multivariate analysis.

[0054] [Embodiment 3] (Configuration of work analysis system 1C according to Embodiment 3) Figure 7 shows the configuration of work analysis system 1C according to Embodiment 3. Work analysis system 1C differs from work analysis system 1B according to Embodiment 2 in that it is connected to an LLM (Large Language Model) server 2C in addition to the VLM server 2. The LLM server 2C includes LLM2L. LLM2L is a type of model that constitutes a generation AI capable of recognizing text and can generate answers to questions based on text.

[0055] Furthermore, the work analysis system 1C differs from the work analysis system 1B according to Embodiment 2 in that it includes a know-how acquisition unit 11d and a report acquisition unit 11e.

[0056] The response acquisition unit 11b outputs the response (analysis result) obtained from the VLM2M based on the prompt 12a and projection vector to the work analysis result display screen 16D1 of the input / output unit 16.

[0057] The know-how acquisition unit 11d generates a know-how question 16D21 (Figure 9) that asks for know-how to improve worker h's work skills and promote proficiency, based on the answer (analysis result) displayed on the work analysis result display screen 16D1 of the input / output unit 16. The know-how acquisition unit 11d then outputs the know-how question 16D21 to the know-how question screen 16D2 of the input / output unit 16. The know-how acquisition unit 11d accepts input of answers (know-how) to the know-how question 16D21 from worker h, other workers, or experts. The know-how acquisition unit 11d stores the input know-how in association with the answers (analysis results) in a predetermined recording area.

[0058] The report acquisition unit 11e generates a report output prompt 12Ca (Figure 10) that instructs the output of a report regarding the response (analysis result) obtained by the VLM2M from the response acquisition unit 11b. The report acquisition unit 11e inputs the generated report output prompt 12Ca to the VLM2M, LLM2L, or other model. The report acquisition unit 11e acquires the report generated and output by the VLM2M, LLM2L, or other model based on the report output prompt 12Ca, and outputs the acquired report via the report output screen 16D3 (Figure 11) of the input / output unit 16.

[0059] Furthermore, when the report acquisition unit 11e outputs a new report via the output unit, it searches the know-how acquisition unit 11d based on the answers (analysis results) related to the report. The report acquisition unit 11e then outputs the know-how that has been stored in association with the answers (analysis results) along with the answers (analysis results) in the report.

[0060] (Work analysis process according to Embodiment 3) Figure 8 is a flowchart showing the work analysis process according to Embodiment 3. The work analysis process according to Embodiment 3 differs from the work analysis process according to Embodiment 2 in that steps S16 and S17 are executed after step S15, and is otherwise the same as Embodiment 2.

[0061] In step S16, the know-how acquisition unit 11d acquires know-how. That is, the know-how acquisition unit 11d generates a know-how question 16D21 (Figure 9) based on the feedback output in step S15 (the answer (analysis result) displayed on the work analysis result display screen 16D1) and outputs it to the know-how question screen 16D2 of the input / output unit 16. The know-how acquisition unit 11d then accepts the input of the answer (know-how) to the know-how question 16D21 and stores it in a predetermined memory area. The answer (know-how) to the know-how question 16D21 acquired by the know-how acquisition unit 11d may be relearned by VLM2M.

[0062] Figure 9 shows the know-how question screen 16D2 according to Embodiment 3. The know-how question 16D21 asks whether there is any further know-how regarding the know-how about "screw tightening work" shown in feedback 16D11, which states, "When tightening screws, if you press the screwdriver a little harder while tightening, you will be less likely to strip the screw threads."

[0063] Next, in step S17, the report acquisition unit 11e reads the report output prompt 12Ca (Figure 10) from the storage unit 12 and inputs it to the VLM2M, LLM2L, or other model to instruct it to create a report. The report acquisition unit 11e then has the VLM2M, LLM2L, or other model generate and acquire a report regarding the response (analysis result) obtained by the response acquisition unit 11b in step S15, and outputs it. In step S17, the report acquisition unit 11e can also include the know-how acquired in step S16 in the report.

[0064] (Report output prompt 12Ca according to Embodiment 3) Figure 10 shows a report output prompt 12Ca according to Embodiment 3. The report output prompt 12Ca has the sections of instruction 12Ca1, question 12Ca2, and answer 12Ca3.

[0065] Instruction 12Ca1 includes the specification of a report generation task. The report generation task allows VLM2M, LLM2L, or other models to generate a report on the responses (analysis results) from VLM2M.

[0066] Question 12Ca2 is a section that contains the specific content of the report to be created for VLM2M, LLM2L, or other models. In the example in Figure 10, the instruction is, "Create a report listing the work processes that need improvement and the know-how for improvement for each of the workers A, B, and C."

[0067] (Report output screen 16D3 according to Embodiment 3) Figure 11 shows the report output screen 16D3 according to Embodiment 3. VLM2M, LLM2L, or other models generate a report 16D31 in response to question 12Ca2 (Figure 10) and output it to the report output screen 16D3 of the input / output unit 16. The report 16D31 has items 311 and 312 for each worker.

[0068] Item 311 is a “work process that needs improvement,” indicating a work process that the worker in question needs to improve. Item 312 is “know-how for improvement,” indicating know-how for improving the work skills in the work process that the worker in question needs improvement. Item 312 may also display know-how related to “screw tightening,” such as “When tightening screws, if you press the screwdriver a little harder while tightening, you are less likely to strip the screw threads,” as shown in feedback 16D11 on the work analysis result display screen 16D1 (Figure 4). Item 312 may also display know-how obtained by asking a know-how question 16D21 on the know-how question screen 16D2 (Figure 9).

[0069] (Modification of Embodiment 3) In Embodiment 3, in step S17 of the work analysis process (Figure 8), a report regarding the response (analysis result) from VLM2M is generated by VLM2M, LLM2L, or another model, acquired, and output. The report is not limited to a comparison list of the response (analysis result) from VLM2M, but may also include advice regarding worker h's future work based on the history of VLM2M's response (analysis result) regarding worker h's past work. In this modification, an example is described in which advice regarding worker h's future work is output based on the history of VLM2M's response (analysis result) regarding worker h's past work.

[0070] (Report output prompt 12Da according to a modified example of Embodiment 3) Figure 12 shows a report output prompt 12Da according to a modified example of Embodiment 3. The report output prompt 12Da has sections for instruction 12Da1, question 12Da2, and answer 12Da3.

[0071] Instruction 12Da1 includes the specification of a report generation task. The report generation task can cause VLM2M, LLM2L, or other models to generate a report (advice) regarding the response (analysis results) from VLM2M.

[0072] Question 12Da2 is a section containing the specific content of the report (advice) that you request to be created for VLM2M, LLM2L, or other models. In the example in Figure 12, the instruction is: "I will perform task X today. Based on my history of performing task X previously, please tell me what points I should be careful about when performing task X today."

[0073] (Report output screen 16D4 according to a modified example of Embodiment 3) Figure 13 shows the report output screen 16D4 according to a modified example of Embodiment 3. VLM2M, LLM2L, or other models generate a report 16D41 in response to question 12Da2 (Figure 12) and output it to the report output screen 16D4 of the input / output unit 16. Report 16D31 includes advice based on the history, such as, "...there was something that needed improvement in the 'screw tightening work' of work process X from three days ago. Specifically, in the 'screw tightening work,' ...you can reduce the risk of stripping the screw threads by pressing the screwdriver a little harder while tightening."

[0074] Report 16D31 may also include know-how obtained by asking questions via know-how question 16D21 on the know-how question screen 16D2 (Figure 9).

[0075] (Effects of Embodiment 3) According to Embodiment 3 described above, the work analysis system generates a second prompt that instructs the output of a report regarding the analysis results output by the model. The work analysis system then inputs the second prompt to the model or another model, retrieves the report output by the model or other model based on the second prompt, and outputs the retrieved report via the output unit. Thus, according to Embodiment 3, it is possible to compare the analysis results of multiple workers or to issue a warning based on the analysis results of a particular worker.

[0076] Furthermore, according to Embodiment 3 described above, the work analysis system generates questions about know-how to improve work skills along with the analysis results and outputs them via the output unit. The work analysis system then accepts input of know-how in response to the questions, accumulates the know-how in relation to the analysis results, and when outputting a new report via the output unit, includes the know-how related to the analysis results relevant to the report in the output. Thus, according to Embodiment 3, when providing feedback on the analysis results to the worker, it is possible to add know-how input by other skilled workers, etc., to the feedback, thereby promoting skill transfer.

[0077] Although several embodiments have been described above, these are merely illustrative examples for explaining the present invention and are not intended to limit the scope of the present invention to these embodiments only. The present invention can also be implemented in various other forms, such as forms in which some of the components of the above embodiments are omitted, forms in which at least some of the components are replaced, forms in which components are added, or forms that combine some or all of the embodiments.

[0078] 1, 1B, 1C: Work analysis system, 2: VLM server, 2M: VLM, 11, 11B: Processor, 12a: Prompt, 13: Memory, 15: Sensor, 15D: Sensor data, 16: Input / Output unit, h: Operator

Claims

1. A work analysis system that outputs analysis results for improving the work skills of a worker when performing a task, wherein the work analysis system is connectable to a model which is generated by pre-training a plurality of modalities including images and text as training data, and takes data recording the worker's actions related to the task as input and outputs the analysis results based on said input, wherein the processor of the work analysis system receives the worker's data as input data, identifies the type of task based on the input data, determines whether the task is a target task based on the type of task, inputs a prompt to the model, along with the input data, instructing the model to output the analysis results based on the input data if it is determined that the task is a target task, acquires the analysis results output by the model based on the input data and the prompt, and outputs the acquired analysis results via an output unit.

2. A work analysis system according to claim 1, wherein the processor receives a plurality of types of data recording the actions of the worker related to the work, integrates the plurality of types of data to generate a latent space, and uses feature data in the latent space corresponding to the plurality of types of data as the input data.

3. A work analysis system according to claim 2, wherein the plurality of types of data include at least one of image data, sensor data, and acoustic data.

4. A work analysis system according to claim 2, characterized in that the latent space is generated based on an autoencoder or generalized multivariate analysis.

5. A work analysis system according to claim 1, wherein the model is characterized in that it has learned the data of skilled workers with a certain level of work skill as the learning data.

6. A work analysis system according to claim 1, wherein the processor inputs the data of skilled workers whose work skills are above a certain level as additional information to the model along with the input data and the prompt.

7. A work analysis system according to claim 1, characterized in that the analysis results include at least one of a change in the procedure of the work, a modification of the operation, a change in the tools used by the worker when performing the work, and an adjustment of the work environment in which the work is performed.

8. A work analysis system according to claim 1, wherein the output unit includes an image output device and / or an audio output device.

9. A work analysis system according to claim 1, wherein the processor visually highlights a portion of the analysis results and outputs them via the output unit.

10. A work analysis system according to claim 1, wherein the processor inputs a second prompt to the model or another model instructing the model to output a report relating to the analysis results output by the model, acquires the report output by the model or the other model based on the second prompt, and outputs the acquired report via the output unit.

11. A work analysis system according to claim 10, wherein the processor generates questions regarding know-how for improving work skills along with the analysis results and outputs them via the output unit, accepts input of the know-how in response to the questions, stores the know-how in association with the analysis results, and when outputting a new report via the output unit, includes the know-how associated with the analysis results related to the report in the output.

12. A work analysis system according to claim 1, wherein the prompt includes specifying a skills transfer task to improve the work skills and requesting the reason for the answer based on the input data.

13. A work analysis method performed by a work analysis system that outputs analysis results for improving the work skills of workers when they perform a task, wherein the work analysis system is connectable to a model that is generated by pre-training a plurality of modalities including images and text as training data, and takes data recording the actions of a worker related to the task as input and outputs the analysis results based on said input, and the processor of the work analysis system is characterized by having the following processes: receiving the worker's data as input data; identifying the type of task based on the input data; determining whether the task is a target task based on the type of task; inputting a prompt to the model, along with the input data, to instruct the output of the analysis results if it is determined that the task is a target task; acquiring the analysis results output by the model based on the input data and the prompt; and outputting the acquired analysis results via an output unit.

14. A work analysis program for causing a computer to function as the work analysis system described in any one of claims 1 to 11.