Attention extraction system and attention extraction method

By acquiring and analyzing image and coordinate information within the demonstrator's field of vision, the fixation pattern is determined, and fixation images are extracted and stored. This solves the problem of difficulty in grasping fixation status in existing technologies, realizes the extraction and provision of appropriate attention information, and improves the work efficiency and quality of operators.

CN119317931BActive Publication Date: 2026-01-09INFORMATION SYST ENG INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202480002623.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-05-11
Filing Date
2024-04-19
Publication Date
2026-01-09
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

Existing technologies struggle to grasp the gaze state based on factors such as gaze duration and gaze shift, making it difficult to effectively extract and provide appropriate attention information to the operator.

Method used

By acquiring image information and viewpoint coordinates within the demonstrator's field of vision, the fixation pattern is determined, the fixation area is set, fixation images are extracted and stored, and the correctness of the task object is determined through a database, providing appropriate attention information.

Benefits of technology

It can accurately grasp the instructor's gaze state, provide appropriate attention information, and improve the operator's work efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119317931B_ABST
    Figure CN119317931B_ABST
Patent Text Reader

Abstract

The present application provides an attention extraction system and method for extracting attention information on a job. In the attention extraction system (100) for extracting attention information on a job, the attention extraction device (1) has: an acquisition unit that acquires image information and coordinate information in time sequence in association with a job; a determination unit that determines a gazing pattern of a demonstrator by calculating a viewpoint shift of the demonstrator within a field of view; an extraction unit that sets a gazing area according to the gazing pattern and extracts a gazing image of a job object gazed at by the demonstrator in the gazing area; and a storage unit that stores the extracted gazing image in a database as attention information in the job in association with the field of view, the viewpoint shift, and the gazing pattern.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to an attention extraction system and an attention extraction method that extract attention information for a work. BACKGROUND

[0002] In the past, as a technology related to attention extraction, for example, a skill transmission system of Patent Literature 1 and a wiring work assistance system disclosed in Patent Literature 2 have been proposed.

[0003] In the skill transmission system disclosed in Patent Literature 1, a technology of the following gist is disclosed: first image data and line-of-sight data of a field of view of a first user are machine-learned based on teacher data, feature data of the image is extracted and information of a focus position of the first user in the first image is input as the teacher data, the information of the focus position is registered in association with the feature data, second image data of a field of view of a second user is transmitted to a server, focus position data associated with the feature data corresponding to the second image is read out, assistance display data is created, and the assistance display data is transmitted to a glasses terminal of the second user, in the glasses terminal of the second user, display is performed so that the information of the focus position coincides with the field of view of the second user based on the assistance display data.

[0004] In addition, in the wiring work assistance system disclosed in Patent Literature 2, a line-of-sight tracking device that tracks a line of sight of a worker, an imaging device that images a wiring and the like of a work object, and image processing of a visual range of the worker are used to recognize a character and a color. A portable work management device is disclosed, which has an image processing section that guides information of the object stored in a data storage management section to the worker, and stores a work end signal from a limit switch provided to a work tool and an image of the work object.

[0005] PRIOR ART DOCUMENTS

[0006] PATENT LITERATURE

[0007] Patent Literature 1: Japanese Patent Application Publication No. 2017-191490

[0008] Patent Literature 2: Japanese Patent Application Publication No. 2015-061339 SUMMARY

[0009] PROBLEMS TO BE SOLVED BY THE INVENTION

[0010] Here, in Patent Literature 1, it is premised that the line of sight is cooperated among users and the skill is transmitted among the users and the like. That is, it is difficult to grasp the state of the gaze from the gaze time and gaze shift and the like of the viewpoint and to extract appropriate attention information to be provided to the worker. Moreover, there is no description or suggestion regarding determination of the attention action that indicates in what state the demonstrator gazes at what.

[0011] In addition, in Patent Literature 2, it is premised that the line of sight of the worker is associated with the work object and the information of the object is guided to the worker. That is, it is difficult to grasp the state of the gaze and extract appropriate attention information to provide to the worker according to the gaze time and gaze shift of the viewpoint and the like. Moreover, there is no description or suggestion regarding determination of the attention action indicating in what state the demonstrator gazes at what.

[0012] Therefore, the present application is proposed in view of the above-described problems, and an object thereof is to provide an attention extraction system and an attention extraction method capable of grasping the state of the gaze of the demonstrator and achieving extraction and provision of appropriate attention information.

[0013] Means for solving the problem

[0014] The attention extraction system of the first application extracts attention information for a work, characterized in that the attention extraction system has: an acquisition unit that acquires, in time series, image information within a field of view of a demonstrator who performs the work and coordinate information indicating a viewpoint at which the demonstrator gazes within the field of view in association with the work; a determination unit that determines a gaze pattern of the demonstrator based on the image information and the coordinate information acquired by the acquisition unit, and calculates a viewpoint shift of the demonstrator within the field of view; an extraction unit that sets a gaze region based on the gaze pattern determined by the determination unit, and extracts a gaze image of a work object at which the demonstrator gazes in the gaze region; and a storage unit that stores the gaze image extracted by the extraction unit in a database in association with the field of view, the viewpoint shift, and the gaze pattern as attention information in the work.

[0015] The attention extraction system of the second application is characterized in that, in the first application, the acquisition unit further has a determination unit that determines correctness of the work object included in the gaze image, and the attention extraction system further has a display unit that displays a determination result of the determination unit.

[0016] The attention extraction system of the third application is characterized in that, in the second application, there is further a database that stores a correlation between past gaze image information acquired in advance and reference information indicating correctness of the work object associated with the gaze image, the determination unit refers to the database to determine the correctness of the work object, and acquires corresponding information corresponding to the result of the determination from the database, and the display unit further outputs the corresponding information acquired by the determination unit.

[0017] The attention extraction system of the fourth invention is characterized in that, in the second invention, an input unit that inputs the correspondence information output by the determination unit is further provided, and the storage unit associates the correspondence information input by the input unit with the gaze image and stores it as an attention dataset that contains a set condition for identifying an object of gaze.

[0018] The attention extraction system of the fifth invention is characterized in that, in the second invention, the image information acquired by the acquisition unit contains recording date and time information, recording position information, and recording control information related to the acquisition operation of the image information, and the display unit displays the recording date and time information, the recording position information, and the recording control information in the center of the field of view range before the acquisition unit acquires the image information, and in the acquisition of the image information, only the recording control information displayed in the corner of the field of view range is switched.

[0019] The attention extraction system of the sixth invention is characterized in that, in the second invention, the display unit further has an attention information display area that switches and displays an acquisition display area and an attention display area, wherein the acquisition display area displays, for a worker who performs the work, a category of the acquisition unit that acquires the image information, a category of the attention dataset stored in the database, and an instruction to select the start of the work, respectively, and the attention display area displays, after the selection, correspondence information corresponding to the gaze image of the worker and attention information based on the attention dataset, the attention information displayed in the attention display area contains boost information that is displayed at at least any of the start of the work, the middle of the work, or the end of the work according to the progress of the work of the worker, and information displayed in the attention information display area, including at least any of the correspondence information, the attention information, or the boost information, is displayed based on the result of the determination by the attention dataset and the determination unit and according to the gaze condition of the worker who performs the work.

[0020] The attention extraction method of the seventh application is characterized by causing a computer to execute the following steps: an acquisition step of acquiring, in association with a job, image information within a field of view of an instructor who performs the job and coordinate information indicating a point of gaze of the instructor within the field of view in a time series; a determination step of determining a mode of gaze of the instructor as a wide field of view mode in a case where the coordinate information indicating a time series of shifts of the point of gaze of the instructor within the field of view is gathered at a center of the field of view and determining the mode of gaze of the instructor as a vigilance mode in a case where the coordinate information indicating a time series of shifts of the point of gaze of the instructor within the field of view is dispersed to a position outside the center of the field of view, based on the image information and the coordinate information acquired by the acquisition step; an extraction step of setting a gaze region in accordance with the mode of gaze determined by the determination step and extracting a gaze image of a job object at which the instructor gazes in the gaze region; and a storage step of storing the gaze image extracted by the extraction step in association with the field of view, the shift of the point of gaze, and the mode of gaze as attention information in the job in a database.

[0021] Effects of the Invention

[0022] According to the first to seventh applications, the determination unit determines the mode of gaze of the instructor based on the image information and the coordinate information, based on the shift of the point of gaze of the instructor within the field of view. Therefore, in the extraction unit, the gaze region can be set in accordance with the determined mode of gaze, and the gaze image of the job object at which the instructor gazes in the gaze region can be extracted. Thus, the gaze image indicating what the instructor gazes at in what situation can be extracted based on the gaze time and the shift of the point of gaze, and the state of the gaze can be accurately grasped, and appropriate attention information can be provided.

[0023] In particular, according to the second application, the gaze region includes the wide field of view mode of the instructor and the vigilance mode. Therefore, the wide field of view mode in which the shift of the point of gaze of the instructor is gathered at the center of the field of view and the vigilance mode in which the shift of the point of gaze of the instructor is dispersed to a position outside the center of the field of view can be set. Thus, the gaze image indicating what the instructor gazes at in what situation can be extracted.

[0024] In particular, according to the third application, the acquisition unit further has a judgment unit. Therefore, the correctness of the job object included in the gaze image can be judged. Thus, the state of the gaze of the instructor and the job performer can be accurately grasped, and appropriate attention information can be provided.

[0025] In particular, according to the fourth application, the database stores a correlation between the past gazing image information acquired in advance and reference information indicating the correctness of the work object associated with the gazing image. Therefore, the database is referred to in the determination unit, the correctness of the work object is determined, and the corresponding information corresponding to the result of the determination can be output by the display unit. Thus, the state of the gazer's gaze can be accurately grasped, and appropriate attention information can be provided.

[0026] In particular, according to the fifth application, the input unit inputs the corresponding information output by the determination unit. Therefore, the storage unit can associate the corresponding information with the gazing image and store it as an attention data set including a set condition for identifying a gazing object. Thus, the state of the gaze can be accurately grasped, and appropriate attention information can be provided.

[0027] In particular, according to the sixth application, the display unit switches the information displayed in the field of view range before and during the acquisition of the image information. Therefore, before the image information is acquired, the record date and time information, the record position information, and the record control information of the work of the demonstrator are displayed in the center of the field of view range, and during the acquisition, only the record control information is displayed in the corner of the field of view range. Thus, the acquisition of the image information of the demonstrator who performs the work and the extraction of the gazing image can be performed without burdening the demonstrator.

[0028] In particular, according to the seventh application, the display unit switches and displays the acquisition display area and the attention display area. Therefore, the corresponding information, the attention information, or the boost information can be allocated and displayed according to the gazing condition of the worker who performs the work according to the attention data set and the determination result. Thus, the state of the gazer's gaze can be accurately grasped, and appropriate attention information can be provided.

[0029] According to the eighth application, the determination step determines the gazing pattern of the demonstrator based on the image information and the coordinate information by calculating the shift of the point of view of the demonstrator in the field of view range. Therefore, in the extraction step, the gazing area can be set according to the determined gazing pattern, and the gazing image of the work object on which the demonstrator gazes in the gazing area can be extracted. Thus, the gazing image indicating what the demonstrator gazes at in what condition can be extracted based on the gazing time and the gazing shift of the point of view, the state of the gaze can be accurately grasped, and appropriate attention information can be provided. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is a schematic diagram showing an example of the structure of the attention extraction system in the present embodiment.

[0031] Figure 2 is a schematic diagram showing an example of the attention extraction of the demonstrator and the work evaluation of the worker of the attention extraction system 100 in the present embodiment.

[0032] Figure 3 (a)~ Figure 3 (c) is a schematic diagram illustrating an example of the attention extraction method in this embodiment.

[0033] Figure 4 (a) is a schematic diagram illustrating an example of the structure of an attention extraction device. Figure 4 (b) is a schematic diagram illustrating an example of the function of the attention extraction device.

[0034] Figure 5 This is a schematic diagram illustrating an example of a database in this embodiment.

[0035] Figure 6 This is a schematic diagram illustrating a first variant of the database in this embodiment.

[0036] Figure 7 This is a schematic diagram illustrating an example of a data table stored in the database in this embodiment.

[0037] Figure 8 This is a schematic diagram illustrating an example of a flowchart of the attention extraction method in this embodiment.

[0038] Figure 9 (a)~ Figure 9 (e) is a schematic diagram illustrating an example of the attention extraction method in this embodiment.

[0039] Figure 10 (a) and Figure 10 (b) is a schematic diagram illustrating an example of the display of the operator's device in this embodiment.

[0040] Figure 11 This is a schematic diagram illustrating an example of the attention extraction device in this embodiment.

[0041] Figure 12 (a)~ Figure 12 (e) is a schematic diagram showing an example of the display of the operator terminal in this embodiment. Detailed Implementation

[0042] Hereinafter, an example of an attention extraction system and attention extraction method according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0043] (Note the extraction system 100 and the extraction method in this implementation.)

[0044] Reference Figure 1 and Figure 2 An example of the structure of the attention extraction system 100 in this embodiment will be described. Figure 1is a diagram showing an example of the structure of the attention extraction system 100 in the present embodiment, Figure 2 is a diagram showing an example of the attention extraction of the instructor and the work evaluation of the worker in the attention extraction system 100 in the present embodiment.

[0045] The attention extraction system 100 is configured to extract an object of work to which the attention of the instructor is paid using image information within the field of view of the instructor who is performing the work and coordinate information indicating a point of view at which the instructor gazes within the field of view. In the attention extraction system 100, the image information and the coordinate information can be acquired within the field of view of the instructor by various conditions with respect to various information including the image information, the coordinate information, and work information related to the work.

[0046] For example, as shown in Figure 1 , the attention extraction system 100 includes an attention extraction device 1, an instructor device 2, a worker device 3, and a server 4, and for example, the instructor device 2 and the worker device 3 can also be provided in plural in the work area 50. The attention extraction system 100 can also perform transmission / reception of various information with respect to the attention extraction device 1, the instructor device 2, the worker device 3, the server 4, and other user devices (not shown) as objects via a known communication network 5, for example.

[0047] The attention extraction system 100 acquires a point of view shift 2b of a field of view destination at which the instructor gazes included in the field of view 2a of the instructor via the instructor device 2 worn by the instructor, for example. Among the acquisition information acquired by the attention extraction system 100 from the instructor, in addition to the image information (image within the field of view of the work being performed), the coordinate information (point of view, position information, and the like), the work information (work date and time, work instruction, process information, and the like), and the instructor information (instructor ID, device ID, and the like), various information of the work performed by the instructor (assignment work, group work, and the like) can also be included.

[0048] For example, as shown in Figure 2 , the attention extraction system 100 acquires the image information within the field of view of the instructor who is performing the work and the coordinate information indicating the point of view at which the instructor gazes within the field of view in time series in association with the work performed by the instructor from the instructor device 2 worn by the instructor. With respect to the acquisition of the image information and the coordinate information, a technology of acquiring a point of view possessed by a known eye tracking technology, a head-mounted display, or smart glasses, and the like can also be used, for example.

[0049] The attention extraction system 100 acquires the image information and the coordinate information by the above-described technique possessed by the instructor device 2, for example, determines the gaze pattern of the instructor in the attention extraction device 1, and extracts the gaze image according to the determined gaze pattern. The attention extraction system 100, for example, associates the extracted gaze image of the instructor with the reference information associated with the gaze image, and stores the same in the database as an attention data set. The attention extraction system 100, for example, can also refer to the database, acquire the attention information associated with the reference information, and display the same on the operator device 3 of the operator who uses the attention data set.

[0050] After that, the attention extraction system 100, for example, evaluates the work of the operator according to the image information of the work object 6 contained in the field of view range 3a of the operator, the numerical information and the coordinate information of the viewpoint shift of the field of view destination, and the like, according to the attention data set selected by the operator, and according to the evaluation result, for example, determines the work object 6 as the gaze information, further acquires the corresponding information such as "(1) confirmation", "(2) adjustment", and the like, and displays the same superimposed on the field of view range 3a of the operator device 3 together with the actual display. The details of each structure of the attention extraction device 1, the instructor device 2, and the operator device 3 are described later.

[0051] Here, with reference to Figure 3 , the gaze pattern of the instructor is described. The attention extraction system 100, for example, determines the gaze pattern of the instructor by the attention extraction method shown in (a) to Figure 3 (c) described above. For example, as shown in (a) of Figure 3 , the attention extraction system 100 determines the gaze pattern in the field of view range of the instructor according to the image information in the field of view range 2a acquired by the instructor device 2 and the coordinate information (x-axis, y-axis) of the viewpoint shift of the instructor. Figure 3 The attention extraction system 100 determines the gaze pattern of the instructor according to the viewpoint shift of the instructor in the field of view range 2a calculated from the coordinate information according to the image information acquired via the instructor device 2. For example, as shown in (b) of

[0052] , in the case where the characteristic of the gaze shift of the instructor is "narrow shift range (concentration)", "few attention points at the viewpoint", the attention extraction system 100 can determine it as "wide field of view pattern". In addition, for example, as shown in (c) of Figure 3 , in the case where the characteristic of the gaze shift of the instructor is "wide shift range (dispersion)", "many attention points at the viewpoint", the attention extraction system 100 can determine it as "alert pattern". Figure 3

[0053] ​Note extraction system 100 sets a gazing region in which the instructor gazes, extracts a gazing image of a work object on which the instructor gazes in the set gazing region, associates the extracted gazing image with the field of view range, the point of view shift, and the gazing pattern, and stores the gazing image in a database of server 4 or the like as attention information in the work of the instructor. Further, the instructor and the worker can each be plural, for example, the same work process can be worked by a plurality of instructors and workers, and note extraction system 100 can, for example, distribute one work process of the instructor in a manner in which a plurality of workers share the work.

[0054] The distribution of the instructor and the worker realized by note extraction system 100 can be, for example, distributed by an evaluator via note extraction device 1 in addition to being decided in accordance with, for example, the number of processes and the degree of difficulty of the work process, the skill of the worker, the work delivery period, and the like, and it is arbitrary to which work to distribute how many workers and how to arrange the workers, and can be appropriately determined.

[0055] <Note extraction device 1>

[0056] Note extraction device 1, for example, determines the gazing pattern of the instructor from the point of view shift of the instructor in the field of view range of the work performed by the instructor acquired by instructor device 2 based on the image information in the field of view range and the coordinate information indicating the point of view on which the instructor gazes in the field of view range.

[0057] The gazing region determined by the determination unit includes, for example, a wide field of view pattern in which the point of view shift of the instructor is included in a field of view radius of 5 to 20 degrees and is concentrated in the center of the field of view range, and a vigilance pattern in which the point of view shift of the instructor is included in a field of view radius of 4 degrees or less and is dispersed to a position outside the center of the field of view range.

[0058] For example, as shown in (b) of the above-described Figure 3 Note extraction device 1 determines a case in which the shift range of the line of sight of the instructor is "narrow (concentrated)" and the number of attention points of the point of view is "few" as a "wide field of view pattern", and on the other hand, for example, as shown in (c) of the above-described Figure 3 Note extraction device 1 determines a case in which the shift range of the line of sight of the instructor is "wide (dispersed)" and the number of attention points of the point of view is "many" as a "vigilance pattern".

[0059] Note extraction device 1, for example, sets a gazing region in accordance with the determined gazing pattern, and extracts a gazing image of a work object on which the instructor gazes in the gazing region. Note extraction device 1, for example, associates the extracted gazing image with the field of view range, the point of view shift, and the gazing pattern, and stores the gazing image in a database as attention information in the work.

[0060] Here, Figure 11One example of the attention extraction device screen la in which the attention extraction device 1 in the attention extraction system 100 is shown is illustrated in FIG. 1. Figure 11 For example, a screen operated by an evaluator, various information acquired by the instructor device 2 is displayed.

[0061] The attention extraction device 1 displays, for example, with reference to a database, on the attention extraction device screen la, a gaze monitoring area including at least: a gaze display area lb (a point of view movement distance graph, a point of view tracking, a gaze area, a gaze object, and the like) that displays coordinate information, a gaze area, a gaze image, and a gaze shift in a time series; an object setting area lc (a gaze object evaluation, a recognizer generation, and the like) that displays setting information for setting with respect to the gaze image displayed in the gaze display area lb; and a determination setting area Id (a corresponding information setting, and the like) that displays a determination condition of the setting information displayed in the object setting area lc.

[0062] The attention extraction device 1 receives, for example, via the setting item menu displayed on the attention extraction device screen la, conditions and values and the like related to various settings and adjustments from the evaluator. The attention extraction device 1 performs, for example, display, condition setting, and adjustment in the gaze display area lb, the object setting area lc, and the determination setting area Id, in accordance with the received various conditions and values.

[0063] The attention extraction device 1 determines, for example, via the setting item menu displayed on the attention extraction device screen la, correctness of a work object included in the gaze image of the instructor in accordance with various conditions received from the evaluator. The attention extraction device 1 can also display the determination result on the instructor device 2 or the worker device 3.

[0064] The attention extraction device 1 receives, for example, via the setting item menu displayed on the attention extraction device screen la, input of corresponding information from the evaluator. The attention extraction device 1 can also perform update and deletion of existing conditions and settings, corresponding information, and the like and store them in the database of the server 4, in addition to receiving input of new corresponding information.

[0065] The attention extraction device 1 can also, for example, associate the input corresponding information with the gaze image, generate an attention data set including a setting condition for identifying a gaze object, and store the generated attention data set in the database of the server 4.

[0066] Figure 4Fig. 1 is a schematic diagram showing an example of the structure of the attention extraction device 1. As the attention extraction device 1, in addition to a single board computer such as Raspberry Pi (registered trademark) or the like, a publicly known electronic device such as a personal computer (PC) or the like can be used. The attention extraction device 1 has, for example, a housing 10, a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a storage 104, and I / F's 105 to 107. The structures 101 to 107 are connected by an internal bus 110.

[0067] The CPU 101 controls the entire attention extraction device 1. The ROM 102 stores the action code of the CPU 101. The RAM 103 is a work area used by the CPU 101 at work. The storage 104 stores various information such as a learning model, a database, and the like. As the storage 104, in addition to an SD memory card, for example, a publicly known data storage medium such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or the like is used.

[0068] The I / F 105 is a publicly known interface for transmitting and receiving various information with the demonstrator device 2, the operator device 3, the server 4, the communication network 5, or the like connected according to the use. The I / F 105 can also be provided with a plurality of I / F's, for example.

[0069] The I / F 106 is a publicly known interface for transmitting and receiving various information with the input section 108 connected according to the use. As the input section 108, a keyboard is used, for example, and a manager or the like who performs management of the attention extraction system 100 or the like inputs or selects various information or a control command of the attention extraction device 1 or the like via the input section 108.

[0070] The I / F 107 is a publicly known interface for transmitting and receiving various information with the display section 109 connected according to the use. The display section 109 outputs various information stored in the storage 104 and a processing status of the attention extraction device 1 or the like. As the display section 109, a display is used, for example, and can be a touch panel type, for example. In this case, the display section 109 can include the input section 108.

[0071] In addition, as the I / F 105 to the I / F 107, the same interface can be used, for example, and a plurality of interfaces can be used, for example, as the respective I / F 105 to the I / F 107. In addition, at least any one of the demonstrator device 2, the operator device 3, the server 4, the communication network 5, the input portion 108, and the display portion 109 can be removed according to the situation.

[0072] <Storage portion 14 (database)>

[0073] The storage portion 14 stores various databases in the holding portion 104, for example. In the storage portion 14, the correlation between the past evaluation target information acquired in advance and the reference information associated with the past evaluation target information is stored, and a learning model having the correlation, for example, is stored. The storage portion 14 can be provided in the demonstrator device 2, the operator device 3, and the server 4, for example.

[0074] In the storage portion 14, the attention action (attention point, point of view, and the like) of the demonstrator at the time of acquisition of the work of the demonstrator is recorded, for example. The attention extraction system 100 can determine the gaze pattern from the information of the gaze image, the field of view range, and the point of view shift acquired via the demonstrator device 2, for example. The attention extraction system 100 can set a case where the characteristics of the eye movement of the demonstrator are “narrow shift range (concentration)” and “few attention points of the point of view” as the “wide field of view mode”, and on the other hand, for example, as described above, a case where the characteristics of the eye movement of the demonstrator are “wide shift range (dispersion)” and “many attention points of the point of view” are determined as the “alert mode”, and stored as the attention information in the work of the demonstrator in the storage portion 14 of the server 4 or the like. Figure 3

[0075] The storage portion 14 stores information related to the work of the demonstrator, which is acquired by gestures, voices, eye movements, and other input devices, for example, through various interfaces (not shown) of the input system possessed by the demonstrator device 2. The attention extraction system 100 stores various information acquired via the input device in the storage portion 14, for example, by respectively associating the information with information about the demonstrator and information about the work of the demonstrator. The attention extraction system 100 can appropriately record various information according to the recording instruction of the recording start, the recording end, and the like by the demonstrator via the demonstrator device 2, for example.

[0076] ​In the storage section 14, for example, the coordinate positions of the viewpoint settings supplemented by the instructor, the changes in the coordinate positions, the change speeds, and the like are sequentially recorded via the instructor device 2, and in addition thereto, for example, the position and orientation of the head in the work space coordinates can be recorded for each spatial layout of the place where each work is performed. Note that the attention extraction system 100 can acquire various information recorded for each spatial layout by, for example, a known 360° camera, a surveillance camera, or a network camera, and store the information in association with information indicating the coordinate positions of the head of the instructor, the changes in the coordinate positions, and the like in the storage section 14.

[0077] In addition, the storage section 14 can record various data recorded in association with the photographed data and the photographed data as an attention data set in the work of the instructor after the recording to the database is ended. The attention data set is evaluated by an evaluator for the attention points or the techniques via the attention extraction device 1, for example, and registered (recorded) in the storage section 14 as the correspondence information in the monitoring of the attention object recognizer and the attention object.

[0078] Here, for example, the attention object recognizer refers to a device generated using the attention object image evaluated by the attention extraction system 100, and for example, can be a known image processing program such as pattern matching or an image recognition model. The image processing program or the image recognition model can be installed in the worker device 3 worn by the worker when performing the work, for example. Thereby, the worker can evaluate the attention information for the work based on the time-series image information and the coordinate information acquired via the worker device 3 when performing the work, for example.

[0079] In addition, the storage section 14 can further have a database in which the correlation between, for example, past gaze image information acquired in advance and reference information indicating the correctness of the work object associated with the gaze image is stored. The storage section 14 can be used when the correctness of the work object in the work of the instructor is referred to and determined by the determination section 15, for example. The determination section 15 can acquire the corresponding correspondence information according to the result of the determination, for example. The display section 16 can further output the corresponding correspondence information from the storage section 14 according to the content determined by the determination section 15, for example.

[0080] In addition, the attention object recognizer and the correspondence information in the monitoring of the attention object evaluated as the attention technique by the instructor device 2 and the determination section 15 can be registered in the storage section 14 as the attention data set of the work in the worker device 3.

[0081] In the storage section 14, for example, past evaluation target information acquired by the instructor device 2 and reference information can be stored. The correlation is, for example, constructed by using machine learning using a plurality of pieces of learning data, with the past evaluation target information and the reference information as one set of learning data. As a learning method, for example, deep learning such as a convolutional neural network is used.

[0082] In this case, for example, the correlation represents a degree of association between a plurality of pairs of information (a plurality of data included in the past evaluation target information, a plurality of data included in the reference information).

[0083] The correlation is appropriately updated in the process of machine learning. That is, the correlation represents, for example, a function optimized in accordance with the past image data A, the corresponding information A, and the reference information. Therefore, using the correlation constructed based on all results of evaluation of the evaluation target information in the past, an evaluation result of the evaluation target information is generated. Thus, even in a case where the evaluation target information has a configuration state combined with other image data A, corresponding information A, or the like, an optimal evaluation result can be generated.

[0084] In addition, the evaluation target information stored in the storage section 14 can be different from the past evaluation target information except in a case where it is the same or similar. Thus, the attention extraction system 100 can quantitatively generate an optimal evaluation result. Furthermore, the attention extraction system 100 can improve a generalization ability when machine learning is performed, and can achieve an improvement in evaluation accuracy for unknown evaluation target information.

[0085] Furthermore, the correlation can have, for example, a plurality of correlation degrees representing degrees of association between a plurality of data included in the past evaluation target information and a plurality of data included in the reference information. For example, in a case where a learning model is constructed by a neural network, the correlation degrees can correspond to weight variables.

[0086] The past evaluation target information represents information of the same kind as the above-described evaluation target information. The past evaluation target information includes, for example, a plurality of pieces of evaluation target information acquired when the image data A is evaluated in the past.

[0087] The reference information is associated with the past evaluation target information, and indicates information related to the image data A, the corresponding information A. The reference information, for example, can include various information related to the progress of the work process, the inspection list, the attention image, the priority of the display information, the combination of display and non-display, the numerical value and data, the distribution, and the like, in addition to indicating the work range, the work instruction, the work step, the content of the subsequent process or the associated work, the content of the shared work or the substitute work of the other worker who performs the work in the same area, the evaluation based on various work relationships and mutual influences (for example, "check", "adjust", "attention (attention!)", "determination", "correspondence", and the like) for the image data A. In addition, the specific content included in the reference information can be arbitrarily set.

[0088] For example, as shown in Figure 5 , the association can also indicate the degree of association between the past evaluation target information and the reference information. In this case, by using the association, it is possible to store the degree of association with respect to the relationship of each of the plurality of data included in the reference information ("Reference A" to "Reference C" in Figure 5 ) to the plurality of data included in the past evaluation target information ("image data A" to "image data C" in Figure 5 ). Therefore, for example, via the association, it is possible to associate one data included in the past evaluation target information with a plurality of data included in the reference information, and it is possible to achieve the generation of a multi-faceted evaluation result.

[0089] The association, for example, has a plurality of association degrees that respectively associate the plurality of data included in the past evaluation target information with the plurality of data included in the reference information. The association degree, for example, is indicated in a manner of three or more levels such as a percentage, ten levels, or five levels, and is indicated by a characteristic of a line (for example, thickness, and the like). For example, "image data A" included in the past evaluation target information shows an association degree AA "80%" between "Reference A" included in the reference information, and an association degree AB "55%" between "Reference B" included in the reference information. That is, the "association degree" indicates the degree of association between each data, and for example, the higher the association degree, the stronger the association of each data. In addition, when the association is constructed by the machine learning described above, it can also be set that the association has three or more levels of association degrees.

[0090] The past evaluation target information, for example, can also be as shown in Figure 6As illustrated, the past evaluation target information is divided into the past work target fixation image and the past work target correspondence information, and stored in the database. In this case, the correlation degree is calculated from the relationship between the combination of the image data of the past fixation image and the past correspondence information and the reference information. Further, the past evaluation target information can be stored in the database by dividing the past correspondence information, for example, in addition to the above information.

[0091] For example, the combination of "image data a" included in the past fixation image and "correspondence information a" included in the past correspondence information shows the correlation degree AAA "85%" with "reference A" and the correlation degree ABA "25%" with "reference B". In this case, the data can be stored in such a manner that the past object data and the past correspondence information are independent of each other. Therefore, when the evaluation result is generated, the precision can be improved and the range of options can be expanded.

[0092] The past evaluation target information can include, for example, synthetic data and similarity. The synthetic data is represented by the similarity of three or more levels between the past object data or the past correspondence information. The synthetic data can be stored in the database in the form of an image or a character string, for example, in addition to being stored in the form of a numerical value, a matrix, or a histogram.

[0093] Here, reference is made to Figure 7 An example of the database in the present embodiment will be described. In the database, in addition to various information related to, for example, an instructor (instructor worker), an evaluator, a worker, and the like who are users of the attention extraction system 100, information related to the contents and procedures of various works and a plurality of data sets for the workers to use are stored as attention data sets.

[0094] The attention data sets, for example, correspond each data table information, work of an instructor, a worker, and the like, evaluation of an evaluator, and the like to each data. The evaluation of the evaluator is input, for example, by the input unit 17. The attention data sets store, for example, at least "work information table", "work step table", "correspondence information table", "instructor work record table", "viewpoint record table", "viewpoint record data table", and "fixation object table".

[0095] < < Work Information Table > >

[0096] In the "work information table", for example, data for identifying a work performed by a worker is stored (saved). In the "work information table", for example, "work ID" and "work name" for identifying a work performed by an instructor or a worker are stored.

[0097] < < Work Step Table > >

[0098] In the "work step table", data relating to the steps of the work performed by the instructor or the worker is stored (saved), for example. In the "work step table", a "work step ID" identifying the step of the work is stored, for example, and in association with the "work step ID", a "work ID", a "work step name", and a "work order" are stored respectively. The "work step table" is associated with the "work information table" by the "work ID", for example.

[0099] <Correspondence information table>

[0100] In the "correspondence information table", data for identifying information corresponding to the work displayed by the display unit after the correctness of the work object is determined by the determination unit based on the image information and the coordinate information of the viewpoint acquired by the acquisition unit when the worker performs the work is stored (saved), for example.

[0101] In the "correspondence information table", a "correspondence information ID", a "work step ID", "preliminary", which indicates a prompt of a nudge-based urging or a prior confirmation information, and a "display type", which indicates a display type such as an "interruption type" displayed according to the work status starting with a missed detection, and the like, for identifying information corresponding to the work are stored, for example.

[0102] In the "display type", a trigger indicating a start timing, a trigger indicating a display time or a display end, and the like are set, for example, as in the animation effect timing setting. In the case where the display type is "interruption", for example, information indicating a determination timing for performing the interruption is stored as a "display / determination start condition". The determination start condition can also be a trigger for specifying the interruption determination, for example.

[0103] Also, in the "correspondence information table", a "determination reference" indicating how the instructor determines "observed" and "not observed" when shifting to "interruption determination" and a display content displayed as the correspondence information are stored as "correspondence information content".

[0104] In the "correspondence information content", the correspondence information produced at the work site can be stored by a mark or the like in addition to the designation of the file name, for example. In the "correspondence information content", a single-shot instruction such as a contact with a superior or a senior can be stored, for example, and can be generated by input through an editing screen.

[0105] Further, in the "correspondence information table", a "work result record display condition" indicating a timing and the like at which the content of the result of each work step is recorded, and a "work result record content" indicating the content of the work result of the report and the check list of the work for recording are stored, for example.

[0106] Note that the attention extraction system 100 can also refer to the "work result record display condition" and determine that the gaze object has become in an appropriate state by image comparison, for example, based on a signal input by a gesture or a voice. The attention extraction system 100 can also store various information and data, for example, as a condition at the time of storing a determination result, in the "work result record display condition".

[0107] In addition, the attention extraction system 100 can also, for example, set the result input of the content in the "work result record content" as necessary. The attention extraction system 100 can also, for example, discriminate the present result input and determine that one work step is completed based on the result of the discrimination. The "correspondence information table" is associated with the "work step table", for example, by the "work step ID".

[0108] < <Teaching work record table> >

[0109] In the "teaching work record table", various information related to the work performed by the teacher is stored (saved), for example. In the "teaching work record table", a "teaching work record ID" for identifying the recorded teaching work, a "teaching work date and time" indicating the date and time at which the teaching work was performed, a "teaching worker" indicating the teacher who performed the teaching work, a "teaching work place" indicating the place and area where the teaching work was performed, and a "work ID" are stored in association with each other, for example. The "teaching work record table" is associated with the "work information table", for example, by the "work ID".

[0110] < <Viewpoint record table> >

[0111] In the "viewpoint record table", information and data related to the viewpoint at which the work was performed by the teacher, which is acquired by the acquisition unit 11, is stored (saved), for example. In the "viewpoint record table", a "viewpoint record ID" for identifying the viewpoint record of the work by the teacher and data of coordinate information of the viewpoint based on eye tracking are stored as a "viewpoint record data file".

[0112] In the "viewpoint record table", stream data acquired in units of milliseconds can be stored, for example. Thus, the attention extraction system 100 can store, for example, actual file names of data such as JavaScript (registered trademark) object notation (JSON: JavaScript Object Notation) without directly saving the data to a DB.

[0113] In addition, in the "viewpoint recording table", for example, a "view range recording image file" in which data indicating image information within the range of the view of the instructor recorded using the external camera is stored. Further, in the "viewpoint recording table", a "teaching operation recording ID" for identifying the history of the teaching operation performed by the instructor is stored. The "viewpoint recording table" is associated with the "teaching operation recording table", for example, by the "teaching operation recording ID".

[0114] < < Viewpoint recording data table > >

[0115] In the "viewpoint recording data table", for example, detailed information of a plurality of pieces of coordinate information of the viewpoint acquired by the acquisition unit 11 when the instructor performs the operation is stored. In the "viewpoint recording data table", a "viewpoint identification data file" identifying a file in which viewpoint recording data is stored, and a "viewpoint recording elapsed time" indicating the time of the viewpoint identification recorded by the instructor are stored. The "viewpoint recording elapsed information" stores (retains) the elapsed time of the recording of the viewpoint as a record (record) with the record start of the operation performed by the instructor set to 00:00:00, for example. The "viewpoint recording elapsed information" can be recorded at intervals of a unit time (for example, 1 second intervals or the like) together with various position information, for example. The "viewpoint recording data table" is associated with the "teaching operation recording table", for example, by the "viewpoint identification data file".

[0116] In addition, in the "viewpoint recording data table", position information of the instructor performing the operation within the operation space is stored as "instructor performing the operation position information (X, Y, Z)". In the "instructor performing the operation position information (X, Y, Z)", for example, in addition to the three-dimensional coordinates of the position information of the head of the HMD worn by the instructor performing the operation, information such as "instructor performing the operation orientation information (rad)", "instructor performing the operation eye position information (x, y, z)", "instructor performing the operation eye angle information (rad)", and the like acquired by the acquisition unit 11 can be stored together. The "viewpoint recording data table" is associated with the "teaching operation recording table", for example, by the "viewpoint recording data file".

[0117] < < Gaze object table > >

[0118] In the "gaze object table", for example, various information related to the object on which the instructor gazes is stored (saved). A plurality of pieces of various information set by the evaluator are stored. In the "gaze object table", for example, a "gaze object ID" identifying the object on which the instructor gazes, a "view range recording image file", a "viewpoint recording data file", a "viewpoint recording elapsed time", and an "operation step ID" are stored.

[0119] In addition, in the "gaze object table", in addition to storing the "gaze-time visual field range image" as a still image cut out from the "visual field range recording image file" at the moment in time of the viewpoint recording, for example, the "gaze object position (x, y)" indicating the position related to the gaze object extracted by the extraction section 13, the "gaze object range (w, h)" indicating the range, and the "gaze object image" indicating the image of the gaze object can be stored as the attention information in the work in accordance with the gaze pattern determined by the determination section 12 from the image information and the coordinate information associated in the "viewpoint recording table".

[0120] With respect to the attention information, for example, the information acquired by the acquisition section 11 can be evaluated, set, and recorded by an evaluator after being generated by the determination section 12 and the extraction section 13. The "gaze object table" is associated with the "viewpoint recording table" by the "gaze object ID", associated with the "work step table" by the "work step ID", and associated with the "viewpoint recording data table" by the "viewpoint recording data file".

[0121] Figure 4 (b) is a schematic diagram indicating an example of the function of the attention extraction apparatus 1. The attention extraction apparatus 1 has, for example, the acquisition section 11, the determination section 12, the extraction section 13, the storage section 14, the determination section 15, the display section 16, the input section 17, and the monitoring display section 18. In addition, Figure 4 Each function shown in (b) of FIG. 1 is realized by the CPU 101 executing the program stored in the storage section 104 or the like with the RAM 103 as a work area.

[0122] <<Acquisition Section 11 (Acquisition Unit)>>

[0123] The acquisition section 11 acquires the image information within the visual field range of the demonstrator performing the work and the coordinate information indicating the viewpoint at which the demonstrator gazes within the visual field range in association with the work in time series. The acquisition section 11 is used, for example, when implementing the acquisition step S110 described later. The timing at which the acquisition section 11 acquires the image information and the coordinate information from the demonstrator apparatus 2 can be arbitrarily set. The acquisition section 11 stores the acquired image information and the coordinate information in the storage section 104 in association with the work performed by the demonstrator in time series, for example, via the storage section 14.

[0124] <<Determination Section 12 (Determination Unit)>>

[0125] The determination section 12 determines the gaze pattern of the demonstrator from the image information and the coordinate information acquired by the acquisition section 11 based on the displacement of the viewpoint of the demonstrator within the visual field range. The determination section 12 is used, for example, when implementing the determination step S120 described later. The determination section 12, for example, as described above, determines the gaze pattern of the demonstrator from the image information and the coordinate information acquired by the acquisition section 11 based on the displacement of the viewpoint of the demonstrator within the visual field range. Figure 3As shown, the gaze pattern within the field of view of the instructor is determined based on the image information in the field of view 2a acquired using the instructor device 2 and the coordinate information (x-axis, y-axis) of the shift of the viewpoint of the instructor.

[0126] Next, if there are features and tendencies in the information acquired by the acquisition section 11 that are different from the stored "wide field of view pattern" and "alert pattern", the determination section 12 can also determine this as a new pattern. The determination section 12 can arbitrarily set various patterns using the attention extraction device 1. The determination section 12, for example, via the storage section 14, saves a new pattern in the storage section 104 in a time series in association with the instructor and the work performed by the instructor in addition to changes to the acquired existing patterns. Furthermore, the determination section 12, for example, can also perform the setting of various patterns by an evaluator based on the information acquired using the acquisition section 11 and stored in the database.

[0127] <<Extraction section 13 (extraction unit)>>

[0128] The extraction section 13 sets the gaze area based on the gaze pattern determined by the determination section 12 and extracts the gaze image of the work object 6 gazed at by the instructor in the gaze area. The extraction section 13, for example, in the case where there are multiple gaze patterns determined by the determination section 12, can also discriminate the gaze image of the work object 6 contained in the set gaze area based on each gaze pattern and perform extraction in a time series.

[0129] The extraction section 13, for example, extracts the gaze image of the work object 6 contained in the set gaze area based on the gaze pattern, and in the case where there is no order, can also extract multiple gaze images in association with one work in addition to extracting in a time series in the order of the order in the work and confirmation of the instructor.

[0130] The extraction section 13, for example, can also output the extraction result to the storage section 14 and the attention extraction device 1 and the like via the communication network 5. The extraction section 13, for example, via the storage section 14, saves the extracted gaze image in the storage section 104 in association with the instructor, the work performed by the instructor, the gaze pattern, and the like. Furthermore, regarding the association based on the gaze image and the gaze pattern and the like of the extraction section 13, this can also be performed by an evaluator based on various information saved in the storage section 104.

[0131] <<Storage section 14 (storage unit, database)>>

[0132] The storage section 14 stores the attention information in the database in association with the field of view range, the viewpoint shift, and the attention mode of the attention image extracted by the extraction section 13 as the attention information in the work of the instructor. The storage section 14 stores or extracts various information in or from the storage section 104.

[0133] The storage section 14 stores the correlation between the plurality of past attention image information acquired in advance and the reference information indicating the correctness of the work object associated with the attention image. The storage section 14 stores or extracts various information, for example, in accordance with the processing contents of the acquisition section 11, the determination section 12, the extraction section 13, the determination section 15, the display section 16, the input section 17, and the monitoring display section 18.

[0134] <<Determination Section 15 (Determination Unit)>>

[0135] The determination section 15 determines the correctness of the work object included in the attention image of the work acquired via the worker's worker device 3 using the attention data set related to the work, for example. The determination section 15 refers to the database stored in the storage section 14 or the storage section 104, and determines the correctness of the work of the attention information (e.g., the work object 6, etc.) performed by the worker based on the image information acquired via the worker's worker device 3, with reference to the database.

[0136] The determination section 15 acquires various corresponding information from the database, for example, in accordance with the result of the determination. The determination section 15 acquires the corresponding information displayed in the field of view range 3a of the worker's worker device 3 (information related to the work that the worker should deal with such as "(1) Confirmation OK: 01234", "(2) Adjustment Step: XXX", etc.), for example. The determination section 15 transmits the acquired corresponding information to the display section 16 of the worker's worker device 3, for example. Figure 2

[0137] With regard to the determination section 15, for example, when the same work is worked by a plurality of workers in the same work area, the correctness determination of the work object of the plurality of workers can also be collectively determined. In this case, for example, the plurality of workers pre-set the attention data set for the common work in each worker device 3 in advance. The determination section 15 can also acquire a plurality of work spaces layouts for each work based on the relationship between the coordinate information indicating the position and orientation of the head of each worker and the coordinate information indicating the work space of the work, for example, in accordance with the coordinate information acquired by the acquisition section 11, and determine the work and the work object of each worker comprehensively or in units of groups.

[0138] ​Further, the determination section 15, for example, acquires information on whether the original operator is facing the correct direction with respect to the work object that the original operator should work on, whether there is another operator who can support in the case where the original operator is not facing the correct direction (in the case where the original operator is not paying attention), whether there is another operator who is facing the work object that the original operator should work on, and the like. The determination section 15, for example, can also transmit corresponding information to the operator device 3 of the other operator who is working together in the surroundings where the work object can be properly worked on, based on the acquired information and the information on the work procedure of each operator, even in the case where the original operator is not paying attention.

[0139] <Display section 16 (display unit)>

[0140] The display section 16 outputs various corresponding information acquired by the determination section 15, for example, to the operator device 3 of the operator who is working on the corresponding work, using the attention data set generated based on the work of the instructor. The display section 16 (display portion 109 of the operator device 3) displays various corresponding information transmitted from the determination section 15 in the field of view 3a of the operator device 3 worn by the operator, for example, superimposed on the actual video.

[0141] With regard to the display section 16, for example, in the case where the same work is shared by a plurality of operators in the same work area, the corresponding information acquired by the determination section 15 is appropriately displayed according to the work situation and the work position of each operator.

[0142] In addition, the display section 16, for example, displays the display content of the attention information extracted with respect to the work of the instructor in the display portion 109 of the operator device 3. Figure 9 The display content of the attention information extracted with respect to the work of the instructor is displayed in the display portion 109 of the operator device 3. Figure 9 is a schematic diagram showing an example of the attention extraction method in the present embodiment. The display section 16, for example, displays the recording date and time information, the recording position information, and the recording control information in the center within the field of view 2a, as shown in (a) of Figure 9 of the attention extraction method in the present embodiment. The display section 16, for example, can also switch to display only the recording control information (for example, "recording temporarily stopped," "recording end," and the like) in the corner of the field of view 2a in the acquisition of the video information by the acquisition section 11.

[0143] In addition, the display section 16, for example, can display the display content indicating each operation in the center within the field of view 2a of the instructor device 2, as shown in (c) to Figure 9 (e) of Figure 9 The display section 16, for example, can also display the display content indicating each operation in the center within the field of view 2a of the instructor device 2, as shown in (c) to

[0144] The display section 16 displays the display contents shown in (a) and (b) of FIG. 9 and (a) to (e) of FIG. 10 on the operator device 3. Figure 10 Figure 10 Figure 12 Figure 12 Figure 10 Figure 10 Figure 12 Figure 12

[0145] The display section 16 displays the determination result of the display determination section 15, for example, in the field of view range 3a of the operator device 3 worn by the operator or the display device held by the operator, and further displays the corresponding information corresponding to the determination result obtained by the display determination section 15 with reference to the database.

[0146] Figure 10 (a) of FIG. 11 is a display of the field of view range 3a in a state in which there is no corresponding information for the operator in the work, for example. In addition, (b) of FIG. 11 is a display of the field of view range 3a in a state in which there is corresponding information for the operator in the work, in which a point that should be gazed at as the corresponding information (correct gaze object of the field of view range 3a: O mark + "Attention!"), a step or associated information that the operator should confirm or refer to (center of the field of view range 3a: "Check-3-A", "Check", "Determination", "Correspondence", "Reference Destination", and the like), and the work status / inspection result of the operator himself or a cooperator (left side of the field of view range 3a: work site, work content, work progress, and the like) are displayed in the field of view range 3a. Figure 10 Further, the display section 16 can further have a notice information display area that switches and displays the acquisition display area and the notice display area, in which the acquisition display area displays the category of the acquisition image information of the acquisition section 11, the category of the attention dataset stored in the database, and an instruction to select the start of the work, for example, at the start of the work by the operator, and in the notice display area, the corresponding information corresponding to the gaze image of the operator and the notice information are displayed based on the attention dataset selected by the operator.

[0147]

[0148] ​​​​​​​​​Further, the display section 16 can appropriately display the attention information displayed in the field of view range 3a of the worker device 3, for example, information displayed at at least any one of the time of the start of work, the middle of work, or the end of work, as boost information according to the progress of work by the worker. The display section 16, for example, according to the attention data set and the determination result of the determination section 15, refers to the attention database, acquires corresponding information, attention information, or boost information, and the like, and allocates display of the acquired various information according to the gazing condition of the worker who is performing work.

[0149] As for the boost information displayed by the display section 16, for example, as an approach to display to the worker, it can be display such as boost default (unconsciously drive), boost push (naturally drive), boost that makes the worker aware but leaves the selection to the worker, boost annotation (intentionally drive), and boost reward (drive with reward), and the like, which are not perceived by the subject. As for the display of boost information, for example, it can also be associated with the work procedure and the progress of work, the personality, temperament, evaluation of the worker, the difficulty of work, the estimated end time, and the like, and appropriately extracted and displayed from the database according to the actual progress of work by the worker.

[0150] In addition, the display section 16, in addition to directly displaying the reference information listed as a display candidate of boost information, for example, as a result of determination by the determination section 15, can also randomly display a type that asks questions such as "Is it okay?", "Wait a moment?", and the like. In addition, the display section 16 can also randomly display in the center of the field of view range 3a of the worker device 3 or at a corresponding position of the gazing object 6a within the field of view range 3a.

[0151] In addition, the display section 16 can appropriately set and display, for example, mix the output method of display or randomly change the output method of display (vision, hearing, vibration feedback (vibration), approach of boost information (annotation), method of boost information (dialogue, directly ask, make think "?", suddenly change), according to the skill of the worker and the progress of work, and the like.

[0152] In addition, the display section 16, for example, can perform heuristics, normality (maintain the status quo), and display of reference information for giving change to prevent overlooking due to habituation and boredom, expectation of maintaining the status quo, as a bias associated with the displayed boost information. The display section 16, for example, as a prerequisite process for displaying boost information, can have a unit (not shown) that grasps the flow and steps of the entire work that the worker is performing work, and can also have a unit (not shown) that detects a deviation from the content that should be paid attention to.

[0153] Further, the display section 16 can display, for example, a speech or the like regarding the labor of the operator, at which timing what kind of assist information is displayed is arbitrary, in the case where the operation of the operator is ended. Thus, for example, the corresponding information displayed to the operator can be displayed variably, and overlooking of the reference information due to the habit and the boredom of the operator, expectation to maintain the status quo can be prevented.

[0154] Further, the display section 16 displays, for example, the display content shown in Figure 11 in the attention extraction apparatus 1. Figure 11 For example, the gaze monitoring area of the attention extraction apparatus 1 in the present embodiment is displayed as the attention extraction apparatus screen la. As described above, in the display section 16, for example, for the operation based on the instructor, the coordinate information, the gaze area, the gaze image, and the gaze shift are displayed in the time series in the gaze display area lb by the evaluator referring to the database via the attention extraction apparatus 1.

[0155] The display section 16 displays, for example, as shown in Figure 11 the setting information for setting the gaze image displayed in the gaze display area lb in the object setting area lc. Further, the display section 16 displays, for example, the determination condition of the setting information displayed in the object setting area lc in the determination setting area Id.

[0156] The display section 16 displays, for example, as shown in

[0157] <<Input section 17 (input unit)>>

[0158] The input section 17 inputs, for example, the corresponding information output by the determination section 15. The input section 17 receives, for example, via the attention extraction apparatus screen la displayed in the attention extraction apparatus 1, for example, in addition to the conditions and values related to various settings and adjustments from the evaluator, for example, via the setting item menu displayed in the attention extraction apparatus screen la. Thus, for example, the acquisition section 11, the determination section 12, the extraction section 13, the storage section 14, the determination section 15, the display section 16, and the monitoring display section 18 perform various processes in accordance with various conditions and values received by the input section 17.

[0159] The corresponding information input by the input section 17 is associated with the gaze image, and is stored in the database of the server 4 as an attention data set including the setting condition for identifying the object of the gaze.

[0160] <<Monitoring display section 18 (monitoring display unit)>>

[0161] The monitoring display section 18 displays, for example, various information for the evaluator to refer to, set the work based on the instructor's work, on the attention extraction device screen la as a gaze monitoring area. The monitoring display section 18 displays, for example, a gaze display area lb that displays the coordinate information stored in the database, the gaze area, the gaze image, and the gaze shift in chronological order, an object setting area lc that displays the setting information for setting the gaze image displayed in the gaze display area, and a determination setting area Id that displays the determination conditions of the setting information displayed in the object setting area.

[0162] The attention extraction device 1 acquires, for example, the image information related to the work via the instructor's instructor device 2. The work performed by the instructor can be preselected, for example, for each of various attention data sets set according to the work of the worker, and the gaze data set of the work of the instructor is generated by the attention data set corresponding to the selected work.

[0163] The attention data set extracted by the attention extraction device 1 is selected, for example, by the workers A to C, and the work of the workers A to C is evaluated by the selected attention data set. In addition, the evaluator can also monitor, for example, the work situation of the workers A to C using the attention data set via the attention extraction device 1. The attention extraction device 1 can acquire, for example, the work status of the workers A to C together with the work using the gaze data set, acquire the image and position information of the work object, for example, by the worker device 3 worn by each worker, and display on the display section 16 of the attention extraction device 1.

[0164] The evaluator can refer to, for example, the work situation of the workers A to C displayed on the attention extraction device screen la displayed on the attention extraction device 1, and in the case where a work different from the attention data set is generated or evaluated, for example, in the case where the worker C is unable to perform the original work (skips the work step and is located before other work objects), in the case where the worker B is able to confirm the position of the original work of the worker C, the work of the gaze information that the worker C should have gazed on is set to "ignore", and the work of the gaze information that the worker B should have gazed on is instructed in real time as "delegation".

[0165] <Instructor device 2>

[0166] The instructor device 2 is worn by, for example, an expert, a qualified person, or the like in addition to the instructor of the work, and acquires the image information observed by the eyes of the instructor through the work of the instructor. The instructor device 2 can be, for example, a publicly known eye tracking technology, a head-mounted display, smart glasses, or the like, and can acquire sound, surrounding sound information, temperature, humidity, position information, spatial information, and the like together with the image information.

[0167] The demonstrator device 2, for example, can be connected to the attention extraction device 1, other demonstrator devices 2, the worker device 3, and the server 4 in a state capable of data communication, and can also be built-in with the attention extraction device 1.

[0168] <Worker device 3>

[0169] The worker device 3 is worn by a person other than a worker who performs a work, and acquires image information observed by the worker's eyes through the worker's work. The worker device 3 can be, for example, a publicly known eye tracking technology, a head-mounted display, smart glasses, or the like, and can acquire sound, surrounding sound information, temperature, humidity, position information, spatial information, and the like together with the image information.

[0170] The worker device 3, for example, can be connected to the attention extraction device 1, the demonstrator device 2, other worker devices 3, and the server 4 in a state capable of data communication, and can also be built-in with the attention extraction device 1.

[0171] <Communication network 5>

[0172] The communication network 5, for example, indicates an Internet to which the attention extraction device 1, the demonstrator device 2, and the worker device 3 are connected via a communication circuit, and can be constituted by an optical fiber communication network. The communication network 5 can be realized by a publicly known communication network such as a wireless communication network in addition to a wired communication network.

[0173] <Learning model>

[0174] The learning model generates a database through machine learning. The learning model acquires a plurality of pairs of learning data each of which includes a learning target image including a gazed image captured using the demonstrator device 2 and reference information indicating correctness of the gazed image captured using the demonstrator device 2 for each work in the demonstrator device 2. The learning model generates a database in which a plurality of learning target images and a plurality of reference information are associated with each other through machine learning using the plurality of learning data.

[0175] Figure 8 is a schematic view indicating an example of an attention extraction method in the present embodiment. The attention extraction method has an acquisition step S110, a determination step S120, an extraction step S130, a storage step S140, and a determination step S150. In addition, the attention extraction method can be implemented using the attention extraction system 100.

[0176] <Acquisition step S110>

[0177] The acquisition step S110 acquires, for example, the image information and the coordinate information in association with the work in chronological order. The acquisition step S110 acquires, for example, using the instructor device 2 having a publicly known camera, a photographing device, or the like. Alternatively, the acquisition step S110 can also acquire, for example, the image information and the coordinate information in association with the work in chronological order from the worker device 3. The acquisition step S110 acquires, for example, the image information and the like using the equipment selected by the instructor and the worker.

[0178] <The determination step S120>

[0179] The determination step S120 determines the gazing pattern of the instructor based on the image information acquired in the acquisition step S110 and the coordinate information described above.

[0180] In the determination step S120, for example, in a case where the point of view shift of the instructor is included in the field of view radius of 5 degrees to 20 degrees of the instructor and has a feature of being concentrated on the center of the field of view range, it is determined to be "wide field of view mode". Alternatively, in the determination step S120, for example, in a case where the point of view shift of the instructor is included in the field of view radius of 4 degrees or less of the instructor and has a feature of being dispersed to a position farther outside than the center of the field of view range, it is determined to be "alert mode".

[0181] The determination step S120 determines the gazing pattern in the field of view range of the instructor to be "wide field of view mode" and "alert mode" based on the image information in the field of view range 2a acquired by the instructor device 2 and the coordinate information (x-axis, y-axis) of the point of view shift of the instructor. The determination step S120 determines the gazing pattern to be "wide field of view mode" and "alert mode" based on the coordinate information of the point of view shift of the instructor, but if other features and tendencies can be confirmed, the confirmed gazing pattern can be re-determined.

[0182] <The extraction step S130>

[0183] The extraction step S130 extracts, for example, a gazing image of a work object on which the instructor gazes, by setting a gazing region. The extraction step S130 sets, for example, a gazing region on which the instructor gazes, based on the gazing pattern (for example, "wide field of view mode", "alert mode") determined by the determination step S120. The extraction step S130 extracts, for example, a gazing image of a work object on which the instructor gazes, from the image information of the instructor in the set gazing region.

[0184] As for the extraction step S130, for example, if there are a plurality of gaze images in the image information of the gaze region of the instructor, the extraction can be performed in the order of the gazes or the time of the gazes. The extraction step S130 extracts a gaze image of the work object in the gaze region of the instructor's gaze from the image information according to the conditions set by the evaluator, for example. Thus, the state of the instructor's gaze can be grasped, and appropriate attention information can be extracted.

[0185] <the storage step S140>

[0186] The storage step S140, for example, associates the gaze image with the field of view range, the point of view shift, and the gaze pattern, and stores it as the attention information. The storage step S140, for example, associates the gaze image extracted by the extraction step S130 with the field of view range, the point of view shift, and the gaze pattern, and stores it as the attention information in the database as described above.

[0187] <the determination step S150>

[0188] The determination step S150 determines the correctness of the work object included in the gaze image acquired via the worker device 3 using the attention data set, for example. The determination step S150 determines the correctness of the work object based on the gaze image acquired by the worker device 3 using the acquisition step S110, for example. The determination step S150 can determine the correctness of the work object by the evaluator who is monitoring the work, for example, via the attention extraction device 1, in addition to the evaluation using the selected attention data set for each work.

[0189] The determination step S150 can determine the correctness of the work object performed by the worker based on the attention data set selected in advance in the worker device 3, for example, in the case where the gaze image is acquired by the worker device 3 using the acquisition step S110. The determination step S150 can determine the work of a plurality of workers for the same work region.

[0190] The determination step S150 can determine the work of a plurality of workers simultaneously using the same or common attention data set, for example. Thus, for example, for a work process that a certain worker has overlooked or skipped or a work process that differs in order, determination corresponding to the work position and the work situation of each worker can be performed.

[0191] The judgment step S150, for example, refers to various data tables stored in the database, such as the "Job Information Table," "Job Step Table," "Corresponding Information Table," "Teaching Job Record Table," "Viewpoint Record Table," "Viewpoint Record Data Table," and "Gaze Object Table," etc., to determine the correctness of the job object based on these data tables and gaze datasets. Based on the judgment result, the corresponding job and various corresponding information stored in the reference link are obtained. The corresponding information obtained through the judgment step S150 can, for example, be displayed on the corresponding teaching operator and the operator's device 3 of other operators.

[0192] Furthermore, the determination step S150 may involve, for example, adding or updating corresponding information from various data tables stored in the database via the input unit 17. The input unit 17 can appropriately input, for example, the work procedures, determination conditions, corresponding work, and reference links of the determined object and the determination object stored in various data tables.

[0193] Therefore, the operation of the attention extraction device 1 in this embodiment ends.

[0194] Furthermore, according to this embodiment, for example, attention is paid to the fact that the image extraction device 1 acquires image information via the acquisition unit 11 of the teacher device 2. The image information may include, for example, recording date and time information related to the teacher's work, recording location information, and recording control information related to the image information acquisition operation.

[0195] Furthermore, according to this embodiment, for example, the display unit 16... Figure 9 The display on the teacher device 2 is switched as shown. For example, before the acquisition unit 11 acquires image information, the display unit 16 displays the recording date and time information, recording location information, and recording control information in the center of the field of view. Furthermore, during image information acquisition, the acquisition unit 11 only switches to displaying the recording control information in a corner of the field of view. This allows for a display that does not obstruct the teacher's work.

[0196] Furthermore, according to this embodiment, the display unit 16 displays attention information. The displayed information, for example, is based on information such as the worker's work progress and status, and is displayed at the start, middle, or end of the work, using a boosting theory. This eliminates the tendency to interfere with previous interactions with computer systems or information, and by displaying boosting information based on the boosting theory, the effectiveness of notifications from the attention extraction device 1 can be improved.

[0197] Embodiments of the present application are described, but the above-described embodiments are presented as examples and are not intended to limit the scope of the application. The above-described novel embodiments can be implemented in other various ways, and various omissions, substitutions, and changes can be made without departing from the scope of the application. The above-described embodiments and modifications thereof are included in the scope and spirit of the application, and are included in the scope of the application and equivalents thereof recited in the claims.

[0198] Label Explanation

[0199] 1: attention extraction device; 1a: attention extraction device screen; 1b: gaze display area; 1c: object setting area; 1d: determination setting area; 1e: work space diagram; 1f: work agent mark; 2: instructor device; 2a: field of view range (instructor); 2b: gaze pattern; 3: worker device; 3a: field of view range (worker); 4: server; 5: communication network; 6: work object; 6a: gaze object; 10: housing; 11: acquisition unit; 12: determination unit; 13: extraction unit; 14: storage unit (database); 15: determination unit; 16: display unit; 17: input unit; 18: monitoring display unit; 50: work area; 100: attention extraction system; S110: acquisition step; S120: determination step; S130: extraction step; S140: storage step; S150: determination step.

Claims

1. An attention extraction system that extracts attention information for a work, characterized by comprising: an acquisition unit that acquires, in association with the work, image information within a field of view of a demonstrator who performs the work and coordinate information that indicates a point of gaze of the demonstrator within the field of view, in a time series; a determination unit that determines a mode of gaze of the demonstrator as a wide field of view mode in a case where the coordinate information that indicates a time series of shifts in the point of gaze of the demonstrator within the field of view is gathered around a center of the field of view, and determines the mode of gaze of the demonstrator as a vigilant mode in a case where the coordinate information that indicates a time series of shifts in the point of gaze of the demonstrator within the field of view is dispersed to a position outside the center of the field of view, based on the image information and the coordinate information acquired by the acquisition unit; an extraction unit that sets a gaze region in accordance with the mode of gaze determined by the determination unit, and extracts a gaze image of a work object on which the demonstrator gazes in the gaze region; and a storage unit that stores the gaze image extracted by the extraction unit in association with the field of view, the shift in the point of gaze, and the mode of gaze in a database as attention information in the work.

2. The attention extraction system according to claim 1, characterized in that the acquisition unit further comprises a judgment unit that acquires a gaze image of a worker who performs a work, and judges a work object performed by the worker as correct or incorrect based on image information included in the gaze image and object information related to the work that the worker should cope with, which is stored in advance in a database, and the attention extraction system further comprises a display unit that displays a result of the judgment by the judgment unit.

3. The attention extraction system according to claim 2, characterized in that the attention extraction system further comprises a database that stores a correlation between past gaze image information acquired in advance and reference information that indicates correctness of the work object associated with the gaze image, the judgment unit refers to the database to judge the work object as correct or incorrect, and acquires corresponding information corresponding to a result of the judgment from the database, and the display unit further outputs the corresponding information acquired by the judgment unit.

4. The attention extraction system according to claim 2, characterized in that the attention extraction system further comprises an input unit that inputs the corresponding information output by the judgment unit, and the storage unit stores the corresponding information input by the input unit in association with the gaze image as an attention data set that includes a set condition for identifying a gaze object.

5. The attention extraction system according to claim 2, characterized in that the image information acquired by the acquisition unit includes recording date and time information of the work performed by the demonstrator, recording position information, and recording control information related to an acquisition operation of the image information. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ The display unit displays the recording date and time information, the recording position information, and the recording control information in the center of the field of view range before the acquisition unit acquires the image information, and in the acquisition of the image information, only the recording control information is displayed in a corner of the field of view range.

6. The attention extraction system according to claim 2, wherein the display unit further has an attention information display area that switches and displays an acquisition display area and an attention display area, wherein the acquisition display area displays, for a worker who performs the work, a category of the acquisition unit that acquires the image information, a category of an attention dataset that includes a set condition for identifying an attention object stored in the database, and an instruction to select a start of the work, respectively, and the attention display area displays, after the selection, corresponding information corresponding to an attention image of the worker based on the attention dataset and attention information, the attention information displayed in the attention display area includes boost information that is displayed at at least any of a work start, a work middle, or a work end, according to a work progress of the worker, information displayed in the attention information display area, including at least any of the corresponding information, the attention information, or the boost information, is allocated to be displayed based on a result of the determination by the determination unit and according to an attention state of the worker who performs the work.

7. An attention extraction method of extracting attention information for a work, comprising: the attention extraction method causing a computer to execute: an acquisition step of acquiring, in time series, image information in a field of view range of a demonstrator who performs the work and coordinate information indicating a point of view at which the demonstrator gazes in the field of view range in association with the work; a determination step of determining a gazing pattern of the demonstrator as a wide field of view pattern in a case where the coordinate information indicating a time series of a point of view shift of the demonstrator in the field of view range is gathered in a center of the field of view range, and determining the gazing pattern of the demonstrator as a vigilant pattern in a case where the coordinate information indicating a time series of a point of view shift of the demonstrator in the field of view range is dispersed to a position farther outside than the center of the field of view range, based on the image information acquired by the acquisition step and the coordinate information; an extraction step of setting a gazing area according to the gazing pattern determined by the determination step, and extracting a gazing image of a work object on which the demonstrator gazes in the gazing area; a storage step of storing the gazing image extracted by the extraction step in association with the field of view range, the point of view shift, and the gazing pattern as attention information in the work in a database. ​

Citation Information

Patent Citations

  • Wire connection work support system

    JP2015061339A

  • Skill transmission system and method

    JP2017191490A

  • Sight capturing method and man-machine interaction method adopting sight capturing

    CN102508551A

  • Alarm device

    CN110348281A