Information processing apparatus and information processing method
By acquiring dynamic images from the production site and using inference models to detect the operator's movement range, emotions, and gaze direction, the problem of the inability to monitor the detailed condition of operators in existing technologies has been solved, enabling detailed monitoring of operators and improving production efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies cannot effectively monitor the detailed condition of workers, such as their range of motion, emotions, and direction of gaze within the monitoring area corresponding to the work process.
By setting up cameras on the production site, dynamic images are captured, and inference models are used to detect the workers' movement range, emotions, and gaze direction, providing detailed detection results.
It enables detailed monitoring of workers' conditions, including accurate detection of their movement range, emotions, and gaze direction, supporting improved productivity and appropriate attention to workers.
Smart Images

Figure CN116486471B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to information processing apparatus and information processing methods. Background Technology
[0002] Previously, attempts have been made to improve work processes at workplaces by using dynamic images captured by cameras installed at work sites such as factories. For example, Japanese Patent Application Publication No. 2020-204819 discloses an information processing apparatus for image analysis of dynamic images captured by ceiling cameras. The information processing apparatus analyzes the dynamic images to determine whether a worker exists in the monitoring area corresponding to each process, and generates data indicating the time period in which a worker was determined to be present in the monitoring area.
[0003] In the technology disclosed in Japanese Patent Application Publication No. 2020-204819, although it is possible to monitor whether an operator exists in the monitoring area corresponding to each process, it is not possible to monitor the detailed condition of the operator. Summary of the Invention
[0004] This disclosure was made in view of the above-mentioned problems, and its purpose is to provide an information processing apparatus and information processing method capable of grasping the detailed situation of the operator.
[0005] According to one example of this disclosure, the information processing apparatus includes an acquisition unit, a motion range detection unit, and a providing unit. The acquisition unit acquires dynamic images from a camera installed at the production site that captures images of the target operator and the area surrounding the operator. The motion range detection unit uses an inference model to detect the motion range of the target operator's work reflected in the first frame of a predetermined number of consecutive frames contained in the dynamic images. The providing unit provides the detection results from the motion range detection unit. The inference model is generated through learning processing using multiple learning datasets. Each of the multiple learning datasets includes: a second frame of a predetermined number of consecutive frames contained in the dynamic images reflecting a specific operator; and labels representing the motion range of the specific operator's work reflected in the second frame of the predetermined number of frames.
[0006] According to one example of this disclosure, the information processing apparatus includes: an acquisition unit that acquires dynamic images from a camera installed at the production site and which captures the face of an operator; an emotion detection unit that detects the emotions of the operator reflected in each frame of the dynamic images; and a provision unit that provides the shifts in emotions detected by the emotion detection unit.
[0007] In the above disclosure, it is preferable that the emotion detection unit outputs scores for each of the multiple emotion categories. Preferably, the providing unit also provides a notification to prompt the operator's attention if the score for the object category among the multiple categories deviates from a predetermined range.
[0008] According to one example of this disclosure, the information processing apparatus includes: an acquisition unit that acquires dynamic images from a camera installed at a production site and which captures the face of a worker; a gaze detection unit that detects the direction of the worker's gaze in each frame of the dynamic image; and a providing unit that provides an image showing an object present in front of the worker. Based on the gaze direction detected by the gaze detection unit, the providing unit determines the position of the worker's viewpoint in the image and displays a mark at the determined position in the image.
[0009] According to one example of this disclosure, the information processing method includes the following steps: acquiring dynamic images from a camera set up at a production site that captures images of the operator and the surrounding area; using an inference model to detect the action range of the operator's work reflected in a predetermined number of consecutive first frames contained in the dynamic images; and providing the detection result. The inference model is generated through learning processing using multiple learning datasets, each containing: a predetermined number of consecutive second frames contained in the dynamic images reflecting a specific operator; and labels representing the action range of the specific operator's work reflected in the predetermined number of second frames.
[0010] According to one example of this disclosure, the information processing method includes the following steps: acquiring dynamic images from a camera installed at the production site and capturing the face of the operator; detecting the operator's emotions reflected in each frame of the dynamic images; and providing a progression of the detected emotions.
[0011] In the above disclosure, preferably, the detection step includes the following steps: outputting scores for each of the multiple categories of emotion. The provisioning step preferably includes the following steps: providing a notification to prompt the operator's attention when the score of the object category among the multiple categories deviates from a predetermined range.
[0012] According to one example of this disclosure, the information processing method includes the following steps: acquiring a dynamic image from a camera installed at the production site that captures the face of a worker; detecting the direction of the worker's gaze in each frame of the dynamic image; and providing an image reflecting an object present in front of the worker. The providing step includes the following steps: determining the position of the worker's viewpoint in the image based on the detected gaze direction; and displaying a mark at the determined position in the image.
[0013] Based on this information, users can gain a detailed understanding of the worker's situation (the range of actions they are performing, the direction of their gaze, and their emotions).
[0014] The above and other objects, features, aspects and advantages of the present invention will become clear from the following detailed description, which is understood in conjunction with the accompanying drawings and is relevant to the present invention. Attached Figure Description
[0015] Figure 1 This is a diagram showing the overall structure of a system using the information processing apparatus of this embodiment.
[0016] Figure 2 This is a schematic diagram illustrating an example of the hardware structure of an information processing device according to an embodiment.
[0017] Figure 3 This is a diagram illustrating an example of the functional structure of an information processing device in an embodiment.
[0018] Figure 4 This is a diagram representing an example of an inference model.
[0019] Figure 5 It is a diagram representing the three frames corresponding to the three action ranges corresponding to the process "welding", and the frame that does not belong to any action range.
[0020] Figure 6 This is a graph representing the verification results of the estimated action range.
[0021] Figure 7 This is an example of a picture that provides a view.
[0022] Figure 8 This is an illustration representing another example of providing a visual representation.
[0023] Figure 9 This is another example of a picture that provides a visual representation.
[0024] Figure 10 It is a graph showing the relationship between workers' emotions and production indicators. Detailed Implementation
[0025] Embodiments of the present invention will be described in detail with reference to the accompanying drawings. Furthermore, the same or corresponding parts in the drawings are labeled with the same reference numerals and their descriptions are not repeated. The various modifications described below can also be selectively combined as appropriate.
[0026] Figure 1 This is a diagram showing the overall structure of a system using the information processing apparatus of this embodiment. (e.g.) Figure 1 As shown, system 1 includes production line 2, information processing device 10, PLC (Programmable Logic Controller) 20, and cameras 30 and 40.
[0027] Production line 2 comprises multiple processes 3_1 to 3_n, producing various products. These processes 3_1 to 3_n include, for example, a "welding" process, a "substrate assembly" process, a "substrate-to-body assembly" process, and an "inspection" process. Various equipment can be installed in each process of the production line. This equipment includes robots, processing devices, inspection devices, and various sensors.
[0028] PLC 20 is a control device that controls the entire production line 2 and can be communicatively connected to the equipment installed on production line 2. Various industrial Ethernet networks can be used as the network for communicatively connecting PLC 20 to the equipment. Examples of industrial Ethernet networks include EtherCAT, Profinet IRT, MECHATROLINK-III, Powerlink, SERCOS-III, and CIP Motion. Furthermore, field networks other than industrial Ethernet networks can also be used. For example, if motion control is not required, DeviceNet, CompoNet / IP, etc., can also be used.
[0029] The PLC 20 acts as the master device in the master-slave control system, acquiring information from the equipment as input data. Following a pre-programmed user program, the PLC 20 performs calculations using the acquired input data. Based on the execution of these calculations, the PLC 20 determines the control content for the master-slave control system and outputs the corresponding control data to the equipment. The PLC 20 repeatedly acquires input data from the equipment and outputs control data to the equipment at a predetermined cycle (control cycle).
[0030] Cameras 30 and 40 are configured to film workers performing tasks on production line 2. Figure 1 In the example shown, cameras 30 and 40 are configured to capture images of worker 4 in process 3-3. Specifically, camera 30 is positioned to capture the face of worker 4 from the front. Camera 40 is positioned to capture images of worker 4 and the workbench in process 3-3. Cameras 30 and 40 output the captured dynamic image data (hereinafter referred to as "dynamic image") to the information processing device 10. Furthermore, cameras 30 and 40 can be installed not only in process 3-3 but also in processes other than process 3-3.
[0031] The information processing device 10, for example, is a general-purpose computer that analyzes the detailed condition of the operator 4 performing the work in process 3-3 based on the dynamic images obtained from cameras 30 and 40. Furthermore, when analyzing the condition of the operator 4, the information processing device 10 can also utilize input data obtained from PLC 20 and control data output from PLC 20.
[0032] <Hardware Structure of Information Processing Device>
[0033] Figure 2 This is a schematic diagram illustrating an example of the hardware structure of an information processing device according to an implementation method. For example... Figure 2 As shown, the information processing device 10 typically has a structure that follows a general computer architecture. Specifically, the information processing device 10 includes a processor 11 such as a CPU (Central Processing Unit) or MPU (Micro-Processing Unit), memory 12, storage 13, a display controller 14, an input interface 15, a camera interface 16, and a communication interface 17. These components are connected to each other via a bus in a manner that enables data communication.
[0034] The processor 11 performs various processes in this embodiment by expanding and executing various programs stored in the memory 13 in the memory 12.
[0035] Memory 12 is typically a volatile storage device such as DRAM, which stores programs and other data read from memory 13.
[0036] The memory 13 is typically a non-volatile magnetic storage device such as a hard disk drive. The memory 13 stores the model generation program 131, the action region detection program 134, the emotion detection program 135, the gaze detection program 136, and the provider program 137 executed by the processor 11. Furthermore, the memory 13 stores multiple learning datasets 132 utilized in the execution of the model generation program 131, and the inference model 133 generated through the execution of the model generation program 131. Various programs installed in the memory 13 are stored in a state that is also stored on a memory card or similar device.
[0037] The display controller 14 is connected to the display device 70 and outputs signals to the display device 70 for displaying various information according to the internal commands from the processor 11.
[0038] Input interface 15 mediates data transmission between processor 11 and input devices 75 such as keyboard, mouse, touch panel, and dedicated console. That is, input interface 15 accepts operation commands given by the user through the input device 75.
[0039] The camera interface 16 mediates data transmission between the processor 11 and the cameras 30 and 40. More specifically, the processor 11 outputs shooting instructions to the cameras 30 and 40 via the camera interface 16. The camera interface 16 outputs the moving images received from the cameras 30 and 40 according to the shooting instructions to the processor 11. The camera interface 16 operates as an acquisition unit for obtaining moving images from the cameras 30 and 40.
[0040] Communication interface 17 mediates data transmission between processor 11 and external devices (such as PLC 20). Communication interface 17 typically includes Ethernet (registered trademark), USB (Universal Serial Bus), etc. Furthermore, various programs stored in memory 13 can also be downloaded from a distribution server, etc., via communication interface 17.
[0041] In the case of using a computer with a structure conforming to a general computer architecture as described above, in addition to the applications used to provide the functions of this embodiment, an OS (Operating System) can also be installed to provide the basic functions of the computer. In this case, the program of this embodiment can also perform processing by calling necessary modules from the program modules provided as part of the OS in a prescribed order and at predetermined times. That is, the program of this embodiment itself may not sometimes contain the modules described above, but performs processing in cooperation with the OS.
[0042] Alternatively, some or all of the functions provided by the execution of the model generation program 131, the action range detection program 134, the emotion detection program 135, the gaze detection program 136, and the provider program 137 can be installed as dedicated hardware circuitry.
[0043] <Functional Structure of Information Processing Device>
[0044] Figure 3 This is a diagram illustrating an example of the functional structure of an information processing apparatus according to an implementation method. For example... Figure 3As shown, the information processing device 10 includes a storage unit 101, a model generation unit 102, an action range detection unit 103, an emotion detection unit 104, a gaze detection unit 105, and a supply unit 106. The storage unit 101 is implemented using a memory 12 and a storage unit 13. The model generation unit 102 is implemented by a processor 11 executing a model generation program 131. The action range detection unit 103 is implemented by a processor 11 executing an action range detection program 134. The emotion detection unit 104 is implemented by a processor 11 executing an emotion detection program 135. The gaze detection unit 105 is implemented by a processor 11 executing a gaze detection program 136. The supply unit 106 is implemented using a display controller 14, an input interface 15, and a processor 11 executing a supply program 137.
[0045] (Structures related to the detection function of the action range)
[0046] Each process involves multiple motion zones. For example, in the case of the "soldering" process, it includes: the motion zone for moving the substrate from the previous process and mounting the substrate onto the fixture; the motion zone for soldering the component onto the substrate; and the motion zone for removing the substrate from the fixture and moving the substrate to the next process.
[0047] The model generation unit 102 generates an inference model 133, which infers the action range of each frame of the dynamic image captured by the camera 40. The model generation unit 102 stores the generated inference model 133 in the storage unit 101.
[0048] The inference model 133 can be appropriately configured to perform computational processing using a prescribed algorithm, prescribed rules, or functional expression, which performs an inference task corresponding to the object data. The output of the inference model 133 can be appropriately configured to determine the result of the inference task performed. In one example of this embodiment, the inference model 133 is composed of a trained machine learning model generated through machine learning. The machine learning model has parameters that can be adjusted through machine learning. The structure and type of the machine learning model can be appropriately selected according to the embodiment.
[0049] Figure 4 This is a diagram representing an example of an inference model. Figure 4 An inference model 133 composed of a neural network is shown.
[0050] like Figure 4As shown, the inference model 133 has an input layer 51, one or more intermediate (hidden) layers 52, and an output layer 53. The number of intermediate layers 52 can be appropriately determined according to the implementation method. Intermediate layers 52 can also be omitted. The number of layers in the neural network constituting the inference model 133 can be appropriately determined according to the implementation method. The input layer 51 can be appropriately configured to receive object data. The output layer 53 can be appropriately configured to output a value corresponding to the inference result. The input layer 51 can also be configured to receive information other than object data, and the output layer 53 can also be configured to output information other than the information corresponding to the inference result.
[0051] The input layer 51, intermediate layer 52, and output layer 53 each have one or more nodes (neurons). The number of nodes in each of the input layer 51, intermediate layer 52, and output layer 53 is not particularly limited and can be appropriately determined according to the implementation method. Each node in each of the input layer 51, intermediate layer 52, and output layer 53 can be coupled to all nodes in adjacent layers. Therefore, the inference model 133 can be constructed as a fully coupled neural network. However, the coupling relationship between nodes is not limited to this example and can be appropriately determined according to the implementation method. For example, each node can also be connected to specific nodes in adjacent layers, or to nodes in layers other than adjacent layers.
[0052] Weights (coupling loads) are assigned to the couplings between nodes. Thresholds are set for each node; essentially, the output of each node is determined by whether the sum of the products of each input and each weight exceeds the threshold. The threshold can also be represented by an activation function. In this case, the sum of the products of each input and each weight is input to the activation function, and the activation function's operation is performed, thereby determining the output of each node. The type of activation function can be arbitrarily chosen. The weights of the couplings between nodes in input layer 51, intermediate layer 52, and output layer 53, as well as the thresholds for each node, are examples of parameters used in the computational processing of inference model 133.
[0053] In machine learning, the values of the parameters of inference model 133 are appropriately tuned to acquire the ability to perform the desired inference task using multiple training datasets 132. Training dataset 132 consists of a combination of training data and positive solution labels. In one example, with respect to training dataset 132, inference model 133 is trained (parameter values are adjusted) so that the execution result of the inference task obtained from inference model 133 by inputting training data into it is suitable for the corresponding positive solution label, thus constituting machine learning. In machine learning methods, based on the machine learning model, known methods such as backpropagation can be employed.
[0054] In this embodiment, the learning dataset 132 is pre-created based on dynamic images acquired by camera 40. The dynamic images reflect a specific operator selected for machine learning. The multiple learning datasets 132 each contain: training data, which is a predetermined number of consecutive frames contained in the dynamic images; and correct labels, which represent the action range of the specific operator's work reflected in the training data. Thus, by inputting the predetermined number of frames, an inference model 133 is generated that outputs labels representing the inferred action range.
[0055] The motion interval detection unit 103 detects the motion interval to which each frame of the dynamic image obtained from the camera 40 belongs. Specifically, the motion interval detection unit 103 inputs a predetermined number of consecutive frames containing the target frame (hereinafter referred to as the "target frame") of the motion interval into the inference model 133. For example, a predetermined number of frames (m+n+1 frames) consisting of m consecutive frames before the target frame, the target frame, and n consecutive frames after the target frame are input into the inference model 133. The motion interval detection unit 103 detects the motion interval represented by the label output from the inference model 133 as the motion interval to which the target frame belongs.
[0056] (Emotion Detection Department)
[0057] The emotion detection unit 104 detects the operator's emotions based on dynamic images acquired from the camera 30. The emotion detection unit 104 can detect emotions using known techniques (e.g., Japanese Patent Application Publication No. 2016-149063).
[0058] For example, the emotion detection unit 104 detects faces and facial features (eyes, eyebrows, nose, mouth, etc.) for each frame of a dynamic image. In face detection and facial feature detection, any algorithm can be used, exemplified by well-known methods; therefore, detailed explanations are omitted.
[0059] The emotion detection unit 104 identifies the operator's emotions (expressions) reflected in the frame based on the detected state of the face and facial organs. In this embodiment, emotions are classified into five types: "neutral," "glad," "angry," "surprise," and "sad." Alternatively, emotions can be classified into seven types, including the above five, plus "disgust" and "fear." As a result of emotion recognition, a score is output, which is a numerical representation of the intensity of each of the five (or seven) emotions, summed to 100. The scores for each emotion are also called expression component values. Emotions (expressions) also depend on the operator's physical and mental condition. Therefore, the scores can be used to estimate the operator's physical and mental condition.
[0060] Furthermore, emotion recognition can utilize any algorithm, including those based on well-known methods. For example, the emotion detection unit 104 extracts features related to the relative position and shape of facial organs based on their location information. These features can include Haar-like features, distances between feature points, Fourier descriptors, etc. Next, the emotion detection unit 104 inputs the extracted features into discriminators for each of the five (or seven) facial expressions, calculating the intensity of each expression. Each discriminator can be generated through learning from sample images. Finally, the emotion detection unit 104 normalizes the output values from the five (or seven) discriminators to a total of 100, outputting scores (expression component values) for the five (or seven) emotions.
[0061] The emotion detection unit 104 stores the emotion recognition results along with timestamp information in the database within the storage unit 101.
[0062] (Gaze Inspection Department)
[0063] The gaze detection unit 105 detects the operator's gaze based on dynamic images acquired from the camera 30. The gaze detection unit 105 uses known techniques (e.g., Japanese Patent Application Publication No. 2009-266086) to detect the gaze.
[0064] For example, the gaze detection unit 105 estimates the facial orientation of the worker reflected in each frame of the dynamic image. Furthermore, the method used to estimate facial orientation is not limited to a specific method, but a method that can perform the estimation more accurately, quickly, and simply is preferred.
[0065] Furthermore, the gaze detection unit 105 detects the outline of the operator's eyes and the pupil reflected in each frame. For example, the gaze detection unit 105 may consider methods such as edge detection or corner detection to detect the inner and outer corners of the eyes. In addition, after detecting the outline of the pupil through edge detection, the gaze detection unit 105 detects the left and right ends of the pupil.
[0066] The gaze detection unit 105 calculates feature parameters based on the eye contour and pupil detection results. These feature parameters represent the relationship between the inner and outer corners of the eye and the left and right ends of the pupil. For example, the feature parameters represent the relative coordinates of the inner corner of the eye relative to the left end of the pupil (in other words, the vector between the left end of the pupil and the inner corner of the eye), and the relative coordinates of the outer corner of the eye relative to the right end of the pupil (in other words, the vector between the right end of the pupil and the outer corner of the eye). Alternatively, the feature parameters can also represent the ratio of the lengths of these two vectors. Any feature parameter represents the position of the pupil relative to the eye contour.
[0067] The gaze detection unit 105 estimates the worker's pupil direction by applying the estimated facial orientation and feature parameters to the correlation between the facial orientation, feature parameters, and pupil direction. The correlation is pre-created. The gaze detection unit 105 then calculates the worker's gaze direction by adding the estimated facial orientation to the estimated pupil direction.
[0068] (Provided by the Department)
[0069] The providing unit 106 provides the detection results of the motion range detection unit 103, the emotion detection unit 104, and the gaze detection unit 105, as well as a screen displaying various information obtained based on the detection results. Specifically, the providing unit 106 displays the screen on the display device 70. The various information can be generated separately based on the detected worker's motion range, emotion, and gaze, or it can be generated by combining multiple items selected from the motion range, emotion, and gaze.
[0070] <Verification of the Estimation Example of Action Range>
[0071] The specific verification results of the estimated motion range for the "welding" process are explained.
[0072] Figure 5 This diagram represents three frames corresponding to the three motion zones of the "welding" process, and frames that do not belong to any motion zone. As described above, the "welding" process includes: the motion zone "first zone" for moving the substrate from the previous process and mounting the substrate onto the fixture; the motion zone "second zone" for welding the component onto the substrate; and the motion zone "third zone" for removing the substrate from the fixture and moving the substrate to the next process. Figure 5 (a), (b), and (c) represent frames belonging to the action intervals "Interval 1", "Interval 2", and "Interval 3", respectively. In a moving image, there exist frames that do not belong to any of the action intervals "Interval 1", "Interval 2", and "Interval 3", that is, frames that do not perform any operation within the action intervals "Interval 1", "Interval 2", and "Interval 3". Therefore, an inference model 133 is generated to classify each frame of the moving image into any of the action intervals "Interval 1", "Interval 2", "Interval 3", or "None". The action interval "None" is an interval where no operation within the action intervals "Interval 1", "Interval 2", and "Interval 3" is performed.
[0073] Figure 6 This is a graph representing the verification results of the estimated action range. In Figure 6 The upper part shows the action zones categorized by human confirmation of the moving images. That is, Figure 6 The upper part represents the correct solution for the action range. On the other hand, in Figure 6The lower part shows the action range inferred using inference model 133.
[0074] Inference was derived using inference model 133 generated under the following conditions. Figure 6 The lower part shows the action range.
[0075] • Model used: 3DResNet (https: / / github.com / kenshohara / 3D-ResNets-PyTorch)
[0076] Input data: 16 frames of images, each pixel representing the RGB intensity, with an image size of 112 pixels × 112 pixels.
[0077] • Learning Rate: 0.1 (or 0.01 if the Validation Loss converges)
[0078] • Data Augmentation:
[0079] 50% Horizontal Flip
[0080] Spatial Crop is randomly selected from the four corners and one center.
[0081] Randomly extract 16 frames from the moving image
[0082] • Transfer learning: using r3d50_K_200
[0083] Depth 50, epoch 200, classes 700, using the Kinectis-700 dataset.
[0084] • Number of data used: Action interval "Interval 1": 10, Action interval "Interval 2": 10, Action interval "Interval 3": 15, Action interval "None": 2
[0085] • Small batch size: 30.
[0086] like Figure 6 As shown, the action range inferred by inference model 133 is similar to the action range classified by human confirmation. Therefore, inference model 133 has high inference accuracy.
[0087] <Example of provided image>
[0088] Figure 7 This is an example of a picture that provides a view. Figure 7The screen 60 shown is provided by the providing unit 106 and includes a graph 61 indicating the shift of the detected motion range. By checking the screen 60, the user can determine whether the operator's actions are appropriate.
[0089] Figure 8 This is an illustration representing another example of providing a visual representation. Figure 9 This is another example of a picture that provides a visual representation. Figure 8 , 9 The image 65 shown is provided by the providing unit 106. For example... Figure 8 , 9 As shown, screen 65 includes areas 66 to 68.
[0090] In area 66, the moving image captured by camera 30 is reproduced. In area 66, frames are displayed based on the operation of operation bar 69. Alternatively, even without operation of operation bar 69, the latest frame captured from camera 30 can be displayed in area 66.
[0091] In region 66, markers 66a to 66d and lines 66e and 66f are overlaid in the dynamic image.
[0092] Label 66a represents the position of the pupil relative to the outline of the worker's right eye as reflected in the moving image. Label 66b represents the position of the pupil relative to the outline of the worker's left eye as reflected in the moving image. Labels 66a and 66b are generated based on the outline of the eyes and the pupil detected from the frame displayed in region 66.
[0093] Line 66e represents the gaze direction of the worker's right eye as reflected in the moving image. Line 66f represents the gaze direction of the worker's left eye as reflected in the moving image. Lines 66e and 66f are generated based on the gaze directions detected from the frames displayed in region 66.
[0094] Therefore, by confirming the markers 66a, 66b and lines 66e, 66f, users can easily grasp the outline of the operator's eyes, the state of the pupils, and the direction of their gaze.
[0095] The marker 66c indicates a negative type of emotion reflected in the moving image. Specifically, the marker 66c represents the emotion with the highest score among "neutral," "surprise," "anger," and "sadness," and is accompanied by a pattern corresponding to the emotion. Figure 8 The marker 66c indicates the emotion "normal". Figure 9 The marker 66c indicates the emotion "sadness". Additionally, an indicator 66g is shown around the marker 66c, indicating the fractional magnitude of the emotion represented by the marker 66c.
[0096] The designation 66d indicates a positive type of emotion reflected in the moving image. Specifically, the designation 66d represents an emotion with a high score between "normal" and "joyful," and is accompanied by a pattern corresponding to the emotion. Figure 8 The marker 66d indicates the emotion "normal". Figure 9 The mark 66d indicates the emotion "joy". Additionally, an indicator 66h is shown around the mark 66d, indicating the fractional magnitude of the emotion represented by the mark 66d.
[0097] Users can grasp the operator's emotions by confirming markers 66c and 66d, and can grasp the degree of emotions by confirming indicators 66g and 66h.
[0098] An image of an object present in front of the operator is displayed in area 67. This image can be prepared in advance or obtained from a camera different from cameras 30 and 40. A marker 67a indicating the operator's viewpoint is overlaid in area 67. The position of marker 67a is determined based on the direction of the gaze detected from the frame displayed in area 66. Figure 8 In the image 65 shown, the operator's gaze is directed towards the upper left; therefore, in area 67, mark 67a is displayed in the upper left portion of the image. Specifically, in the image of area 67, mark 67a is superimposed on the standard book A projected in the upper left. Figure 9 In the image 65 shown, the operator's line of sight is downward, therefore, in area 67, mark 67a is displayed at the bottom of the image. Specifically, in the image of area 67, mark 67a is superimposed on the component box reflected on the lower side.
[0099] By confirming area 67, users can easily determine where the operator is looking.
[0100] Area 68 displays a chart representing the worker's emotional progression. Specifically, the chart shows the progression of scores for five emotions: "neutral," "glad," "surprise," "angry," and "sad." Line 68a is displayed in area 68, representing the time corresponding to the frame displayed in area 66. Therefore, by observing the scores of each emotion overlapping with line 68a, the user can grasp the worker's emotions reflected in the frame displayed in area 66.
[0101] <Examples of how test results are used>
[0102] Figure 10 This is a graph showing the relationship between worker morale and production targets. Figure 10 The upper part shows the shifts in production quantity per unit time and defect rate as production indicators. Figure 10 The lower part shows the progression of the operator's scores for each emotion. Figure 10 In the example shown, as the score for "sad" increased, a decrease in the number of products produced per unit of time and an increase in the defect rate were observed.
[0103] Therefore, managers confirm Figure 8 , 9 In area 68, it is possible to identify workers exhibiting emotions that lead to decreased productivity and provide appropriate attention to them. Furthermore, as mentioned above, emotions depend on physical and mental condition. Therefore, managers can identify... Figure 8 , 9 Area 68 can detect changes in the worker's physical and mental condition, allowing the worker to rest.
[0104] Furthermore, the provision unit 106 can also be based on Figure 10 The relationship shown illustrates that when the score of the object category among multiple emotion categories deviates from the prescribed range, a notification prompting the operator to pay attention is provided. Specifically, the providing unit 106 can also compare the score of the emotion "sadness" with a threshold, and if the score of the emotion "sadness" exceeds the threshold, a notification prompting appropriate attention is provided. By having the operator receive the above notification, managers can provide appropriate attention in advance. As a result, a decline in productivity can be suppressed.
[0105] Workers are encouraged to review the work standards while performing their tasks. Therefore, managers can review the work standards by confirming them. Figure 8 , 9 The system uses area 67 to determine whether the operator's viewpoint shifts in the desired sequence. This allows the manager to determine whether the work has been performed in the appropriate steps.
[0106] Furthermore, the providing unit 106 can also store reference information indicating changes in viewpoint during standard operations, and calculate the similarity between the reference information and changes in the marker 67a displayed in area 67. The reference information is pre-created. The providing unit 106 can also provide notifications indicating different work steps if the similarity between the reference information and the changes in the marker 67a displayed in area 67 is less than a threshold. This allows managers to easily identify which operators need to be trained on the required work steps.
[0107] Managers confirm Figure 7 The screen 60 shown can create an ideal work procedure based on the shift in motion range detected from dynamic images obtained by filming skilled operators. Alternatively, the providing unit 106 can automatically create work standards based on the detected shift in motion range and provide the created work standards.
[0108] <Variation Example>
[0109] The memory 13 of the information processing device 10 may not store the model generation program 131. That is, the information processing device 10 may not have a model generation unit 102. In this case, the information processing device 10 can obtain the inferred model 133 from another device that has the model generation program 131 installed. The processor of the other device implements the model generation unit 102 by executing the model generation program 131.
[0110] The memory 13 of the information processing device 10 may not store one or two of the motion range detection program 134, emotion detection program 135, and gaze detection program 136. That is, the information processing device 10 may not have one or two of the functional blocks of the motion range detection unit 103, emotion detection unit 104, and gaze detection unit 105. For example, if the information processing device 10 only has the emotion detection unit 104, the providing unit 106 can provide a screen 65 that includes areas 66 and 68 but does not include area 67. If the information processing device 10 only has the gaze detection unit 105, the providing unit 106 can provide a screen 65 that includes areas 66 and 67 but does not include area 68. If the information processing device 10 only includes the motion range detection unit 103, the providing unit 106 can provide... Figure 7 The image shown is 60, but no image is provided. Figure 8 and Figure 9 The image shown is 65. In the case where the information processing device 10 only has an emotion detection unit 104 and a gaze detection unit 105, the providing unit 106 provides... Figure 8 , 9 The image shown is 65, but no further information is provided. Figure 7 The image shown is 60. In the case where the information processing device 10 only has an action range detection unit 103 and an emotion detection unit 104, the providing unit 106 provides... Figure 7 The screen shown is 60, and a screen 65 that includes areas 66 and 68 but does not include area 67 is also provided. When the information processing device 10 only includes the motion range detection unit 103 and the gaze detection unit 105, the providing unit 106 provides... Figure 7 The screen shown is 60, and a screen 65 is provided that includes areas 66 and 67 but does not include area 68.
[0111] Embodiments of the present invention have been described, but should be considered as illustrative rather than restrictive in all respects. The scope of the invention is defined by the technical solutions of the invention and is intended to include all modifications equivalent to or within the scope of the technical solutions of the invention.
Claims
1. An information processing device, wherein, The information processing device has a first acquisition unit that acquires a first dynamic image from a first camera located at a production site where an operation involving multiple action zones is performed, and which captures images of the operator and the surrounding area of the operator. The first dynamic image comprises multiple frames. The information processing device also has: The action region detection unit uses an inference model to detect the action region to which each of the plurality of frames belongs; and The providing unit provides the detection results of the action range detection unit. The inference model is generated through learning processing using multiple learning datasets, each containing: a predetermined number of consecutive second frames within dynamic images depicting a specific worker; and labels representing the action range of the specific worker's task depicted in the predetermined number of second frames. The motion interval detection unit inputs the first frame, which contains a predetermined number of consecutive object frames from the plurality of frames, into the inference model, and detects the motion interval represented by the label output from the inference model as the motion interval to which the object frame belongs.
2. The information processing apparatus according to claim 1, wherein, The information processing device has: The second acquisition unit acquires a second dynamic image from a second camera installed at the production site that captures the face of the operator. as well as An emotion detection unit detects the emotions of the object operator as reflected in each frame of the second dynamic image. The providing unit also provides the emotional shifts detected by the emotion detection unit.
3. The information processing apparatus according to claim 1, wherein, The information processing device has: The second acquisition unit acquires a second dynamic image from a second camera installed at the production site that captures the face of the worker; and A gaze detection unit detects the gaze direction of the operator as reflected in each frame of the second dynamic image. The providing unit also provides an image of the object present in front of the operator. The providing unit determines the position of the operator's viewpoint in the image based on the gaze direction detected by the gaze detection unit, and displays a mark at the determined position in the image.
4. The information processing apparatus according to claim 1, wherein, The information processing device has: The second acquisition unit acquires a second dynamic image from a second camera installed at the production site that captures the face of the operator. An emotion detection unit detects the emotions of the object operator as reflected in each frame of the second dynamic image; and A gaze detection unit detects the gaze direction of the operator as reflected in each frame of the second dynamic image. The providing unit also provides the emotion shift detected by the emotion detection unit. The providing unit also provides an image of the object present in front of the operator. The providing unit determines the position of the operator's viewpoint in the image based on the gaze direction detected by the gaze detection unit, and displays a mark at the determined position in the image.
5. An information processing method, wherein, The information processing method includes the following steps: acquiring dynamic images from cameras installed at the production site where a work involving multiple action zones is being performed, and capturing images of the operator and the surrounding area of the operator; the dynamic images contain multiple frames. The information processing method also includes the following steps: Using an inference model, detect the action region to which each of the multiple frames belongs within the multiple action regions; and Provide test results. The inference model is generated through learning processing using multiple learning datasets, each containing: a predetermined number of consecutive second frames within dynamic images depicting a specific worker; and labels representing the action range of the specific worker's task depicted in the predetermined number of second frames. The detection step includes the following steps: inputting the first frame, which contains a specified number of consecutive object frames from the plurality of frames, into the inference model, and detecting the action interval represented by the label output from the inference model as the action interval to which the object frame belongs.
Citation Information
Patent Citations
Line of sight detecting device and method and program
JP2009266086A
Emotion estimation system and emotion estimation method
JP2016149063A
Device, method, and program for processing information, and recording medium
JP2020204819A
Emotion estimation apparatus, method, and program
US20190239795A1
Training system and data collection device
US20230360437A1