Device for determining whether action is necessary, program for determining whether action is necessary
The device and program use skeletal information to accurately assess operator needs by analyzing joint points and postures, overcoming the limitations of facial expression capture, thereby improving service quality and customer impression.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- KONICA MINOLTA INC
- Filing Date
- 2021-12-10
- Publication Date
- 2026-04-14
AI Technical Summary
Existing systems struggle to accurately determine the need for staff intervention when operators wear masks or sunglasses, as they cannot effectively capture facial expressions or body contours, leading to inaccurate psychological state detection.
A device and program that utilize skeletal information generation from captured images to identify joint points suitable for determining intervention needs, excluding irrelevant points, and analyze postures and movements to determine if staff response is necessary.
Accurately determines the need for staff intervention by reflecting the operator's psychological state through skeletal information, enhancing service quality and customer impression.
Smart Images

Figure 0007844857000004 
Figure 0007844857000005 
Figure 0007844857000006
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a device for determining whether or not it is necessary to take any action when an operator performs an operation on a target device such as an electronic device installed in a store or the like. [Background technology]
[0002] In recent years, stores have installed various electronic devices, such as image forming machines, ATMs, ticket reservation machines, vending machines, and amusement machines, within their limited space. If these electronic devices are not operated correctly, customers will not be able to purchase goods or receive services. Furthermore, ensuring that customers use these electronic devices correctly is a critical issue for businesses. Just as with the turnover rate of seats in restaurants, profit margins will not improve unless the turnover rate of electronic device usage increases. For these reasons, it has become common practice in recent years for customer service staff to receive appropriate training and guidance on how to handle individual customers operating electronic devices. Such individual support for customers includes providing assistance to those who are confused about how to operate the electronic devices, and reporting those who attempt to damage the devices or use them for fraudulent purposes.
[0003] The problem here lies in determining the timing of providing assistance to the operator. In recent years, the repercussions of staff reductions have been significant, and it has become a heavy burden for customer service staff to keep an eye on operators using electronic devices while handling various store tasks. To alleviate this burden, user operation support devices have existed for some time that can determine whether individual assistance is necessary.
[0004] The user operation support device described in Patent Document 1 captures the user's gaze and facial expressions from images captured by a camera. If the user's gaze indicates that they are looking for a store clerk or checking the contents of an instruction manual, the device notifies the store clerk that assistance is needed. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2019-101775 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] When using image processing techniques to have a computer determine whether or not staff intervention is necessary, a problem arises as to how to detect the psychological state of the operator attempting to operate the target device.
[0007] The operator's gaze and facial expressions succinctly reflect the operator's psychological state when attempting to operate the target device. Therefore, as described in Patent Document 1, the first approach is to determine whether or not staff intervention is necessary based on the operator's gaze and facial expressions. When determining whether or not intervention is necessary based on the operator's facial expressions and gaze, the operator's face, including facial features such as the eyes, nose, and mouth, must be captured by the camera. However, in recent years, wearing a mask when entering a store has become commonplace, and some operators even wear sunglasses in addition to masks when operating the target device, making it impossible to capture the operator's facial expressions and gaze.
[0008] If it is difficult to capture facial features with a camera, a second approach is to take an external view image of the operator's body and determine the operator's psychological state based on the outline shape of the operator's body as seen in this external view image. This is because the operator's psychological state is also reflected in their body.
[0009] However, the contour shape in external view images changes depending on differences in body shape, such as being overweight or underweight, and differences in clothing, such as wearing thick or thin clothes or a skirt. Since the contour shape includes extraneous shapes due to differences in body shape and clothing, even if image recognition is applied to the contour shape, it cannot be said that it is possible to detect the fundamental psychological state of the operator who is about to operate the target device. Therefore, even if one tries to determine whether or not an action is necessary based on the contour shape recognition results, there is a problem in that the accuracy of the determination cannot be improved.
[0010] The purpose of this disclosure is to provide a device for determining whether or not an operator needs to respond to a target device such as an electronic device. [Means for solving the problem]
[0011] The above problem relates to a device for determining whether staff need to respond to an operator operating a target device, comprising an acquisition means for acquiring a captured image of the location where the target device is installed, The system includes a skeletal information generation means that generates skeletal information from captured images in which the operator is visible, and a necessity determination means that determines whether or not staff intervention is necessary based on the skeletal information. The necessity determination means selects joint points of the operator's body that appear in captured images taken around the target device that are suitable for determining whether or not intervention is necessary, and excludes other joint points from the skeletal information. Furthermore, if the system determines from the skeletal information that the operator's face is rotated at a predetermined angle relative to the lower body and is in a turning posture, it will determine that action is required. This problem is solved by a device for determining whether or not action is necessary, characterized by the features described above. Furthermore, the need for staff to respond to an operator operating a target device may be resolved by a response necessity determination device comprising: acquisition means for acquiring a photograph of the location where the target device is installed; skeletal information generation means for generating skeletal information from the photograph, in which the operator is visible; and necessity determination means for determining whether staff response is necessary if the operator's posture corresponds to a confused pose, based on the skeletal position information shown in the skeletal information, wherein the necessity determination means determines from the skeletal information that the operator's face is at a predetermined rotation angle relative to the lower body and is in a turning posture, and the operator's posture corresponds to a confused pose. Furthermore, the need for staff to respond to an operator operating a target device may be resolved by a response necessity determination device comprising: acquisition means for acquiring a photographed image of the location where the target device is installed; skeletal information generation means for generating skeletal information from the photographed image in which the operator is visible; and necessity determination means for determining whether staff response is necessary based on the skeletal information, wherein the skeletal information generation means generates skeletal information for each of a plurality of frames and identifies the operator's movements from the skeletal information generated for each frame, and the necessity determination means calculates the rotation angle between the plane passing through the joint points of the face shown in the skeletal information generated from one frame and the plane passing through the joint points of the lower body, and if the calculated rotation angle exceeds a predetermined threshold, the operator's movements are determined to be confusing and a response is deemed necessary.
[0012] The aforementioned necessity determination means may identify the operator's posture from the positional relationship of the skeleton shown in the skeletal information.
[0013] The aforementioned necessity determination means may determine that action is necessary if the operator's posture corresponds to a posture of confusion.
[0014] When the position of the wrist joint point shown in the skeletal information is higher than the shoulder joint point, the posture of the operator may be regarded as a confused pose.
[0015] The necessity determination means Operator When it is determined from the skeletal information that the face makes a predetermined rotation angle with respect to the lower body and is in a turning posture, the posture of the operator may be regarded as a confused pose.
[0016] The skeletal information generation means may generate skeletal information for each of a plurality of frames and identify the operation of the operator from the skeletal information generated for each frame.
[0017] The necessity determination means may determine whether the operation of the operator is the operation of a person who is confused about the operation, and if so, issue a determination result that corresponding measures are necessary.
[0018] When the position of the wrist joint point shown in the skeletal information transitions from a position lower than the shoulder joint point to a higher position, it may be regarded as corresponding to the operation of a person who is confused about the operation.
[0019] The necessity determination means calculates the rotation angle formed by the plane passing through the face joint point shown in the skeletal information generated from one frame with respect to the plane passing through the lower body joint points, and if the calculated rotation angle exceeds a predetermined threshold value, the operation of the operator may be regarded as a confused operation.
[0020] Furthermore, the device for determining whether staff need to respond to an operator operating a target device comprises: acquisition means for acquiring a photograph of the location where the target device is installed; skeletal information generation means for generating skeletal information from the photographed image in which the operator is visible; and necessity determination means for determining whether staff response is necessary based on the skeletal information. The necessity determination means From the captured images of the area surrounding the target device, select the joint points of the operator's body that are suitable for determining whether or not a response is necessary, and exclude the other joint points from the skeletal information. calculates the degree of necessity for corresponding measures by the staff from the rotation angle formed by the operator's face with respect to the lower body, and notifies the staff Characterized by .
[0021] The acquisition means may acquire captured images at predetermined intervals, the skeletal information generation means may generate skeletal information for each captured image, and the notification by the necessity determination means may be made by displaying the continuous time change of the degree of need for response calculated from the skeletal information generated in the order of shooting.
[0022] The necessity determination means determines whether an action is necessary based on the calculated degree of need for action, and the notification by the necessity determination means may be made by plotting the necessity of individual actions at each of multiple points in time on the time axis, thereby displaying the discrete time changes in the necessity of an action.
[0023] The necessity determination means may calculate the amount of movement of the skeletal information generated for one frame compared to the skeletal information generated for the previous frame, and if the calculated amount of movement is small and the period of stillness with a small amount of movement continues for a predetermined threshold or longer, it may determine that action is necessary.
[0024] Furthermore, the device for determining whether staff need to respond to an operator operating a target device comprises: acquisition means for acquiring captured images of the location where the target device is installed; skeletal information generation means for generating skeletal information from captured images in which the operator is visible; and necessity determination means for determining whether staff response is necessary based on the skeletal information. The necessity determination means selects joint points of the operator's body that appear in captured images of the area around the target device that are suitable for determining whether response is necessary, and does not include other joint points in the skeletal information. The skeletal information generation means generates skeletal information for each of multiple frames and identifies the operator's movements from the skeletal information generated for each frame. The aforementioned necessity determination means is Furthermore, the system calculates the amount of movement of the skeletal information generated for a given frame compared to the skeletal information generated for the previous frame. If the calculated amount of movement is small, and the period of stillness with a small amount of movement continues for a predetermined threshold or longer, the system determines that action is required. While a specific part or all parts of the operator's body are moving, the system calculates the total movement of the joints in that specific part or all parts of the operator's body. The system then compares this total movement to a threshold value to determine the operator's skill level. If the determined skill level is that of a beginner, the system changes the predetermined threshold value to a shorter value. Characterized by .
[0025] The aforementioned necessity determination means may select joint points of the operator's body that appear in the captured image taken around the target device and are suitable for determining whether a response is necessary, while excluding other joint points from the skeletal information.
[0026] The aforementioned skeletal information may represent the positions of joint points connecting body segments using coordinates in a coordinate system with the neck as the origin.
[0027] The aforementioned skeletal information may represent the positions of the joints connecting the body segments in coordinates of the camera's image plane coordinate system.
[0028] The acquisition means may include a recording means for capturing multiple frames of images obtained by a camera and recording the image data of the multiple frames captured by the acquisition means onto a recording medium, and the skeletal information generated by the skeletal information generation means may be recorded on the recording medium together with the image data.
[0029] The recording means may identify periods when the operator's body is still and periods when the operator's body is in motion from multiple frames of image data recorded on the recording medium, and the skeletal information generation means may generate skeletal information from frames in which the operator's body is in motion.
[0030] The skeletal information generation means may detect the speed of the operator's body movements from multiple frames recorded on the recording medium, and generate skeletal information from image data of frames in which the speed of the operator's body movements meets a predetermined standard.
[0031] The skeletal information generation means may generate skeletal information from the image data of each frame if the frame immediately preceding the current time, or multiple frames immediately preceding the current time, constitute the operating period.
[0032] If the captured image includes the bodies of multiple operators operating each of the multiple target devices, the skeletal information generation means may generate skeletal information corresponding to each operator, and the necessity determination means may determine whether or not a response is necessary for each operator.
[0033] When the aforementioned skeletal information generation means generates skeletal information corresponding to each operator, the joint point information included in the skeletal information may include information that uniquely identifies the corresponding operator among multiple operators.
[0034] The aforementioned response may be defined as providing appropriate support for the operation of the target device.
[0035] The aforementioned response may be defined as taking appropriate measures against torts against the target device or torts using the target device.
[0036] The aforementioned target device may be either an image forming apparatus or an automated teller machine.
[0037] The skeletal information generation means may extract feature quantities from the captured image, estimate the relationships between the feature quantities and estimate the locations that appear to be joint points of the human body based on the feature quantities, and the generation of the skeletal information may include a process to clarify which part of the human body each joint point corresponds to by combining the estimated joint point locations with the estimated relationships.
[0038] The estimation of the relationship between the joint points may be an inference operation using a convolutional neural network that targets information indicating the relationship between features, and the estimation of a likely joint point location may be an inference operation using a convolutional neural network that targets probability information indicating how likely it is to be a joint point.
[0039] The above problem relates to a program for determining whether staff need to respond to an operator operating a target device, which involves a computer performing an acquisition step to acquire a photograph of the location where the target device is installed, a skeletal information generation step to generate skeletal information from the photographed image in which the operator is visible, and a necessity determination step to determine whether staff response is necessary based on the skeletal information, wherein in the necessity determination step, the computer selects joint points of the operator's body that appear in the photographed image taken around the target device that are suitable for determining whether response is necessary, and the other joint points are not included in the skeletal information. Furthermore, if the system determines from the skeletal information that the operator's face is rotated at a predetermined angle relative to the lower body and is in a turning posture, it will determine that action is required. This can also be resolved by a program that determines whether or not action is necessary, characterized by the above features. Furthermore, the program determines whether staff need to respond to an operator operating a target device, and involves a computer performing an acquisition step to acquire a photographic image of the location where the target device is installed, a skeletal information generation step to generate skeletal information from the photographic image in which the operator is captured, and a necessity determination step to determine whether staff response is necessary based on the skeletal information, wherein in the necessity determination step, the program selects joint points of the operator's body that appear in the photographic image taken around the target device and is suitable for determining whether response is necessary, and other joint points are not included in the skeletal information, in the skeletal information generation step, skeletal information is generated for each of multiple frames, the operator's movements are identified from the skeletal information generated for each frame, and in the necessity determination step, the program calculates the rotation angle between the plane passing through the joint points of the face shown in the skeletal information generated from one frame and the plane passing through the joint points of the lower body, and if the calculated rotation angle exceeds a predetermined threshold, the program determines that response is necessary. Furthermore, the program for determining whether staff need to respond to an operator operating a target device involves a computer that performs the following steps: an acquisition step of acquiring a photograph of the location where the target device is installed; a skeletal information generation step of generating skeletal information from the photographed image in which the operator is visible; and a necessity determination step of determining whether staff response is necessary based on the skeletal information. In the necessity determination step, the program selects joint points of the operator's body that appear in the photographed image of the area around the target device that are suitable for determining whether staff response is necessary, and excludes other joint points from the skeletal information. In addition, it calculates the degree to which staff response is necessary from the rotation angle of the operator's face relative to the lower body and notifies the staff. Furthermore, the program determines whether staff need to respond to an operator operating a target device, and involves the following steps: an acquisition step to acquire a photographic image of the location where the target device is installed; a skeletal information generation step to generate skeletal information from the photographic image in which the operator is visible; and a necessity determination step to determine whether staff response is necessary based on the skeletal information. In the necessity determination step, the program selects joint points of the operator's body that appear in the photographic image taken around the target device that are suitable for determining whether response is necessary, and excludes other joint points from the skeletal information. In the skeletal information generation step, the program generates skeletal information for each of multiple frames. The system identifies the operator's movements from the skeletal information generated for each frame. In the necessity determination step, it further calculates the amount of movement of the skeletal information generated for one frame compared to the skeletal information generated for the previous frame. If the calculated amount of movement is small and the period of stillness with small movement continues for a predetermined threshold or longer, it determines that action is necessary. Furthermore, while a specific part or all parts of the operator's body are moving, the system calculates the total amount of movement of the joint points included in that specific part or all parts of the operator's body. By comparing the total amount of movement with a threshold, the system determines the operator's skill level. If the determined skill level is that of a beginner, the predetermined threshold is changed to a shorter value. Alternatively, the issue may be resolved by a response necessity determination program that causes a computer to determine whether staff need to respond to an operator operating a target device, the program comprising: an acquisition step of acquiring a photograph of the location where the target device is installed; a skeletal information generation step of generating skeletal information from the photograph, in which the operator is visible; and a necessity determination step of causing the computer to make a determination that staff response is necessary if the operator's posture corresponds to a confused pose, based on the skeletal position information shown in the skeletal information, and wherein in the necessity determination step, if the skeletal information determines that the operator's face is at a predetermined rotation angle relative to the lower body and is in a turning posture, the operator's posture is deemed to correspond to a confused pose. Furthermore, the issue may be resolved by a response necessity determination program that causes a computer to determine whether staff need to respond to an operator operating a target device, the program comprising: an acquisition step of acquiring a captured image of the location where the target device is installed; a skeletal information generation step of generating skeletal information for each of multiple frames of the captured image in which the operator is captured, and identifying the operator's movements from the skeletal information generated for each frame; and a necessity determination step of determining whether staff response is necessary based on the skeletal information, wherein in the necessity determination step, the program calculates the rotation angle between the plane passing through the joint points of the face shown in the skeletal information generated from one frame and the plane passing through the joint points of the lower body, and if the calculated rotation angle exceeds a predetermined threshold, the program determines that the operator's movements constitute confusing movements and that response is necessary. [Effects of the Invention]
[0040] The skeletal information generated by the skeletal information generation means represents the skeletal structure of the operator attempting to operate the target device. Since the skeleton represents the fundamental state of the operator's body, regardless of differences in body shape (fat or thin) or clothing (thick or thin clothing), it is highly likely that the operator's psychological state at the time of operation is clearly reflected. Therefore, by determining whether intervention is necessary based on the skeletal information, it is possible to accurately determine whether staff intervention is required.
[0041] By being able to appropriately determine whether or not action is necessary, it is possible to improve the impression of the store and enhance the quality of service. [Brief explanation of the drawing]
[0042] [Figure 1] This disclosure shows the configuration of the store system. [Figure 2] This is a hardware configuration diagram of the device 30 for determining whether action is necessary. [Figure 3] The software configuration of the device 30 for determining whether action is necessary is shown. [Figure 4] Figure 4(a) shows the data structure of skeletal information. Figure 4(b) shows the anthropometric body label 2010, including the anthropometric body number. Figure 4(c) shows the anatomical body label 2020, including the joint point number. [Figure 5] This shows a three-dimensional coordinate system that represents the joint points included in the skeletal information. [Figure 6] This flowchart shows the procedure for determining whether the operator's actions / posture indicate confusion. [Figure 7] Figure 7(a) shows the user touching their hair with their right hand, and Figure 7(b) shows the user touching their chin with their right hand. [Figure 8] This flowchart shows the procedure for calculating the degree of response required based on the operator's turning direction. [Figure 9] Figure 9(a) shows a situation where a user who was standing in front of the image forming apparatus 10 and operating the image forming apparatus 10 turns around. Figure 9(b) shows an example of a three-dimensional vector indicating the direction of the user's face. [Figure 10] This flowchart shows the calculation procedure for determining the amount of movement of a specific body part or the entire body of the operator. [Figure 11] This shows that between frames, the right wrist joint point 3004 moves from 3D position 3104 to 3204, and the right elbow joint point 3003 moves from 3D position 3103 to 3203. [Figure 12] This flowchart shows the procedure for determining whether the operator's behavior pattern has switched from a stationary state to a confused state. [Figure 13] This flowchart shows the procedure for determining whether the confused state continued for a predetermined time, or whether the cycle of confused state and still state repeated more than a predetermined number of times. [Figure 14] This flowchart shows the procedure for displaying the level of action required. [Figure 15] An example of a graph showing the change in the level of support needed over time is shown. [Figure 16] Here is an example of a basic emotion model. [Figure 17] This is a flowchart showing the processing steps of the program that creates the teacher model. [Figure 18] This is a flowchart showing the emotion model determination procedure according to the second embodiment. [Modes for carrying out the invention]
[0043] The embodiments of the action requirement determination device according to this disclosure will be described below with reference to the drawings. The following describes one embodiment of a store system including the action requirement determination device.
[0044] [1] Store system configuration Figure 1 shows the configuration of the store system described herein. As shown in Figure 1, the store system 1 includes an image forming apparatus 10 that performs jobs such as copying, scanning, printing, and faxing in response to requests from the operator, a camera 20 installed on the ceiling of the store, and a response requirement determination device 30 installed on the cash register counter 4000.
[0045] The image forming apparatus 10 is a Multifunction Peripheral (MFP) that performs image formation using an electrophotographic method, and consists of a document transport unit 11, a scanner unit 12, a printer unit 13, a paper feed unit 14, and an operation unit 15. The document transport unit 11 feeds documents placed on a tray at the top of the device one by one to the scanner unit 12, and the scanner unit 12 optically reads the image recorded on the document. The printer unit 13 forms an image on the paper supplied from the paper feed unit 14 and outputs the printed paper from the output port. The operation unit 15 includes a touch panel display 17.
[0046] The touch panel display 17 is located on the front side of the image forming apparatus and displays job-related information and other information to the user. It also outputs the coordinates of the location touched by the user.
[0047] The key unit 18 includes keys for start and stop operations, keys for job selection, and a keyboard for inputting letters from A to Z, and accepts instructions to start and stop jobs, as well as input of usernames and email addresses. The camera 20 is mounted on the ceiling of the store so as to capture the area where the image forming apparatus 10 is installed and to capture the operator operating the image forming apparatus 10. The camera 20 can shoot in video mode or still image mode. In video mode, for example, it outputs moving image data at a frame rate of 1 / 40th of a second. In still image mode, it takes pictures of the operator located near the image forming apparatus 10 at longer intervals, i.e., every few seconds or minutes, and outputs multiple still image data that intermittently captures the posture and movement of the operator located near the image forming apparatus 10.
[0048] The response necessity determination device 30 determines whether individual assistance is needed for the operator based on the video or still image data captured by the camera 20. Individual assistance for the operator includes providing support for the operation of the image forming apparatus 10, reporting illegal activities using the image forming apparatus 10, and reporting illegal activities against the image forming apparatus 10. Explaining all the sub-concepts of individual assistance would be complicated, so we will explain the representative sub-concept, which is providing support for the operation of the image forming apparatus 10. In other words, when an operator attempts to operate the image forming apparatus 10, the response necessity determination device 30 determines whether support is needed for the operator's operation. If it determines that support is needed, it notifies the store clerk that support is needed for the operator through a display on the monitor 31, flashing of the call lamp 32, and sounding of the speaker 33.
[0049] [2] Configuration of the device 30 for determining whether action is necessary (2-1) Hardware configuration of the device 30 for determining whether action is required Figure 2 is a hardware configuration diagram of the response necessity determination device 30. The response necessity determination device 30 is equipped with a CPU 201, boot ROM 202, HDD 203, RAM 204, interrupt control circuit 205, image recording unit 206, and peripheral control circuit 207, and processes data according to the operation output from the image forming apparatus 10 and the video output from the camera 20. The operation output from the image forming apparatus 10 is the output from the operation detection unit 16 (see Figure 2) provided in the image forming apparatus 10, which indicates whether an operation by the operator has started or ended. Specifically, when a document is set, the number of copies is set, or money is inserted into the coin vending machine, the operation detection unit 16 raises the operation output to indicate that an operation by the operator has started. On the other hand, when the start key is pressed, the operation detection unit 16 lowers the operation output to indicate that an operation by the operator has ended.
[0050] When power is turned on, the response necessity determination device 30 performs a boot operation based on the bootstrap program stored in the boot ROM 202, loads the programs installed on the HDD 203 into the RAM 204, and starts the programs.
[0051] The interrupt control circuit 205 notifies the CPU 201 of the start and end of operation of the image forming apparatus 10 by outputting an interrupt signal based on the operation output of the operation detection unit 16 of the image forming apparatus 10.
[0052] The image recording unit 206 encodes the video signal captured from the camera 20 and writes it to the HDD 203. Information generated from this video signal (skeletal information generated by the skeleton detection application 110 described later) is recorded in the HDD 203 in association with the video signal.
[0053] The peripheral control circuit 207 has multiple I / O ports and connectors assigned to the address space of RAM 204, and by setting control values for the multiple I / O ports and connectors, it performs actions such as displaying information on the monitor 31, blinking the call lamp 32, and sounding the speaker 33.
[0054] The GPU208 is composed of multiple computing units, which are used to build a neural network and perform the processing necessary for skeleton detection.
[0055] (2-2) Software configuration of the device 30 for determining whether action is required Figure 3 shows the software configuration of the response necessity determination device 30. As shown in this figure, the software configuration of the response necessity determination device 30 has a layered structure in which the skeleton detection application 110 and the response necessity determination application 120 are placed on top of the operating system 100.
[0056] [3] Skeleton detection application 110 The skeleton detection application 110 generates skeletal information using images captured by the camera 20. Skeletal information generation involves extracting features, estimating the relationships between features and the locations of possible joints in the human body based on these features, and then deriving the locations of the joints and the relationships between them to generate skeletal information of the human body. Skeletal information is information that simplifies the human body by connecting multiple joints, and has the characteristic of having a high degree of freedom in the camera's shooting conditions.
[0057] Estimating the relationship between joint points is an inference operation that targets information (Part Affinity Field) that indicates the relationship (affinity) between features in a captured image. Here, features are calculated for each of several key locations in the captured image, and estimating the relationship between joint points determines how these key locations are related.
[0058] Estimating a possible joint point is an inference operation that targets probability information (Confidence Map) indicating how likely a point is to be a joint point.
[0059] These inference operations can be performed, for example, by using a CNN (Convolutional Neural Network) inference operation with a GPU208.
[0060] Furthermore, OpenPose is a well-known software that can perform skeletal detection by estimating the relationships between joint points and the locations that appear to be joint points.
[0061] (3-1) Data structure of skeletal information Figure 4(a) shows the data structure of skeletal information generated by the skeletal detection application 110 for a given person. As shown in this figure, the skeletal information 2000 is composed of multiple joint point information 2100, 2200, 2300...2400. The number of joint point information in the skeletal information is variable and is composed of joint point information that represents each of the joint points of the human body that are captured in the image data. The skeletal information shows a partial skeleton with one joint point information, and the entire skeleton of the operator is shown with all 18 joint point information. The entire skeleton is a skeleton represented by multiple joint points and multiple body segments connected via each joint point. The partial skeleton is a skeleton represented by one joint point and two body segments connected via that joint point. Furthermore, the skeletal information generated by the skeletal detection application 110 can represent the shape of joints that are not visible in the image captured by the camera.
[0062] Joint point information has a common structure and, as shown in Figure 4(a), consists of a human body label 2010, a body part label 2020, coordinate information 2030, and confidence level 2040.
[0063] The anthropometric body label 2010 includes an anthropometric body number. The anthropometric body number indicates which joint of the human body is captured in the image, starting from the side closest to the camera, as shown in Figure 4(b).
[0064] The body part label 2020 includes joint point numbers. As shown in Figure 4(c), the joint point numbers, ranging from 0 to 17, indicate which joint in the human body the joint point information corresponds to. The body parts identified by these numbers include the nose, neck, right shoulder, right elbow, right wrist, left shoulder, left elbow, left wrist, right hip, right knee, right ankle, left hip, left knee, left ankle, right eye, left eye, right ear, and left ear.
[0065] The coordinate information 2030 indicates the location of the corresponding joint point information using the coordinates of the image sensor coordinate system (camera coordinate system) in the camera 20.
[0066] Skeletal information allows for the identification of the posture and movements of multiple operators. Therefore, whether there is one image forming machine or multiple machines are installed, it is possible to determine from a single captured image whether individual attention is required for each operator when each machine is being operated by a different operator.
[0067] A confidence score of 2040 indicates how likely the corresponding joint point is.
[0068] (3-2) Coordinate System To determine the positional relationship of joint points from skeletal information, it is desirable to convert the coordinates of the joint points, expressed in the camera coordinate system, into a three-dimensional coordinate system. In this embodiment, the coordinates of the joint points shown in the skeletal information are converted from the camera coordinate system to a three-dimensional coordinate system as shown in Figure 5. The three-dimensional coordinate system in Figure 5 is a Cartesian coordinate system with the neck joint point 3001 as the origin, and is called the world coordinate system. The direction from the left shoulder joint point coordinate towards the right shoulder joint point coordinate is the X-axis direction, and the direction from the neck joint point towards the top of the head is the Z-axis direction. By using this coordinate system, if the wrist is raised, the Z-coordinate of that wrist becomes positive, making it possible to determine whether the posture is one in which the hand is raised.
[0069] The transformation from the camera coordinate system to the world coordinate system is achieved through a transformation via a weak calibration coordinate system. The weak calibration coordinate system is a three-dimensional coordinate system that establishes a correspondence between common feature points on the image and the three-dimensional position of the camera in the store. The three-dimensional coordinate M in the weak calibration coordinate system is shown in equation 1 below. sfm When observed in the camera coordinate system at 2D coordinate m, the projection relationship between the weak calibration coordinate system and the camera coordinate system is given by the camera's projection transformation matrix P and scale factor λ in the weak calibration coordinate system. m It can be expressed as shown in the following equation 1 using .
[0070]
number
[0071] On the other hand, any point M in the world coordinate systemworld It is represented by the following Equation (2) using a rigid body transformation with a rotation matrix R and a translation vector t.
[0072]
Equation
[0073] Assuming a 3D rigid body transformation matrix D using a rotation matrix R and a translation vector t, M world is expressed as shown in the following Equation (3).
[0074]
Equation
[0075] The rotation matrix R is calculated from the orthonormal basis vectors e i of the weakly calibrated coordinate system.
[0076] The orthonormal basis vectors e i are obtained from the points S x , S y , S z in the weakly calibrated coordinate system and a point S0 in the weakly calibrated coordinate system.
[0077] The translation vector t is given by the amount of translation from a point S0 in the weakly calibrated coordinate system to the origin O sfm .
[0078] In this way, by expressing the 三维 coordinates of the operator's joint points in the world coordinate system with the origin at the coordinates of the neck, the position of the joint points can be represented by relative values regardless of whether the left and right wrists are in the lower right, upper right, upper left, or lower left of the camera coordinate system of the captured image, and the positional relationship with other joint points can be determined.
[0079] [4] Corresponding Necessity Judgment Application 120 The response necessity determination application 120 is composed of multiple manual logic programs and determines whether individual responses are necessary for the operator operating the image forming apparatus 10 by referring to the positional relationships of joint points shown in the skeletal information. Manual logic refers to logic constructed by program developers using if statements, switch statements, for statements, etc. of high-level programming languages, and is in contrast to logic constructed by applying machine learning to a neural network (machine learning logic). As an example of such manual logic programs, the response necessity determination application 120 includes a confusion detection program 121, a turning detection program 122, a movement amount calculation program 123, a stationary pattern determination program 124, a confusion pattern determination program 125, and a response necessity display program 126, as shown in Figure 3.
[0080] Of these programs, the confusion detection program 121, the turn detection program 122, and the movement amount calculation program 123 identify the operator's movement or posture from the positional relationship of joint points shown in the skeletal information. On the other hand, the static pattern determination program 124, the confusion pattern determination program 125, and the response necessity display program 126 determine whether individual assistance by a store employee is necessary based on the identified operator's movement or posture. Whether the need for individual assistance is determined by movement or posture is determined by the shooting mode of the camera 20. If the shooting mode is video mode, it is important that the same movement continues across multiple frames, so the operator's movement appearing in the images across multiple frames must be one that requires individual assistance.
[0081] If the shooting mode is still image mode, it is important that the operator is in a specific posture in the still image obtained in a single shot. Therefore, it is sufficient if the operator's posture appearing in a single still image requires individual adjustment.
[0082] When the operator starts taking images, the loader 102 of the operating system 100 loads these programs into RAM 22 and provides the tasks, which are instances of these programs, to the multitasking execution environment of the kernel 101 to determine whether it is necessary to provide individual support to the operator operating the image forming apparatus 10.
[0083] (4-1) Confusion Detection Program 121 When the skeletal detection application 110 generates skeletal information, the confusion detection program 121 instructs the CPU 201 to determine whether the operator's movements and posture indicate confusion based on the positional relationships of the joint points shown in the skeletal information. Specifically, the confusion detection program 121 consists of program code that represents the steps in the flowchart in Figure 6.
[0084] Figure 6 is a flowchart showing the procedure for determining whether the operator's movements / posture indicate confusion. The procedure involves waiting for a period of time during which skeletal information can be generated (in this embodiment, this is 1 frame, but can be changed to 2 to 10 frames) to elapse (step S101).
[0085] If the shooting mode is video mode, step S101 becomes Yes at predetermined intervals of a certain number of frames. If the shooting mode is still image mode, step S101 becomes Yes each time a shooting point arrives at predetermined intervals. Once either period has elapsed (step S101 is Yes), the height of the left and right wrists is calculated from the joint point information of the left and right wrists (step S102), and the height of the left and right shoulders is calculated from the joint point information of the left and right shoulders (step S103).
[0086] Next, it is determined whether the height of either the left or right wrist is higher than the height of the left or right shoulder (step S104). If either the left or right wrist is higher than the height of the left or right shoulder (Yes in step S104), a result of determination that it is a confused movement is made (step S105), and the process returns to step S101.
[0087] In the world coordinate system, a situation where the height of either the left or right wrist is higher than the height of the left or right shoulders is highly likely to indicate that the operator is touching a part of their body unrelated to the operation of the image forming apparatus 10, such as their hair or chin, and is in a state of confusion.
[0088] Figure 7(a) shows the user touching their hair with their right hand, and Figure 7(b) shows the user touching their chin with their right hand. The fact that both wrists are not on the operating section 15 of the image forming apparatus 10, but on the chin and hair, suggests that the user is in a state of confusion. In this case, in the world coordinate system, the right wrist joint point 3004 is above the right shoulder joint point 3002, and step S104 becomes Yes. In this case, the confusion detection program 121 determines that the operator is in a state of confusion (step S105).
[0089] If the shooting mode is still image mode, a confused posture is indicated if step S104 is Yes for a single still image. If the shooting mode is video mode, a confused behavior is indicated if step S104 is Yes for multiple frames.
[0090] (4-2) Turn detection program 122 The turning detection program 122 is an application program that determines whether the operator's movement and posture are turning movements based on the positional relationships of joint points shown in the skeletal information generated for the current frame, and outputs the angle of the turn as the degree of need for response.
[0091] Figure 8 is a flowchart showing the procedure for calculating the degree of response required according to the operator's turning direction. After waiting for the period during which skeletal information can be generated to elapse (step S111), when the period during which skeletal information can be generated has elapsed (Yes in step S111), a plane vector is calculated from the image data of the current frame that passes through three points: the joint points of the left and right ears, the joint points of the left and right shoulders, and the neck joint, and the normal of this plane vector is taken as the vector in the direction of the face (step S112). A plane vector is calculated from the image data of the current frame that passes through three points: the joint points of the left and right hips, the joint points of the left and right knees, and the neck joints of the left and right ankles, and the normal of this plane vector is taken as the vector in the direction of the lower body (step S113). The rotation angle between the vector in the direction of the face and the vector in the direction of the lower body is determined, and the degree of response required is calculated from this rotation angle (step S114).
[0092] When an operator encounters difficulties operating the image forming apparatus 10 and needs assistance, they will likely turn their head to look around and try to find a store employee located at the cash register counter 4000 or elsewhere. Because of this tendency, the system calculates the rotation angle of the face vector relative to the front of the image forming apparatus 10 to determine the degree of the turn.
[0093] Figure 9(a) shows a situation where a user who was standing in front of the image forming apparatus 10 and operating the image forming apparatus 10 turns around. In this figure, the joint points 3001 of the neck, 3016 of the right ear, and 3017 of the left ear are located on the operator's face. Therefore, in step S112, as shown in Figure 9(b), a vector 8001 that forms the normal to the plane 8000 passing through these points is calculated, and the degree of assistance required is calculated from the amount of rotation of the normal vector 8001. If the shooting mode is still image mode, the rotation angle calculated for a single still image is used as the rotation angle of the operator's turning posture to calculate the degree of assistance required. If the shooting mode is video mode, the rotation angles calculated for multiple frames are used as the rotation angle of the operator's turning motion to calculate the degree of assistance required.
[0094] (4-3) Movement calculation program 123 The movement amount calculation program 123 is an application program that determines whether a specific part of the operator or a part of the whole body moves based on the positional relationship of joint points shown in the skeletal information generated for the current frame, and consists of program code that represents the state procedure in the flowchart of Figure 10.
[0095] Figure 10 is a flowchart showing the calculation procedure for determining the amount of movement of a specific body part or body part of the operator. Variable i is a subscript variable that indicates each of the 18 joint points of the human body, and D(i) is the amount of movement of the joint point indicated by subscript i, indicating how far that joint point has moved from the previous pixel. sumD is an integration variable for accumulating the amount of movement of the joint point.
[0096] First, the variable sumD for accumulating the amount of change is set to 0 (step S121), and it is determined whether a frame in which skeletal information can be detected has passed (step S122). If a frame in which skeletal information can be detected has passed (Yes in step S122), the target joint is determined according to the mode setting for the whole body or a specific body part (step S123), the variable i is initialized to 1 (step S124), the amount of change D(i) of joint point i from the previous frame is calculated (step S125), and the amount of change D(i) of joint point i is added to sumD (step S126).
[0097] Step S127 determines whether the variable i is less than the total number of joint points N, and depending on the result of the determination, it switches between continuing the loop from steps S125 to S128 or exiting the loop. If the variable i is less than N, the variable i is incremented (step S128), and the process returns to step S125 to continue the loop. If the variable i is greater than or equal to N, the loop from steps S125 to S128 is exited.
[0098] Assume that the right wrist joint point 3004, the right elbow joint point 3003, and the right shoulder joint point 3002 shown in Figure 5 constitute the specific area in step S123. Then, as shown in Figure 11, assume that between frames, the right wrist joint point 3004 moves from 3D position 3104 to 3204, the right elbow joint point 3003 moves from 3D position 3103 to 3203, and the right shoulder joint point 3002 moves from 3D position 3102 to 3202. In this case, in step S125 in Figure 10, the change amounts D(3) from 3D position 3104 to 3204, the change amounts D(2) from 3D position 3103 to 3203, and the change amounts D(1) from 3D position 3102 to 3202 are calculated, respectively, and in step S126, a sum-of-products operation is performed to sum these change amounts D(1), D(2), and D(3).
[0099] After exiting the above loop, the system proceeds to the judgment step sequence of steps S129 and S130. Step S129 determines whether the sum value sumD is below the quiescent state criterion. If step S129 is Yes, the system is in a quiescent state (step S131). Step S130 determines whether the sum value sumD is below the proficiency threshold D. T This determines whether the value falls below a certain threshold. If it does (Yes in step S130), the operator is considered a beginner, and the threshold for the idle period is set to a beginner level (step S132). Proficiency threshold D T If the above conditions are met (No in step S130), the operator is designated as a skilled operator, and the threshold for the idle period is set for skilled operators (step S133).
[0100] The threshold here indicates the standard for how long an operator standing in front of the image forming apparatus 10 must remain still before a store employee should speak to them. However, even if the total amount of movement is large, if the confusion detection program 121 detects that the operator is touching their hair or chin and exhibiting a confused posture, they will not be classified as a beginner.
[0101] The sum of the movement amounts sumD is large, and the proficiency threshold D TIf the above conditions are met, the operator is assumed to be standing in front of the copier and operating it smoothly, and the threshold for the period of inactivity is increased. If the operator of the image forming apparatus 10 is an expert, even if they stand in front of the image forming apparatus 10 without moving their hands for an extended period, it is highly likely that they are simply reading the on-screen manual. They are likely just trying to understand an unfamiliar function and will hardly need any assistance from a store employee. Therefore, in this case, the threshold for the period of inactivity is increased. The sum of the movement amounts sumD is small, and the proficiency threshold D T If the time falls below a certain threshold, the operator is assumed to be standing in front of the copier and operating it slowly, and the threshold for the period of inactivity is shortened. If the operator of the image forming apparatus 10 is a novice, there is a high probability that they will stand in front of the image forming apparatus 10 for several minutes and occupy it. In this case, the threshold for the period of inactivity is shortened, and it is assumed that individual attention is needed, so a store employee is asked to speak to the operator earlier.
[0102] The judgment in steps S132 to S133 can be changed according to the function settings for the application 120 that determines whether or not action is required. The square root of the sum of the squares of the absolute values of the change in skeletal coordinates and the square root of the sum of the squares of the differences in the change in skeletal coordinates can be calculated, and if these values are above a certain level, the user can be judged as an expert; otherwise, they can be judged as a beginner. By dividing the amount of change by the unit time (preferably an integer multiple of the frame duration), the speed of a specific part or a whole body part can be determined, and the level of skill in operation can be judged to determine whether individual action is required. The amount of change between frames is accumulated over the unit time, and the speed is calculated by taking the square root of the square of the value obtained from the accumulation.
[0103] Furthermore, in the flowchart of Figure 10, a threshold for the resting period is set according to the amount of movement and the total amount of movement. However, the handling of the amount of movement and the total amount of movement can be changed by configuring the functions of the action necessity determination application 120. Specifically, the amount of movement calculated by the amount of movement calculation program 123 is compared with a threshold, and if the total amount of movement is lower than the threshold, it can be determined that individual action is necessary. It can also be determined that individual action is necessary if the total amount of movement of joint points included in a specific part of the user's body or all parts is lower than the threshold.
[0104] If the total value of the movement calculated by the movement calculation program 123 is below a threshold, the total value of the movement can be displayed on the monitor 31 as an indication of the degree to which action is needed, and the store staff can be notified.
[0105] The confusion detection program 121, the turning detection program 122, and the movement amount calculation program 123 are examples of manual logic programs. In addition, the system can determine postures and movements such as sitting or squatting for extended periods based on the positional relationships of joint points in the lower body, and determine whether individual assistance by a store employee is necessary.
[0106] (4-4) Stationary pattern determination program 124 The static pattern determination program 124 is a program that determines whether the operator's action pattern has switched from a static state to a confused action, and consists of program code that causes the CPU 201 to execute the procedure shown in Figure 12. Figure 12 is a flowchart showing the procedure for determining whether the operator's action pattern has switched from a static state to a confused action. T is a variable for accumulating the elapsed time of the static period, and first this variable T is initialized to 0 (step S201).
[0107] Step S202 determines whether the operator has entered a stationary state, and step S203 determines whether the operator has entered a confused movement or posture. If no movement is observed in the operator's whole body or specific parts, and it is determined that the operator is in a stationary state (Yes in step S202), time T is incremented (step S204), and the loop of steps S202 and S203 is returned.
[0108] If it is determined that the operator is in a confused state or posture (Yes in step S203), it is determined whether the elapsed time T of the stationary state exceeds a threshold (step S205). If the elapsed time T exceeds the threshold (Yes in step S205), it is determined that individual attention is required (step S206), and the loop returns to steps S202 and S203. If the elapsed time T is below the threshold (No in step S205), the loop returns to steps S202 and S203.
[0109] (4-5) Confusion Pattern Determination Program 125 The confusion pattern determination program 125 is a program that determines whether the operator's behavior pattern indicates whether the confused state continued for a predetermined time or whether the confused state and the resting state were repeated several times. It consists of program code that causes the CPU 201 to execute the procedure shown in Figure 13. Figure 13 is a flowchart showing the determination procedure for determining whether the confused state continued for a predetermined time or whether the confused state and the resting state were repeated more than a predetermined number of times. In this flowchart, Tm is a variable indicating the elapsed time of the confused state, and Cm is a counter variable indicating how many times the confused state occurred. First, these variables are initialized to 0 (step S211), and then the program proceeds to the loop from steps S212 to S215.
[0110] Step S212 determines whether the operator has entered a confused state, step S213 determines whether the operator has entered a stationary state, step S214 determines whether the counter variable Cm is greater than or equal to the threshold Ct, and step S215 determines whether the elapsed time Tm of the stationary state is greater than 0. It is desirable to set the threshold Cm in step S214 to approximately 3 times. If the operating state continues, steps S212 and S213 are repeated, and if the stationary state continues, steps S212 to S215 are repeated. If the operator enters a confused state in this state (Yes in step S212), the elapsed time Tm of confusion is incremented (step S216), and it is determined whether the elapsed time Tm exceeds the threshold Tt (step S217). It is desirable to set the threshold Tt to approximately 3 seconds.
[0111] If the result is below the threshold (No in step S217), the process returns to the loop of steps S212-S215. If the elapsed time Tm exceeds the threshold Tt (Yes in step S217), the process determines that individual action is required (step S218), and the process returns to the loop of steps S212-S215.
[0112] If the system is in a stationary state (Yes in step S213), but the counter variable Cm falls below the threshold Ct (No in step S214), and the elapsed time Tm of the confused state exceeds 0 (Yes in step S215), the counter variable Cm indicating that the system has entered a confused state is incremented (step S219), the elapsed time Tm of the confused state is reset to 0 (step S220), and the loop returns to steps S212-S215. In other words, if a confused state occurs with a length less than the threshold for the confused state, Tm will exceed 0 in step S216, so step S215 is Yes, and the counter variable Cm, which indicates the number of times the system has entered a confused state, is incremented in step S219. If a confused state with a length less than the threshold for the confused state occurs many times, and Cm exceeds the threshold Ct, then step S214 becomes Yes. The system determines that individual attention is required (step S221), resets Cm to 0 (step S222), and resets Tm to 0 (step S224).
[0113] (4-6) Program 126 for displaying the degree of need for support The Action Needed Display Program 126 is a program that displays the action needed, and consists of program code that executes the flowchart in Figure 14. Figure 14 is a flowchart of the procedure for displaying the action needed. In this flowchart, it is determined whether the action needed has been calculated (step S231), and if the action needed has been calculated (Yes in step S231), if the action needed exceeds a threshold (Yes in step S232), it is determined that individual action is needed (step S233), and if the action needed falls below a threshold (No in step S232), it is determined that individual action is not needed (step S234). Each time, the action needed and whether action is needed are plotted at the current point in the action needed graph (step S235).
[0114] The support need graph is a graph with the horizontal axis representing time and the vertical axis representing the support need, and an example is shown in Figure 15. The support need graph in Figure 15 was obtained by repeatedly plotting the support need at step S235 and interpolating it with a curve over 800 frames. In this graph, if the threshold is set to 0.4, the support need calculated for periods P1, P2, and P3 exceeds the threshold. In that case, a call to the store clerk at step S110 is made during these periods.
[0115] [5] Summary As described above, according to this embodiment, if a part of the operator's body is captured in the image output by the camera 20 and at least two joint points are extracted, the skeleton detection application 110 generates skeleton information indicating the joint points captured in the image, and determines whether individual attention is needed for the operator based on the positional relationship of the joint points shown in the skeleton information. Therefore, even if facial features are not captured in the image, it is possible to determine whether individual attention is needed for the operator based on the skeleton information. Even when the camera shooting conditions are limited, the need for individual attention can be determined with high accuracy, which can improve the impression of the store and enhance the quality of service.
[0116] [6] Second embodiment In the first embodiment, the confusion detection program 121, the turning detection program 122, and the movement amount calculation program 123 determined the operator's confused movements, turning, and body movements using the logic (called manual logic) shown in the flowcharts of Figures 6, 8, and 10. However, this logic can only cover a limited range of operator emotions. Therefore, in the second embodiment, a teacher model is used to determine whether assistance is needed.
[0117] This section describes the process for determining the necessity of individualized support using a teacher model. To perform this support need determination, it is necessary to create a teacher model in advance and install the created teacher model and the program for determining the necessity of individualized support on HDD203.
[0118] The former teacher model consists of feature templates that enable classification using the basic emotion model shown in Figure 16. The basic emotion model in Figure 16 classifies eight emotions: joy, acceptance, fear, surprise, sadness, disgust, and anger.
[0119] Of the eight emotions thus classified, those marked with hatching (such as cowardice, apprehension, worry, fear, and confusion) are negative, and if an operator standing in front of the image forming machine exhibits these behaviors, some kind of action must be taken.
[0120] Figure 17 is a flowchart showing the processing steps of the program that creates the teacher model.
[0121] u is a variable that indicates each of the multiple negative emotions to be recognized. In step S1, the variable u is initialized to 1. In the following step S2, the performer is instructed to perform an action that represents the u-th emotion from among negative emotions such as cowardice, anxiety, worry, fear, and terror. During this time, joint points are extracted from the performer to generate skeletal information, and the body movements that characterize the information action are obtained from the skeletal information to be used as action data. The period from when the moving part starts moving until it returns to its original position is considered one action, and corresponding action data is generated. In step S3, feature quantities of a feature space with 7 dimensions, such as moving part p, frequency f, amplitude am, amount of movement dx, velocity v, acceleration α, and attribute e, are extracted from the skeletal information, and in step S4, the extracted feature quantities are used as templates for each emotional action.
[0122] The movement site p is one of the joint points shown in the skeletal information and is a key point. Frequency f is the fundamental frequency of a single operation, and is the reciprocal of the time it takes for the moving part to return to its original position after it has started moving.
[0123] The amplitude am is the amount of displacement of the moving part when it is at its most displaced state, from the time it starts moving until it returns to its original position. The motion amount dx is a value that indicates how much the moving part has displaced between frames, and is calculated for each of the x, y, and z axes. Velocity v is the value obtained by dividing the amount of motion in a frame by time, and is calculated for each of the X, y, and z components.
[0124] The acceleration α is the derivative of the velocity in each frame.
[0125] Attribute e indicates the emotion corresponding to the action and its intensity.
[0126] Step S5 determines whether the variable u is less than the total number of negative emotions v. If it is (Yes in step S5), the variable u is incremented (step S6), and the process returns to step S2. The loop continuation requirement in step S5 is that u is less than v, so the process returns to step S2 until u is between 1 and v-1, and then exits the loop when the variable u becomes v.
[0127] Through the above process, a training model is created, which is a set of templates corresponding to each of the multiple emotions. The set of templates is T{T j This is represented by the notation |j=0,1,2··N}.
[0128] Furthermore, in the second embodiment, instead of the confusion detection program 121 to the movement amount calculation program 123, a determination program is installed on the HDD 203 that determines which of the basic emotion models shown in Figure 16 corresponds to the emotion expressed by the operator's body. This determination program consists of program code that causes the CPU 201 to execute the procedure shown in Figure 18.
[0129] Figure 18 is a flowchart showing the emotion model determination procedure according to the second embodiment.
[0130] i is an index variable that indicates the motion features captured by the camera. k is an index variable that indicates each of the multiple templates, showing which of the multiple templates obtained through supervised learning is being processed.
[0131] The user's movement characteristics {Pi, fi, ami, dxi, vi, αi, ei} are extracted from the input sequence of video footage of the user taken during the monitoring period (step S11), and the variable k is initialized to 1 (step S12), and the input sequence Ti{P i ,f i ,am i ,dx i ,v i ,α i ,e i} and the template group T{T jThe k-th template T among |j=0,1,2··N} k The similarity level with the k-th template is calculated (step S13). Then, the similarity levels calculated for the k-th template are aggregated (step S14). Step S15 determines whether k is less than the total number of templates Tmax. If it is less than Tmax (Yes in step S15), k is incremented (step S16) and the process returns to step S13. The loop continuation requirement for step S15 is that k falls below kmax. Therefore, the loop continues from step S13 until i becomes kmax-1, and the loop from steps S13 to S16 continues. When k reaches kmax, step S15 becomes No, and the loop from steps S13 to S16 is exited. Next, the template with the highest aggregated similarity level is selected (step S17), and the emotion corresponding to the selected template is set as the operator's emotion (step S18). If this emotion is negative, the system determines that individual intervention is necessary.
[0132] [7] Variant Although the present invention has been described above based on embodiments, it goes without saying that the present invention is not limited to the embodiments described above, and the following modifications are possible.
[0133] (1) This embodiment provides a separate device for determining whether individual action is required, but is not limited to this. That is, the image forming apparatus may be equipped with a camera 20 for photographing the operator, and the skeleton detection application 110 and the action requirement determination application 120 may be installed on the image forming apparatus to determine whether individual action is required. Alternatively, the skeleton detection application 110 and the action requirement determination application 120 may be installed on a server computer or edge computer to determine whether individual action is required.
[0134] (2) For skeletal detection, it is desirable to use the GPU208 to construct a circuit block (PAF block) that processes PAF (PartAffiliationFields) which shows the relationships between joint points, and a circuit block (CM block) that processes CM (CofidenceMap) which shows the likelihood of joint points, and then perform CNN inference operations on CM and PAF.
[0135] A PAF block consists of multiple stages, and a CMAP block also consists of multiple stages. Each stage contains multiple arithmetic units, and the arithmetic units in one stage are connected to the arithmetic units in the preceding stage by edges, forming a many-to-one connection. Weight coefficients are also assigned to these edges.
[0136] In the PAF block, the first-stage arithmetic unit performs CNN inference operations only on the feature quantities F extracted from the image, and outputs the calculation result P1.
[0137] The t-th stage arithmetic unit from the second stage onward performs CNN inference on the feature quantities F extracted from the image and the result Pt-1 from the previous inference operation, and outputs the result Pt. By repeating this CNN inference operation, the PAF is refined.
[0138] In the CMAP block, the first stage arithmetic unit takes the feature vector F extracted from the image and the output of the final stage of the PAF block as input, performs a CNN inference operation, and outputs the calculation result C1.
[0139] The arithmetic unit in the sth stage (from the second stage onward) takes the feature vector F extracted from the image, the result Cs-1 from the previous stage's inference operation, and the output of the final stage of the PAF block as input, performs the sth stage CNN inference operation, and outputs the result Cs.
[0140] The CNN inference operation involves performing a many-to-one sum-of-products operation by multiplying the values of multiple arithmetic units from the previous stage by edge-specific weight coefficients and summing them up. A bias value is then added to the sum-of-products value, and the result is substituted into the activation function. The PAF and CMap are refined by performing CNN inference operations in both the CMAP and PAF blocks. Gaussian functions and other similar functions can be used as activation functions.
[0141] (3) Skeletal information is generated by extracting joint points from the entire body, but is not limited to this. The joint points that are subject to the determination of whether individual adjustments are necessary are the wrist joint points and the facial joint points. Of the joint points of the operator's body that appear in the captured images taken around the target device, the joint points that are used to determine whether individual adjustments are necessary are these wrist joint points and facial joint points. Therefore, it is possible to select which joint points to include in the joint point information and exclude the other joint points.
[0142] Furthermore, the reliability of the joint point information in the skeletal data is shown in the reliability score for joint point information, or Cmap. Joint points with a low probability shown in the Cmap may be excluded from the determination of whether individual attention is necessary.
[0143] (4) If multiple target devices are installed in a store and the same operator is operating each of them, the skeletal detection application 110 extracts only the joint points from each person and connects the joint points with edges. Prior to this connection, the joint points are grouped based on the distance between them, etc., to prevent confusion between the joint points of one person and the joint points of another person. At this time, the human body label in the joint point information indicates which target device the operator is operating. The posture and movements of each operator operating the target device may be represented according to the human body label in the joint point information, and the need for individual handling may be determined.
[0144] (5) Since the image recording unit 206 takes pictures of a fixed location within the store, it is acceptable to perform differential extraction recording and extract only the parts in which the operator is present. Differential extraction recording is a recording method implemented in surveillance cameras, etc., which extracts and records only the parts of the video signal that have changed from the background image when compared with the previous frame. Then, only the period of action is extracted, and the period for which skeletal information should be extracted is selected. Which period of action should be selected? Firstly, the period of action before the static period (called the pre-freeze period of action) is selected. This is because the pre-freeze period of action is the period in which the operator is performing the preceding operations, and is considered to represent the operator's skill level. If the operator's actions are stopped for a long time after such an action period, it can be inferred that the operator is in a frozen state, that is, that the operator has forgotten how to operate and is frozen. Therefore, the operator's skill level is calculated during the pre-static period of action, and if the skill level is low, a judgment is made that individual attention is needed for the operator due to the progression of the frozen state.
[0145] Secondly, we select the period of action that follows the static period (referred to as the post-freeze action period). The post-freeze action period is the period during which the operator, having encountered difficulties in the operation, finally expresses feelings of confusion. Therefore, we determine whether the operator exhibits a confused posture or behavior during the post-freeze action period.
[0146] Furthermore, in differential extraction recording, it is desirable to calculate the speed of the operator appearing in multiple frames and select frames where that speed is above a predetermined lower limit and below a predetermined upper limit for detection of skeletal information. Since the target of skeletal information detection can be narrowed down at the stage when performing recording by differential extraction recording, the operating load of the response necessity determination device 30 can be reduced.
[0147] (6) The response necessity determination device 30 has determined whether individual response is necessary, but it is not necessary to actually call a store employee. It may be sufficient to simply generate statistical data indicating that individual response was necessary. By referring to such statistical data, the business operator running the store can analyze the operator's behavior to determine under what circumstances individual response by a store employee became necessary. The statistical data should preferably include the basic information that formed the basis of the determination that individual response was necessary, and a time code indicating the date and time when individual response was determined to be necessary.
[0148] Depending on whether the priority is on collecting statistical data or on the speed of individual responses, it is desirable to change the threshold for the quiescent period as shown in the first embodiment.
[0149] (7) The program for determining whether individual intervention is necessary may determine whether the user is experiencing difficulties based on the user's posture and movements captured by the camera, such as by comparing them with a human body model of the user experiencing difficulties. The human body model is a composite body in which rigid body segments are connected at joint points. The learned human body model for each posture shows the joint angles of each of the 18 joint points when the human body model assumes a specific posture. When skeletal information is generated by the skeletal detection application 110, the joint angles are calculated from the joint point information in the skeletal information, and an input system consisting of those joint angles is obtained. Then, the learned human body model for each posture that is closest to the input system is selected and set as the corresponding posture.
[0150] (8) The device for determining whether action is required in relation to this disclosure is to be installed at the counter of a store, but is not limited to that. It may also be installed in the staff waiting room, break room, or security room of a store. It may also be installed in any facility other than a store that is visited by an unspecified number of people. It may be installed in a company's service center, waiting room, the counter of a financial institution, the counter of a government office or administrative agency, etc.
[0151] (9) The application 120 for determining whether assistance is needed made a determination, but is not limited to, whether the person was in a confused posture or movement, turning around, etc. It also determined whether there had been an illegal act against the image forming machine, and if there had been an illegal act, it determined that individual assistance was needed and instructed the store clerk to contact the police or the security company.
[0152] The system detects crouching or squatting in front of the image forming apparatus 10, kicking or hitting the image forming apparatus 10, and if any of these actions occur, it is highly likely that an illegal act such as property damage has been committed. Whether or not a person is crouching in front of the image forming apparatus 10 can be determined by checking whether the joint points of the left and right knees are higher than the joint points of the left and right hips. This allows the system to detect acts such as removing paper from the paper feed unit 14 of the image forming apparatus 10 or replacing paper without permission.
[0153] Whether or not the motion is a kicking motion of the image forming apparatus 10 can be determined by whether one of the left or right ankles is swinging downwards and whether the amount of movement is large. Whether or not the action is a tapping motion on the image forming apparatus 10 can be determined by swinging either the left or right wrist downwards and observing the amount of movement.
[0154] (10) In stores, the target devices subject to the determination of whether individual staff intervention is necessary are not limited to image forming machines. They may also be automated teller machines (ATMs). This can also be applied to determinations regarding operational support and crime prevention in ATMs. Appropriate measures may also be taken against illegal acts against ATMs or illegal acts using the target devices. Such measures may include reporting to the police or security company, or contacting financial institutions. Furthermore, in the case of ATMs, the determination may be made that individual intervention is necessary regarding the act of making a phone call on a mobile phone in front of the ATM, because there is a high possibility that a special fraud is occurring. [Industrial applicability]
[0155] The device for determining whether action is necessary, as disclosed herein, can create a new type of service that integrates artificial intelligence-based skeletal detection with a device, and has the potential to be used in various industrial fields, including office automation equipment and information equipment, as well as retail, rental, real estate, advertising, transportation, and publishing. [Explanation of symbols]
[0156] 10 Image forming apparatus 30. Device for determining whether action is necessary. 100 Operating Systems 110 Skeleton Detection Applications 120 Application for determining whether action is required 121 Confusion Detection Program 122 Direction Detection Program 123 Movement Calculation Program 124 Stationary Pattern Detection Program 125 Confusion Pattern Determination Program 126 Program to display the degree of need for support 201 CPU 202 Boot ROM 203 HDD 204 RAM 205 Interrupt control circuit 206 Image Recording Unit 207 Peripheral control circuits
Claims
1. A device for determining whether staff need to respond to an operator operating a target device, An acquisition means for acquiring a captured image of the location where the target device is installed, A skeletal information generation means that generates skeletal information from a captured image in which the operator is visible, A means of determining whether or not staff intervention is necessary based on skeletal information. Equipped with, The aforementioned necessity determination means is From the captured images of the area surrounding the target device, select the joint points of the operator's body that are suitable for determining whether or not a response is necessary, and do not include the other joint points in the skeletal information. Furthermore, if the skeletal information determines that the operator's face is rotated at a predetermined angle relative to their lower body, resulting in a turning posture, the system will determine that action is required. A device for determining whether or not a response is necessary, characterized by the above.
2. A device for determining whether staff need to respond to an operator operating a target device, An acquisition means for acquiring a captured image of the location where the target device is installed, A skeletal information generation means that generates skeletal information from a captured image in which the operator is visible, A means of determining whether or not staff intervention is necessary based on skeletal information. Equipped with, The aforementioned necessity determination means is From the captured images of the area surrounding the target device, select the joint points of the operator's body that are suitable for determining whether or not a response is necessary, and do not include the other joint points in the skeletal information. The aforementioned skeletal information generation means generates skeletal information for each of the multiple frames, and identifies the operator's actions from the skeletal information generated for each frame. The aforementioned necessity determination means is The system calculates the rotation angle between the plane passing through the joint points of the face, as shown in the skeletal information generated from frame 1, and the plane passing through the joint points of the lower body. If the calculated rotation angle exceeds a predetermined threshold, the system determines that action is required. A device for determining whether or not a response is necessary, characterized by the above.
3. A device for determining whether staff need to respond to an operator operating a target device, An acquisition means for acquiring a captured image of the location where the target device is installed, A skeletal information generation means that generates skeletal information from a captured image in which the operator is visible, A means of determining whether or not staff intervention is necessary based on skeletal information. Equipped with, The aforementioned necessity determination means is From the captured images of the area surrounding the target device, select the joint points of the operator's body that are suitable for determining whether or not a response is necessary, and do not include the other joint points in the skeletal information. Furthermore, the system calculates the degree to which staff intervention is necessary based on the rotation angle between the operator's face and their lower body, and notifies the staff accordingly. A device for determining whether or not a response is necessary, characterized by the above.
4. A device for determining whether staff need to respond to an operator operating a target device, An acquisition means for acquiring a captured image of the location where the target device is installed, A skeletal information generation means that generates skeletal information from a captured image in which the operator is visible, A means of determining whether or not staff intervention is necessary based on skeletal information. Equipped with, The aforementioned necessity determination means is From the captured images of the area surrounding the target device, select the joint points of the operator's body that are suitable for determining whether or not a response is necessary, and do not include the other joint points in the skeletal information. The aforementioned skeletal information generation means generates skeletal information for each of the multiple frames, and identifies the operator's actions from the skeletal information generated for each frame. The aforementioned necessity determination means is Furthermore, the system calculates the amount of movement of the skeletal information generated for a given frame compared to the skeletal information generated for the previous frame. If the calculated amount of movement is small, and the period of stillness with a small amount of movement continues for a predetermined threshold or longer, the system determines that action is required. Furthermore, while a specific part or all parts of the operator's body are moving, the system performs a process to sum the movement amounts of the joint points included in that specific part or all parts of the operator's body. The operator's skill level is determined by comparing the total amount of movement with a threshold value, and if the determined skill level is that of a beginner, the predetermined threshold value is changed to a shorter value. A device for determining whether or not a response is necessary, characterized by the above.
5. A device for determining whether staff need to respond to an operator operating a target device, An acquisition means for acquiring a captured image of the location where the target device is installed, A skeletal information generation means that generates skeletal information from a captured image in which the operator is visible, Based on the skeletal position information shown in the skeletal information, if the operator's posture corresponds to a confused pose, a determination means is made to determine whether staff intervention is necessary. Equipped with, The necessity determination means determines from the skeletal information that the operator's face is rotated at a predetermined angle relative to the lower body and is in a turning posture, and if so, the operator's posture corresponds to a confused pose. A device for determining whether or not a response is necessary, characterized by the above.
6. A device for determining whether staff need to respond to an operator operating a target device, An acquisition means for acquiring a captured image of the location where the target device is installed, A skeletal information generation means that generates skeletal information from a captured image in which the operator is visible, A means of determining whether or not staff intervention is necessary based on skeletal information. Equipped with, The aforementioned skeletal information generation means generates skeletal information for each of the multiple frames, and identifies the operator's actions from the skeletal information generated for each frame. The aforementioned necessity determination means is The system calculates the rotation angle between the plane passing through the joint points of the face, as shown in the skeletal information generated from frame 1, and the plane passing through the joint points of the lower body. If the calculated rotation angle exceeds a predetermined threshold, the system determines that the operator's actions constitute confusing actions and that action is required. A device for determining whether or not a response is necessary, characterized by the above.
7. The aforementioned necessity determination means calculates the degree to which staff intervention is necessary, based on the rotation angle of the operator's face relative to their lower body, and notifies the staff. A device for determining whether a response is necessary, as described in any of claims 1, 2, or 4 to 6.
8. The acquisition means acquires captured images at predetermined intervals, and the skeletal information generation means generates skeletal information for each captured image. The notification by the aforementioned necessity determination means is made by displaying the continuous time change in the degree of need for action calculated from the skeletal information generated in the order of shooting. The device for determining whether a response is necessary, according to claim 3 or 7.
9. The aforementioned necessity determination means determines whether or not a response is necessary based on the calculated degree of need for a response, The notification by the aforementioned necessity determination means is made by plotting the necessity of individual action at each of multiple points in time on the time axis, thereby displaying the discrete time changes in the necessity of action. The device for determining whether a response is necessary, as described in claim 3 or 7.
10. The aforementioned skeletal information represents the positions of the joints connecting the body segments, using coordinates in a coordinate system with the neck as the origin. A device for determining whether a response is necessary, as described in any one of claims 1 to 9.
11. The aforementioned skeletal information represents the positions of the joints connecting the body segments, using coordinates in the coordinate system of the camera's imaging plane. A device for determining whether a response is necessary, as described in any one of claims 1 to 9.
12. The acquisition means captures multiple frames of images obtained by the camera, The system includes a recording means for recording image data of multiple frames acquired by the acquisition means onto a recording medium, The skeletal information generated by the skeletal information generation means is recorded on the recording medium together with the image data. A device for determining whether or not a response is necessary, as described in any of claims 1 to 11.
13. The recording means identifies from multiple frames of image data recorded on the recording medium the period during which the operator's body is still and the period during which the operator's body is in motion. The aforementioned skeletal information generation means generates skeletal information from frames in which the operator's body is in motion. The device for determining whether or not a response is necessary, as described in feature 12.
14. The speed of the operator's body movements is detected from multiple frames recorded on the recording medium. The aforementioned skeletal information generation means generates skeletal information from image data of frames in which the speed of the operator's body movements meets a predetermined standard. The device for determining whether or not a response is necessary, as described in feature 13.
15. The aforementioned skeletal information generation means generates skeletal information from the image data of each frame if the frame immediately preceding the current time, or multiple frames immediately preceding the current time, constitute the operating period. The device for determining whether or not a response is necessary, as described in feature 14.
16. If the captured image includes the bodies of multiple operators operating each of the multiple target devices, the skeletal information generation means generates skeletal information corresponding to each operator, and the necessity determination means determines whether or not a response is necessary for each operator. A device for determining whether or not a response is necessary, as described in any of claims 1 to 15.
17. When the aforementioned skeletal information generation means generates skeletal information corresponding to each operator, the joint point information included in the skeletal information includes information that uniquely identifies the corresponding operator among multiple operators. The device for determining whether or not a response is necessary, as described in feature 16.
18. The aforementioned response refers to the act of providing appropriate support for the operation of the target device. A device for determining whether or not a response is necessary, as described in any one of claims 1 to 17.
19. The aforementioned response refers to taking appropriate measures against torts against the target device or torts using the target device. A device for determining whether or not a response is necessary, as described in any one of claims 1 to 17.
20. The aforementioned target device is either an image forming apparatus or an automated teller machine (ATM). A device for determining whether a response is necessary, as described in any one of claims 1 to 19.
21. The aforementioned skeletal information generation means extracts features from the captured image, and based on these features, estimates the relationships between the features and estimates the locations that appear to be joint points of the human body. The generation of the aforementioned skeletal information is This process includes combining the estimated joint point locations with the estimated correlations to determine which part of the human body each joint point corresponds to. A device for determining whether or not a response is necessary, as described in any of claims 1 to 20.
22. The estimation of the relationships between the aforementioned joint points is an inference operation using a convolutional neural network that targets information indicating the relationships between features. The estimation of the aforementioned joint point-like locations is an inference operation using a convolutional neural network that targets probability information indicating how likely it is to be a joint point. The device for determining whether or not a response is necessary, as described in feature 21.
23. A program for determining whether staff need to respond to an operator operating a target device, which causes a computer to determine whether staff need to respond to the operator operating the device. The acquisition step involves acquiring an image of the location where the aforementioned target device is installed, A skeletal information generation step that generates skeletal information from a captured image in which the operator is visible, A necessity determination step that determines whether or not staff intervention is necessary based on skeletal information. A program that causes a computer to execute a program to determine whether or not action is necessary, In the aforementioned necessity determination step, among the joint points of the operator's body that appear in the captured images taken around the target device, those suitable for determining the necessity of action are selected, and the other joint points are not included in the skeletal information. Furthermore, if the skeletal information determines that the operator's face is rotated at a predetermined angle relative to their lower body, resulting in a turning posture, the system will determine that action is required. A program for determining whether action is necessary, characterized by the following features.
24. A program for determining whether staff need to respond to an operator operating a target device, wherein the program causes a computer to determine whether staff need to respond to an operator operating the target device, The acquisition step involves acquiring an image of the location where the aforementioned target device is installed, A skeletal information generation step that generates skeletal information from a captured image in which the operator is visible, A necessity determination step that determines whether or not staff intervention is necessary based on skeletal information. A program that causes a computer to execute a program to determine whether or not action is necessary, In the aforementioned necessity determination step, among the joint points of the operator's body that appear in the captured images taken around the target device, those suitable for determining the necessity of action are selected, and the other joint points are not included in the skeletal information. In the aforementioned skeletal information generation step, skeletal information is generated for each of the multiple frames, and the operator's actions are identified from the skeletal information generated for each frame. In the aforementioned necessity determination step, The system calculates the rotation angle between the plane passing through the joint points of the face, as shown in the skeletal information generated from frame 1, and the plane passing through the joint points of the lower body. If the calculated rotation angle exceeds a predetermined threshold, the system determines that action is required. A program for determining whether action is necessary, characterized by the following features.
25. A program for determining whether staff need to respond to an operator operating a target device, wherein the program causes a computer to determine whether staff need to respond to an operator operating the target device, The acquisition step involves acquiring an image of the location where the aforementioned target device is installed, A skeletal information generation step that generates skeletal information from a captured image in which the operator is visible, A necessity determination step that determines whether or not staff intervention is necessary based on skeletal information. A program that causes a computer to execute a program to determine whether or not action is necessary, In the aforementioned necessity determination step, among the joint points of the operator's body that appear in the captured images taken around the target device, those suitable for determining the necessity of action are selected, and the other joint points are not included in the skeletal information. Furthermore, the system calculates the degree to which staff intervention is necessary based on the rotation angle between the operator's face and their lower body, and notifies the staff accordingly. A program for determining whether action is necessary, characterized by the following features.
26. A program for determining whether staff need to respond to an operator operating a target device, wherein the program causes a computer to determine whether staff need to respond to an operator operating the target device, The acquisition step involves acquiring an image of the location where the aforementioned target device is installed, A skeletal information generation step that generates skeletal information from a captured image in which the operator is visible, A necessity determination step that determines whether or not staff intervention is necessary based on skeletal information. A program that causes a computer to execute a program to determine whether or not action is necessary, In the aforementioned necessity determination step, among the joint points of the operator's body that appear in the captured images taken around the target device, those suitable for determining the necessity of action are selected, and the other joint points are not included in the skeletal information. In the aforementioned skeletal information generation step, skeletal information is generated for each of the multiple frames, and the operator's actions are identified from the skeletal information generated for each frame. In the aforementioned necessity determination step, Furthermore, the system calculates the amount of movement of the skeletal information generated for a given frame compared to the skeletal information generated for the previous frame. If the calculated amount of movement is small, and the period of stillness with a small amount of movement continues for a predetermined threshold or longer, the system determines that action is required. Furthermore, while a specific part or all parts of the operator's body are moving, the system performs a process to sum the movement amounts of the joint points included in that specific part or all parts of the operator's body. The operator's skill level is determined by comparing the total amount of movement with a threshold value, and if the determined skill level is that of a beginner, the predetermined threshold value is changed to a shorter value. A program for determining whether action is necessary, characterized by the following features.
27. A program for determining whether staff need to respond to an operator operating a target device, which causes a computer to determine whether staff need to respond to the operator operating the device. The acquisition step involves acquiring an image of the location where the aforementioned target device is installed, A skeletal information generation step that generates skeletal information from a captured image in which the operator is visible, Based on the skeletal position information shown in the skeletal information, if the operator's posture corresponds to a confused pose, a determination step is made to determine whether staff intervention is necessary. A program that causes a computer to execute a program to determine whether or not action is necessary, The program for determining whether action is necessary is characterized in that, in the aforementioned necessity determination step, if it is determined from skeletal information that the operator's face is at a predetermined rotation angle relative to the lower body and the operator is in a turning posture, the operator's posture is deemed to correspond to a confused pose.
28. A program for determining whether staff need to respond to an operator operating a target device, which causes a computer to determine whether staff need to respond to the operator operating the device. The acquisition step involves acquiring an image of the location where the aforementioned target device is installed, A skeletal information generation step involves generating skeletal information for each of the multiple frames of the captured image in which the operator is visible, and identifying the operator's movements from the skeletal information generated for each frame. A necessity determination step that determines whether or not staff intervention is necessary based on skeletal information. A program that causes a computer to execute a program to determine whether or not action is necessary, In the aforementioned necessity determination step, the program calculates the rotation angle between the plane passing through the joint points of the face, as shown in the skeletal information generated from one frame, and the plane passing through the joint points of the lower body. If the calculated rotation angle exceeds a predetermined threshold, the program determines that the operator's actions constitute confusing actions and that action is necessary.
Citation Information
Patent Citations
Device and program for estimating confusion
JP2010117964A
Behavior recognition device and behavior recognition program
JP2017228100A
Customer service monitoring device, customer service monitoring system, and customer service monitoring method
JP2018022284A
Customer service necessity determination apparatus, customer service necessity determination method, and program
JP2018190012A
User operation support device and user operation support program
JP2019101775A