Image processing device, image processing method and program
The image processing device enhances fraud detection by identifying user areas and changes in users using skeletal feature points and posture analysis, addressing inaccuracies in low-performance surveillance cameras to improve fraud detection accuracy.
Patent Information
- Application Number
- JP2024203782
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-19
- Filing Date
- 2024-11-22
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-12-15
AI Technical Summary
Surveillance cameras with low performance or inappropriate installation locations struggle to accurately record facial details and user actions, leading to inaccuracies in fraud detection, particularly in bank transfer fraud scenarios.
An image processing device that identifies a user area within images and detects changes in users using skeletal feature points and posture analysis, even with low-performance cameras, to accurately calculate usage time and detect potential fraud.
Accurately detects fraud victims or potential fraud victims with high accuracy despite camera limitations, preventing false alarms and improving fraud detection efficiency.
Smart Images

Figure 0007810239000001 
Figure 0007810239000002 
Figure 0007810239000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, and a program. [Background technology]
[0002] There is a demand for technology to prevent the damage caused by bank transfer fraud. Related technologies are disclosed in Patent Documents 1 and 2. Patent Documents 1 and 2 disclose technology that analyzes images generated by a surveillance camera installed on an operating terminal such as an ATM (automated teller machine) to identify a person and determine whether the identified person is talking on a mobile phone. Patent Document 1 also discloses technology that determines whether a person who has been using an operating terminal for a long period of time is a victim of bank transfer fraud or is likely to be a victim of bank transfer fraud. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-238204 [Patent Document 2] Japanese Patent Application Laid-Open No. 2010-218392 Summary of the Invention [Problem to be solved by the invention]
[0004] It is known that victims of frauds such as bank transfer fraud tend to operate their operating terminals while talking on a mobile phone, or use the terminal for a long period of time. As disclosed in Patent Documents 1 and 2, detecting people engaging in these behaviors through image analysis can reduce fraud damage. However, the inventors have discovered the following new problem with this technology.
[0005] Some surveillance cameras installed to film users of operation terminals may have low performance (e.g., low frame rate, low resolution, etc.) and may not be able to clearly record the facial details or actions of the user of the operation terminal. In addition, some cameras may be installed in a position and orientation that films the user of the operation terminal from above or diagonally above, and may not be able to record the facial details of the user of the operation terminal to an extent that facial recognition technology can accurately recognize them. Because the number of installed operation terminals is enormous, the work of replacing all surveillance cameras with high-performance surveillance cameras or changing their installation positions and orientations is a heavy burden.
[0006] The present invention aims to provide a technology that can accurately detect fraud victims or potential fraud victims based on images generated by surveillance cameras that have limitations in performance and installation location. [Means for solving the problem]
[0007] According to the present invention, a user area specifying means for specifying a user area, which is an area where a user of the operation terminal exists, from within the image to be processed; a user change detection means for detecting a change in the user of the operation terminal based on the image of the user area; An image processing apparatus is provided, comprising:
[0008] Further, according to the present invention, The computer A user area is identified from the image to be processed, which is an area where the user of the operating terminal is present; An image processing method is provided for detecting a change in the user of the operation terminal based on an image of the user area.
[0009] Further, according to the present invention, Computer, a user area specifying means for specifying a user area, which is an area where a user of the operation terminal exists, from within the image to be processed; a user change detection means for detecting a change in the user of the operation terminal based on the image of the user area; A program is provided to function as a [Effects of the Invention]
[0010] According to the present invention, a technology is realized that can detect fraud victims or potential fraud victims with high accuracy based on images generated by surveillance cameras that have limitations in performance and installation location. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram for explaining an image generated by a surveillance camera according to an embodiment of the present invention; [Figure 2] 1 is a diagram for explaining an image generated by a surveillance camera according to an embodiment of the present invention; [Figure 3] 1 is a diagram for explaining an image generated by a surveillance camera according to an embodiment of the present invention; [Figure 4] FIG. 2 is a functional block diagram of an image processing apparatus according to an embodiment of the present invention; [Figure 5] FIG. 10 is a diagram illustrating an example of a result of detecting a person area. [Figure 6] 10A and 10B are diagrams for explaining a process for identifying a user area according to the present embodiment. [Figure 7] 10A and 10B are diagrams for explaining a process for detecting a change of a user according to the present embodiment. [Figure 8] 10 is a flowchart showing an example of a processing flow of the image processing apparatus of the present embodiment. [Figure 9] FIG. 1 is a diagram illustrating an example of a hardware configuration of an image processing apparatus according to an embodiment of the present invention. [Figure 10] FIG. 2 is a functional block diagram of an image processing apparatus according to an embodiment of the present invention; [Figure 11] 10A and 10B are diagrams for explaining a call posture detection process according to the present embodiment; [Figure 12] 10A and 10B are diagrams for explaining a process of calculating a certainty factor that a user is taking a predetermined posture according to the present embodiment. [Figure 13] 10 is a flowchart showing an example of a processing flow of the image processing apparatus of the present embodiment. [Figure 14] 10 is a flowchart showing an example of a processing flow of the image processing apparatus of the present embodiment. [Figure 15] FIG. 2 is a functional block diagram of an image processing apparatus according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0013] First Embodiment "Images generated by surveillance cameras" The image processing device of this embodiment analyzes images generated by a surveillance camera installed to capture images of users of operation terminals such as ATMs, and accurately detects fraud victims or potential victims. Fraud includes, but is not limited to, bank transfer fraud.
[0014] Here, we will explain the images generated by a surveillance camera. Surveillance cameras have limitations in performance and installation location. For this reason, surveillance cameras cannot clearly record the facial details and behavior of the user of the operation terminal.
[0015] For example, cameras with low performance (e.g., low frame rate, low resolution, etc.) are used as surveillance cameras, and in this case, the surveillance cameras cannot clearly record the facial details and actions of the user of the operation terminal.
[0016] Furthermore, for example, as shown in FIG. 1, the surveillance camera 100 is installed in a position and orientation that allows it to capture an image of the user 101 of the operation terminal 102 from above or diagonally above. That is, the surveillance camera 100 may be installed in a position and orientation that prevents it from capturing an image of the face of the user 101 of the operation terminal 102 from the front. In this case, the surveillance camera 100 cannot record enough details of the face of the user 101 of the operation terminal 102 to allow accurate recognition by facial recognition technology. Furthermore, the generated image may include people and objects other than the user 101. For example, as shown in FIGS. 2 and 3, the image F generated by the surveillance camera 100 may include, in addition to the user 101 in front of the operation terminal 102, the operation terminal 102, other people 103 such as people waiting in line or passersby, and other objects 104 such as walls and partitions.
[0017] "Image Processing Device Overview" Next, an overview of the image processing device of this embodiment will be described. The image processing device of this embodiment executes a process for detecting fraud victims or potential fraud victims with high accuracy based on images generated by a surveillance camera. Specifically, the image processing device of this embodiment executes, based on the images generated by the surveillance camera, "a process for identifying a user area, which is an area where the user of the operating terminal exists, from the image to be processed" and "a process for detecting that the user of the operating terminal has changed, based on the image of the user area."
[0018] One of the main features of the image processing device of this embodiment is that it executes a process to detect when the user of the operating terminal has changed. Based on the detection result, the usage time of the operating terminal for each user can be calculated.
[0019] If the surveillance camera has high performance (e.g., high frame rate, high resolution, etc.) and is installed in an appropriate position and orientation, it can use well-known tracking technology or face recognition technology to distinguish multiple users detected in an image from one another with high accuracy. Therefore, there is no need to go to the trouble of executing a process to detect a change in the user of the operating terminal. In fact, the technologies described in Patent Documents 1 and 2, which are considered to be based on the premise that the surveillance camera has high performance and is installed in an appropriate position and orientation, do not execute a "process to detect a change in the user of the operating terminal."
[0020] However, if the surveillance camera has low performance (e.g., low frame rate, low resolution, etc.) or is installed in an inappropriate location, as in this embodiment, it becomes difficult to accurately distinguish between multiple users detected in an image using well-known tracking or facial recognition technology. As a result, different users may be mistaken for the same user, or the same user appearing in multiple images may be mistaken for different users. As a result, the accuracy of the calculation results for the usage time of each user's operating device may deteriorate.
[0021] Therefore, the image processing device of this embodiment performs a process for detecting a change in the user of an operating terminal, which is not performed in conventional technology, and calculates the usage time of each user's operating terminal based on the detection results. As a result, even if the surveillance camera has limitations in performance or installation location, it becomes possible to accurately calculate the usage time of each user's operating terminal.
[0022] Another main feature of the image processing device of this embodiment is that all of the above processes performed by the image processing device are characterized by content suitable for processing images generated by surveillance cameras that have limitations in performance and installation location. Therefore, the accuracy of the above processes is high even when the surveillance cameras have limitations in performance and installation location. As a result, the usage time of each user's operating terminal can be calculated with high accuracy.
[0023] "Functional configuration of image processing device" Next, the functional configuration of the image processing device will be described in detail. Fig. 4 shows an example of a functional block diagram of the image processing device 10 of this embodiment. As shown in the figure, the image processing device 10 has a user area identification unit 11, a user switching detection unit 12, and an output unit 13.
[0024] The user area specifying unit 11 specifies a user area, which is an area in which the user of the operating terminal is present, from within the image to be processed.
[0025] The "image to be processed" is an image generated by the above-mentioned surveillance camera. The surveillance camera is installed in a position and orientation that captures the user of the operating terminal, and generates a video. The multiple images contained in this video are the images to be processed in chronological order (in the order they were captured).
[0026] Next, a process for identifying a user area from within an image to be processed will be described. The user area identification unit 11 detects a person area from within the image to be processed, and identifies one of the detected person areas as the user area. This will be described in detail below.
[0027] -Processing to detect human regions from within the image to be processed- The process of detecting a person region from an image to be processed can be realized using any known person detection technology. For example, it may be realized using a model for detecting people generated by machine learning, or it may be realized by other means. In known person detection technology, for example, a rectangular region including a person is detected as a person region.
[0028] As mentioned above, surveillance cameras have limitations in performance and installation location. Therefore, the accuracy of the process for detecting the person region may be insufficient. As a result, a part of a person or an object 104 other than a person may be mistakenly recognized as a person. Furthermore, as mentioned above, when a surveillance camera captures an image from above or diagonally above, it generates an image that includes other people 103. As a result, the process for detecting the person region may also detect other people 103. As a result, in the process for detecting a person region from the image to be processed, as shown in FIG. 5, in addition to the person region W1 including the user 101, a person region W2 including other people 103, a person region W3 including only a part of a person, or a person region W4 including other objects 104 may be detected. Therefore, the user region identification unit 11 executes a "process for identifying one of one or more detected person regions as a user region," which will be described in detail below, and identifies an appropriate user region (a region including the user 101) from among the detected person regions.
[0029] -Process of identifying one of the detected person areas as the user area- In this process, the user area identification unit 11 uses a first detection result that detects a person area from the image to be processed, and a second detection result that detects feature points of the person's skeleton from the image to be processed.
[0030] The first detection result includes the size and certainty of each detected person region. The size of the person region is the size of the area occupied by the person region and can be expressed, for example, by the number of pixels. The certainty is a value indicating the degree of certainty that the person region is an area containing a person (a measure of how certain the result is). Techniques for calculating such certainty are widely known in well-known person detection technologies.
[0031] The second detection result includes coordinates (information indicating a position within the image) of each of a plurality of feature points of the human skeleton detected from the image to be processed. The detection of the feature points of the human skeleton is realized using a well-known technique such as OpenPose.
[0032] The user area identification unit 11 uses the first detection result and the second detection result to identify one of the detected person areas that is most suitable as one that includes the user of the operating terminal. Specifically, the user area identification unit 11 applies the determination method shown in Fig. 6 to each of the detected person areas as a processing target, and determines whether or not each person area is a user area.
[0033] In S100, the user area identification unit 11 determines whether the size of the person area to be processed is larger than a preset threshold. If the size of the person area to be processed is smaller than the threshold ("small" in S100), the user area identification unit 11 determines that the person area to be processed is not a user area (S102).
[0034] If the size of the person region to be processed is larger than the threshold value ("large" in S100), the user region identification unit 11 determines whether at least one person's skeletal feature point has been detected in the image to be processed that includes the person region to be processed (S101). If any person's skeletal feature points have been detected ("detected" in S101), the user region identification unit 11 determines whether the number of person's skeletal feature points included in each of the one or more detected person regions is the largest (S103).
[0035] If it is the largest ("YES" in S103), the user area identification unit 11 determines that the person area to be processed is a user area (S110). On the other hand, if it is not the largest ("NO" in S103), the user area identification unit 11 determines that the person area to be processed is not a user area (S105).
[0036] If no skeletal feature points of a person are detected in the image to be processed that includes the person area to be processed ("Not detected" in S101), the user area identification unit 11 determines whether at least one skeletal feature point of a person is detected in the reference image (S104).
[0037] The "reference image" is an image generated before the image to be processed that includes the person area to be processed, and includes the user of the operating terminal included in the image to be processed. The reference image changes dynamically. An example of the process for determining the reference image is described below.
[0038] If human skeletal feature points are detected in the reference image ("Yes" in S104), the user region identification unit 11 determines the inclusion relationship between the human region detected in the processing target image and the human skeletal feature points detected in the reference image, based on the area occupied by each of the detected one or more human regions in the image and the coordinates within the image of each of the human skeletal feature points detected in the reference image.The user region identification unit 11 then determines whether the human region to be processed contains the largest number of human skeletal feature points among the detected one or more human regions (S106).
[0039] If it is the largest ("YES" in S106), the user area identification unit 11 determines that the person area to be processed is a user area (S110). On the other hand, if it is not the largest ("NO" in S106), the user area identification unit 11 determines that the person area to be processed is not a user area (S107).
[0040] If no skeletal feature points of a person are detected in the reference image ("No" in S104), the user area identification unit 11 determines whether the person area to be processed has the highest confidence level (confidence level indicated by the first detection result) among the one or more detected person areas (S108).
[0041] If it is the largest ("YES" in S108), the user area identification unit 11 determines that the person area to be processed is a user area (S110). On the other hand, if it is not the largest ("NO" in S108), the user area identification unit 11 determines that the person area to be processed is not a user area (S109).
[0042] In this way, the user area identification unit 11 can identify a user area from the image to be processed based on the first detection result of detecting a person area from the image to be processed and the second detection result of detecting skeletal feature points of a person from the image to be processed. Furthermore, if skeletal feature points of a person are not detected from the image to be processed, the user area identification unit 11 can identify a user area from the image to be processed based on skeletal feature points of a person detected from a reference image that was generated before the image to be processed.
[0043] 4, the user switching detection unit 12 detects that the user of the operating terminal has switched, based on the image of the user area identified by the user area identification unit 11. The user switching detection unit 12 detects that the user of the operating terminal has switched (that the user in the image to be processed and the user in the image to be compared are different people) based on the result of comparing feature data extracted from the image of the user area in the image to be processed with feature data extracted from the image of the user area in the image to be compared.
[0044] The "feature data extracted from the image of the user area" is data that indicates the external characteristics of the user shown in the image. Examples include, but are not limited to, clothing characteristics, hairstyle characteristics, facial characteristics, whether or not the user is wearing glasses or a hat, and the characteristics of those.
[0045] The "comparison image" is an image generated before the processing target image. The comparison image changes dynamically. An example of a process for determining the comparison image is described below.
[0046] The above-mentioned "reference image" and "comparison image" have in common the fact that they are both "images generated before the image to be processed." However, they differ in that the "reference image" is an image to be used in the above-mentioned "processing to determine the user area," while the "comparison image" is an image to be used in the above-mentioned "processing to detect a change in the user of the operating terminal."
[0047] As mentioned above, surveillance cameras have limitations in performance and installation location. Therefore, the accuracy of the process of determining whether or not the images represent the same person by comparing the feature data is insufficient. Therefore, the user change detection unit 12 determines that the user of the operating terminal has changed when it is determined that the images are not the same person in M or more consecutive images to be processed (M is an integer equal to or greater than 2). Then, the unit determines that the user of the operating terminal has changed based on the timing when the earliest chronologically processed image among the M consecutive images to be processed was generated.
[0048] This process will be explained below using the specific example shown in Fig. 7. The above-mentioned reference image and comparison image will also be explained. Note that M is set to "16".
[0049] First, the image of frame number 1 becomes the image to be processed, and processing is executed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "◯"). The image processing device 10 sets a person ID (identifier) of "1" for the person (user of the operation terminal) included in the user area, and sets "1" as the number of frames stayed. In addition, the image of frame number 1 is set as the image to be referenced and the image to be compared. Note that since there are no previous frames, processing is not executed by the user switching detection unit 12.
[0050] Next, the image of frame number 2 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "◯"). Then, the user switching detection unit 12 compares feature data extracted from the image of the user area in the image to be processed with feature data extracted from the image of the user area in the image to be compared. If the similarity is equal to or greater than a reference value, the user switching detection unit 12 determines that the person included in the two images is the same person, and if it is less than the reference value, it determines that the person is not the same person. Here, it is assumed that the person is determined to be the same person (identity determination "◯"). The image processing device 10 leaves the person ID as "1" and updates the number of stay frames to "2". Then, the image processing device 10 updates the reference image and the comparison image to the image of frame number 2.
[0051] Next, the image with frame number 3 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has not been identified (person detection "X"). A situation in which the user area is not identified is, for example, when a user area with a size equal to or larger than the threshold is not detected ("small" in S100 of FIG. 6). In this case, processing by the user switching detection unit 12 is not performed. The image processing device 10 leaves the person ID at "1", the number of staying frames at "2", and sets the number of consecutive failures to "1". Then, the image processing device 10 leaves the image with frame number 2 as the reference image and the comparison image.
[0052] Next, the image with frame number 4 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has not been identified (person detection "X"). In this case, processing is not performed by the user switching detection unit 12. The image processing device 10 leaves the person ID as "1", the number of frames stayed as "2", and updates the number of consecutive failures to "2". Then, the image processing device 10 leaves the image with frame number 2 as the reference image and the comparison image.
[0053] Next, the image of frame number 5 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "O"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image of frame number 2. Here, it is assumed that it has been determined that it is the same person (identity determination "O"). The image processing device 10 leaves the person ID as "1" and updates the number of frames stayed to "5". In other words, it is assumed that the person stayed during frames 3 and 4, where person detection failed. Then, the image processing device 10 updates the number of consecutive failures to "0". Furthermore, the image processing device 10 updates the reference image and the comparison image to the image of frame number 5.
[0054] Next, the image with frame number 6 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "◯"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image with frame number 2. Here, it is assumed that it has not been determined that they are the same person (identity determination "X"). The image processing device 10 leaves the person ID as "1", the number of frames stayed as "5", and updates the number of consecutive failures to "1". Furthermore, the image processing device 10 leaves the image with frame number 5 as the reference image and the image to be compared.
[0055] Next, the image of frame number 7 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "◯"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image of frame number 2. Here, it is assumed that it has not been determined that they are the same person (identity determination "X"). The image processing device 10 leaves the person ID as "1", the number of frames stayed as "5", and updates the number of consecutive failures to "2". Furthermore, the image processing device 10 leaves the image of frame number 5 as the reference image and the image to be compared.
[0056] Next, the image with frame number 8 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has not been identified (person detection "X"). In this case, processing is not performed by the user switching detection unit 12. The image processing device 10 leaves the person ID at "1", the number of frames stayed at "5", and updates the number of consecutive failures to "3". Then, the image processing device 10 leaves the image with frame number 5 as the reference image and the comparison image.
[0057] Similar processing is performed thereafter, but in all of the images of frame numbers 9-20, the user area was identified (person detection: "Yes"), but it was not determined that they were the same person (identity determination: "No"). When processing of frame number 20 is completed, the person ID is "1", the number of frames stayed is "5", the number of consecutive failures is "15", and the reference image and comparison image are the images of frame number 5.
[0058] Next, the image with frame number 21 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "◯"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image with frame number 2. Here, it is assumed that it has not been determined that the person is the same (identity determination "X"). As a result, the number of consecutive failures becomes "16 (M or more)". Therefore, the image processing device 10 updates the person ID to "2". Furthermore, since the timing of the user switching is determined to be the first failure of these 16 consecutive failures (the timing when frame number 6 was generated), the image processing device 10 updates the number of stayed frames to "16". Furthermore, the image processing device 10 updates the number of consecutive failures to "0". Furthermore, the image processing device 10 updates the reference image and the comparison image to the image with frame number 21.
[0059] In the above example, the reference image and the comparison image are the most recent images among the images determined to be of the same person in the same person determination process by the user switching detection unit 12.
[0060] 4, the output unit 13 outputs information about the detection result by the user switching detection unit 12. This output can be realized by using a dedicated system, email, an app, or the like.
[0061] For example, the output unit 13 may calculate the usage time of each user in real time based on the detection result by the user switching detection unit 12, and output the calculation result. The output destination is a display or the like that can be viewed by a monitor.
[0062] As another example, the output unit 13 may calculate the usage time of each user in real time based on the detection results of the user switching detection unit 12 and monitor whether the usage time exceeds a reference value. If the usage time exceeds the reference value, warning information may be output. The output destination may be, for example, the operating terminal or a display or speaker installed near the operating terminal. In this case, the warning information may be a warning message such as, "Beware of bank transfer fraud." Alternatively, the output destination may be a display or speaker viewed by a monitor or an administrator of the operating terminal, or a mobile device carried by such a person. In this case, the warning information may be a warning message such as, "The usage time of the customer using operating terminal No. 3 has exceeded the reference value. This may be a bank transfer fraud. Please check."
[0063] As another example, the output unit 13 may output the processing result of the user switching detection unit 12 as is. The output processing result includes the determination result of whether or not the user has been switched, and, if it is determined that the user has been switched, the timing at which the user of the operating terminal was switched (in the example of FIG. 7, the date and time at which the image of frame number 6 was generated). Note that, only when it is determined that the user has been switched, the output unit 13 may output information indicating that the user has been switched and the timing at which the user of the operating terminal was switched. The output destination is a device that executes predetermined processing. For example, the device monitors the usage time of each user based on the input information and performs warning processing according to the usage time.
[0064] Next, an example of the processing flow of the image processing device 10 will be described using the flowchart in Fig. 8. Details of each process have been described above, so explanations here will be omitted as appropriate. In real-time processing, the following processes are executed on images generated by the surveillance camera.
[0065] First, the image processing device 10 acquires one image as an image to be processed (S10). Then, the image processing device 10 executes a process for detecting a person area (S11) and a process for detecting feature points of the person's skeleton (S12) for the image to be processed.
[0066] Next, the image processing device 10 identifies a user area from within the image to be processed based on the first detection result that detects the person area and the second detection result that detects the feature points of the person's skeleton (S13).
[0067] Next, the image processing device 10 determines whether the person included in the user area in the image to be processed is the same person based on the comparison result between the feature data extracted from the image of the user area in the image to be compared with the feature data extracted from the image of the user area in the image to be compared (S14). Note that if there is no image generated before the image to be processed, this process may be skipped.
[0068] Next, the image processing device 10 determines whether the user of the operating terminal has changed based on the determination result of S14 and the history of determination results up to this point (S15). Note that if there is no image generated before the image to be processed, this process may be skipped.
[0069] Then, the image processing device 10 outputs the determination result of S15 (S16).
[0070] "Hardware configuration of image processing device" An example of the hardware configuration of the image processing device 10 will be described. Fig. 9 is a diagram showing an example of the hardware configuration of the image processing device 10. Each functional unit of the image processing device 10 is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, programs loaded into the memory, a storage unit such as a hard disk that stores the programs (this can store programs that are pre-loaded when the device is shipped, as well as programs downloaded from storage media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface. Those skilled in the art will understand that there are many variations in the implementation methods and devices.
[0071] As shown in FIG. 9, the image processing device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The image processing device 10 does not necessarily have to have the peripheral circuit 4A. The image processing device 10 may be configured as multiple physically and / or logically separated devices, or may be configured as a single device that is physically and logically integrated. In the former case, each of the multiple devices that make up the image processing device 10 may have the above hardware configuration.
[0072] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data to and from each other. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, etc., and an interface for outputting information to an output device, an external device, an external server, etc. Examples of input devices include a keyboard, a mouse, a microphone, etc. Examples of output devices include a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0073] "Effects of image processing devices" As described above, the image processing device 10 executes the "processing for detecting a change in the user of the operating terminal."
[0074] If the surveillance camera has high performance (e.g., high frame rate, high resolution, etc.) and is installed in an appropriate position and orientation, it can use well-known tracking technology or face recognition technology to distinguish multiple users detected in an image from one another with high accuracy. Therefore, there is no need to go to the trouble of executing a process to detect a change in the user of the operating terminal. In fact, the technologies described in Patent Documents 1 and 2, which are considered to be based on the premise that the surveillance camera has high performance and is installed in an appropriate position and orientation, do not execute a "process to detect a change in the user of the operating terminal."
[0075] However, if the surveillance camera has low performance (e.g., low frame rate, low resolution, etc.) or is installed in an inappropriate location, as in this embodiment, it becomes difficult to accurately distinguish between multiple users detected in an image using well-known tracking or facial recognition technology. As a result, different users may be mistaken for the same user, or the same user appearing in multiple images may be mistaken for different users. As a result, the accuracy of the calculation results for the usage time of each user's operating device may deteriorate.
[0076] Therefore, the image processing device 10 executes a process that detects a change in the user of the operating terminal, which is not performed in conventional technology, and calculates the usage time of the operating terminal for each user based on the detection result. As a result, even if the surveillance camera has limitations in performance or installation location, it becomes possible to accurately calculate the usage time of the operating terminal for each user.
[0077] Furthermore, all of the processes executed by the image processing device 10 are characterized by content suitable for processing images generated by surveillance cameras that have limitations in performance and installation location. Therefore, even if the surveillance cameras have limitations in performance and installation location, the accuracy of the processing by the image processing device 10 is high. As a result, the usage time of each user's operating terminal can be calculated with high accuracy.
[0078] Furthermore, as described with reference to FIG. 7, the image processing device 10 does not increment the number of frames stayed when person detection or identity determination fails. In cases where the surveillance camera has low performance (e.g., low frame rate, low resolution, etc.) or is installed in an inappropriate location, as in this embodiment, there is a relatively high possibility that person detection or identity determination will fail. For this reason, if the number of frames stayed is incremented when person detection or identity determination fails, a false alarm may be output. The image processing device 10 can prevent this inconvenience by being configured to "not increment the number of frames stayed when person detection or identity determination fails."
[0079] <Second embodiment> "Image Processing Device Overview" The image processing device 10 of this embodiment analyzes images generated by a surveillance camera to detect that a user of an operating terminal is in a call posture. The surveillance camera used in this embodiment also has limitations in performance and installation location, as in the first embodiment. For this reason, the process of detecting a call posture executed by the image processing device 10 of this embodiment has distinctive content that is suitable for processing images generated by surveillance cameras that have limitations in performance and installation location. For this reason, the accuracy of the process is high even when the surveillance camera has limitations in performance and installation location.
[0080] "Functional configuration of image processing device" 10 shows an example of a functional block diagram of the image processing device 10 of this embodiment. As shown in the figure, the image processing device 10 differs from the image processing device 10 of the first embodiment in that it has a posture determination unit 14 instead of the user switching detection unit 12.
[0081] The configuration of the user area identification unit 11 is the same as that in the first embodiment, and therefore a description thereof will be omitted here.
[0082] The posture determination unit 14 calculates a degree of certainty that the user of the operation terminal is in a predetermined posture, based on the image of the user area. The predetermined posture is a call posture. The posture determination unit 14 detects the call posture based on skeletal feature points of a person detected in the image of the user area. Specifically, when a target feature point, which is part of the skeletal feature points of a person, is in a predetermined state, that is, when a target feature point in a predetermined state is detected in the image of the user area, the posture determination unit 14 determines that the user of the operation terminal is in a call posture.
[0083] Here, a description will be given of feature points of interest that have reached a predetermined state. For example, as shown in Fig. 11, feature points corresponding to the wrist, elbow, and shoulder are feature points of interest. The predetermined state is when the angle θ formed by the wrist feature point, elbow feature point, and shoulder feature point is equal to or less than a threshold value. Note that the feature points of interest that have reached the predetermined state illustrated here are merely examples, and are not limited to these.
[0084] As mentioned above, surveillance cameras have limitations in performance and installation location. When images generated by such surveillance cameras are processed and detected as a predetermined state, the detection accuracy of a feature point of interest deteriorates. Therefore, when a call posture is determined based on the detection of a feature point of interest that has reached a predetermined state, the accuracy of the determination deteriorates. Therefore, the posture determination unit 14 calculates a certainty factor that the user of the operating terminal is in a predetermined posture based on the history of detection results of feature points of interest that have reached a predetermined state. If this certainty factor exceeds a reference value, the system determines that the user of the operating terminal is in a predetermined posture. A method for calculating this certainty factor will be described below.
[0085] The posture determination unit 14 executes a process of detecting a feature point of interest that has attained a predetermined state in time-series order for each of a plurality of images to be processed (a process of detecting a predetermined posture).The posture determination unit 14 then determines a confidence level according to the number of times a predetermined posture is detected consecutively. The more times a predetermined posture is detected consecutively, the higher the confidence level becomes.
[0086] Here, an example of a process for determining the confidence level according to the number of times a predetermined posture is detected consecutively will be described. In this example, the posture determination unit 14 updates the confidence level based on the following rules.
[0087] (Rule 1) When a feature point of interest that is in a predetermined state is detected from the image to be processed, the confidence level is increased by a predetermined value. (Rule 2) When a feature point of interest that is not in a predetermined state is detected in the image to be processed, the certainty factor is reset to the initial value. (Rule 3) If no feature point of interest is detected in the image to be processed, the confidence level is maintained as it is.
[0088] A specific example will be described with reference to Fig. 12. The horizontal axis represents the number (frame number) of the image to be processed, and the vertical axis represents the confidence level.
[0089] Detection result (1) "Call posture detected" occurs when a feature point of interest that is in a predetermined state is detected from the image to be processed. In this case, rule 1 above is applied. Detection result (2) "Other posture detection" occurs when a feature point of interest that is not in a predetermined state is detected in the image to be processed. In this case, rule 2 above is applied. Detection result (3) "No target feature point detected" occurs when no target feature point is detected in the image to be processed. In this case, rule 3 above is applied.
[0090] 12, the detection result for frame 1 is (1). Therefore, rule 1 is applied, and the confidence level is increased by a predetermined value. Next, the detection result for frame 2 is also (1), so rule 1 is applied and the confidence level is further increased by a predetermined value. Next, the detection result for frame 3 is (3). Therefore, rule 3 is applied and the confidence level remains the same. Next, the detection result for frame 4 is (2), so rule 2 is applied and the confidence is reset to the initial value. Next, the detection result for frame 5 is (2) or (3), so either rule 2 is applied and the confidence is reset to the initial value, or rule 3 is applied and the confidence is maintained at the initial value. Next, the detection results for frames 6 to 9 are all (1), so rule 1 is applied and the confidence level is increased by a predetermined value.
[0091] Returning to FIG. 10 , the output unit 13 outputs warning information when the confidence calculated by the posture determination unit 14 exceeds a reference value. The output destination may be, for example, the operation terminal or a display or speaker installed near the operation terminal. In this case, the warning information may be a warning message such as "Beware of bank transfer fraud." Alternatively, the output destination may be a display or speaker viewed by a monitor or an administrator of the operation terminal, or a mobile terminal carried by such a person. In this case, the warning information may be a warning message such as "The customer using operation terminal No. 3 is operating while talking on the phone. There is a possibility of bank transfer fraud. Please check."
[0092] Next, an example of the processing flow of the image processing device 10 will be described using the flowchart in Fig. 13. Details of each process have been described above, so explanations here will be omitted as appropriate. In real-time processing, the following processes are executed on images generated by the surveillance camera.
[0093] First, the image processing device 10 acquires one image as an image to be processed (S20). Then, the image processing device 10 executes a process for detecting a person area (S21) and a process for detecting feature points of the person's skeleton (S22) for the image to be processed.
[0094] Next, the image processing device 10 identifies a user area from within the image to be processed based on the first detection result that detects the person area and the second detection result that detects the feature points of the person's skeleton (S23).
[0095] Next, the image processing device 10 executes a process of detecting a feature point of interest that has reached a predetermined state from the image of the user area (a process of detecting a predetermined posture) (S24).Then, the image processing device 10 updates the confidence that the user of the operation terminal is taking the predetermined posture based on the detection result of S24 (S25).
[0096] If the confidence level exceeds the reference value (Yes in S26), the image processing device 10 outputs warning information (S27). Note that if the confidence level does not exceed the reference value (No in S26), the image processing device 10 does not output warning information.
[0097] "Variations" Here, a modified example of the image processing device 10 of this embodiment will be described. The posture determination unit 14 calculates a certainty that the right half of the person's body is assuming a predetermined posture based on the feature points of the right half of the body among the feature points of the person's skeleton detected from the image to be processed. Also, the posture determination unit 14 calculates a certainty that the left half of the person's body is assuming a predetermined posture based on the feature points of the left half of the body among the feature points of the person's skeleton detected from the image to be processed. The method of calculating the predetermined posture and the certainty is as described above.
[0098] Then, the posture determination unit 14 calculates the larger of the degree of certainty that the right half of the person's body is taking a predetermined posture and the degree of certainty that the left half of the person's body is taking a predetermined posture as the degree of certainty that the user of the operating terminal is taking a predetermined posture.
[0099] The output unit 13 outputs warning information when the certainty determined as described above (the greater of the certainty that the right half of the person's body is in a predetermined posture and the certainty that the left half of the person's body is in a predetermined posture) exceeds a reference value.
[0100] Next, an example of the processing flow of the image processing device 10 in this modified example will be described using the flowchart in Fig. 14. Since the details of each process have been described above, the description here will be omitted as appropriate. In real-time processing, the following processes are executed on images generated by the surveillance camera.
[0101] First, the image processing device 10 acquires one image as an image to be processed (S30). Then, the image processing device 10 executes a process for detecting a person area (S31) and a process for detecting feature points of the person's skeleton (S32) for the image to be processed.
[0102] Next, the image processing device 10 identifies a user area from within the image to be processed based on the first detection result that detects the person area and the second detection result that detects the feature points of the person's skeleton (S33).
[0103] Next, the image processing device 10 executes a process of detecting a target feature point that has reached a predetermined state (a process of detecting a predetermined posture) based on the feature points of the left half of the body among the skeletal feature points of the person detected from the image of the user area (S34).Then, based on the detection result of S34, the image processing device 10 updates the confidence that the user of the operation terminal is assuming a predetermined posture with the left half of the body (S35).
[0104] Furthermore, the image processing device 10 executes a process of detecting a target feature point that has reached a predetermined state (a process of detecting a predetermined posture) based on the feature points of the right half of the body among the skeletal feature points of the person detected from the image of the user area (S36).Then, based on the detection result of S36, the image processing device 10 updates the confidence that the user of the operation terminal is assuming a predetermined posture with the right half of the body (S37).
[0105] Next, the image processing device 10 selects the greater of the degree of certainty that the user of the operation terminal is assuming a predetermined posture with the left half of the body and the degree of certainty that the user of the operation terminal is assuming a predetermined posture with the right half of the body (S38).
[0106] If the selected certainty exceeds the reference value (Yes in S39), the image processing device 10 outputs warning information (S40). If the selected certainty does not exceed the reference value (No in S39), the image processing device 10 does not output warning information.
[0107] "Hardware configuration of image processing device" The hardware configuration of the image processing device 10 is the same as that of the first embodiment.
[0108] "Effects of image processing devices" The image processing device 10 of this embodiment analyzes images generated by a surveillance camera and detects whether the user of the operating terminal is in a call posture. The process of detecting the call posture executed by the image processing device 10 of this embodiment is characterized by content suitable for processing images generated by surveillance cameras that have limitations in performance and installation location. Therefore, the accuracy of the process is high even when the surveillance camera has limitations in performance and installation location.
[0109] <Third embodiment> The image processing device 10 of this embodiment has the functions described in the first embodiment and the functions described in the second embodiment.
[0110] 15 shows an example of a functional block diagram of the image processing device 10. As shown in the figure, the image processing device 10 includes a user area identification unit 11, a user switching detection unit 12, an output unit 13, and a posture determination unit 14.
[0111] The functional configurations of the user area specifying unit 11, the user switching detection unit 12, and the posture determination unit 14 are the same as those described in the first and second embodiments.
[0112] The output unit 13 may have both the output process described in the first embodiment and the output process described in the second embodiment. That is, the output process may be performed separately in accordance with the detection result by the user switching detection unit 12 and the determination result by the posture determination unit 14.
[0113] Additionally, the output unit 13 may perform output processing that integrates the detection result by the user switching detection unit 12 and the determination result by the posture determination unit 14. For example, the output unit 13 may output warning information when the usage time of the user of the operation terminal calculated based on the detection result by the user switching detection unit 12 exceeds a reference value and the confidence calculated by the posture determination unit 14 exceeds a reference value.
[0114] The output destination may be, for example, the operation terminal or a display or speaker installed near the operation terminal. The warning information in this case may be a warning message such as "Beware of bank transfer fraud." The output destination may also be a display or speaker viewed by a monitor or the administrator of the operation terminal, or a mobile terminal carried by such a person. The warning information in this case may be a warning message such as "The usage time of the customer on operation terminal No. 3 has exceeded the standard value. Furthermore, this customer is operating the terminal while talking on the phone. There is a possibility of bank transfer fraud. Please check."
[0115] The hardware configuration of the image processing device 10 is the same as that of the first embodiment.
[0116] According to the image processing device 10 of this embodiment, the same effects as those of the first and second embodiments are realized.
[0117] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations can also be adopted.
[0118] In this specification, "acquisition" includes at least one of the following: "the device retrieves data stored in another device or storage medium (active acquisition)" based on user input or program instructions, such as receiving data by making a request or inquiry to another device, or accessing and reading out another device or storage medium; "the device inputs data output from another device (passive acquisition)" based on user input or program instructions, such as receiving data that is distributed (or transmitted, push notification, etc.), and selecting and acquiring data from received data or information; and "the device generates new data by editing data (converting it to text, rearranging data, extracting some data, changing the file format, etc.), and then acquires the new data."
[0119] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes. 1. A user area specifying means for specifying a user area, which is an area where a user of an operating terminal exists, from within an image to be processed; a user change detection means for detecting a change in the user of the operation terminal based on the image of the user area; An image processing device having: 2. The image processing device described in 1, wherein the user switching detection means detects a change in the user of the operating terminal based on the comparison result between feature data extracted from an image of the user area in the image to be processed and feature data extracted from an image of the user area in a comparison image generated before the image to be processed. 3. The user switching detection means a process of determining whether or not a person included in the image of the user region in the image to be processed and a person included in the image of the user region in the image to be compared are the same person based on a comparison result between the feature data extracted from the image to be processed and the feature data extracted from the image to be compared, the process being repeated in order for a plurality of images to be processed; If it is determined that M or more consecutive images to be processed (M is an integer of 2 or more) are not of the same person, it is determined that the user of the operation terminal has changed; 3. An image processing device according to claim 2, wherein the timing at which the earliest chronologically ordered image to be processed among the M consecutive images to be processed was generated is determined to be the timing at which the user of the operating terminal was switched. 4. The user area specifying means 4. An image processing device according to any one of claims 1 to 3, which identifies the user area from the image to be processed based on a first detection result that detects a person area from the image to be processed and a second detection result that detects feature points of the person's skeleton from the image to be processed. 5. The user area specifying means An image processing device described in 4, which identifies the person area in which the user of the operating terminal is located based on the size of the person area and the characteristic points of the person's skeleton for the multiple person areas indicated in the first detection result. 6. The user area specifying means 6. An image processing device according to claim 4 or 5, which, if no skeletal feature points of the person are detected in the image to be processed, identifies the user area from the image to be processed based on skeletal feature points of the person detected in a reference image generated before the image to be processed. 7. An image processing device described in any one of 1 to 6, further comprising an output means for outputting warning information when the usage time of the user of the operating terminal calculated based on the detection result of the user switching detection means exceeds a reference value. 8. An image processing device according to any one of 1 to 6, further comprising a posture determination means for calculating a degree of certainty that the user of the operation terminal is taking a predetermined posture based on the image of the user area. 9. An image processing device as described in 8, further comprising an output means for outputting warning information when the usage time of the user of the operating terminal calculated based on the detection result of the user switching detection means exceeds a reference value and the certainty calculated by the posture determination means exceeds a reference value. 10. The plurality of images are processed in chronological order; The posture determination means performing a process of detecting the predetermined orientation from the image to be processed for each of the plurality of images to be processed; 10. The image processing device according to 8 or 9, wherein the certainty factor is determined according to the number of times the predetermined posture is detected consecutively. 11. The posture determination means execute a process of detecting the predetermined posture based on a target feature point among feature points of a person's skeleton detected from the image to be processed; As the process of detecting the predetermined posture, the feature point of interest that has reached a predetermined state is detected; When the feature point of interest that is in the predetermined state is detected from the image to be processed, the certainty factor is increased; resetting the certainty factor to an initial value when the feature point of interest that is not in the predetermined state is detected from the image to be processed; 11. The image processing device according to 10, wherein the certainty factor is maintained as it is if the feature point of interest is not detected from the image to be processed. 12. The posture determination means calculating a degree of certainty that the right half of the person is taking the predetermined posture based on feature points of the right half of the person's body among feature points of the person's skeleton detected from the image to be processed; calculating a degree of certainty that the left half of the person is taking the predetermined posture based on feature points of the left half of the person's body among feature points of the person's skeleton detected from the image to be processed; 12. An image processing device according to any one of claims 8 to 11, wherein the larger of the degree of certainty that the right half of the person is assuming the specified posture and the degree of certainty that the left half of the person is assuming the specified posture is calculated as the degree of certainty that the user of the operating terminal is assuming the specified posture. 13. The computer A user area is identified from the image to be processed, which is an area where the user of the operating terminal is present; An image processing method for detecting a change in the user of the operation terminal based on an image of the user area. 14. Computer, a user area specifying means for specifying a user area, which is an area where a user of the operation terminal exists, from within the image to be processed; a user change detection means for detecting a change in the user of the operation terminal based on the image of the user area; A program that functions as a
[0120] This application claims priority based on Japanese Patent Application No. 2021-006332, filed on January 19, 2021, the disclosure of which is incorporated herein in its entirety. [Explanation of symbols]
[0121] 10 Image processing device 11 User area identification section 12 User switching detection unit 13 Output section 14 Posture determination section 1A processor 2A Memory 3A input / output I / F 4A peripheral circuit 5A Bus
Claims
1. A program to be executed by a computer, The program A process of acquiring an image from a camera installed in a position capable of capturing an image of a person operating an ATM (automated teller machine); A process of detecting that a person is talking based on information about the skeleton of the person appearing in the image; and outputting a message including information for identifying the ATM to a predetermined output destination according to the result of the detection, The information about the skeleton includes angles based on a first feature point, a second feature point, and a third feature point.
2. 2. The program according to claim 1, wherein the process of detecting whether a person is talking detects whether a person is talking based on a positional relationship between a plurality of parts of the person's body.
3. The program according to claim 2 , wherein the plurality of body parts of the person includes at least a shoulder.
4. The program according to claim 1 , which causes the computer to execute a process of calculating a score indicating that the person has a posture regarding a call.
5. The program according to claim 4 , wherein the score is higher the more times the person is detected to be in a posture related to a call.
6. 6. The program according to claim 1, wherein the information for identifying the ATM includes a number assigned to the ATM.
7. 7. The program according to claim 1, wherein the predetermined output destination includes a terminal for viewing by a monitor or a manager of the ATM.
8. A program described in any one of claims 1 to 7, wherein the angle based on the first feature point, the second feature point, and the third feature point includes the angle formed by a first line segment connecting the first feature point and the second feature point and a second line segment connecting the second feature point and the third feature point.
9. The program described in Claim 8, wherein the second feature point indicates the position of the person's elbow.
10. The program comprises: a process of detecting whether the person's usage time of the ATM has exceeded a preset time; A program described in any one of claims 1 to 9, wherein the process of outputting a message including information for identifying the ATM to a predetermined output destination is executed when it is detected that the person is making a call and that the person's usage time at the ATM has exceeded a predetermined time.
11. The computer An image is acquired from a camera installed in a position where it can capture an image of a person operating the ATM, Detecting that the person is talking based on information about the skeleton of the person appearing in the image; outputting a message including information for identifying the ATM to a predetermined output destination according to the result of the detection; An image processing method in which the information about the skeleton includes angles based on a first feature point, a second feature point, and a third feature point.
12. means for acquiring an image from a camera installed in a position capable of photographing a person operating the ATM; means for detecting that a person is making a call based on information about the skeleton of the person appearing in the image; a means for outputting a message including information for identifying the ATM to a predetermined output destination in response to the result of the detection; and An image processing device in which the information about the skeleton includes angles based on a first feature point, a second feature point, and a third feature point.
Citation Information
Patent Citations
Transaction monitoring device and transaction monitoring system
JP2010079744A
Transaction monitoring device
JP2010170275A
Phone call decision device, its method, and program
JP2010218392A
Monitoring device and monitoring method
JP2010238204A
Device for estimation of occupant posture
JP2011123733A