Image processing program, image processing method, and image processing system
The image processing device enhances fraud detection by identifying user areas and switching events in low-performance or improperly installed surveillance cameras, ensuring accurate terminal usage time calculation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-14
AI Technical Summary
Surveillance cameras with low performance or improper installation orientations struggle to accurately record user details for facial recognition, leading to misidentification of users and inaccurate calculation of terminal usage time.
An image processing device that identifies the user area and detects user switching based on image analysis, using techniques like human detection and skeletal feature point analysis to enhance accuracy despite camera limitations.
Accurately detects fraud victims or potential victims by identifying user areas and switching events, improving the calculation of terminal usage time even with low-performance or improperly installed cameras.
Smart Images

Figure 2026065103000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing program, an image processing method, and an image processing system.
Background Art
[0002] Techniques for suppressing transfer fraud damage are desired. Related techniques are disclosed in Patent Documents 1 and 2. Patent Documents 1 and 2 disclose a technique for analyzing an image generated by a surveillance camera installed in an operation terminal such as an ATM (automated teller machine) to identify a person and determining whether the identified person is making a call on a mobile phone. Further, Patent Document 1 discloses a technique for determining that a person who has used an operation terminal for a long time is or may be a victim of transfer fraud.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] As behavioral tendencies during operation of an operation terminal by a fraud victim such as transfer fraud, "operating while talking on a mobile phone", "using for a long time", etc. are known. By detecting a person performing these behaviors through image analysis as disclosed in Patent Documents 1 and 2, it becomes possible to suppress fraud damage. However, the present inventors newly found the following problems in this technology.
[0005] Some surveillance cameras installed to film users of control terminals may have low performance (e.g., low frame rate, low resolution, etc.) and may not be able to clearly record the details of users' faces or their actions. Furthermore, some cameras may be installed in positions and orientations that film users from above or at an angle above, and may not record details of their faces accurately enough for facial recognition technology to recognize them properly. Given the enormous number of control terminals, replacing all surveillance cameras with high-performance ones or changing their positions and orientations would be a significant burden.
[0006] The present invention aims to provide a technology that accurately detects fraud victims or potential fraud victims based on images generated by surveillance cameras that have limitations in performance and installation location. [Means for solving the problem]
[0007] According to the present invention, A user area identification means for identifying the user area, which is the area where the user of the operating terminal is located, from within the image to be processed, A user switching detection means for detecting that the user of the operating terminal has switched based on the image of the user area, An image processing device having the following is provided.
[0008] Furthermore, according to the present invention, Computers From the image to be processed, the user area, which is the area where the user of the operating terminal is located, is identified. An image processing method is provided that detects a change in the user of the operating terminal based on the image of the user area.
[0009] Furthermore, according to the present invention, Computers, A user area identification means for identifying the user area, which is the area where the user of the operating terminal is located, from within the image to be processed. A user switching detection means that detects when the user of the operating terminal has switched based on the image of the user area. A program that functions as such is provided.
Advantages of the Invention
[0010] According to the present invention, a technique for highly accurately detecting a fraud victim or a person who may be a fraud victim based on an image generated by a surveillance camera having limitations in performance and installation position is realized.
Brief Description of the Drawings
[0011] [Figure 1] It is a diagram for explaining an image generated by the surveillance camera of the present embodiment. [Figure 2] It is a diagram for explaining an image generated by the surveillance camera of the present embodiment. [Figure 3] It is a diagram for explaining an image generated by the surveillance camera of the present embodiment. [Figure 4] It is an example of a functional block diagram of the image processing apparatus of the present embodiment. [Figure 5] It is a diagram showing an example of the result of detecting a person area. [Figure 6] It is a diagram for explaining the process of specifying the user area of the present embodiment. [Figure 7] It is a diagram for explaining the process of detecting the switching of users in the present embodiment. [Figure 8] It is a flowchart showing an example of the process flow of the image processing apparatus of the present embodiment. [Figure 9] It is a diagram showing an example of the hardware configuration of the image processing apparatus of the present embodiment. [Figure 10] It is an example of a functional block diagram of the image processing apparatus of the present embodiment. [Figure 11] It is a diagram for explaining the call posture detection process of the present embodiment. [Figure 12] It is a diagram for explaining the process of calculating the confidence level that the user of the present embodiment is taking a predetermined posture. [Figure 13] It is a flowchart showing an example of the process flow of the image processing apparatus of the present embodiment. [Figure 14] It is a flowchart showing an example of the processing flow of the image processing apparatus according to the present embodiment. [Figure 15] It is an example of the functional block diagram of the image processing apparatus according to the present embodiment.
Mode for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.
[0013] <First Embodiment> 「Regarding the image generated by the surveillance camera」 The image processing apparatus according to the present embodiment analyzes an image generated by a surveillance camera installed to photograph a user of an operation terminal such as an ATM, and highly accurately detects a fraud victim or a person who may be a fraud victim. Fraud includes, for example, transfer fraud, etc., but is not limited thereto.
[0014] Here, the image generated by the surveillance camera will be described. The surveillance camera has limitations in performance and installation location. Therefore, the surveillance camera cannot clearly record the details of the face and actions of the user of the operation terminal.
[0015] For example, a camera with low performance (e.g., low frame rate, low resolution, etc.) is used as the surveillance camera. In this case, the surveillance camera cannot clearly record the details of the face and actions of the user of the operation terminal.
[0016] Furthermore, as shown in Figure 1, for example, the surveillance camera 100 is installed in a position and orientation that photographs the user 101 of the operating terminal 102 from above or diagonally above. In other words, the surveillance camera 100 may be installed in a position and orientation that does not photograph the face of the user 101 of the operating terminal 102 from the front. In this case, the surveillance camera 100 cannot record the details of the face of the user 101 of the operating terminal 102 to a degree that allows for accurate recognition using facial recognition technology. In addition, the generated image may include people or objects other than the user 101. For example, as shown in Figures 2 and 3, the image F generated by the surveillance camera 100 may include the operating terminal 102, other people 103 such as people waiting in line or passersby, and other objects 104 such as walls or partitions, in addition to the user 101 standing in front of the operating terminal 102.
[0017] "Overview of Image Processing Equipment" Next, an overview of the image processing device of this embodiment will be described. The image processing device of this embodiment performs processing to detect fraud victims or potential fraud victims with high accuracy based on images generated by a surveillance camera. Specifically, the image processing device of this embodiment performs the following based on images generated by a surveillance camera: "processing to identify the user area, which is the area where the user of the operating terminal is located, from the image to be processed" and "processing to detect that the user of the operating terminal has changed, based on the image of the user area."
[0018] One of the main features of the image processing device of this embodiment is that it performs a process to detect when the user of the operating terminal changes. Based on this detection result, the usage time of each user's operating terminal can be determined.
[0019] If the surveillance camera is high-performance (e.g., high frame rate, high resolution) and installed in the appropriate position and orientation, well-known tracking and facial recognition technologies can be used to accurately identify multiple users detected in the image. Therefore, there is no need to perform any processing to detect when the user of the operating terminal changes. In fact, the technologies described in Patent Documents 1 and 2, which are thought to be based on the premise that the surveillance camera is high-performance and installed in the appropriate position and orientation, do not perform any processing to detect when the user of the operating terminal changes.
[0020] However, if the surveillance camera is of low performance (e.g., low frame rate, low resolution, etc.) or installed in an inappropriate location, as in this embodiment, it becomes difficult to accurately identify multiple users detected in the image using well-known tracking or facial recognition technologies. As a result, there is a risk of misidentifying different users as the same user, or misidentifying the same user appearing across multiple images as different users. Consequently, the accuracy of the calculation of each user's terminal usage time deteriorates.
[0021] Therefore, the image processing device of this embodiment performs a process that does not occur in conventional technology, namely "a process to detect when the user of the operating terminal has changed," and calculates the usage time of each user's operating terminal based on the detection result. As a result, even if the surveillance camera has limitations in performance or installation location, it becomes possible to accurately calculate the usage time of each user's operating terminal.
[0022] Another key feature of the image processing device of this embodiment is that all of the processes performed by the image processing device are specifically designed to be suitable for processing images generated by surveillance cameras that have limitations in performance and installation location. Therefore, even when surveillance cameras have limitations in performance and installation location, the accuracy of the above processes is improved. As a result, the usage time of each user's terminal can be calculated with high accuracy.
[0023] "Functional Configuration of Image Processing Devices" Next, the functional configuration of the image processing device will be described in detail. Figure 4 shows an example of a functional block diagram of the image processing device 10 of this embodiment. As shown in the figure, the image processing device 10 has a user area identification unit 11, a user switching detection unit 12, and an output unit 13.
[0024] The user area identification unit 11 identifies the user area from the image to be processed, which is the area where the user of the operating terminal is located.
[0025] The "images to be processed" are the images generated by the aforementioned surveillance cameras. The surveillance cameras are installed in a position and orientation that captures the users of the operating terminals, and they generate video footage. Multiple images included in this video footage, in chronological order (in the order they were captured), become the images to be processed.
[0026] Next, we will explain the process of identifying the user area from the image to be processed. The user area identification unit 11 detects the person area from the image to be processed and identifies one of the one or more detected person areas as the user area. This will be explained in detail below.
[0027] -Processing to detect areas containing people within the image to be processed- The process of detecting human regions within an image to be processed can be implemented using any known human detection technique. For example, it may be implemented using a human detection model generated by machine learning, or by other means. In known human detection techniques, for example, a rectangular region containing a person is detected as a human region.
[0028] As mentioned above, surveillance cameras have limitations in terms of performance and installation location. Therefore, the accuracy of the process for detecting the person's area is insufficient. As a result, there is a possibility that parts of a person or other objects 104 other than people may be mistakenly identified as people. Also, as mentioned above, when the surveillance camera shoots from above or diagonally above, it generates images that include other people 103. As a result, there is a possibility that the process for detecting the person's area may also detect other people 103. Consequently, in the process of detecting the person's area from the image to be processed, as shown in Figure 5, in addition to the person's area W1 that includes the user 101, there is a possibility that person's area W2 that includes other people 103, person's area W3 that includes only a part of a person, person's area W4 that includes other objects 104, etc. may be detected. Therefore, the user area identification unit 11 executes the "process of identifying one of the one or more detected person's areas as the user's area," which will be explained in detail below, and identifies an appropriate user area (the area including user 101) from among the detected person's areas.
[0029] - A process to identify one of the detected person areas as the user area - The user area identification unit 11 utilizes, in this process, a first detection result in which a human area is detected from the image to be processed, and a second detection result in which characteristic points of a person's skeleton are detected from the image to be processed.
[0030] The first detection result includes the size of each detected person region and the confidence level. The size of the person region is the size of the area occupied by the person region, and can be expressed, for example, in pixels. The confidence level is a value that indicates the degree of confidence that the person region is an area containing a person (a measure of how certain the result is). Techniques for calculating such confidence levels are widely known in well-known person detection technologies.
[0031] The second detection result includes the coordinates (information indicating the position within the image) of each of the multiple feature points of the human skeleton detected from the image being processed. The detection of the human skeleton feature points is achieved using well-known techniques such as OpenPose.
[0032] The user area identification unit 11 uses the first detection result and the second detection result to identify the one or more person areas that are most appropriate to contain the user of the operating terminal. Specifically, the user area identification unit 11 applies the judgment method shown in Figure 6 to each of the one or more person areas that were detected as a processing target to determine whether or not each person area is a user area.
[0033] In S100, the user area identification unit 11 determines whether the size of the person area to be processed is greater than a preset threshold. If the size of the person area to be processed is smaller than the threshold ("smaller" in S100), the user area identification unit 11 determines that the person area to be processed is not a user area (S102).
[0034] If the size of the person region to be processed is greater than the threshold ("great" in S100), the user region identification unit 11 determines whether at least one feature point of the person's skeleton has been detected in the image to be processed, which includes the person region to be processed (S101). If it has been detected ("detected" in S101), the user region identification unit 11 determines which person region to be processed has the largest number of feature points of the person's skeleton, based on the number of feature points of the person's skeleton contained in each of the detected person regions (S103).
[0035] If it is the largest value (YES in S103), the user area identification unit 11 determines that the person area to be processed is a user area (S110). On the other hand, if it is not the largest value (NO in S103), the user area identification unit 11 determines that the person area to be processed is not a user area (S105).
[0036] If no feature points of the human skeleton are detected in the image to be processed, including the area of the person being processed (S101, "Not detected"), the user area identification unit 11 determines whether at least one feature point of the human skeleton is detected in the reference image (S104).
[0037] The "referenced image" is an image generated before the image containing the person region being processed, and includes the user of the operating terminal contained within the image being processed. The referenced image changes dynamically. An example of the process for determining the referenced image is described below.
[0038] If feature points of a person's skeleton are detected in the reference image (S104, "Yes"), the user area identification unit 11 determines the inclusion relationship between the person area detected in the image and the feature points of a person's skeleton detected in the reference image, based on the area occupied within each of the detected person areas and the coordinates within each of the feature points of a person's skeleton detected in the reference image. The user area identification unit 11 then determines which of the detected person areas contains the largest number of feature points of a person's skeleton (S106).
[0039] If it is the largest value (YES in S106), the user area identification unit 11 determines that the person area to be processed is a user area (S110). On the other hand, if it is not the largest value (NO in S106), the user area identification unit 11 determines that the person area to be processed is not a user area (S107).
[0040] If no skeletal feature points of a person are detected in the reference image (S104, "None"), the user area identification unit 11 determines which of the one or more detected person areas has the highest confidence level (confidence level shown in the first detection result) and is the person area to be processed (S108).
[0041] If it is the largest value (YES in S108), the user area identification unit 11 determines that the person area to be processed is a user area (S110). On the other hand, if it is not the largest value (NO in S108), the user area identification unit 11 determines that the person area to be processed is not a user area (S109).
[0042] In this way, the user area identification unit 11 can identify the user area from the image to be processed based on a first detection result in which a human area is detected from the image to be processed, and a second detection result in which characteristic points of a human skeleton are detected from the image to be processed. Furthermore, if characteristic points of a human skeleton are not detected from the image to be processed, the user area can be identified from the image to be processed based on characteristic points of a human skeleton detected from a reference image that was generated before the image to be processed.
[0043] Returning to Figure 4, the user switching detection unit 12 detects that the user of the operating terminal has switched based on the image of the user region identified by the user region identification unit 11. The user switching detection unit 12 detects that the user of the operating terminal has switched (that the user in the image to be processed and the user in the image to be compared are different people) based on the comparison result between the feature data extracted from the image of the user region in the image to be processed and the feature data extracted from the image of the user region in the image to be compared.
[0044] "Feature data extracted from user area images" refers to data that describes the user's appearance characteristics as seen in the image. Examples include, but are not limited to, clothing characteristics, hairstyle characteristics, facial features, and the presence or absence and characteristics of items such as glasses or hats.
[0045] The "comparison image" is an image created before the image being processed. The comparison image changes dynamically. An example of the process for determining the comparison image is explained below.
[0046] It should be noted that both the "referenced image" and the "comparison image" mentioned above are images that were generated before the "image being processed." However, they differ in that the "referenced image" is used for the "process to determine the user area" mentioned above, while the "comparison image" is used for the "process to detect a change in the user of the operating terminal" mentioned above.
[0047] As mentioned above, surveillance cameras have limitations in terms of performance and installation location. Therefore, the accuracy of the process for determining whether or not the same person is captured by comparing the feature data is insufficient. Accordingly, the user switching detection unit 12 determines that the user of the operating terminal has switched if it is determined that M (where M is an integer of 2 or more) or more consecutive images to be processed are not of the same person. The timing at which the earliest chronologically ordered image among those M consecutive images to be processed was generated is then determined to be the timing at which the user of the operating terminal switched.
[0048] The process will be explained below using the specific example shown in Figure 7. The reference image and comparison image mentioned above will also be explained. Note that M will be set to "16".
[0049] First, the image with frame number 1 becomes the image to be processed, and processing by the user area identification unit 11 is executed. Here, it is assumed that the user area has been identified (person detection "○"). The image processing device 10 sets the person ID (identifier) "1" for the person (user of the operating terminal) included in that user area, and sets the number of frames stayed to "1". It also sets the image with frame number 1 as the image to be referenced and the image to be compared. Note that there are no frames prior to this, so processing by the user switching detection unit 12 is not executed.
[0050] Next, the image with frame number 2 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, we assume that the user area has been identified (person detection "○"). Then, the user switching detection unit 12 compares the feature data extracted from the image of the user area in the image to be processed with the feature data extracted from the image of the user area in the image to be compared. If the similarity is greater than or equal to the threshold value, the user switching detection unit 12 determines that the person in those two images is the same person; if it is less than the threshold value, it determines that they are not the same person. Here, we assume that they have been determined to be the same person (identity determination "○"). The image processing device 10 keeps the person ID as "1" and updates the number of frames stayed to "2". Then, the image processing device 10 updates the reference image and the comparison image to the image with the image with frame number 2.
[0051] Next, the image with frame number 3 becomes the image to be processed, and processing by the user area identification unit 11 is executed. Here, we assume that the user area was not identified (person detection "×"). Situations in which the user area is not identified include, for example, when no user area larger than the threshold size is detected ("small" in S100 of Figure 6). In this case, processing by the user switching detection unit 12 is not executed. The image processing device 10 keeps the person ID at "1", the number of frames stayed at "2", and sets the number of consecutive failures to "1". The image processing device 10 then keeps the reference image and the comparison image at the image with frame number 2.
[0052] Next, the image with frame number 4 becomes the image to be processed, and processing by the user area identification unit 11 is executed. Here, we assume that the user area was not identified (person detection "×"). In this case, processing by the user switching detection unit 12 is not executed. The image processing device 10 keeps the person ID at "1", the number of frames stayed at "2", and updates the number of consecutive failures to "2". Then, the image processing device 10 keeps the referenced image and the comparison target image as the image with frame number 2.
[0053] Next, the image with frame number 5 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "○"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image with frame number 2. Here, it is assumed that it has been determined to be the same person (same person determination "○"). The image processing device 10 keeps the person ID as "1" and updates the number of frames stayed to "5". That is, it is assumed that the person stayed even during frames 3 and 4, when person detection failed. Then, the image processing device 10 updates the number of consecutive failures to "0". In addition, the image processing device 10 updates the reference image and the comparison image to the image with frame number 5.
[0054] Next, the image at frame number 6 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "○"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image at frame number 2. Here, it is assumed that it was not determined to be the same person (same person determination "×"). The image processing device 10 keeps the person ID at "1", the number of frames stayed at "5", and updates the number of consecutive failures to "1". Also, the image processing device 10 keeps the reference image and the comparison image at frame number 5.
[0055] Next, the image with frame number 7 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "○"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image with frame number 2. Here, it is assumed that it was not determined to be the same person (same person determination "×"). The image processing device 10 keeps the person ID at "1", the number of frames stayed at "5", and updates the number of consecutive failures to "2". Also, the image processing device 10 keeps the reference image and the comparison image at the same as the image with frame number 5.
[0056] Next, the image at frame number 8 becomes the image to be processed, and processing by the user area identification unit 11 is executed. Here, we assume that the user area was not identified (person detection "×"). In this case, processing by the user switching detection unit 12 is not executed. The image processing device 10 keeps the person ID at "1", the number of frames stayed at "5", and updates the number of consecutive failures to "3". The image processing device 10 then keeps the reference image and the comparison image at frame number 5.
[0057] The same processing is performed thereafter, but for images with frame numbers 9-20, the user area was identified (person detection "○"), but it was not determined to be the same person (identity determination "×"). When the processing for frame number 20 is completed, the person ID will be "1", the number of frames stayed will be "5", the number of consecutive failures will be "15", and the reference image and the image to be compared will be the image with frame number 5.
[0058] Next, the image with frame number 21 becomes the image to be processed, and processing is performed by the user area identification unit 11. Here, it is assumed that the user area has been identified (person detection "○"). Then, the user switching detection unit 12 performs the same person determination process described in the processing of the image with frame number 2. Here, it is assumed that it was not determined to be the same person (same person determination "×"). As a result, the number of consecutive failures becomes "16 (M or more)". Therefore, the image processing device 10 updates the person ID to "2". Also, since the timing of the user switching is determined to be the very first failure among these 16 consecutive failures (the timing when frame number 6 is generated), the image processing device 10 updates the number of frames stayed to "16". The image processing device 10 also updates the number of consecutive failures to "0". Finally, the image processing device 10 updates the reference image and the comparison image to the image with the image with frame number 21.
[0059] The same process is repeated thereafter. In the example above, the reference image and the comparison image are the most recent images among those determined to be of the same person by the user switching detection unit 12's same person determination process.
[0060] Returning to Figure 4, the output unit 13 outputs information related to the detection result by the user switching detection unit 12. This output can be achieved using a dedicated system, email, application, etc.
[0061] For example, the output unit 13 may calculate the usage time for each user in real time based on the detection result from the user switching detection unit 12 and output the calculation result. The output destination is a display or the like that can be viewed by a supervisor.
[0062] As another example, the output unit 13 may calculate the usage time of each user in real time based on the detection result of the user switching detection unit 12, and monitor whether the usage time exceeds a standard value. If the usage time exceeds the standard value, warning information may be output. The output destination may be, for example, the operating terminal or a display or speaker installed near the operating terminal. In this case, the warning information may be a cautionary message such as, "Please be careful of wire fraud." Alternatively, the output destination may be a display or speaker viewed by a monitor or the administrator of the operating terminal, or a mobile device carried by such a person. In this case, the warning information may be a cautionary message such as, "The usage time of the customer on operating terminal number 3 has exceeded the standard value. This may be wire fraud. Please check."
[0063] As another example, the output unit 13 may output the processing result from the user switching detection unit 12 as is. The output processing result includes the determination result of whether or not the user has switched, and, if it is determined that the user has switched, the timing of the user switching on the operating terminal (in the example of Figure 7, the date and time when the image with frame number 6 was generated). In addition, the output unit 13 may output information indicating that the user has switched and the timing of the user switching on the operating terminal only if it is determined that the user has switched. The output destination is a device that performs a predetermined process. This device, for example, monitors the usage time of each user based on the input information and performs warning processing according to the usage time.
[0064] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 8. Since the details of each process have been described above, explanations will be omitted here as appropriate. In real-time processing, the following processes are performed on the images generated by the surveillance camera.
[0065] First, the image processing device 10 acquires one image as the image to be processed (S10). Then, the image processing device 10 performs the following processes on the image to be processed: detecting the region of a person (S11) and detecting the feature points of the person's skeleton (S12).
[0066] Next, the image processing device 10 identifies the user area from the image to be processed based on the first detection result in which the human area was detected and the second detection result in which the characteristic points of the human skeleton were detected (S13).
[0067] Next, the image processing device 10 determines whether the person contained in the user region of the image to be processed is the same person based on the comparison result between the feature data extracted from the user region of the image to be compared and the feature data extracted from the user region of the image to be compared (S14). Note that if there are no images generated before the image to be processed, this process may be skipped.
[0068] Next, the image processing device 10 determines whether the user of the operating terminal has changed based on the determination result in S14 and the history of determination results to date (S15). Note that if there are no images generated before the image to be processed, this process may be skipped.
[0069] Then, the image processing device 10 outputs the judgment result from S15 (S16).
[0070] "Hardware configuration of image processing equipment" An example of the hardware configuration of the image processing device 10 will be described. Figure 9 shows an example of the hardware configuration of the image processing device 10. Each functional unit of the image processing device 10 is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, a program loaded into memory, a storage unit such as a hard disk that stores that program (which can store programs that are pre-installed at the time of shipment, as well as programs downloaded from storage media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface. It will be understood by those skilled in the art that there are various modifications to the implementation method and the device.
[0071] As shown in Figure 9, the image processing device 10 includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. Peripheral circuitry 4A includes various modules. The image processing device 10 does not necessarily have peripheral circuitry 4A. The image processing device 10 may be composed of multiple physically and / or logically separate devices, or it may be composed of a single physically and logically integrated device. In the former case, each of the multiple devices constituting the image processing device 10 may have the above hardware configuration.
[0072] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuits 4A, and input / output interface 3A to send and receive data to and from each other. Processor 1A is a processing unit such as a CPU or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). Input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. Input devices include, for example, keyboards, mice, microphones, etc. Output devices include, for example, displays, speakers, printers, mailers, etc. Processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0073] "Effects and Effects of Image Processing Devices" As explained above, the image processing device 10 performs a process to detect that the user of the operating terminal has changed.
[0074] If the surveillance camera is high-performance (e.g., high frame rate, high resolution) and installed in the appropriate position and orientation, well-known tracking and facial recognition technologies can be used to accurately identify multiple users detected in the image. Therefore, there is no need to perform any processing to detect when the user of the operating terminal changes. In fact, the technologies described in Patent Documents 1 and 2, which are thought to be based on the premise that the surveillance camera is high-performance and installed in the appropriate position and orientation, do not perform any processing to detect when the user of the operating terminal changes.
[0075] However, if the surveillance camera is of low performance (e.g., low frame rate, low resolution, etc.) or installed in an inappropriate location, as in this embodiment, it becomes difficult to accurately identify multiple users detected in the image using well-known tracking or facial recognition technologies. As a result, there is a risk of misidentifying different users as the same user, or misidentifying the same user appearing across multiple images as different users. Consequently, the accuracy of the calculation of each user's terminal usage time deteriorates.
[0076] Therefore, the image processing device 10 performs a process that detects when the user of the operating terminal changes, which is not performed in conventional technology, and calculates the usage time of each user's operating terminal based on the detection result. As a result, even if the surveillance camera has limitations in performance or installation location, it becomes possible to accurately calculate the usage time of each user's operating terminal.
[0077] Furthermore, all the processes performed by the image processing device 10 are characterized in a way that makes them suitable for processing images generated by surveillance cameras that have limitations in performance and installation location. Therefore, even when surveillance cameras have limitations in performance and installation location, the accuracy of the processing performed by the image processing device 10 is high. As a result, the usage time of each user's terminal can be calculated with high accuracy.
[0078] Furthermore, as explained with reference to Figure 7, the image processing device 10 does not add the number of frames spent when person detection or identity determination fails. In cases where the surveillance camera is low-performance (e.g., low frame rate, low resolution, etc.) or the installation location is inappropriate, as in this embodiment, the possibility of person detection or identity determination failing is relatively high. Therefore, if the number of frames spent is added when person detection or identity determination fails, there is a risk of outputting a false alarm. The image processing device 10 can suppress this problem by having a configuration that "does not add the number of frames spent when person detection or identity determination fails."
[0079] <Second Embodiment> "Overview of Image Processing Equipment" The image processing device 10 of this embodiment analyzes images generated by a surveillance camera to detect whether the user of the operating terminal is in a calling position. The surveillance camera used in this embodiment, like in the first embodiment, has limitations in performance and installation location. Therefore, the process for detecting the calling position performed by the image processing device 10 of this embodiment has a distinctive structure suitable for processing images generated by surveillance cameras with performance and installation location limitations. As a result, the accuracy of the above process is improved even when the surveillance camera has performance and installation location limitations.
[0080] "Functional Configuration of Image Processing Devices" Figure 10 shows an example of a functional block diagram of the image processing device 10 of this embodiment. As shown in the figure, the image processing device 10 differs from the image processing device 10 of the first embodiment in that it has a posture determination unit 14 instead of a user switching detection unit 12.
[0081] The configuration of the user area identification unit 11 is the same as in the first embodiment, so its explanation is omitted here.
[0082] The posture determination unit 14 calculates the degree of confidence that the user of the operating terminal is in a predetermined posture based on the image of the user's area. The predetermined posture is the phone-talk posture. The posture determination unit 14 detects the phone-talk posture based on the characteristic points of the person's skeleton detected within the image of the user's area. Specifically, the posture determination unit 14 determines that the user of the operating terminal is in a phone-talk posture if a particular characteristic point, which is a part of the characteristic points of the person's skeleton, is in a predetermined state, that is, if a particular characteristic point in a predetermined state is detected within the image of the user's area.
[0083] Here, we will explain the feature points of interest that have reached a predetermined state. For example, as shown in Figure 11, the feature points corresponding to the wrist, elbow, and shoulder are the feature points of interest. The predetermined state is when the angle θ formed by the feature points of the wrist, elbow, and shoulder is below a threshold. Note that the feature points of interest that have reached a predetermined state as exemplified here are merely examples and are not limited to these.
[0084] However, as mentioned above, surveillance cameras have limitations in terms of performance and installation location. Processing images generated by such surveillance cameras results in poor detection accuracy for specific feature points that have reached a predetermined state. Therefore, if the system determines that a user is in a calling posture based on the detection of specific feature points that have reached a predetermined state, the accuracy of that determination will be poor. Accordingly, the posture determination unit 14 calculates the confidence level that the user of the operating terminal is in a predetermined posture based on the history of detection results for specific feature points that have reached a predetermined state. If this confidence level exceeds a standard value, the system determines that the user of the operating terminal is in a predetermined posture. The method for calculating this confidence level will be explained below.
[0085] The posture determination unit 14 performs a process to detect a feature point of interest that has reached a predetermined state in chronological order for each of the multiple images to be processed in a time series (a process to detect a predetermined posture). The posture determination unit 14 then determines the confidence level according to the number of times the predetermined posture has been detected consecutively. The more times it has been detected consecutively, the higher the confidence level.
[0086] Here, we will describe an example of a process that determines the confidence level according to the number of times a predetermined posture has been detected consecutively. In this example, the posture determination unit 14 updates the confidence level based on the following rules.
[0087] (Rule 1) If a feature point of interest that is in a predetermined state is detected in the image to be processed, the confidence level is increased by a predetermined value. (Rule 2) If a feature point of interest that is not in the specified state is detected in the image to be processed, the confidence level is reset to the initial value. (Rule 3) If no feature point of interest is detected in the image to be processed, the confidence level is maintained as is.
[0088] Let's illustrate with a specific example using Figure 12. The horizontal axis represents the image number (frame number) of the image being processed, and the vertical axis represents the confidence level.
[0089] Detection result (1) "Call posture detection" occurs when a feature point of interest in a predetermined state is detected within the image being processed. In this case, Rule 1 above is applied. Detection result (2) "Other posture detection" occurs when a feature point of interest that is not in the predetermined state is detected in the image being processed. In this case, Rule 2 above is applied. Detection result (3) "Feature point of interest not detected" indicates that the feature point of interest was not detected in the image being processed. In this case, Rule 3 above applies.
[0090] As shown in Figure 12, the detection result for frame 1 is (1). Therefore, Rule 1 is applied, and the confidence level increases by a predetermined value. Next, the detection result for frame 2 is also (1). Therefore, Rule 1 is applied, and the confidence level increases further by a predetermined value. Next, the detection result for frame 3 is (3). Therefore, rule 3 is applied, and the confidence level remains the same. Next, the detection result for frame 4 is (2). Therefore, rule 2 is applied, and the confidence level is reset to its initial value. Next, the detection result for frame 5 is either (2) or (3). Therefore, rule 2 is applied and the confidence level is reset to its initial value, or rule 3 is applied and the confidence level remains at its initial value. Next, the detection results for frames 6 through 9 are all (1). Therefore, Rule 1 is applied, and the confidence level increases by a predetermined value.
[0091] Returning to Figure 10, the output unit 13 outputs warning information if the confidence level calculated by the posture determination unit 14 exceeds a standard value. The output destination is, for example, the operating terminal itself or a display or speaker installed near the operating terminal. In this case, the warning information could be a cautionary message such as, "Please be careful of wire fraud." Alternatively, the output destination may be a display or speaker viewed by a supervisor or the administrator of the operating terminal, or a mobile device carried by such a person. In this case, the warning information could be a cautionary message such as, "The customer using operating terminal number 3 is operating it while on a call. This may be wire fraud. Please check."
[0092] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 13. Since the details of each process have been described above, explanations will be omitted here as appropriate. In real-time processing, the following processes are performed on the images generated by the surveillance camera.
[0093] First, the image processing device 10 acquires one image as the image to be processed (S20). Then, the image processing device 10 performs a process to detect the region of a person (S21) and a process to detect the feature points of the person's skeleton (S22) on the image to be processed.
[0094] Next, the image processing device 10 identifies the user area from the image to be processed based on the first detection result in which the human area was detected and the second detection result in which the characteristic points of the human skeleton were detected (S23).
[0095] Next, the image processing device 10 performs a process to detect a feature point of interest that is in a predetermined state from the image of the user area (a process to detect a predetermined posture) (S24). Then, based on the detection result from S24, the image processing device 10 updates the confidence level that the user of the operating terminal is in a predetermined posture (S25).
[0096] If the confidence level exceeds the threshold value (Yes in S26), the image processing device 10 outputs warning information (S27). If the confidence level does not exceed the threshold value (No in S26), the image processing device 10 does not output warning information.
[0097] "Variations" Here, a modified version of the image processing apparatus 10 of this embodiment will be described. The posture determination unit 14 calculates the confidence level that a predetermined posture is being adopted by the right half of a person based on the feature points of the right half of the person's skeleton detected in the image to be processed. The posture determination unit 14 also calculates the confidence level that a predetermined posture is being adopted by the left half of a person based on the feature points of the left half of the person's skeleton detected in the image to be processed. The method for calculating the predetermined posture and confidence level is as described above.
[0098] The posture determination unit 14 then calculates the confidence level that the user of the operating terminal is in the predetermined posture, whichever is greater of the confidence level that the right half of the person's body is in the predetermined posture and the confidence level that the left half of the person's body is in the predetermined posture.
[0099] The output unit 13 outputs warning information if the confidence level determined as described above (the greater of the confidence level that the right half of the person's body is in the predetermined posture and the confidence level that the left half of the person's body is in the predetermined posture) exceeds a standard value.
[0100] Next, using the flowchart in Figure 14, an example of the processing flow of the image processing device 10 in this modified example will be explained. Since the details of each process have been described above, explanations will be omitted here as appropriate. In real-time processing, the following processes are performed on the image generated by the surveillance camera.
[0101] First, the image processing device 10 acquires one image as the image to be processed (S30). Then, the image processing device 10 performs a process to detect the region of a person (S31) and a process to detect the feature points of the person's skeleton (S32) on the image to be processed.
[0102] Next, the image processing device 10 identifies the user area from the image to be processed based on the first detection result in which the human area was detected and the second detection result in which the characteristic points of the human skeleton were detected (S33).
[0103] Next, the image processing device 10 performs a process to detect a feature point of interest that is in a predetermined state (a process to detect a predetermined posture) based on the feature points of the left half of the human skeleton detected from the image of the user area (S34). Then, based on the detection result of S34, the image processing device 10 updates the confidence level that the user of the operating terminal is in a predetermined posture with the left half of their body (S35).
[0104] Furthermore, the image processing device 10 performs a process to detect a feature point of interest that is in a predetermined state (a process to detect a predetermined posture) based on the feature points of the right half of the human skeleton detected from the image of the user area (S36). Then, based on the detection result from S36, the image processing device 10 updates the confidence level that the user of the operating terminal is in a predetermined posture with the right half of their body (S37).
[0105] Next, the image processing device 10 selects the larger of the two confidence levels: the confidence level that the user of the operating terminal is in a predetermined posture with the left half of their body, and the confidence level that the user of the operating terminal is in a predetermined posture with the right half of their body (S38).
[0106] If the selected confidence level exceeds the threshold value (Yes in S39), the image processing device 10 outputs warning information (S40). If the selected confidence level does not exceed the threshold value (No in S39), the image processing device 10 does not output warning information.
[0107] "Hardware configuration of image processing equipment" The hardware configuration of the image processing device 10 is the same as in the first embodiment.
[0108] "Effects and Effects of Image Processing Devices" The image processing device 10 of this embodiment analyzes images generated by a surveillance camera to detect whether the user of the operating terminal is in a calling position. The process for detecting the calling position performed by the image processing device 10 of this embodiment has a distinctive structure that is suitable for processing images generated by surveillance cameras that have limitations in performance and installation location. Therefore, even when the surveillance camera has limitations in performance and installation location, the accuracy of the above process is improved.
[0109] <Third Embodiment> The image processing apparatus 10 of this embodiment has the functions described in the first embodiment and the functions described in the second embodiment.
[0110] Figure 15 shows an example of a functional block diagram of the image processing device 10. As shown in the figure, the image processing device 10 includes a user area identification unit 11, a user switching detection unit 12, an output unit 13, and a posture determination unit 14.
[0111] The functional configurations of the user area identification unit 11, the user switching detection unit 12, and the posture determination unit 14 are as described in the first and second embodiments.
[0112] The output unit 13 may have both the output processing described in the first embodiment and the output processing described in the second embodiment. That is, output processing may be performed separately according to the detection result by the user switching detection unit 12 and the determination result by the posture determination unit 14.
[0113] In addition, the output unit 13 may perform output processing that integrates the detection results from the user switching detection unit 12 and the judgment results from the posture determination unit 14. For example, the output unit 13 may output warning information if the user's usage time of the operating terminal, calculated based on the detection results from the user switching detection unit 12, exceeds a standard value, and the confidence level calculated by the posture determination unit 14 also exceeds a standard value.
[0114] The output destination could be, for example, the operating terminal itself or a display or speaker installed near the operating terminal. In this case, the warning information could be a cautionary message such as, "Please be careful of wire fraud." Alternatively, the output destination could be a display or speaker viewed by a supervisor or the administrator of the operating terminal, or a mobile device carried by such a person. In this case, the warning information could be a cautionary message such as, "The usage time of customer using operating terminal number 3 has exceeded the limit. Furthermore, this customer is operating the terminal while on a call. This may be wire fraud. Please check."
[0115] The software configuration of the image processing device 10 is the same as in the first embodiment.
[0116] The image processing apparatus 10 of this embodiment achieves the same effects and advantages as those of the first and second embodiments.
[0117] The embodiments of the present invention have been described above with reference to the drawings, but these are merely examples of the present invention, and various other configurations can also be adopted.
[0118] In this specification, "acquisition" includes at least one of the following: "active acquisition" based on user input or program instructions, such as "the device retrieving data stored in another device or storage medium," for example, receiving data by requesting or querying another device, accessing and reading data by another device or storage medium; "passive acquisition" based on user input or program instructions, such as "inputting data output from another device into the device," for example, receiving data that is distributed (or transmitted, push notification, etc.), selecting and acquiring data from the received data or information; and "generating new data by editing data (converting to text, rearranging data, extracting some data, changing file format, etc.) and acquiring said new data."
[0119] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. User area identification means for identifying the user area, which is the area where the user of the operating terminal is located, from within the image to be processed. A user switching detection means for detecting that the user of the operating terminal has switched based on the image of the user area, An image processing device having 2. The image processing apparatus according to claim 1, wherein the user switching detection means detects that the user of the operating terminal has switched based on the result of comparing feature data extracted from the image of the user region in the image to be processed with feature data extracted from the image of the user region in a comparison target image that was generated before the image to be processed. 3. The user switching detection means is: Based on the comparison result between the feature data extracted from the image to be processed and the feature data extracted from the image to be compared, the process of determining whether the person included in the user area image in the image to be processed and the person included in the user area image in the image to be compared are the same person is repeatedly performed on multiple images in sequence as the image to be processed. If it is determined that M or more consecutive images to be processed are not of the same person, it is determined that the user of the operating terminal has changed. The image processing apparatus according to claim 2, wherein the timing at which the earliest chronologically ordered image among the M consecutive images to be processed is generated is determined to be the timing at which the user of the operating terminal switches. 4. The user area identification means is: An image processing apparatus according to any one of 1 to 3, which identifies the user area from the image to be processed based on a first detection result in which a human area is detected from the image to be processed and a second detection result in which characteristic points of a human skeleton are detected from the image to be processed. 5. The user area identification means is: The image processing apparatus according to 4, which identifies a person region where a user of the operating terminal is located, based on the size of the person region and the characteristic points of the person's skeleton, with respect to the plurality of person regions indicated in the first detection result. 6. The user area identification means is: If feature points of the person's skeleton are not detected in the image to be processed, the image processing apparatus according to 4 or 5 identifies the user area from the image to be processed based on feature points of the person's skeleton detected in a reference image generated before the image to be processed. 7. An image processing apparatus according to any one of 1 to 6, further comprising an output means for outputting warning information when the user usage time of the operation terminal, calculated based on the detection result of the user switching detection means, exceeds a standard value. 8. An image processing apparatus according to any one of 1 to 6, further comprising posture determination means for calculating the degree of confidence that the user of the operating terminal is in a predetermined posture based on the image of the user area. 9. The image processing apparatus according to 8, further comprising an output means for outputting warning information when the user usage time of the operating terminal calculated based on the detection result of the user switching detection means exceeds a standard value, and the confidence level calculated by the posture determination means exceeds a standard value. 10. Multiple images become the images to be processed in chronological order. The posture determination means is The process of detecting the predetermined posture from the image to be processed is performed for each of the multiple images to be processed. The image processing apparatus according to 8 or 9, which determines the confidence level according to the number of times the predetermined posture is detected consecutively. 11. The posture determination means is: Based on the feature points of interest among the skeletal feature points of a person detected from the image to be processed, the process of detecting the predetermined posture is executed. As a process for detecting the predetermined posture, the feature point of interest that has reached a predetermined state is detected, If the feature point of interest that is in the predetermined state is detected from the image to be processed, the confidence level is increased. If a feature point of interest that is not in the predetermined state is detected in the image to be processed, the confidence level is reset to its initial value. The image processing apparatus according to claim 10, which maintains the confidence level if the feature point of interest is not detected in the image to be processed. 12. The posture determination means is: Based on the feature points of the right half of the human skeleton detected from the image to be processed, the degree of confidence that the right half of the human is in the predetermined posture is calculated. Based on the feature points of the left half of the human skeleton detected from the image to be processed, the degree of confidence that the left half of the human is in the predetermined posture is calculated. An image processing apparatus according to any one of 8 to 11, wherein the greater of the confidence that the right half of the person is in the predetermined posture and the confidence that the left half of the person is in the predetermined posture is used as the confidence that the user of the operating terminal is in the predetermined posture. 13. Computers, From the image to be processed, the user area, which is the area where the user of the operating terminal is located, is identified. An image processing method for detecting that the user of the operating terminal has changed based on the image of the user area. 14. Computers, A user area identification means for identifying the user area, which is the area where the user of the operating terminal is located, from within the image to be processed. A user switching detection means that detects when the user of the operating terminal has switched based on the image of the user area. A program that makes it function as such.
[0120] This application claims priority based on Japanese Patent Application No. 2021-006332, filed on 19 January 2021, and incorporates all of its disclosures herein. [Explanation of symbols]
[0121] 10 Image Processing Device 11 User area identification section 12 User switching detection unit 13 Output section 14 Posture determination section 1A Processor 2A Memory 3A input / output I / F 4A Peripheral Circuits 5A bus
Claims
1. The process involves acquiring images from a camera positioned to capture a person operating an ATM (automated teller machine), and The process of obtaining the usage time of the aforementioned ATM, Based on the aforementioned image, a process is performed to detect the posture of the person, If the aforementioned usage time and the detection result of the person's posture each meet predetermined criteria, a process is performed to output a warning about fraud. An image processing program equipped with the following features.
2. The process that outputs a warning about the aforementioned fraud is: The image processing program according to claim 1, which outputs a warning about fraud when the usage time exceeds a standard value and the user is shown to be in a predetermined posture.
3. The image processing program according to claim 2, wherein the predetermined posture includes a telephone posture.
4. The process for detecting the posture of the person is as follows: The image processing program according to claim 3, comprising: a process for determining the characteristic points of a person's skeleton; and a process for determining posture based on the angle formed by a line segment connecting a first characteristic point and a second characteristic point, and a line segment connecting the second characteristic point and a third characteristic point.
5. The process for obtaining the ATM usage time is the image processing program according to claim 1, which is required for each user of the ATM.
6. The image processing program according to claim 5, further comprising a process for determining whether the user of the ATM has changed based on the aforementioned image.
7. The image processing program according to claim 1, wherein the process of outputting the aforementioned fraud warning is performed by the ATM.
8. The image processing program according to claim 7, which includes outputting a warning about the aforementioned fraud to an information processing device different from the ATM, and outputting a message of a different length than the message output by the ATM.
9. One or more computers, The process involves acquiring images from a camera positioned to capture a person operating an ATM, and The process of obtaining the usage time of the aforementioned ATM, Based on the aforementioned image, a process is performed to detect the posture of the person, If the aforementioned usage time and the detection result of the person's posture each meet predetermined criteria, a process is performed to output a warning about fraud. An image processing method that performs this task.
10. A means of acquiring images from a camera installed in a position that allows for the capture of a person operating an ATM, A means for obtaining the usage time of the aforementioned ATM, A means for detecting a person's posture based on the aforementioned image, If the aforementioned usage time and the detection result of the person's posture each meet predetermined criteria, a means for outputting a warning about fraud, An image processing system equipped with the following features.
Citation Information
Patent Citations
Phone call decision device, its method, and program
JP2010218392A
Monitoring device and monitoring method
JP2010238204A