Image processing device, image processing method, and program
The image processing device accurately estimates the purpose of individuals in surveillance videos by analyzing clothing, possessions, posture, and movement, improving surveillance accuracy and distinguishing between legitimate and suspicious activities.
Patent Information
- Application Number
- JP2024507288
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2042-03-16
AI Technical Summary
Existing surveillance technologies lack the ability to accurately estimate the purpose of individuals in captured images, leading to inaccuracies when determining purposes based on appearance or frequency, and fail to distinguish between legitimate activities and suspicious behavior.
An image processing device that acquires video, generates appearance information on clothing, possessions, posture, and movement, and appearance degree information, and estimates the purpose of individuals in a target area using predefined criteria.
Accurately estimates the purpose of individuals in a video by combining appearance and frequency information, distinguishing between workers, passersby, and suspicious activities, enhancing surveillance effectiveness.
Smart Images

Figure 0007722554000001 
Figure 0007722554000002 
Figure 0007722554000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, and a recording medium. [Background technology]
[0002] Techniques related to the present invention are disclosed in Patent Documents 1 to 3 and Non-Patent Document 1.
[0003] Patent Document 1 discloses a technology that uses image analysis to recognize uniforms, caps, company logos, possession of cardboard boxes, etc., and determines whether a delivery person has arrived based on the recognition results.
[0004] Patent Document 2 discloses a technique for extracting people whose appearance frequency satisfies a predetermined condition from a moving image.
[0005] Patent Document 3 discloses a technology that calculates the feature values of each of multiple key points of a human body contained in an image, searches for images containing human bodies with similar postures or movements based on the calculated feature values, and classifies images with similar postures or movements together.
[0006] Non-Patent Document 1 discloses a technique related to human skeleton estimation. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Patent Publication No. 2021-022767 [Patent Document 2] International Publication No. 2017 / 077902 [Patent Document 3] International Publication No. 2021 / 084677 [Non-patent literature]
[0008] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 Summary of the Invention [Problem to be solved by the invention]
[0009] Surveillance cameras are becoming widespread. Rather than simply capturing images with a surveillance camera, the ability to estimate the purpose of people in the captured images will not only increase crime prevention effectiveness, but also enable the results of the estimation to be used in a variety of applications.
[0010] The technology disclosed in Patent Document 1 is a technology for determining whether a delivery person has arrived, and cannot determine other purposes. Furthermore, when determining based solely on the appearance of a person, as in the technology disclosed in Patent Document 1, the accuracy of the determination is poor. For example, if the person is disguised (e.g., disguised as a delivery person), an incorrect determination will be made.
[0011] The technology disclosed in Patent Document 2 is a technology for extracting people whose appearance frequency satisfies a predetermined condition from a video, but is not a technology for estimating the purpose of the people in the video. Furthermore, when extraction is based solely on appearance frequency, as in the technology disclosed in Patent Document 2, the accuracy of extraction decreases.
[0012] Patent Document 3 and Non-Patent Document 1 are technologies for estimating the posture and movement of a person, but are not technologies for estimating the purpose for which a target person is in a target area.
[0013] In view of the above-mentioned problems, one example of the object of the present invention is to provide an image processing device, an image processing method, and a recording medium that solve the problem of accurately estimating the purpose of a person included in a moving image. [Means for solving the problem]
[0014] According to one aspect of the present invention, an acquisition means for acquiring a moving image of a target area; an analysis means for generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; an estimation means for estimating a purpose for which the target person is in the target area based on the appearance information and the appearance degree information; An image processing apparatus is provided, comprising:
[0015] According to one aspect of the present invention, The computer Acquire a video image of the target area, generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; estimating a purpose for which the target person is in the target area based on the appearance information and the appearance degree information; An image processing method is provided.
[0016] According to one aspect of the present invention, Computer, an acquisition means for acquiring a moving image of the target area; an analysis means for generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; an estimation means for estimating a purpose for which the target person is in the target area based on the appearance information and the appearance degree information; A recording medium is provided on which a program for causing the device to function as the above is recorded. [Effects of the Invention]
[0017] According to one aspect of the present invention, an image processing device, an image processing method, and a recording medium are realized that solve the problem of accurately estimating the purpose for which a person included in a moving image is present. [Brief explanation of the drawings]
[0018] The above and other objects, features and advantages are described below. Suitable This will become more apparent from the following embodiments and the accompanying drawings.
[0019] [Figure 1] FIG. 1 is a diagram illustrating an example of a functional block diagram of an image processing apparatus. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of an image processing apparatus. [Figure 3] FIG. 10 is a diagram for explaining the processing of an analysis unit. [Figure 4] FIG. 10 is a diagram illustrating an example of information output by the image processing apparatus. [Figure 5] 10 is a flowchart illustrating an example of a processing flow of the image processing device. [Figure 6] FIG. 1 is a diagram illustrating an example of a functional block diagram of an image processing apparatus. [Figure 7] FIG. 10 is a diagram illustrating an example of information output by the image processing apparatus. DETAILED DESCRIPTION OF THE INVENTION
[0020] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0021] First Embodiment 1 is a functional block diagram showing an overview of an image processing device 10 according to the first embodiment. The image processing device 10 includes an acquisition unit 11, an analysis unit 12, and an estimation unit 13.
[0022] The acquisition unit 11 acquires a video of a target area. The analysis unit 12 generates, based on the video, appearance information indicating at least one of the type of clothing, the type of belongings, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video. The estimation unit 13 estimates the purpose of the target person being in the target area based on the appearance information and the appearance degree information.
[0023] The image processing device 10 having such a configuration solves the problem of accurately estimating the purpose for which a person included in a moving image is there.
[0024] <Second embodiment> "overview" Second embodiment image The processing device 10 is a more specific embodiment of the image processing device 10 of the first embodiment. The image processing device 10 estimates the purpose of the target person's presence in the target area based on appearance information indicating at least one of the target person's type of clothing, type of belongings, posture, and movement, and appearance frequency information indicating the frequency with which the target person appears in the video. The configuration of such an image processing device 10 will be described in more detail below.
[0025] "Hardware Configuration" Next, an example of the hardware configuration of image processing device 10 will be described. Each functional unit of image processing device 10 is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, programs loaded into memory, a storage unit such as a hard disk that stores the programs (this can store programs that are pre-loaded when the device is shipped, as well as programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface. Those skilled in the art will understand that there are many variations in the realization methods and devices.
[0026] FIG. 2 is a block diagram illustrating an example of the hardware configuration of an image processing device 10. As shown in FIG. 2, the image processing device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The image processing device 10 does not necessarily have to have the peripheral circuit 4A. Note that the image processing device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration.
[0027] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data among them. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. Examples of input devices include a keyboard, a mouse, a microphone, physical buttons, a touch panel, etc. Examples of output devices include a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0028] "Function Configuration" Next, the functional configuration of the image processing device 10 according to the second embodiment will be described in detail. An example of a functional block diagram of the image processing device 10 according to the second embodiment is shown in FIG. 1. As shown in the figure, the image processing device 10 includes an acquisition unit 11, an analysis unit 12, and an estimation unit 13.
[0029] The acquisition unit 11 acquires video images of a target area. The "target area" is an area that is captured for purposes such as surveillance. For example, the target area may be near the entrance of a building or near an entrance / exit to a site, but is not limited to these. A camera is installed at a position and orientation that captures the video images of the target area.
[0030] The acquisition unit 11 acquires moving images captured by the camera. For example, the camera and the image processing device 10 may be configured to be able to communicate with each other. The acquisition unit 11 may then acquire moving images transmitted by the camera. Alternatively, a moving image file generated by the camera may be input to the image processing device 10 by any means, such as user input. The acquisition unit 11 may then acquire the moving image file input in this manner.
[0031] The analysis unit 12 generates appearance information and appearance frequency information based on the moving image acquired by the acquisition unit 11.
[0032] The "appearance information" indicates at least one of the type of clothing, the type of possessions, the posture, and the movement of a person (hereinafter referred to as "target person") included in the video.
[0033] The "type of clothing" indicated by the appearance information is one of a plurality of predefined clothing types. The plurality of clothing types are predefined so as to include the clothing types of people who may come into the target area. The people who may come into the target area include, for example, workers who perform various tasks and other people. The workers may include at least one of a delivery person who makes a delivery, a flyer distributor who posts flyers, a gas inspector who checks the gas meter, a waterworks employee who checks the water meter, and a garbage collector who collects garbage. Note that the examples given here are merely examples and are not limited to these. Other examples will be described in the following embodiments.
[0034] Assuming that the above-mentioned people will come into the target area, multiple types of clothing may be defined in advance to include a category corresponding to various uniforms and a category called "other clothing (casual clothing)." The various uniforms are the uniforms of various workers who may come into the target area, and examples include the uniforms of delivery people from a first company, delivery people from a second company, gas inspector uniforms, waterworks employee uniforms, and garbage collector uniforms. There are an enormous number of variations in casual clothing. Therefore, they can be grouped into a single category called "other clothing (casual clothing)."
[0035] The "type of possessions" indicated by the appearance information is one of a plurality of predefined types of possessions. The plurality of types of possessions are predefined so as to include types of possessions that are typically possessed by people who may come into the target area. Examples of people who may come into the target area are as described above.
[0036] Assuming that the above-mentioned people will come to the target area, multiple types of possessions may be defined in advance to include deliveries, flyers, dedicated terminals carried by delivery personnel, dedicated terminals carried by gas inspectors, dedicated terminals carried by waterworks personnel, etc. Note that the examples given here are merely examples and are not limiting. Other examples will be described in the following embodiments.
[0037] The "pose" and "movement" indicated by the appearance information are any of a plurality of predefined poses and movements. The plurality of poses and movements are predefined so as to include poses and movements typically performed by people who may be in the target area. Examples of people who may be in the target area are as described above.
[0038] Assuming that the above-mentioned person will come into the target area, a plurality of postures and movements may be predefined, including, for example, postures and movements of standing and waiting in front of the door with a delivery, postures and movements of handing over a delivery, postures and movements of dropping off a flyer in a mailbox, postures and movements of checking a gas meter, postures and movements of checking a water meter, postures and movements of collecting garbage, etc. Note that the examples given here are merely examples and are not limiting. Other examples will be described in the following embodiments.
[0039] "Appearance degree information" indicates the degree to which the target person appears in the video. The appearance degree information may, for example, indicate the length of time that the target person continuously appears in the video. Alternatively, the appearance degree information may indicate the frequency with which the target person appears in the video within a predetermined period. The predetermined period can be set arbitrarily, such as one day, one week, or one month. The frequency may be indicated by the number of times the target person appears within the predetermined period, or by the length of time the target person appears within the predetermined period. There are various ways to count the number of times. For example, the time from when the target person appears in the video until they are no longer visible may be counted as one occurrence, or other methods may be used.
[0040] The analysis unit 12 can generate the appearance information and the appearance degree information based on the analysis results of the video. The analysis of the video is performed by a pre-prepared image analysis system 20. As shown in FIG. 3, the analysis unit 12 inputs the video to the image analysis system 20. Then, the analysis unit 12 obtains the analysis results of the video from the image analysis system 20. The image analysis system 20 may be a part of the image processing device 10, or may be an external device that is physically and / or logically independent from the image processing device 10.
[0041] Here, we will explain the image analysis system 20. The image analysis system 20 has at least one of a face recognition function, a human shape recognition function, a posture recognition function, a movement recognition function, an appearance attribute recognition function, an image gradient feature detection function, an image color feature detection function, an object recognition function, a character recognition function, and a gaze detection function.
[0042] The face recognition function extracts facial features of a person. It may also compare and calculate the similarity between facial features (to determine whether they are the same person, etc.). It may also compare the extracted facial features with facial features of people who have previously been in the target area and are registered in a database, to determine whether the person in the image is a person who has previously been in the target area.
[0043] The human recognition function extracts a person's physical features (e.g., overall features such as whether or not they are fat or thin, height, clothing, etc.). Furthermore, the similarity between the physical features may be compared and calculated (e.g., to determine whether they are the same person). The extracted physical features may also be compared with the physical features of people who have previously visited the target area and are registered in a database, to determine whether the person in the image is a person who has previously visited the target area.
[0044] The posture recognition function and movement recognition function detect the joint points of a person and connect the joint points to create a stick figure model. Then, based on the stick figure model, the person is detected, the person's height is estimated, posture features are extracted, and movements are identified based on changes in posture. Specifically, postures and movements typically performed by people who may enter the target area are defined in advance, and these postures and movements are detected. Furthermore, similarities between posture features and movement features may be compared and calculated (e.g., determining whether the postures or movements are the same). The estimated height may also be compared with the heights of people who have previously entered the target area and are registered in a database, to identify whether the person in the image has previously entered the target area. The posture recognition function and movement recognition function may be implemented using the technology disclosed in Patent Document 3 and Non-Patent Document 1.
[0045] The appearance attribute recognition function recognizes appearance attributes associated with a person (for example, type of clothing, shoe color, hairstyle, whether a hat or tie is worn, etc., for a total of more than 100 types of appearance attributes). For example, the types of clothing of people who may enter the target area described above are defined in advance, and the type of clothing of people appearing in an image is recognized. Furthermore, the similarity of the recognized appearance attributes may be compared and calculated (it is possible to determine whether the attributes are the same). The recognized appearance attributes may also be compared with the appearance attributes of people who have previously entered the target area and are registered in a database, to determine whether the person appearing in the image is a person who has previously entered the target area.
[0046] Image gradient feature detection functions include SIFT, SURF, RIFF, ORB, BRISK, CARD, HOG, etc. According to these functions, the gradient features of each frame image are detected.
[0047] The image color feature detection function generates data indicating the color features of the image, such as a color histogram, and detects the color features of each frame image.
[0048] The object recognition function is realized using an engine such as YOLO (which can extract general objects (such as tools and equipment used in sports and other performances) and people). By using the object recognition function, various predefined objects can be detected from images. Specifically, items typically carried by people who may enter the target area described above are predefined, and these items are detected.
[0049] The character recognition function recognizes numbers, letters, etc.
[0050] The gaze detection function detects the gaze direction of people in the image.
[0051] The analysis unit 12 generates the appearance information and appearance degree information based on the analysis results received from the image analysis system 20 described above.
[0052] Returning to FIG. 1, the estimation unit 13 estimates the purpose of the target person being in the target area based on the appearance information and appearance degree information generated by the analysis unit 12.
[0053] For example, the estimation unit 13 may estimate the purpose for which the target person is in the target area based on the following criteria.
[0054] If the content of the appearance information satisfies a first condition regarding the worker, and the length of time that the target person continuously appears in the video in the appearance degree information is equal to or greater than a first reference value and less than a second reference value, it is estimated that the target person is in the target area to be working. If the content of the appearance information satisfies a first condition regarding the worker, and the length of time that the target person continuously appears in the video in the appearance degree information is less than a first reference value, it is presumed that the target person is in the target area simply passing through. If the content of the appearance information satisfies a first condition regarding the worker, and the length of time that the target person continuously appears in the video in the appearance degree information is equal to or exceeds a second reference value, it is presumed that the target person's purpose in being in the target area is to engage in suspicious behavior.
[0055] The content of the first condition, the first reference value, and the second reference value differ for each worker.
[0056] For example, multiple first conditions are predefined for each worker, such as a first condition for a delivery person from a first company, a first condition for a delivery person from a second company, a first condition for a gas inspector, a first condition for a waterworks employee, a first condition for a garbage collector, and a first condition for a flyer distributor.
[0057] The first condition specifies at least one of the type of clothing, the type of belongings, the posture, and the movement. For example, in the first condition for a delivery person of a first company, the type of clothing is specified as the uniform of the delivery person of the first company, the type of belongings is specified as the delivery item, and the posture and movement are specified as the posture and movement of standing in front of the door while holding the delivery item. When the content of the appearance information matches the content specified in the first condition, the estimation unit 13 determines that the content of the appearance information satisfies the first condition.
[0058] The time range between the first and second reference values indicates an estimate of the time required for each task performed by each worker in the target area. For example, a plurality of first and second reference values are predefined for each worker, such as a first and second reference value for a delivery person from a first company, a first and second reference value for a delivery person from a second company, a first and second reference value for a gas inspector, a first and second reference value for a waterworks employee, a first and second reference value for a garbage collector, and a first and second reference value for a flyer distributor.
[0059] According to the above criteria, if the time that a target person continuously appears in the target area is within the estimated time, it is presumed that the target person is in the target area to be working. If the time that a target person continuously appears in the target area is less than the estimated time, it is presumed that the target person is simply passing by (for example, if the target person is captured in a video image while passing in front of a house). If the time that a target person continuously appears in the target area is longer than the estimated time, it is presumed that the target person is in the target area to be engaging in suspicious behavior. A specific example of suspicious behavior could be someone dressed as a worker and scouting the area for a crime.
[0060] First, the estimation unit 13 determines whether the appearance information of the target person satisfies the first condition of any worker. Then, if the appearance information of the target person satisfies the first condition of any worker, the estimation unit 13 determines which of the three conditions (greater than or equal to the first reference value and less than the second reference value, less than the first reference value, or greater than or equal to the second reference value) the appearance frequency information of the target person satisfies based on the first and second reference values of that worker.
[0061] The estimation unit 13 can output the estimation result. The output may be displayed on a display or a projection device, printed on a printer, or transmitted to an external device. For example, as shown in FIG. 4, the purpose of each of multiple target persons detected in a video may be displayed in a list in chronological order along with the date and time of detection.
[0062] Next, an example of the processing flow of the image processing device 10 will be described with reference to the flowchart of FIG.
[0063] When the image processing device 10 acquires a moving image of a target area (S10), it generates, based on the acquired moving image, appearance information indicating at least one of the type of clothing, the type of belongings, the posture, and the movement of the target person included in the moving image, and appearance degree information indicating the degree to which the target person appears in the moving image (S11).Then, the image processing device 10 estimates the purpose of the target person being in the target area based on the appearance information and the appearance degree information (S12).
[0064] "Action and effect" According to the image processing device 10 of the second embodiment, the purpose of a person in a video can be estimated based on both appearance information indicating at least one of the type of clothing, the type of belongings, the posture, and the movement of the person, and appearance degree information indicating the degree to which the person appears in the video. Estimating the purpose based on both the appearance information and the appearance degree information improves the accuracy of estimation.
[0065] Furthermore, by estimating the purpose of a person in an image using the characteristic criteria described above, it is possible to accurately distinguish between multiple purposes, such as "workers working," "workers passing by," and "suspicious behavior by a person dressed as a worker."
[0066] <Third embodiment> In the third embodiment, the appearance information of the target person further indicates the appearance characteristics of the vehicle used by the target person. For example, the analysis unit 12 can identify the vehicle in which the target person was riding as the vehicle used by the target person. The vehicle in which the target person was riding can be identified by detecting the behavior of getting on and off the vehicle in the image. The first condition specifies at least one of the type of clothing, the type of belongings, posture, movement, and vehicle appearance characteristics. The vehicle appearance characteristics relate to the design characteristics of the vehicle, such as a company logo, company name, or company-specific pattern displayed on the exterior of the vehicle. For example, the first condition for a delivery person of a first company specifies the first company's delivery person's uniform as the type of clothing, the type of belongings as the delivery item, the posture and movement of standing in front of the door while holding the delivery item, and the appearance characteristics of the vehicle used for delivery by the first company.
[0067] Other configurations of the image processing device 10 of the third embodiment are similar to those of the image processing device 10 of the first and second embodiments.
[0068] The image processing device 10 of the third embodiment achieves the same effects as the image processing device 10 of the first and second embodiments. In addition, workers may travel using vehicles designed with their company logo, company name, company-specific patterns, etc. By utilizing the external features of the vehicle used by the target person, it becomes possible to more accurately estimate the purpose of the target person being in the target area.
[0069] <Fourth embodiment> In the fourth embodiment, another specific example of processing for estimating the purpose for which a target person is in a target area based on appearance information and appearance degree information will be described.
[0070] For example, if a person is wearing the uniform of a delivery person from a first company, carrying a delivery item, and the frequency of appearance is above a reference value (for example, more than five times in the past week), the estimation unit 13 can estimate that the person is a delivery person from the first company who is trying to deliver a delivery item.
[0071] In addition, when the person is wearing other clothes (plain clothes), carrying a stack of papers (flyers), and the frequency of appearance is equal to or less than a reference value (for example, less than twice in the last week), the estimation unit 13 determines that the person is a flyer intended to deliver a flyer. distribution It can be presumed that the
[0072] In addition, if the person is wearing other clothes (plain clothes), carrying a specified case, and the frequency of appearance is below a reference value (for example, less than four times in the past week), the estimation unit 13 can estimate that the person is a delivery person whose purpose is to deliver food.
[0073] Furthermore, if the person assumes a posture or movement appropriate for checking the water meter and the frequency of occurrence is within a predetermined reference range (for example, 1 to 2 times in the past month), the estimation unit 13 can estimate that the person is a waterworks employee whose purpose is to check the water meter.
[0074] Furthermore, if the person assumes a posture or movement appropriate for checking a gas meter and the frequency of occurrence is within a predetermined reference range (for example, 1 to 2 times in the last month), the estimation unit 13 can estimate that the person is a gas inspector whose purpose is to check the gas meter.
[0075] Furthermore, if the person is wearing a garbage collector's uniform, carrying an arbitrary object, using a garbage truck, and the frequency of appearance is below a reference value (for example, less than five times in the last week), the estimation unit 13 can estimate that the person is a garbage collector whose purpose is to collect garbage.
[0076] Furthermore, if the person is wearing other clothes (plain clothes) and carrying pocket tissues, and the length of time that they appear continuously in the video meets a predetermined condition (for example, between 3 and 8 hours), the estimation unit 13 can estimate that they are a pocket tissue distributor who intends to distribute pocket tissues.
[0077] Furthermore, if the person is wearing the uniform of a specified moving company, carrying luggage, using a truck, and the time that they continuously appear in the video meets a specified condition (for example, between 30 minutes and 2 hours), the estimation unit 13 can estimate that they are a moving company whose purpose is to carry out moving work.
[0078] Furthermore, if the person is wearing a predetermined uniform, holding a stick or flag in their hand and assuming a traffic control pose, and the time during which they continuously appear in the video satisfies a predetermined condition (for example, between 1 hour and 8 hours), the estimation unit 13 can estimate that they are a traffic controller whose purpose is to direct traffic.
[0079] Furthermore, if the person is wearing a specified uniform, holding construction tools for road repair or the like, appears near a road, and the time that they continuously appear in the video satisfies a specified condition (for example, between 3 and 8 hours), the estimation unit 13 can estimate that they are road construction workers whose purpose is to maintain the road.
[0080] Furthermore, if the person is wearing a predetermined uniform, holding a device such as a counter, appears near a road or intersection, and the time that they continuously appear in the video satisfies a predetermined condition (for example, one hour or more), the estimation unit 13 can estimate that they are a traffic surveyor whose purpose is to survey traffic volume.
[0081] Furthermore, if the person is wearing a police uniform, using a police motorcycle or patrol car, and appears continuously in the video for a period of time that satisfies a predetermined condition (for example, one hour or more), the estimation unit 13 can estimate that the person is a traffic police officer whose purpose is to manage traffic or investigate traffic accidents.
[0082] Furthermore, if the person is wearing a predetermined uniform, holding a sign, and appears continuously in the video for a period of time that satisfies a predetermined condition (for example, between 1 hour and 8 hours), the estimation unit 13 can estimate that the person is a sign-holding staff member whose purpose is to introduce a signboard advertisement.
[0083] Furthermore, if the person is wearing a specified uniform, holding cleaning tools, appears around the stormwater tanks on both sides of the road, and the time that they continuously appear in the video satisfies a specified condition (for example, between 2 minutes and 2 hours), the estimation unit 13 can estimate that they are road sweepers whose purpose is to clean the stormwater tanks on the road.
[0084] Furthermore, if the person is wearing a specified uniform, holding flowers, plants, pots, watering tools, etc., appears near a flower bed or planter, and the time that they continuously appear in the video meets a specified condition (for example, between 30 minutes and 5 hours), the estimation unit 13 can estimate that they are a flower planter whose purpose is to plant flowers.
[0085] Furthermore, if the person is wearing other clothes (plain clothes), has a lunch box and a desk (belongings) placed nearby, and the time that the person appears continuously in the video meets a predetermined condition (for example, between 30 minutes and 3 hours), the estimation unit 13 can estimate that the person is a street bento vendor intending to sell bento boxes on the street.
[0086] Furthermore, if the person is wearing other clothes (plain clothes), has a microphone or musical instrument (possession) nearby, is walking or dancing, and the time that they continuously appear in the video meets a predetermined condition (for example, between 30 minutes and 5 hours), the estimation unit 13 can estimate that they are a street performer who is aiming to perform a street live show.
[0087] Other configurations of the image processing device 10 of the fourth embodiment are similar to those of the image processing device 10 of the first to third embodiments.
[0088] According to the image processing device 10 of the fourth embodiment, the same effects as those of the image processing device 10 of the first to third embodiments are realized. Furthermore, according to the image processing device 10 of the fourth embodiment, it is possible to estimate various purposes for which various people are present at a location.
[0089] <Fifth embodiment> The image processing device 10 of the fifth embodiment has a function of statistically processing the estimated results of the purpose for which the target person is in the target area by time period.
[0090] 6 shows an example of a functional block diagram of an image processing device 10 according to the fifth embodiment. As shown in the figure, the image processing device 10 includes an acquisition unit 11, an analysis unit 12, an estimation unit 13, and a statistics unit 14.
[0091] The statistical unit 14 statistically processes the estimated results of the purpose of the target person being in the target area by time period. The statistical unit 14 can then output the results of the statistical processing. The output can be displayed on a display or projection device, printed on a printer, sent to an external device, etc.
[0092] 7 shows an example of the results of statistical processing by the statistics unit 14. In the example shown, the number of times each purpose was estimated, that is, the number of times a person was detected in the target area for each purpose, is shown for each time period.
[0093] Other configurations of the image processing device 10 of the fifth embodiment are similar to those of the image processing device 10 of the first to fourth embodiments.
[0094] The image processing device 10 of the fifth embodiment achieves the same effects as the image processing device 10 of the first to fourth embodiments. Furthermore, the image processing device 10 of the fifth embodiment can statistically process and output the estimated results of the purpose of the target person being in the target area by time period. The user can use this information for crime prevention purposes or marketing purposes.
[0095] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations may be adopted. The configurations of the above-described embodiments may be combined with each other, or some of the configurations may be replaced with other configurations. Furthermore, various modifications may be made to the configurations of the above-described embodiments without departing from the spirit of the invention. Furthermore, the configurations and processes disclosed in the above-described embodiments and modified examples may be combined with each other.
[0096] In addition, although the flowcharts used in the above explanations show multiple steps (processes) in a sequential order, the order of steps executed in each embodiment is not limited to the order shown. In each embodiment, the order of the steps shown in the drawings can be changed as long as it does not cause any problems in terms of content. Furthermore, the above-mentioned embodiments can be combined as long as the content is not contradictory.
[0097] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. 1. An acquisition means for acquiring a moving image of a target area; an analysis means for generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; an estimation means for estimating a purpose for which the target person is in the target area based on the appearance information and the appearance degree information; An image processing device having: 2. The image processing device described in 1, wherein the estimation means estimates that the purpose is work if the appearance information satisfies a first condition regarding the worker and the appearance degree information indicates that the length of time during which the target person continuously appears in the video is equal to or greater than a first reference value and less than a second reference value. 3. The image processing device described in 2, wherein the estimation means estimates that the object is passing if the appearance information satisfies the first condition and the length of time that the target person continuously appears in the video in the appearance degree information is less than the first reference value. 4. An image processing device as described in 2 or 3, wherein the estimation means estimates that the purpose is suspicious behavior if the appearance information satisfies the first condition and the length of time that the target person continuously appears in the video in the appearance degree information is equal to or longer than the second reference value. 5. An image processing device according to any one of 2 to 4, wherein the worker includes at least one of a delivery person whose job is to make deliveries, a flyer distributor whose job is to post flyers, a gas inspector whose job is to check gas meters, a waterworks employee whose job is to check water meters, and a garbage collector whose job is to collect garbage. 6. An image processing device described in any one of 1 to 5, wherein the appearance information further indicates the appearance characteristics of the vehicle used by the target person. 7. An image processing device according to any one of 1 to 6, further comprising statistical means for statistically processing the estimated results of the purpose for which the target person is in the target area by time period. 8. The computer Acquire a video image of the target area, generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; estimating a purpose for which the target person is in the target area based on the appearance information and the appearance degree information; Image processing methods. 9. Computer an acquisition means for acquiring a moving image of the target area; an analysis means for generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; an estimation means for estimating a purpose for which the target person is in the target area based on the appearance information and the appearance degree information; A recording medium on which a program that functions as a [Explanation of symbols]
[0098] 10 Image processing device 11 Acquisition Department 12 Analysis Department 13 Estimation part 14 Statistics Department 1A processor 2A Memory 3A input / output I / F 4A peripheral circuit 5A Bus
Claims
1. an acquisition means for acquiring a moving image of a target area; an analysis means for generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; an estimation means for, when it is determined based on the appearance information that the target person has the appearance of a first type of worker, comparing a standard for the degree to which the first type of worker appears in the video with the degree to which the target person appears in the video, which is indicated by the appearance degree information, and estimating, based on the comparison result, whether the target person is a suspicious person disguised as the first type of worker; An image processing device having:
2. 2. The image processing device according to claim 1, wherein the estimation means estimates that the purpose of the target person being in the target area is work if the appearance information satisfies a first condition regarding the worker and the appearance degree information indicates that the length of time the target person continuously appears in the moving image is equal to or greater than a first reference value and less than a second reference value.
3. The image processing device according to claim 2, wherein the estimation means estimates that the object is passing if the appearance information satisfies the first condition and the appearance degree information indicates that the length of time during which the target person continuously appears in the moving image is less than the first reference value.
4. The image processing device described in claim 2 or 3, wherein the estimation means estimates that the purpose is suspicious behavior if the appearance information satisfies the first condition and the appearance degree information indicates that the length of time the target person continuously appears in the moving image is equal to or longer than the second reference value.
5. The image processing device according to any one of claims 2 to 4, wherein the worker includes at least one of a delivery person whose job is to deliver, a flyer distributor whose job is to post flyers, a gas inspector whose job is to check gas meters, a waterworks employee whose job is to check water meters, and a garbage collector whose job is to collect garbage.
6. The image processing device according to claim 1 , wherein the appearance information further indicates an appearance feature of a vehicle used by the target person.
7. The image processing device according to claim 1 , further comprising statistical means for statistically processing, by time period, the results of estimation of the purpose for which the target person is in the target area.
8. The computer Acquire a video image of the target area, generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; If it is determined based on the appearance information that the target person has the appearance of a first type of worker, the method compares a standard for the degree to which the first type of worker appears in the video with the degree to which the target person appears in the video, as indicated by the appearance degree information, and, based on the comparison result, estimates whether the target person is a suspicious person pretending to be the first type of worker. Image processing methods.
9. Computer, an acquisition means for acquiring a moving image of the target area; an analysis means for generating, based on the video, appearance information indicating at least one of the type of clothing, the type of possessions, the posture, and the movement of the target person included in the video, and appearance degree information indicating the degree to which the target person appears in the video; an estimation means for, when it is determined based on the appearance information that the target person has the appearance of a first type of worker, comparing a standard for the degree to which the first type of worker appears in the video with the degree to which the target person appears in the video, which is indicated by the appearance degree information, and estimating based on the comparison result whether the target person is a suspicious person disguised as the first type of worker; A program that functions as a
Citation Information
Patent Citations
Image recognition method, device, equipment and medium
CN112380892A
Confirmation action detecting device and alarm system
JP2004334784A
Device and method for recognizing motion
JP2010123019A
Safety measure system and safety measure method
JP2020086482A
Vehicle crime prevention device
JP2020150295A