Processing apparatus, mobile body, processing method, and program
The processing device addresses the complexity and load issues in existing image processing systems by converting high-resolution images to lower resolution and using positional-based target area recognition, enabling efficient and accurate target identification.
Patent Information
- Application Number
- JP2025067424
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing image processing systems for tracking objects are often complex and burdensome in terms of processing load, which can hinder efficient and accurate target identification.
A processing device that captures a first image, converts it to a lower resolution second image, and uses a second processing unit to identify target areas in both images based on relative positional changes, distinguishing between different areas based on distance from the imaging unit to reduce processing load.
Accurate target identification is achieved while significantly reducing processing load by employing image conversion and positional-based target area recognition.
Smart Images

Figure 2025108605000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a processing device, a moving body, a processing method, and a program.
Background Art
[0002] Conventionally, an information processing device that analyzes images captured by two cameras to track an object has been disclosed (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the above technology, the configuration of the equipment to be used may be complicated or the processing load may be high.
[0005] The present invention has been made in consideration of such circumstances, and one of the objectives is to provide a processing device, a moving body, a processing method, and a program that can identify a target accurately while reducing the processing load.
Means for Solving the Problems
[0006] The processing device, the moving body, the processing method, and the program according to this invention adopt the following configuration. (1): The processing device according to an embodiment of the present invention includes an imaging unit that captures a first image, a first processing unit that converts the first image into a second image with a lower resolution than the first image, and a second processing unit that identifies a target area including a predetermined target in the second image and identifies a target area including the target in the first image based on the target area in the second image. The specific area, which is an area for recognizing the target in the target area, changes according to the relative position between the imaging unit and the target.
[0007] (2): In the aspect of (1) above, the relative position is the distance from the imaging unit to the target.
[0008] (3): In the aspect of (2) above, an area where the distance between the imaging unit and the target is greater than or equal to a predetermined value, or exceeds the predetermined value, is defined as a first area, and an area where the distance between the imaging unit and the target is less than or equal to a predetermined value, or less than the predetermined value, is defined as a second area. In the case of the first area, the specific area is defined as the first specific area of the target, and in the case of the second area, the specific area is defined as the second specific area of the target.
[0009] (4): In the aspect of (1) above, the second processing unit analyzes the second image obtained by converting the first image captured at a first time and the second image obtained by converting the first image captured at a second time after the first time, and tracks the target included in the target area of the second image corresponding to the first time in the second image corresponding to the second time.
[0010] (5): In the aspect of (1) above, the second processing unit tracks the target in the second image based on the change in the position of the target in a time-series of second images obtained by converting each of the time-series of first images captured.
[0011] (6): In the aspect of (1) above, the recognition of the target includes a process of recognizing the gesture of the target based on the information of the specific area in the first image.
[0012] (7): In the aspect of (1) above, the recognition of the target object mark includes a process of recognizing the specific region based on recognizing the skeleton or joint points with respect to the object mark region in the first image.
[0013] (8): In the aspect of (7) above, the recognition of the target object mark includes a process of setting, as the specific region, a region including the arm or hand of the target object mark based on the recognition result of the skeleton or joint points.
[0014] (9): In the aspect of (3) above, the first region includes the hand of the target object mark, and the second region is a region including the arm of the target object mark.
[0015] (10): The processing device according to any one of the aspects (1) to (9) above is mounted on a moving body.
[0016] (11): A processing method according to one aspect of the present invention is such that a computer converts a first image captured by an imaging unit into a second image having a lower resolution than the first image, specifies an object mark region including a predetermined target object mark in the second image, specifies an object mark region including the target object mark in the first image based on the object mark region in the second image, and a specific region, which is a region for recognizing the target object mark in the object mark region, changes according to the relative position between the imaging unit and the target object mark.
[0017] (12): A program according to one aspect of the present invention causes a computer to convert a first image captured by an imaging unit into a second image having a lower resolution than the first image, specify an object mark region including a predetermined target object mark in the second image, specify an object mark region including the target object mark in the first image based on the object mark region in the second image, and a specific region, which is a region for recognizing the target object mark in the object mark region, changes according to the relative position between the imaging unit and the target object mark.
Effects of the Invention
[0018] (1)-(12) can identify a target accurately while reducing the processing load.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Embodiments for Carrying Out the Invention
[0020] Hereinafter, with reference to the drawings, a processing device, a moving body, a processing method, and a program according to an embodiment of the present invention will be described.
[0021] <First Embodiment> [Overall Configuration] FIG. 1 is a diagram showing an example of a moving body 10 including a processing device according to an embodiment. The moving body 10 is an autonomous mobile robot. The moving body 10 supports the actions of users. For example, the moving body 10 supports customer shopping or customer service in response to instructions from store clerks, customers, facility staff (hereinafter, these persons are referred to as "users"), etc., or supports the work of the staff.
[0022] The moving body 10 includes a main body 20, a storage container 92, and one or more wheels 94 (wheels 94A and 94B in the figure). The moving body 10 moves in response to instructions based on user gestures, voices, operations on the input unit of the moving body 10 (a touch panel described later), and operations on a terminal device (e.g., a smartphone). The moving body 10 recognizes gestures based on, for example, an image captured by a camera 22 provided on the main body 20.
[0023] For example, the moving body 10 drives the wheels 94 to move so as to follow the customer according to the movement of the user or to lead the customer. At this time, the moving body 10 explains products or work to the user and guides the products or objects that the user is looking for. Also, the user can store the products or luggage to be purchased in the storage container 92 that accommodates them.
[0024] In this embodiment, the mobile body 10 will be described as including the storage container 92. However, instead of (or in addition to) these, the mobile body 10 may be provided with a seating part for the user to sit on in order to move with the mobile body 10, a housing into which the user gets in, steps for the user to place their feet on, etc.
[0025] FIG. 2 is a diagram showing an example of the functional configuration included in the main body 20 of the mobile body 10. The main body 20 includes a camera 22, a communication unit 24, a position specifying unit 26, a speaker 28, a microphone 30, a touch panel 32, a motor 34, and a control device 50 (an example of a "processing device").
[0026] The camera 22 images the periphery of the mobile body 10. The camera 22 is, for example, a fisheye camera capable of imaging the periphery of the mobile body 10 at a wide angle (e.g., 360 degrees). The camera 22 is, for example, attached to the upper part of the mobile body 10 and images the periphery of the mobile body 10 at a wide angle in the horizontal direction. The camera 22 may be realized by combining a plurality of cameras (a plurality of cameras that image a range of 120 degrees or 60 degrees in the horizontal direction). The camera 22 is not limited to one, and a plurality of cameras may be provided on the mobile body 10.
[0027] The communication unit 24 is a communication interface for communicating with other devices using a cellular network, Wi-Fi network, Bluetooth (registered trademark), DSRC (Dedicated Short Range Communication), etc.
[0028] The position specifying unit 26 specifies the position of the mobile body 10. The position specifying unit 26 acquires the position information of the mobile body 10 by means of a GPS (Global Positioning System) device (not shown) built into the mobile body 10. The position information may be, for example, two-dimensional map coordinates or latitude and longitude information.
[0029] The speaker 28 outputs a predetermined sound, for example. The microphone 30 receives the input of the sound uttered by the user, for example.
[0030] The touch panel 32 is configured by superimposing a display unit such as an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) and an input unit capable of detecting the touch position of an operator by a coordinate detection mechanism. The display unit displays a GUI (Graphical User Interface) switch for operation. When the input unit detects a touch operation, a flick operation, a swipe operation, etc. on the GUI switch, it generates an operation signal indicating that a touch operation has been performed on the GUI switch and outputs it to the control device 50. The control device 50 causes the speaker 28 to output sound or causes the touch panel 32 to display an image according to the operation. Further, the control device 50 may move the moving body 10 according to the operation.
[0031] The motor 34 drives the wheels 94 to move the moving body 10. The wheels 94 include, for example, drive wheels driven in the rotational direction by the motor 34 and steering wheels which are non-drive wheels driven in the yaw direction. By adjusting the angle of the steering wheels, the moving body 10 can change its course or rotate.
[0032] In the present embodiment, the moving body 10 is provided with wheels 94 as a mechanism for realizing movement, but the present embodiment is not limited to this configuration. For example, the moving body 10 may be a multi-legged walking robot.
[0033] The control device 50 includes, for example, an acquisition unit 52, a recognition unit 54, a trajectory generation unit 56, a travel control unit 58, an information processing unit 60, and a storage unit 70. Some or all of the acquisition unit 52, the recognition unit 54, the trajectory generation unit 56, the travel control unit 58, and the information processing unit 60 are realized, for example, by a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these functional units may be realized by hardware (including a circuit unit; circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or may be realized by the cooperation of software and hardware. The program may be stored in advance in the storage unit 70 (a storage device having a non-transitory storage medium) such as an HDD (Hard Disk Drive) or a flash memory, or may be stored in a removable storage medium (non-transitory storage medium) such as a DVD or a CD-ROM, and may be installed by mounting the storage medium on a drive device. The acquisition unit 52, the recognition unit 54, the trajectory generation unit 56, the travel control unit 58, or the information processing unit 60 may be provided in a device different from the control device 50 (the moving body 10). For example, the recognition unit 54 may be provided in another device, and the control device 50 may control the moving body 10 based on the processing result of the other device. Also, some or all of the information stored in the storage unit 70 may be stored in another device. A configuration including one or more functional units among the acquisition unit 52, the recognition unit 54, the trajectory generation unit 56, the travel control unit 58, or the information processing unit 60 may be configured as a system.
[0034] The storage unit 70 stores map information 72, gesture information 74, and user information 80. The map information 72 is information in which the shape of a road or passage is represented by, for example, links indicating roads and passages within a facility and nodes connected by the links. The map information 72 may include the curvature of a road, POI (Point Of Interest) information, and the like.
[0035] The gesture information 74 is information in which information regarding a gesture (feature amount of a template) and the operation of the moving body 10 are associated with each other. The gesture information 74 includes first gesture information 76 and second gesture information 78. The user information 80 is information indicating the feature amount of a user. Details of the gesture information 74 and the user information 80 will be described later.
[0036] The acquisition unit 52 acquires an image captured by the camera 22 (hereinafter referred to as a "peripheral image"). The acquisition unit 52 holds, for example, the acquired peripheral image as pixel data in a fisheye camera coordinate system.
[0037] The recognition unit 54 includes, for example, a first processing unit 55A and a second processing unit 55B. The first processing unit 55A converts a first image (for example, a high-resolution image) captured by the camera 22 into a second image (low-resolution image) having a lower resolution than the first image. The second processing unit 55B identifies a target area including a predetermined target in the second image based on the second image, and identifies a target area including the target in the first image based on the target area in the identified second image. The target is, for example, a target to be tracked. Details of the processing of the first processing unit 55A and the second processing unit 55B will be described later.
[0038] In addition, the second processing unit 55B recognizes a body movement (hereinafter referred to as "gesture") by the user based on one or more peripheral images. The recognition unit 54 recognizes the gesture by comparing the feature amount of the user's gesture extracted from the peripheral image with the feature amount of the template (the feature amount indicating the gesture). The feature amount is, for example, data representing feature points such as a person's finger, finger joints, wrist, arm, skeleton, etc., links connecting them, and the inclination and position of the links.
[0039] The trajectory generation unit 56 generates a trajectory that the moving body 10 should travel in the future based on the user's gesture, the destination set by the user, surrounding objects, the position of the user, the map information 72, etc. The trajectory generation unit 56 generates a trajectory by combining a plurality of arcs so that the moving body 10 can move smoothly to the target point. The trajectory generation unit 56 generates a trajectory by, for example, combining three arcs. The trajectory generation unit 56 may generate a trajectory by fitting the state to a geometric model such as a Bézier curve. The trajectory is generated, for example, as a collection of a finite number of trajectory points in actuality.
[0040] The trajectory generation unit 56 performs coordinate conversion between the orthogonal coordinate system and the fisheye camera coordinate system. A one-to-one relationship holds between the orthogonal coordinate system and the fisheye camera coordinate system, and this relationship is stored in the storage unit 70 as correspondence information. The trajectory generation unit 56 generates a trajectory in the orthogonal coordinate system (orthogonal coordinate system trajectory), and converts this trajectory into a trajectory in the fisheye camera coordinate system (fisheye camera coordinate system trajectory). The trajectory generation unit 56 calculates the risk of the fisheye camera coordinate system trajectory. The risk is an index value indicating the likelihood of the moving body 10 approaching an obstacle. The risk tends to be higher as the distance between the trajectory (the trajectory points of the trajectory) and the obstacle is smaller, and lower as the distance between the trajectory and the obstacle is larger.
[0041] When the total value of the risks and the risks of each trajectory point satisfy a preset criterion (for example, when the total value is less than or equal to the threshold Th1 and the risk of each trajectory point is less than or equal to the threshold Th2), the trajectory generation unit 56 adopts the trajectory that satisfies the criterion as the trajectory for the moving body to move.
[0042] If the above-mentioned trajectory does not meet the preset criteria, the following processing may be performed. The trajectory generation unit 56 detects the drivable space in the fisheye camera coordinate system, and converts the detected drivable space in the fisheye camera coordinate system into the drivable space in the orthogonal coordinate system. The drivable space is the space excluding the obstacles and the areas around the obstacles (the areas where risks are set or the areas where the risks are equal to or higher than the threshold value) in the area of the moving direction of the moving body 10. The trajectory generation unit 56 corrects the trajectory so that the trajectory fits within the drivable space transformed into the orthogonal coordinate system. The trajectory generation unit 56 converts the orthogonal coordinate system trajectory into a fisheye camera coordinate system trajectory, and calculates the risk of the fisheye camera coordinate system trajectory based on the surrounding image and the fisheye camera coordinate system trajectory. This process is repeated to search for a trajectory that meets the above-mentioned preset criteria.
[0043] The travel control unit 58 causes the moving body 10 to travel along a trajectory that meets the preset criteria. The travel control unit 58 outputs a command value for causing the moving body 10 to travel along the trajectory to the motor 34. The motor 34 rotates the wheels 94 according to the command value, and moves the moving body 10 along the trajectory.
[0044] The information processing unit 60 controls various devices and equipment included in the main body 20. The information processing unit 60 controls, for example, the speaker 28, the microphone 30, and the touch panel 32. Further, the information processing unit 60 recognizes the voice input to the microphone 30 and the operations performed on the touch panel 32. The information processing unit 60 operates the moving body 10 based on the recognition result.
[0045] In the above example, the recognition unit 54 has been described as using the image captured by the camera 22 provided on the moving body 10 for various processes. However, the recognition unit 54 may perform various processes using an image captured by a camera not provided on the moving body 10 (a camera provided at a position different from the moving body 10). In this case, the image captured by the camera is transmitted to the control device 50 via communication, and the control device 50 acquires the transmitted image and executes various processes based on the acquired image. Further, the recognition unit 54 may execute various processes using a plurality of images. For example, the recognition unit 54 may execute various processes based on an image captured by the camera 22 and a plurality of images captured by a camera provided at a position different from the moving body 10.
[0046] [Support Process] The moving body 10 executes a support process for supporting the user's shopping. The support process includes a process related to tracking and a process related to behavior control.
[0047] [Process Related to Tracking (Part 1)] FIG. 3 is a flowchart showing an example of the flow of the tracking process. First, the control device 50 of the moving body 10 receives the registration of the user (step S100). Next, the control device 50 tracks the user registered in step S100 (step S102). Next, the control device 50 determines whether the tracking was successful (step S104). If the tracking is successful, the process proceeds to step S200 in FIG. 10 described later. If the tracking is not successful, the control device 50 identifies the user (step S106).
[0048] (Process of Registering a User) The process of registering the user in step S100 will be described. The control device 50 of the moving body 10 confirms the user's intention to register based on a specific gesture, voice, or operation on the touch panel 32 of the user (for example, a customer who has come to the store). When the user's intention to register can be confirmed, the recognition unit 54 of the control device 50 extracts the feature amount of the user and registers the extracted feature amount.
[0049] FIG. 4 is a diagram for explaining a process of extracting user feature amounts and a process of registering the feature amounts. The second processing unit 55B of the control device 50 identifies a user from an image IM1 in which the user is imaged, and recognizes joint points and a skeleton of the identified user (executes skeleton processing). For example, the second processing unit 55B estimates a user's face, face parts, neck, shoulders, elbows, wrists, waist, ankles, etc. from the image IM1, and executes skeleton processing based on the positions of the estimated parts. For example, the second processing unit 55B executes skeleton processing using a known method (such as a method like OpenPose) for estimating a user's joint points and skeleton using deep learning. Next, based on the result of the skeleton processing, the second processing unit 55B identifies a user's face, upper body, lower body, etc., extracts feature amounts for each of the identified face, upper body, and lower body, and registers the extracted feature amounts in the storage unit 70 as the user's feature amounts. The feature amounts of the face are, for example, male, female, hairstyle, and face feature amounts. The feature amounts indicating male and female are feature amounts indicating the shape of the head, etc., and the hairstyle is information indicating the length of the hair (such as short hair or long hair) obtained from the shape of the head. The feature amount of the upper body is, for example, the color of the upper body part. The feature amount of the lower body is, for example, the color of the lower body part.
[0050] (Process of tracking the user) The first processing unit 55A converts each of the high-resolution images captured every unit time into a low-resolution image. Having a high resolution means, for example, that the number of pixels per unit area in the image is larger than the number of pixels per unit area in the low-resolution image (high dpi). The first processing unit 55A performs a process of thinning out the pixels of the high-resolution image IM to convert the high-resolution image into a low-resolution image, or applies a predetermined algorithm to convert the high-resolution image into a low-resolution image.
[0051] The second processing unit 55B analyzes the low-resolution image obtained by converting the high-resolution image captured at the first time and the low-resolution image obtained by converting the high-resolution image captured at the second time after the first time, and tracks the object target included in the object target area of the object to be tracked in the low-resolution image corresponding to the first time in the low-resolution image corresponding to the second time. The second processing unit 55B tracks the object target in the low-resolution image based on the change in the position of the object target in the time-series low-resolution images obtained by converting each of the high-resolution images captured in time series. The low-resolution image used for this tracking is, for example, the low-resolution image obtained by converting the most recently captured high-resolution image. This will be specifically described below.
[0052] The process of tracking the user in step S102 will be described. FIG. 5 is a diagram for explaining the process of the recognition unit 54 tracking the user (the process of step S102 in FIG. 3). The first processing unit 55A of the recognition unit 54 acquires an image captured at time T. This image is an image captured by the camera 22 (hereinafter, high-resolution image IM2).
[0053] The first processing unit 55A of the recognition unit 54 converts the high-resolution image IM2 into a low-resolution image IM2# with a lower resolution than the high-resolution image IM2. Next, the second processing unit 55B detects a person and a person detection area including the person from the low-resolution image IM2#.
[0054] The second processing unit 55B estimates the position (person detection area) of the user at time T based on the position of the person detected at time T-1 (before time T) (the person detection area of the user being tracked at time T-1) and the moving direction of the person. If the user detected in the low-resolution image IM2 obtained at time T exists near the position estimated from the position or moving direction of the user who was the tracking target before time T-1, the second processing unit 55B identifies the user detected at time T as the user to be tracked (tracking target). When the user can be identified, it is considered that the tracking has succeeded.
[0055] As described above, since the control device 50 tracks the user using the low-resolution image IM2#, the processing load is reduced.
[0056] In the tracking process, the second processing unit 55B may track the user using not only the positions of the user at time T and time T-1 as described above, but also the feature amounts of the user. FIG. 6 is a diagram for explaining the tracking process using the feature amounts. For example, the second processing unit 55B estimates the position of the user at time T, identifies the user existing near the estimated position, and further extracts the feature amounts of the user. When the extracted feature amounts match the registered feature amounts by a threshold value or more, the control device 50 estimates that the identified user is the user to be tracked and determines that the tracking has succeeded.
[0057] For example, when extracting the feature amounts of the user, the second processing unit 55B extracts the region including the person, and performs skeleton processing on the image (high-resolution image) of the extracted region to extract the feature amounts of the person. Thereby, the processing load is reduced.
[0058] Note that, instead of the feature amounts obtained from the high-resolution image, when the feature amounts obtained from the low-resolution image match the registered feature amounts by a threshold value or more, the second processing unit 55B may estimate that the identified user is the user to be tracked. In this case, the feature amounts for comparison with the feature amounts obtained from the low-resolution image are stored in the storage unit 70 in advance, and these feature amounts are used. Further, instead of (or in addition to) the registered feature amounts, the second processing unit 55B may compare the feature amounts extracted from the image obtained when tracking with the feature amounts obtained from the image captured this time to identify the user.
[0059] For example, even when the user to be tracked overlaps or intersects with another person, the user can be tracked more accurately based on the change in the position of the user and the feature amounts of the user as described above.
[0060] (Processing for identifying a user) The processing for identifying the user in step S106 will be described. When the second processing unit 55B fails to track the user, as shown in FIG. 7, the second processing unit 55B collates the feature amounts of the people around with the feature amounts of the registered users to identify the user to be tracked. For example, the second processing unit 55B extracts the feature amounts of each person included in the image. The second processing unit 55B collates the feature amounts of each person with the feature amounts of the registered users, and identifies the person whose feature amounts match the registered users' feature amounts by a threshold value or more. The second processing unit 55B sets the identified user as the user to be tracked. At this time, the feature amounts used may be the feature amounts obtained from a low-resolution image or the feature amounts obtained from a high-resolution image.
[0061] Through the above processing, the second processing unit 55B of the control device 50 can track the user with higher accuracy.
[0062] [Processing related to tracking (Part 2)] In the above example, the user has been described as a customer who visited the store. However, when the user is a store clerk or a facility staff member (for example, a person engaged in medical care within the facility), the following processing may be performed.
[0063] (Processing for registering a user) The processing for tracking the user in step S102 may be performed as follows. FIG. 8 is a diagram for explaining another example of the processing (the processing in step S102 of FIG. 3) in which the second processing unit 55B tracks the user. The second processing unit 55B extracts the area including a person from the low-resolution image, and extracts the area corresponding to the area extracted from the high-resolution image (the area including the person). The second processing unit 55B further extracts the area including the face part of the person from the area extracted from the high-resolution image, and extracts the feature amounts of the face part of the person. The second processing unit 55B collates the extracted feature amounts of the face part with the feature amounts of the face part of the user to be tracked registered in advance in the user information 80, and when they match, determines that the person included in the image is the user to be tracked.
[0064] (Processing for identifying the user) The processing for identifying the user in step S106 may be performed as follows. When the second processing unit 55B fails to track the user, as shown in FIG. 9, the second processing unit 55B extracts a region including a person in the vicinity from the high-resolution image. The second processing unit 55B extracts a region including the face portion of the person from the extracted region, extracts the feature amount of the face portion of the person, collates the feature amount of the face of the person in the vicinity with the feature amount of the registered user, and determines that the person having a feature amount that matches the threshold value or more is the user to be tracked.
[0065] As described above, the control device 50 can track the user with higher accuracy. In addition, since the control device 50 extracts a person using a low-resolution image and further extracts a person using a high-resolution image as necessary, the processing load can be reduced.
[0066] [Processing related to action control] FIG. 10 is a flowchart showing an example of the flow of the action control process. This process is a process executed after the process of step S104 in FIG. 3. The control device 50 recognizes the gesture of the user (step S200), and controls the action of the moving body 10 based on the recognized gesture (step S202). Next, the control device 50 determines whether to end the service (step S204). If the service is not ended, the process returns to the process of step S102 in FIG. 3, and tracking is continued. When the service is ended, the control device 50 deletes the registered registration information related to the user, such as the feature amount of the user (step S206). For example, when the user makes a gesture indicating an intention to end the service, performs an operation, etc., or inputs a voice, the service ends. Also, when the user or the moving body 10 reaches the boundary with the area outside the area where the service is provided, the provision of the service ends. Thereby, one routine of this flowchart ends.
[0067] The process of step S200 will be described. FIG. 11 is a diagram (part 1) for explaining the process of recognizing a gesture. The second processing unit 55B identifies the same person detection area (target area) as the person detection area including the tracked user detected in the low-resolution image IM2# corresponding to time T in the high-resolution image IM corresponding to time T. Then, the second processing unit 55B cuts out (extracts) the person detection area (target area) in the identified high-resolution image IM. The identified or cut-out person detection area (target area) is not limited to the same person detection area as the person detection area including the above-mentioned tracked user, but may be a person detection area (target area) including the person detection area including the above-mentioned user. For example, in addition to the person detection area including the above-mentioned user, an area including another area may be identified, cut out, and used as the target area.
[0068] The second processing unit 55B executes image recognition processing on the cut-out person detection area. The image recognition processing includes processing for recognizing a person's gesture, skeleton processing, processing for identifying an area including a person's arm or hand, or processing for extracting an area where the degree of change in the user's movement (for example, an arm or a hand) is large. These will be described below.
[0069] FIG. 12 is a diagram (part 2) for explaining the process of recognizing gestures. The second processing unit 55B performs skeleton processing on the image of the user included in the cut-out person detection area. The second processing unit 55B extracts an area (hereinafter, the target area) including one or both of the arms and / or hands from the result of the skeleton processing, and extracts a feature amount indicating the state of one or both of the arms and / or hands in the extracted target area. The target area (an example of the "specific area") is, for example, an area used for gesture recognition. The second processing unit 55B specifies a feature amount that matches the feature amount indicating the above state from the feature amounts included in the gesture information 74. The control device 50 causes the moving body 10 to execute the operation of the moving body 10 associated with the specified feature amount in the gesture information 74. Note that whether to extract an area including the hand or an area including the arm is determined by the position of the user with respect to the moving body 10. For example, when the user is not separated from the moving body 10 by a predetermined distance or more, an area including the hand is extracted, and when the user is separated from the moving body 10 by a predetermined distance or more, an area including the arm is extracted.
[0070] FIG. 13 is a diagram (part 3) for explaining the process of recognizing gestures. The second processing unit 55B may recognize gestures by preferentially using information on an area (an area including parts with a large degree of change among the respective parts) where the degree of change in the movement of a person in a time series is large. Based on the result of skeleton processing on the high-resolution images captured in a time series, the second processing unit 55B extracts an area (specific area) including the arm or hand with a large degree of change among the first degree of change of the user's left arm or left hand and the second degree of change of the user's right arm or right hand, and recognizes the gesture of the user performed by the arm or hand included in the extracted area. That is, the second processing unit 55B recognizes gestures by preferentially using information on an area (specific area) where the degree of change in a time series (for example, the degree of change of the arm or hand) is large among two or more areas. The two or more areas include at least a specific area specified as an area including the right arm or right hand of the object target and a specific area specified as an area including the left arm or left hand of the object target.
[0071] For example, as shown in FIG. 13, the second processing unit 55B extracts, as a target area, an area including the left arm or left hand with a larger degree of change over time from among the degree of change over time of the user's right arm or right hand and the degree of change over time of the left arm or left hand. The second processing unit 55B recognizes, for example, a gesture of the left arm or left hand with a larger degree of change over time.
[0072] Alternatively, the second processing unit 55B may determine whether the user's right arm or right hand is making a gesture for controlling the moving body 10 and whether the user's left arm or left hand is making a gesture for controlling the moving body 10, and recognize a gesture based on the result of the determination.
[0073] In the above example, the tracking target has been described as a person. However, the tracking target may instead (or in addition) be an object that is the subject of an action, such as a robot or an animal. In this case, the second processing unit 55B recognizes a gesture of an object such as a robot or an animal.
[0074] (Processing for recognizing a gesture) The control device 50 determines whether to refer to the first gesture information 76 or the second gesture information 78 of the gesture information 74 based on the relative position between the moving body 10 and the user. As shown in FIG. 14, when the user is not at a predetermined distance from the moving body 10, in other words, when the user is within the first area AR1 set with respect to the moving body 10, the control device 50 determines whether the user is making the same gesture as the first gesture included in the first gesture information 76.
[0075] FIG. 15 is a diagram showing an example of the first gesture included in the first gesture information 76. The first gesture is, for example, a gesture using the hand without using the arm as shown below. - Gesture for moving the moving body 10 forward: This gesture is a gesture of protruding the hand forward. - Gesture for stopping the moving body 10 that is moving forward: This gesture is a gesture of facing the palm of the hand in the forward direction of the user. ·Gesture to move the moving body 10 to the left: This gesture is a gesture of moving the hand to the left. ·Gesture to move the moving body 10 to the right: This gesture is a gesture of moving the hand to the right. ·Gesture to move the moving body 10 backward: This gesture is a gesture of repeatedly moving the fingertips (moving the fingertips closer to the palm) so that the palm faces the vertically opposite direction and the fingertips face the user's direction (a gesture of beckoning). ·Gesture to rotate the moving body 10 to the left: This gesture is a gesture of protruding the index finger and thumb (or a predetermined finger) and rotating the protruding finger to the left. ·Gesture to rotate the moving body 10 to the right: This gesture is a gesture of protruding the index finger and thumb (or a predetermined finger) and rotating the protruding finger to the right.
[0076] As shown in FIG. 16, when the user is at a predetermined distance from the moving body 10, in other words, when the user is present in the second region AR2 set with respect to the moving body 10 (when not present in the first region AR1), the control device 50 determines whether the user is performing the same gesture as the second gesture included in the second gesture information 78.
[0077] The second gesture is a gesture using the arm (the arm between the elbow and the hand) and the hand. Note that the second gesture may be a body movement such as a larger body movement or a larger hand movement than the first gesture. A larger body movement means that when causing a movement (such as straight movement) in the moving body 10, the body movement of the second gesture is larger than the body movement of the first gesture. For example, the first movement may be a gesture using the hand or finger, and the second gesture may be a gesture using the arm. For example, the first movement may be a gesture using the leg below the knee, and the second gesture may be a gesture using the lower body. For example, the first movement may be a gesture using the hand or foot, etc., and the second gesture may be a gesture using the whole body such as jumping.
[0078] When the camera 22 of the moving body 10 images a user existing in the first region AR1 as shown in FIG. 14 described above, the arm portion is difficult to fit into the image, while the hand and fingers can fit into the image. The first region AR1 is a region where the recognition unit 54 cannot recognize or can hardly recognize the user's arm from the image of the user imaged in the first region AR1. When the camera 22 of the moving body 10 images a user existing in the second region AR2 as shown in FIG. 16, the arm portion fits into the image. Therefore, as described above, when a user exists in the first region AR1, the recognition unit 54 recognizes a gesture using the first gesture information 76, and when a user exists in the second region AR2, the recognition unit 54 can recognize the user's gesture with higher accuracy by recognizing a gesture using the second gesture information 78.
[0079] FIG. 17 is a diagram showing an example of the second gesture included in the second gesture information 78. - Gesture for moving the moving body 10 located behind the user in front of the user: This gesture is a gesture in which the user pushes out the arm and hand from near the body to in front of the body. - Gesture for moving the moving body 10 forward: This gesture is a gesture in which the arm and hand protrude forward. - Gesture for stopping the moving body 10 that is moving forward: This gesture is a gesture in which the palm of the hand is facing directly forward among the arm and hand protruding forward. - Gesture for moving the moving body 10 to the left: This gesture is a gesture in which the arm and hand are moved to the left. - Gesture for moving the moving body 10 to the right: This gesture is a gesture in which the arm and hand are moved to the right. - Gesture for moving the moving body 10 backward: This gesture is a gesture that repeats an operation of moving the arm or wrist so that the palm faces in the vertically opposite direction and the fingertips face the user's direction (a waving gesture). - Gesture for rotating the moving body 10 to the left: This gesture is a gesture in which the index finger (or a predetermined finger) protrudes and the protruding finger is rotated to the left. ·Gesture to rotate the mobile body 10 to the right: This gesture is a gesture of protruding the index finger (or a predetermined finger) and rotating the protruding finger to the right.
[0080] [Flowchart] FIG. 18 is a flowchart showing an example of a process in which the control device 50 recognizes a gesture. First, the control device 50 determines whether the user is present in the first area (step S300). If the user is present in the first area, the control device 50 recognizes the user's behavior based on the acquired image (step S302). The behavior is, for example, the movement of the user recognized from images acquired continuously over time.
[0081] Next, the control device 50 refers to the first gesture information 76 and identifies a gesture that matches the behavior recognized in step S302 (step S304). If the gesture that matches the behavior recognized in step S302 is not included in the first gesture information 76, it is determined that no gesture for controlling the movement of the mobile body 10 has been performed. Next, the control device 50 performs an action corresponding to the identified gesture (step S306).
[0082] If the user is not present in the first area (if present in the second area), the control device 50 recognizes the user's behavior based on the acquired image (step S308), refers to the second gesture information 78, and identifies a gesture that matches the behavior recognized in step S308 (step S310). Next, the control device 50 performs an action corresponding to the identified gesture (step S312). Thereby, the processing of one routine of this flowchart is completed.
[0083] For example, in the above processing, the recognition unit 54 does not need to perform the process of recognizing the gesture of the user being tracked and the gesture of a person not being tracked. Thereby, the control device 50 can control the mobile body based on the gesture of the user being tracked with a reduced processing load.
[0084] As described above, the control device 50 can recognize the gesture of the user with higher accuracy and operate the moving body 10 according to the will of the user by switching the gesture to be recognized based on the area where the user exists. As a result, the convenience of the user is improved.
[0085] According to the first embodiment described above, the control device 50 converts the first image into a second image with a lower resolution than the first image, obtains a target area including the target object to be tracked in the second image, and based on the obtained target area in the second image, obtains the target area including the target object in the first image, so that the target object can be accurately specified while reducing the processing load.
[0086] The embodiment described above can be expressed as follows. A storage device storing a program, A hardware processor, and By the hardware processor executing the program stored in the storage device, Convert the first image into a second image with a lower resolution than the first image, Based on the second image, identify a target area including a predetermined target object in the second image, and based on the identified target area in the second image, identify the target area including the target object in the first image, Processing device.
[0087] As described above, the embodiments for carrying out the present invention have been described using the embodiments. However, the present invention is not limited to such embodiments, and various modifications and substitutions can be made without departing from the gist of the present invention.
Explanation of Reference Numerals
[0088] 10‥Moving body, 20‥Main body, 22‥Camera, 50‥Control device, 52‥Acquisition unit, 54‥Recognition unit, 55A‥First processing unit, 55B‥Second processing unit, 56‥Trajectory generation unit, 58‥Travel control unit, 60‥Information processing unit, 70‥Memory unit, 74‥Gesture information, 76‥First gesture information, 78‥Second gesture information, 80‥User information
Claims
1. An imaging unit that captures a first image; A first processing unit that converts the first image into a second image with a lower resolution than the first image; A second processing unit that identifies a target area including a predetermined target mark in the second image and identifies a target area including the target mark in the first image based on the target area in the second image; and A specific area that is an area for recognizing the target mark in the target area changes according to the relative position between the imaging unit and the target mark. A processing device.
2. The relative position is the distance from the imaging unit to the target mark. The processing device according to Claim 1.
3. An area where the distance between the imaging unit and the target mark is equal to or greater than a predetermined value, or greater than a predetermined value is defined as a first area; An area where the distance between the imaging unit and the target mark is equal to or less than a predetermined value, or less than a predetermined value is defined as a second area; In the case of the first area, the specific area is defined as a first specific area of the target mark; In the case of the second area, the specific area is defined as a second specific area of the target mark. The processing device according to Claim 2.
4. The second processing unit: Analyzes the second image obtained by converting the first image captured at a first time and the second image obtained by converting the first image captured at a second time after the first time, and tracks the target mark included in the target area of the second image corresponding to the first time in the second image corresponding to the second time; The processing device according to Claim 1.
5. The second processing unit: Tracks the target mark in the second image based on the change in the position of the target mark in a time series of second images obtained by converting each of the first images captured in a time series. The processing device according to Claim 1.
6. The recognition of the target mark includes a process of recognizing the gesture of the target mark based on the information of the specific area in the first image. The processing device according to Claim 1.
7. The recognition of the target mark includes a process of recognizing the specific area based on recognizing a skeleton or joint points with respect to the target area in the first image. The processing device according to Claim 1.
8. The recognition of the target mark includes a process of defining an area including the arm or hand of the target mark as the specific area based on the recognition result of the skeleton or joint points. The processing device according to Claim 7.
9. The first area includes the hand of the target mark. The second area is an area including the arm of the object target. The processing device according to claim 3.
10. A moving body equipped with the processing device according to any one of claims 1 to 9.
11. A computer converts a first image captured by an imaging unit into a second image with a lower resolution than the first image, identifies a target area including a predetermined object target in the second image, identifies a target area including the object target in the first image based on the target area in the second image, the specific area, which is an area for recognizing the object target in the target area, changes according to the relative position between the imaging unit and the object target. Processing method.
12. Cause a computer to convert a first image captured by an imaging unit into a second image with a lower resolution than the first image, identify a target area including a predetermined object target in the second image, identify a target area including the object target in the first image based on the target area in the second image, the specific area, which is an area for recognizing the object target in the target area, changes according to the relative position between the imaging unit and the object target. Program.
Citation Information
Patent Citations
Gesture recognition device, its method and gesture recognition program
JP2004303014A
Self-propelled teleconferencing platform
JP2014197403A
Distance scalable no touch computing
US20150100926A1
Image processing device, image processing method, and program
WO2020100664A1
Information processing device, imaging device, apparatus control system, movable body, information processing method, and program
JP2018088234A