Operation detection device
The operation detection device enhances vehicle door control by accurately identifying user gestures through defined areas and movement analysis, reducing erroneous openings and processing load.
Patent Information
- Application Number
- JP2024077871
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2025-11-26
AI Technical Summary
Existing vehicle door control systems inaccurately determine user gestures, leading to erroneous opening of sliding doors.
An operation detection device that utilizes a camera to define areas around a door opening, identifies a user based on movement direction and speed within these areas, and determines a valid opening gesture through machine-learned models to accurately detect user operations.
Accurately detects user opening operations, reducing erroneous door openings and minimizing CPU load and power consumption by focusing processing on likely users.
Smart Images

Figure 2025172386000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an operation detection device. [Background technology]
[0002] Patent Document 1 describes a vehicle including a vehicle body having a door opening, a door that opens and closes the door opening, an opening / closing motor that drives the door, a laser radar that detects the area around the door, and a door ECU that controls the opening / closing motor based on the detection results of the laser radar. The door ECU determines whether a user has made a gesture to open the door based on the detection results of the laser radar. If the door ECU determines that the user has made the gesture, it starts the opening operation of the sliding door. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-7171 Summary of the Invention [Problem to be solved by the invention]
[0004] It is desirable for the door ECU described above to improve the accuracy of gesture determination so as not to cause the sliding door to open due to an erroneous gesture determination. [Means for solving the problem]
[0005] An operation detection device that solves the above problem is an operation detection device that detects a user's opening operation to open a vehicle door from a fully closed position where the door opening is fully closed, and is equipped with: a user determination unit that, when an area of the area surrounding the door opening that is closer to the door opening is defined as a first area and an area that is farther from the door opening than the first area is defined as a second area, determines that the moving object moving in the second area toward the first area is the user based on the detection result of a moving object detection unit that is capable of detecting a moving object present in at least the second area, and determines that a moving object that does not move in the second area toward the first area is not the user; and an operation determination unit that determines whether the moving object determined by the user determination unit to be the user has performed the opening operation in the first area. [Effects of the Invention]
[0006] The operation detection device can accurately detect whether or not a user has performed an opening operation to open the vehicle door. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a schematic diagram of a vehicle equipped with an operation detection device. [Figure 2] FIG. 2 shows an image at a specific frame captured by the camera. [Figure 3] FIG. 3 shows an image of the frame captured by the camera next to the frame above. [Figure 4] FIG. 4 is a flowchart showing the flow of processing executed by the operation detection device to detect a user's gesture. [Figure 5] FIG. 5 is a flowchart illustrating the flow of the tracking user setting process. [Figure 6] FIG. 6 is a flowchart illustrating the flow of the user determination process. [Figure 7] FIG. 7 is a flowchart illustrating the flow of the gesture determination process. DETAILED DESCRIPTION OF THE INVENTION
[0008] An embodiment of a vehicle equipped with an operation detection device will be described below. <Configuration of this embodiment> As shown in FIG. 1, a vehicle 10 includes a vehicle body 20, a front door 30, a rear door 40, a door drive unit 50, a camera 60, a door control device 70, and an operation detection device 80.
[0009] In the following description, the front-rear direction of the vehicle 10, the width direction of the vehicle 10, and the up-down direction of the vehicle 10 will be referred to as the front-rear direction, the width direction, and the up-down direction, respectively. In addition, in the width direction, the direction toward the center of the vehicle will be referred to as the inward direction, and the direction away from the center of the vehicle will be referred to as the outward direction.
[0010] <Body 20> The vehicle body 20 has a front opening 21 and a rear opening 22 that open to the side of the vehicle body 20. The front opening 21 and the rear opening 22 are adjacent to each other in the front-to-rear direction. The front opening 21 is located forward of the rear opening 22. The front opening 21 and the rear opening 22 are areas through which a user of the vehicle 10 passes when getting in and out of the vehicle 10. In this embodiment, the user is a driver who possesses a portable device linked to the vehicle 10. In other words, the user is the owner of the vehicle 10. The user may also include passengers associated with the vehicle 10, such as the driver's family members. In this embodiment, the rear opening 22 corresponds to a "door opening."
[0011] <Front Door 30> The front door 30 is supported by the vehicle body 20 so as to be rotatable about an axis extending in the vertical direction. The front door 30 rotates between a fully closed position where the front opening 21 is fully closed and a fully open position where the front opening 21 is fully opened. The front door 30 has side mirrors 31. Although not shown, the vehicle 10 preferably includes a door lock device that locks the front door 30, which is located in the fully closed position, to the vehicle body 20, and a door lock drive unit that switches the state of locking the front door 30 by the door lock device.
[0012] <Rear door 40 / door drive unit 50> The rear door 40 is supported by the vehicle body 20 so as to be movable in the fore-and-aft direction relative to the vehicle body 20. The rear door 40 slides between a fully closed position where the rear opening 22 is fully closed and a fully open position where the rear opening 22 is fully opened. The door drive unit 50 includes, for example, an electric motor and a transmission mechanism that transmits power from the electric motor to the rear door 40. The door drive unit 50 drives the rear door 40 to open and close it. In this respect, the rear door 40 is a so-called power sliding door. Although not shown, the vehicle 10 preferably includes a door lock device that locks the rear door 40 to the vehicle body 20 when it is in the fully closed position, and a door lock drive unit that switches the locking state of the rear door 40 by the door lock device. In this embodiment, the rear door 40 corresponds to the "vehicle door."
[0013] <Camera 60> The camera 60 is installed on the side mirror 31 so that the lens faces outward in the width direction. As shown in Fig. 1, the photographing area of the camera 60 is the area around the rear opening 22. In other words, the photographing area of the camera 60 is the area on the side of the vehicle 10. The photographing area of the camera 60 includes, within the area around the rear opening 22, a first area A1 that is close to the rear opening 22 and a second area A2 that is farther from the rear opening 22 than the first area A1.
[0014] Each time the camera 60 captures an image, it outputs the image to the operation detection device 80. The camera 60 may be a camera dedicated to the operation detection device 80, or may be a camera mounted on the vehicle 10 for another purpose. The frame rate of the camera 60 may be, for example, about 5 to 10 fps. FIGS. 2 and 3 show examples of images captured by the camera 60. The camera 60 corresponds to a "moving object detection unit" capable of detecting a person as a moving object present in the first area A1 and the second area A2.
[0015] <Door control device 70> The door control device 70 is composed of a processing circuit including a CPU and a memory. The door control device 70 controls the door drive unit 50 based on a command signal input to the door control device 70. More specifically, when an opening command signal for opening the rear door 40 is input, the door control device 70 opens the rear door 40 toward the fully open position. On the other hand, when a closing command signal for closing the rear door 40 is input, the door control device 70 closes the rear door 40 toward the fully closed position.
[0016] <Operation detection device 80> The operation detection device 80 is composed of a processing circuit including a CPU and a memory. The operation detection device 80 determines whether the user has performed a gesture as an opening operation to open the rear door 40. In this embodiment, the user's gesture to be detected by the operation detection device 80 is a kick gesture made by the user using the foot. For example, the user's kick gesture may be an action in which the user swings their foot forward and then returns it to its original position. In this regard, it generally takes about one second for the user to make such a kick gesture. In other embodiments, the kick gesture may be a gesture in which the user moves their foot sideways or swings their toe around the heel. The user's gesture may also be a gesture in which the user moves their hand, head, or the like in a predetermined direction.
[0017] As shown in FIG. 1, the operation detection device 80 includes, as functional units, a memory unit 81, an image acquisition unit 82, a person extraction unit 83, a user determination unit 84, a skeleton estimation unit 85, a gesture determination unit 86, and a command unit 87.
[0018] <Storage section 81> The storage unit 81 stores a person extraction model, a skeleton estimation model, and a gesture determination model as models trained by machine learning. The storage unit 81 may be part of the memory of the operation detection device 80 described above.
[0019] The person extraction model is a machine-learned model that, in response to an input image captured by a camera 60, outputs the positions of people appearing in the image. The image input to the person extraction model is an image such as that shown in FIGS. 2 and 3. As shown in FIG. 2, if a person is included in the image, the person extraction model outputs the position of a bounding box that defines the minimum size of the person. As indicated by the dashed-dotted lines in FIGS. 2 and 3, the bounding box is a rectangle consisting of two sides extending horizontally and two sides extending vertically. The position of the bounding box may be any information that can identify the position of the bounding box. For example, when an arbitrary position in the image is taken as the origin, the position of the bounding box may be the coordinates of the two corners located diagonally across the bounding box. If multiple people are included in the input image, the person extraction model outputs the positions of multiple bounding boxes. On the other hand, if no people are included in the input image, the person extraction model does not output the position of the bounding box.
[0020] The skeleton estimation model is a machine-learned model that outputs the positions of a person's skeleton points in response to an input image of the person. The image input to the skeleton estimation model is not the image shown in FIG. 2 itself, but an image of the image shown in FIG. 2 separated by a bounding box. In other words, the image input to the skeleton estimation model generally contains a single person. Furthermore, the size of the image input to the skeleton estimation model is smaller than the size of the image input to the person extraction model. The positions of the skeleton points output by the skeleton estimation model include, for example, the coordinates of the neck, the coordinates of the left and right shoulders, the coordinates of the left and right elbows, the coordinates of the left and right wrists, the coordinates of the left and right hips, the coordinates of the left and right knees, the coordinates of the left and right ankles, and the coordinates of the left and right toes. In other words, the position of a skeleton point does not refer to the coordinates of a single skeleton point of a person, but to the coordinates of multiple skeleton points of the person.
[0021] The gesture determination model is a machine-learned model that outputs whether a person has made a gesture in response to input of chronologically-series positions of skeleton points. The input data for the gesture determination model are the positions of skeleton points of a person appearing in images continuously captured by camera 60 over several seconds. For example, the image within a bounding box obtained by cutting out a person from an image in an Nth frame is defined as the Nth image, and the image within a bounding box obtained by cutting out a person from an image in an Mth frame several seconds after the Nth frame is defined as the Mth image. In this case, the gesture determination model receives input of the positions of skeleton points of the person appearing in the Nth image, the positions of skeleton points of the person appearing in the N+1th image, ..., the positions of skeleton points of the person appearing in the M-1th image, and the positions of skeleton points of the person appearing in the Mth image.
[0022] As described above, in this embodiment, the time required for a user to perform a gesture is approximately one second. In order to improve the accuracy of gesture determination by the gesture determination model, it is preferable to include not only the positional changes of the skeleton points when the user performs a gesture, but also the positional changes of the skeleton points before and after the user performs the gesture in the input data of the gesture determination model. For example, it is preferable to use the positions of the user's skeleton points for three seconds, which is obtained by adding one second before and after the one second required for the user to perform the gesture, as input data for the gesture determination model. For example, if a camera 60 with a frame rate of 10 fps captures images for three seconds, 30 frames of images can be obtained. Therefore, in this case, the input data for the gesture determination model is the positions of the user's skeleton points for 30 frames.
[0023] <Image acquisition unit 82> Image acquisition unit 82 acquires an image each time camera 60 captures the image. In the following description, the image currently acquired by image acquisition unit 82 will be referred to as the current image, the image one frame before the current image will be referred to as the previous image, and the image one frame after the current image will be referred to as the next image.
[0024] <Person extraction part 83> The person extraction unit 83 inputs an image captured by the camera 60 into a person extraction model to acquire the positions of people appearing in the image. Specifically, the person extraction unit 83 acquires the coordinates of the bounding boxes of people appearing in the image. When multiple people appear in the image, the person extraction unit 83 acquires the coordinates of the bounding boxes of the multiple people.
[0025] When the person extraction unit 83 acquires a bounding box, it sets a reference point Pb in the bounding box. As shown in FIG. 2, in this embodiment, the position of the reference point Pb is the coordinate of the center of the bottom side of the bounding box. In other embodiments, the position of the reference point Pb may be the coordinate of the center of the bounding box, or the coordinate of one of the four corners of the bounding box. Furthermore, the position of the reference point Pb may be the coordinate of the center of any side of the bounding box other than the bottom side.
[0026] Next, the person extraction unit 83 calculates the distance from the reference point Pb of the bounding box to a representative point Pr set within the image (hereinafter referred to as the "reference distance Db"). The representative point Pr is a point at a fixed position within the image. The representative point Pr is preferably set at any position within the image on the rear door 40 that is positioned in the fully closed position. In this embodiment, the representative point Pr is set near the center in the front-to-rear direction of the lower edge of the rear door 40. Then, the person extraction unit 83 stores the position of the acquired bounding box, the position of the reference point Pb of the bounding box, and the reference distance Db in memory in association with each other.
[0027] <User Determination Unit 84> The user determination unit 84 determines whether a bounding box corresponding to the bounding box acquired from the current image exists in the previous image. In other words, the user determination unit 84 determines whether the same person as the person appearing in the current image also appears in the previous image. Due to the person's movement speed and the frame rate of the camera 60, the person does not move significantly between the previous image shown in FIG. 2 and the current image shown in FIG. 3. Therefore, if the reference point Pb of the bounding box acquired from the current image is located close to the reference point Pb of the bounding box acquired from the previous image, these bounding boxes can be said to be bounding boxes that separate the same person. In other words, these bounding boxes can be said to be corresponding bounding boxes. Note that the bounding box in FIG. 2, which is the previous image, is illustrated by a dashed line in FIG. 3, which is the current image. As shown in FIG. 3, there are two partially overlapping bounding boxes on the left side of the image, and there are two partially overlapping bounding boxes on the right side of the image. In this way, the two partially overlapping bounding boxes have a correspondence relationship.
[0028] If a bounding box corresponding to the bounding box acquired from the current image exists in a previous image acquired so far, the user determination unit 84 calculates a first velocity V1 and a second velocity V2 as the movement speed of the reference point Pb. Because the reference point Pb also indicates the position of a person, the movement speed of the reference point Pb can be said to be the movement speed of the person corresponding to the reference point Pb.
[0029] When a bounding box corresponding to the bounding box acquired from the current image exists in a previous image that goes back the first period T1 from the current image, the user determination unit 84 calculates a first speed V1, which is the movement speed of the reference point Pb. Specifically, the user determination unit 84 calculates the first speed V1 by subtracting the reference distance Db of the bounding box in the current image from the reference distance Db of the bounding box in the previous image that goes back the first period T1 from the current image, and dividing the result by the first period T1. Incidentally, if the value obtained by multiplying the first period T1 by the frame rate of the camera 60 is the first frame number, the image that goes back the first period T1 from the current image can also be said to be an image that goes back the first number of frames from the current image.
[0030] If the reference point Pb of the bounding box approaches the representative point Pr as the first period T1 passes, the first velocity V1 takes a positive value. If the reference point Pb of the bounding box moves away from the representative point Pr as the first period T1 passes, the first velocity V1 takes a negative value. Furthermore, if the position of the reference point Pb of the bounding box does not change over time, the first velocity V1 becomes "0."
[0031] Furthermore, if a bounding box corresponding to the bounding box acquired from the current image exists in a previous image that predates the current image by the second period T2, the user determination unit 84 calculates a second velocity V2, which is the movement velocity of the reference point Pb. Specifically, the user determination unit 84 calculates the second velocity V2 by subtracting the reference distance Db of the bounding box in the current image from the reference distance Db of the bounding box in the previous image that predates the current image by the second period T2, and dividing the result by the second period T2. The second period T2 is shorter than the first period T1. Therefore, the first velocity V1 can be considered the average velocity of the reference point Pb of the bounding box, and the second velocity V2 can be considered the instantaneous velocity of the reference point Pb of the bounding box. Therefore, when the movement velocity of the reference point Pb of the bounding box changes suddenly, the change in the first velocity V1 is small, while the change in the second velocity V2 is large.
[0032] Then, the user determination unit 84 determines whether the person corresponding to the reference point Pb is a user based on the first speed V1. Specifically, the user determination unit 84 determines whether the first speed V1 is equal to or greater than a first lower limit speed VLth1. A person moving away from the first area A1 or a person standing still in the second area A2 is likely not a user. Furthermore, even if a person is approaching the first area A1, a person moving at an extremely slow speed is likely not a user. The first lower limit speed VLth1 is a threshold for determining that such a person is not a user. On the other hand, the user determination unit 84 determines whether the first speed V1 is less than a first upper limit speed VHth1. Even if a person is moving through the second area A2 toward the first area A1, a person moving through the second area A2 at high speed is likely not a user. The first upper limit speed VHth1 is a threshold for determining that such a person is not a user. Then, when the first speed V1 is equal to or greater than the first lower limit speed VLth1 and less than the first upper limit speed VHth1, the user determination unit 84 determines that the person corresponding to the reference point Pb is a user. On the other hand, when the first speed V1 is less than the first lower limit speed VLth1 or is equal to or greater than the first upper limit speed VHth1, the user determination unit 84 determines that the person corresponding to the reference point Pb is not a user.
[0033] The first lower limit speed VLth1 is a value set to a value greater than "0," and the first speed V1 is a speed relative to the representative point Pr set on the rear door 40. Therefore, when the first speed V1 is equal to or greater than the first lower limit speed VLth1, it means that the reference point Pb of the bounding box is approaching the rear door 40. In other words, when the first speed V1 is equal to or greater than the first lower limit speed VLth1, it means that a person is moving through the second area A2 toward the first area A1. Therefore, determining whether the first speed V1 is equal to or greater than the first lower limit speed VLth1 is equivalent to determining whether the corresponding person is moving through the second area A2 toward the first area A1.
[0034] Furthermore, the user determination unit 84 determines whether or not the person corresponding to the reference point Pb is a user based on the second speed V2. Specifically, the user determination unit 84 determines whether or not the second speed V2 is equal to or greater than a second lower limit speed VLth2. Even if a person is moving through the second area A2 toward the first area A1, if the person suddenly slows down or stops, it is highly likely that the person is not a user. The second lower limit speed VLth2 is a threshold for determining that such a person is not a user. On the other hand, the user determination unit 84 determines whether or not the second speed V2 is less than a second upper limit speed VHth2. Even if a person is moving through the second area A2 toward the first area A1, if the person suddenly accelerates, it is highly likely that the person is not a user. The second upper limit speed VHth2 is a threshold for determining that such a person is not a user. Then, if the second speed V2 is equal to or greater than the second lower limit speed VLth2 and less than the second upper limit speed VHth2, the user determination unit 84 determines that the person corresponding to the reference point Pb is a user. On the other hand, if the second speed V2 is less than the second lower limit speed VLth2 or is equal to or greater than the second upper limit speed VHth2, the user determination unit 84 determines that the person corresponding to the reference point Pb is not a user.
[0035] In the following description, the condition indicating whether the first speed V1 and the second speed V2 are within an appropriate range is referred to as the “user applicable condition.” Furthermore, the first lower limit speed VLth1, the first upper limit speed VHth1, the second lower limit speed VLth2, and the second upper limit speed VHth2 may be fixed values or may be variable values that can be adjusted by the user.
[0036] <Bone Estimation Unit 85> The skeleton estimation unit 85 acquires skeleton points of a user by inputting an image of the person determined by the user determination unit 84 to the skeleton estimation model. The image input by the skeleton estimation unit 85 to the skeleton estimation model is an image of the person determined by the user determination unit 84 to be the user, and is an image separated by a bounding box. The skeleton estimation unit 85 then associates the acquired information on the skeleton points of the user with the user and stores it in memory. Note that processing based on the skeleton estimation model is more demanding than processing based on a person extraction model.
[0037] <Gesture determination unit 86> The gesture determination unit 86 determines whether the user has made a gesture in the first area A1. Specifically, the gesture determination unit 86 determines whether the user has made a gesture by inputting the time-series positions of the user's skeleton points into a gesture determination model. Note that, like the skeleton estimation model, processing based on the gesture determination model is more demanding than processing based on the person extraction model.
[0038] <Command section 87> When the gesture determination unit 86 determines that the user has made a gesture, the command unit 87 outputs an opening operation command signal to the door control device 70.
[0039] <Processing Contents of Operation Detection Device 80> 4 to 7, the flow of processing performed by the operation detection device 80 to detect whether or not the user has made a gesture will be described. This processing is performed in a predetermined control cycle when the vehicle 10 is stopped and the rear door 40 is in the fully closed position. Therefore, it is assumed that the camera 60 is activated before starting this processing.
[0040] As shown in FIG. 4, the operation detection device 80 acquires an image captured by the camera 60 at the latest timing (S11). Subsequently, the operation detection device 80 determines whether a tracking user has been set (S12). The tracking user is a person who is likely to be the user of the vehicle 10 and is the subject of skeleton estimation and gesture determination. In other words, the operation detection device 80 does not subject all people appearing in the image to skeleton estimation and gesture determination, but subjects only those people appearing in the image who are likely to be the user. In this regard, a tracking user is set for one person.
[0041] If a tracking user has not been set (S12: NO), the operation detection device 80 performs a user setting process to set a tracking user (S13). For example, in the user setting process, the operation detection device 80 determines whether a person is present in the image acquired in step S11, and whether the person is a user. Next, the operation detection device 80 determines whether a tracking user has been set based on the result of the user setting process (S14). If a tracking user has not been set (S14: NO), the operation detection device 80 ends this process. In this case, it is determined that a person corresponding to the tracking user is not included in the current image, and the process from step S11 is performed on the next image.
[0042] On the other hand, if a tracked user has been set (S14: YES), the operation detection device 80 inputs an image of the tracked user, specifically, an image within a bounding box that delimits the tracked user, into the skeleton estimation model (S15).The operation detection device 80 then acquires the positions of the skeleton points of the tracked user and stores the positions of the skeleton points of the tracked user in memory in association with the tracked user (S16).Then, the operation detection device 80 determines whether the reference point Pb of the bounding box that delimits the tracked user is present in the first area A1 (S17).
[0043] If the reference point Pb of the bounding box that separates the tracked user does not exist in the first area A1 (S17: NO), the operation detection device 80 ends this process. In this case, although the tracked user is captured in the current image, it is determined that the tracked user has not moved to the first area A1, which is used to determine whether the tracked user has made a gesture, and the process from step S11 is performed on the next image.
[0044] On the other hand, if the reference point Pb of the bounding box that separates the tracked user is present in the first area A1 (S17: YES), the operation detection device 80 performs a gesture determination process (S18) to determine whether or not the tracked user has made a gesture. That is, in the gesture determination process, the operation detection device 80 determines whether or not the tracked user has made a gesture in the first area A1. Thereafter, the operation detection device 80 ends this process.
[0045] In step S12, if a tracking user has been set (S12: YES), the operation detection device 80 inputs the image acquired in step S11 into the person extraction model (S19). Subsequently, the operation detection device 80 determines whether or not a bounding box has been acquired from the output of the person extraction model (S20). If a bounding box cannot be acquired (S20: NO), that is, if the tracking user no longer exists in the second area A2, the operation detection device 80 resets the tracking user setting (S26). Thereafter, the operation detection device 80 ends this process.
[0046] In step S20, if one or more bounding boxes are acquired (S20: YES), the operation detection device 80 determines whether or not a bounding box corresponding to the currently set tracked user exists among the acquired one or more bounding boxes (S21). If a bounding box corresponding to the tracked user does not exist (S21: NO), that is, if the tracked user no longer exists in the second area A2, the operation detection device 80 proceeds to step S26. On the other hand, if a bounding box corresponding to the tracked user exists (S21: YES), the operation detection device 80 sets a reference point Pb of the bounding box (S22).
[0047] Next, the operation detection device 80 calculates a reference distance Db, which is the distance between the reference point Pb and the representative point Pr of the bounding box (S23). Thereafter, the operation detection device 80 performs a user determination process to determine whether the tracked user continues to satisfy the user matching condition (S24). The operation detection device 80 determines whether the tracked user satisfies the user matching condition based on the result of the user determination process (S25). If the tracked user does not satisfy the user matching condition (S25: NO), the operation detection device 80 proceeds to step S26. If the tracked user does not satisfy the user matching condition, this occurs when the tracked user suddenly stops or leaves the first area A1. In this case, the setting of the tracked user is reset to set another person as the tracked user. On the other hand, if the tracked user satisfies the user matching condition (S25: YES), the operation detection device 80 proceeds to step S15.
[0048] The user setting process will be described with reference to FIG. As shown in FIG. 5, in the user setting process, the operation detection device 80 inputs an acquired image into the person extraction model (S31). Next, the operation detection device 80 determines whether or not a bounding box has been acquired from the output of the person extraction model (S32). If a bounding box cannot be acquired (S32: NO), the operation detection device 80 ends this process. On the other hand, if a bounding box has been acquired (S33: YES), the operation detection device 80 sets the position of the reference point Pb of the N-th bounding box out of the Sn bounding boxes acquired from the current image (S33). Here, "Sn" is the total number of acquired bounding boxes, and "N" is a variable that is initialized to "1" at the start of this process. Then, the operation detection device 80 calculates a reference distance Db between the reference point Pb of the N-th bounding box and the representative point Pr (S34).
[0049] Next, it is determined whether a bounding box corresponding to the N-th bounding box of the Sn bounding boxes obtained from the current image exists in the images obtained so far (S35). More precisely, the operation detection device 80 determines whether information on the bounding box corresponding to the N-th bounding box is stored in memory to the extent that the first velocity V1 and second velocity V2 of the reference point Pb of the N-th bounding box can be calculated.
[0050] If a bounding box corresponding to the Nth bounding box does not exist in the images acquired so far (S35: NO), the operation detection device 80 increments "N" by 1 (S39). Next, the operation detection device 80 determines whether "N" is greater than "Sn" (S40). If "N" is greater than "Sn" (S40: YES), the operation detection device 80 ends this processing. On the other hand, if "N" is equal to or less than "Sn" (S40: NO), the operation detection device 80 proceeds to step S35.
[0051] In step S35, if a bounding box corresponding to the Nth bounding box exists in the images acquired so far (S35: YES), the operation detection device 80 performs a user determination process on the person separated by the Nth bounding box (S36). Next, the operation detection device 80 determines whether the person separated by the Nth bounding box satisfies the user matching condition based on the result of the user determination process (S37). If the person separated by the Nth bounding box does not satisfy the user matching condition (S37: NO), the operation detection device 80 proceeds to step S39.
[0052] On the other hand, if the person separated by the Nth bounding box satisfies the user matching condition (S37: YES), the operation detection device 80 sets the user separated by the Nth bounding box as a tracking user (S38). After that, the operation detection device 80 ends this process. In this case, this process ends even if "N" is equal to or less than "Sn". In other words, the operation detection device 80 ends this process when the first user who is suitable as a tracking user is found among the Sn bounding boxes.
[0053] The user determination process will be described with reference to FIG. 6, in the user determination process, the operation detection device 80 calculates a first velocity V1 of the reference point Pb of the bounding box and a second velocity V2 of the reference point Pb based on information stored in memory (S51). In detail, the operation detection device 80 calculates the first velocity V1, which is the movement velocity of the reference point Pb in a first period T1 that is a relatively long period, and the second velocity V2, which is the movement velocity of the reference point Pb in a second period T2 that is a relatively short period.
[0054] Then, the operation detection device 80 determines whether the first speed V1 is equal to or greater than the first lower limit speed VLth1 and less than the first upper limit speed VHth1 (S52). That is, the operation detection device 80 determines whether the person delimited by the bounding box is moving through the second area A2 toward the first area A1 at an appropriate average speed.
[0055] If the first speed V1 is outside the appropriate speed range (S52: NO), the operation detection device 80 ends this process. In this case, it is determined that the person separated by the bounding box is not the user. On the other hand, if the first speed V1 is within the appropriate speed range (S52: YES), the operation detection device 80 determines whether the second speed V2 is equal to or greater than the second lower limit speed VLth2 and less than the second upper limit speed VHth2 (S53). In other words, the operation detection device 80 determines whether the person separated by the bounding box is moving through the second area A2 toward the first area A1 at an appropriate instantaneous speed.
[0056] If the second speed V2 is outside the appropriate speed range (S53: NO), the operation detection device 80 ends this process. In this case, it is determined that the person separated by the bounding box is not the user. On the other hand, if the second speed V2 is within the appropriate speed range (S53: YES), the operation detection device 80 stores information in memory indicating that the person separated by the bounding box satisfies the user matching condition (S54). Thereafter, the operation detection device 80 ends this process.
[0057] The gesture determination process will be described with reference to FIG. 7, the operation detection device 80 determines whether a termination condition for this process is met (S71). The termination condition is met when the elapsed time since the start of this process is equal to or greater than a predetermined determination time, or when the number of images acquired since the start of this process is equal to or greater than a predetermined determination count. The termination condition is also met when the reference point Pb of the bounding box that delimits the tracked user moves from the first area A1 to outside the first area A1.
[0058] If the termination condition is met (S71: YES), the operation detection device 80 resets the setting of the tracked user (S82). In other words, the current tracked user is deemed not to be the user of the vehicle 10, and the next user who satisfies the user matching condition is set as the tracked user. If the termination condition is not met (S71: NO), the operation detection device 80 acquires a new image captured by the camera 60 (S72). The operation detection device 80 inputs the image acquired in step S72 into the person extraction model (S73). Subsequently, the operation detection device 80 determines whether or not a bounding box has been acquired from the output of the person extraction model (S74). If a bounding box cannot be acquired (S74: NO), the operation detection device 80 proceeds to the process of step S82.
[0059] If a bounding box is acquired (S74: YES), the operation detection device 80 determines whether or not a bounding box corresponding to the currently set tracking user exists among the one or more acquired bounding boxes (S75). If a bounding box corresponding to the tracking user does not exist (S75: NO), the operation detection device 80 proceeds to step S82.
[0060] On the other hand, if a bounding box corresponding to the tracked user exists (S75: YES), the operation detection device 80 inputs an image in which the tracked user appears into a skeleton estimation model (S76). Then, the operation detection device 80 acquires the positions of the skeleton points of the tracked user and stores the positions of the skeleton points of the tracked user in memory in association with the tracked user (S77). Thereafter, the operation detection device 80 determines whether a necessary number of skeleton point positions have been accumulated in memory as input data for the gesture determination model (S78). If the accumulation of skeleton point positions is insufficient (S78: NO), the operation detection device 80 proceeds to step S71. On the other hand, if the accumulation of skeleton point positions is sufficient (S78: YES), the operation detection device 80 inputs the time-series skeleton point positions into the gesture determination model (S79). Next, the operation detection device 80 determines whether the tracked user has made a gesture based on the output of the gesture determination model (S80). If it cannot be determined that the tracked user has made a gesture (S81: NO), the operation detection device 80 proceeds to step S71. On the other hand, if it can be determined that the tracked user has made a gesture (S80: YES), the operation detection device 80 outputs an opening operation command signal to the door control device 70 (S81). In this case, the rear door 40 is opened by the door control device 70.
[0061] <Operation of this embodiment> When a person is approaching the rear opening 22 of the vehicle 10 at an appropriate speed, the operation detection device 80 sets the person as a tracked user who is likely to be the user of the vehicle 10. Next, the operation detection device 80 accumulates the positions of the skeleton points of the tracked user and determines whether the tracked user has made a gesture in the first area A1. Then, when it is determined that the tracked user has made a gesture in the first area A1, the operation detection device 80 outputs an open operation command signal to the door control device 70. As a result, the rear door 40 is opened.
[0062] On the other hand, if there is a person moving away from the rear opening 22 of the vehicle 10, the operation detection device 80 does not set the person as a tracking user. Therefore, the operation detection device 80 is prevented from storing the positions of skeleton points for such a person or determining whether or not a gesture has been made in the first area A1. Furthermore, if the speed at which the person set as a tracking user is moving is no longer appropriate, the operation detection device 80 determines that the person is no longer a tracking user. Examples of such cases include when the direction of movement of the person set as a tracking user changes away from the rear opening 22 of the vehicle 10 or when the person set as a tracking user suddenly stops. The operation detection device 80 is also prevented from storing the positions of skeleton points for such a person or determining whether or not a gesture has been made in the first area A1.
[0063] <Effects of this embodiment> (1) The operation detection device 80 determines whether a person is a user based on the direction of movement of the person in the second area A2. The operation detection device 80 then determines whether a gesture has been made in the first area A1 for a person determined to be a user, but does not determine whether a gesture has been made in the first area A1 for a person determined not to be a user. Therefore, the operation detection device 80 determines whether a gesture has been made for a person who is likely to be a user. As a result, the operation detection device 80 can improve the gesture detection accuracy.
[0064] (2) Furthermore, the operation detection device 80 does not perform processes for extracting skeleton points and determining whether a gesture has been made, which impose a load on a person who is not a user. Therefore, the operation detection device 80 can suppress an increase in the load on the CPU and can suppress power consumption.
[0065] (3) The operation detection device 80 determines both the direction of movement of the person in the second area A2 and whether the user has made a gesture in the first area A1 based on multiple images captured by the camera 60. Therefore, the number of sensors required to implement the operation detection device 80 is reduced compared to when detection results from separate sensors are used for both determinations.
[0066] (4) If the person's moving speed deviates from the moving speed of the user when approaching the vehicle, there is a high possibility that the person is not the user. In this regard, the operation detection device 80 determines whether the person is the user based on the first speed V1. Therefore, the operation detection device 80 can suppress determination of whether a person who is likely not the user has made a gesture.
[0067] (5) Because the first speed V1 is the speed for the relatively long first period T1, the first speed V1 can be an appropriate speed even if the person's speed changes suddenly. On the other hand, because the second speed V2 is the speed for the relatively short second period T2, the second speed V2 is unlikely to be an appropriate speed if the person's speed changes suddenly. As described above, the operation detection device 80 determines whether or not a person is a user based on the first speed V1 as the average speed and the second speed V2 as the instantaneous speed. Therefore, the operation detection device 80 can further suppress determination of whether or not a person who is likely not a user has made a gesture.
[0068] <Example of change> This embodiment can be modified as follows: This embodiment and the following modifications can be combined and implemented within the scope of technical compatibility.
[0069] The operation detection device 80 may determine whether or not a person moving through the second area A2 toward the first area A1 is a user based only on the first speed V1 of the first speed V1 and the second speed V2. Similarly, the operation detection device 80 may determine whether or not a person moving through the second area A2 toward the first area A1 is a user based only on the second speed V2 of the first speed V1 and the second speed V2.
[0070] The operation detection device 80 does not have to determine whether the person moving through the second area A2 toward the first area A1 is the user based on the first speed V1 and the second speed V2. For example, the operation detection device 80 may determine whether the person moving through the second area A2 toward the first area A1 is the user based on whether the reference distance Db calculated when the current image was acquired is shorter than the reference distance Db calculated when the previous image was acquired.
[0071] The operation detection device 80 may determine whether or not a person is the user based on the acceleration of the person moving through the second area A2 toward the first area A1. When an image shows only one person, the operation detection device 80 may acquire the position of the skeleton points of the person and determine whether the person has made a gesture, regardless of whether the person in the image is the user or not.
[0072] The operation detection device 80 does not need to use a trained model based on machine learning when identifying the position of a person, acquiring the positions of the person's skeletal points, or determining whether the person has made a gesture. For example, the operation detection device 80 may use image processing to identify the position of a person, acquire the positions of the person's skeletal points, or determine whether the person has made a gesture.
[0073] The operation detection device 80 may detect an opening operation that indicates the user's intention to open the rear door 40. In other words, the opening operation does not have to be a gesture. The opening operation referred to here may be, for example, a user causing a sensor provided on the rear door 40 or in the vicinity of the rear door 40 to read user information. Alternatively, the opening operation may be the user standing still within the first area A1. In this case, the operation detection device 80 may determine that the tracked user has performed an opening operation when the position of the skeleton point of the tracked user does not change significantly over a predetermined period of time.
[0074] The vehicle 10 may include a notification unit that notifies the user of various information. The notification unit may be capable of transmitting information to the user by, for example, sound, light, or display. For example, the notification unit may be a buzzer or siren for answerback, or a light such as a turn signal lamp. The operation detection device 80 preferably causes the notification unit to notify the user that a gesture is being urged when the user enters the first area A1, in other words, when a reference point corresponding to the tracked user is present in the first area A1 (step S17: YES). This allows the operation detection device 80 to prompt the user to start a gesture at an appropriate position.
[0075] The operation detection device 80 may measure the number of times one or more people enter the first area A1. If the number of times in a predetermined period of time reaches or exceeds a predetermined number, the operation detection device 80 may temporarily suspend the series of processes for determining whether or not the user has performed an opening operation. This reduces the possibility that the operation detection device 80 will erroneously determine whether or not the user has performed a gesture in a situation where people frequently pass by around the vehicle.
[0076] In the above embodiment, the operation detection device 80 determines whether the tracked user has entered the first area A1, and then determines whether the tracked user has made a gesture. In a modified example, the operation detection device 80 may determine whether the tracked user has entered the first area A1, and then determine whether the tracked user has made a gesture. That is, even if the operation detection device 80 determines that the tracked user has made a gesture, if the tracked user has not entered the first area A1, the operation detection device 80 does not output an open operation command signal. On the other hand, if the operation detection device 80 determines that the tracked user has made a gesture and if the tracked user has entered the first area A1, the operation detection device 80 outputs an open operation command signal. In other words, the processing of step S17 in the flowchart shown in FIG. 4 may be processing performed after step S80 and before step S81 in the flowchart shown in FIG. 7.
[0077] The image acquisition unit 82 does not have to acquire all images output from the camera 60. The image acquisition unit 82 may acquire images output from the camera while thinning them out in accordance with the frame rate of the camera 60. This prevents an increase in the load on the operation detection device 80 even when a camera 60 with a high frame rate is used.
[0078] The rear door 40 may be a swing door that opens and closes the rear opening 22 by rotating about an axis that extends in the vertical direction. The camera 60 does not have to be mounted on the side mirror 31. The camera 60 may be mounted in any position that allows it to capture images of the first area A1 and the second area A2.
[0079] The vehicle 10 may include a first moving object detector that detects a moving object in the first area A1 and a second moving object detector that detects a moving object in the second area A2. In this case, the second moving object detector may be a millimeter wave radar.
[0080] The vehicle 10 may include a front door drive unit that drives the front door 30. In this case, the door control device 70 may open the front door 30 when an opening operation command signal based on a user gesture is input from the operation detection device 80.
[0081] The vehicle 10 may include a back door and a back door drive unit that drives the back door. In this case, the door control device 70 may open the back door when an open operation command signal based on a user gesture is input from the operation detection device 80. However, it is preferable that the camera 60 is a rear camera provided in the back door. It is also preferable that the first area A1 and the second area A2 are areas set behind the vehicle 10.
[0082] The door control device 70 and the operation detection device 80 are not limited to processing circuits equipped with a CPU and ROM and executing software processing. For example, the door control device 70 and the operation detection device 80 may be equipped with dedicated hardware circuits that execute at least some of the various processes executed in the above-described embodiment. An example of a dedicated hardware circuit is an ASIC. ASIC is an abbreviation for "Application Specific Integrated Circuit." In other words, the door control device 70 and the operation detection device 80 may have any of the following configurations (a) to (c):
[0083] (a) A processing circuit comprising a processing device that executes all of the above processes according to a program, and a program storage device such as a ROM that stores the program. (b) A processing circuit comprising a processing device and a program storage device that execute part of the above processing according to a program, and a dedicated hardware circuit that executes the remaining processing.
[0084] (c) A processing circuit having dedicated hardware circuitry for performing all of the above processes. Here, there may be a plurality of software execution devices each having a processing device and a program storage device, and a plurality of dedicated hardware circuits.
[0085] <Summary of this embodiment> The operation detection device detects a user's opening operation to open a vehicle door from a fully closed position where the door opening is fully closed, and includes: a user determination unit that, when an area around the door opening that is closer to the door opening is defined as a first area and an area farther from the door opening than the first area is defined as a second area, determines that a moving object moving through the second area toward the first area is the user based on the detection result of a moving object detection unit that is capable of detecting a moving object present in at least the second area, and determines that a moving object that does not move through the second area toward the first area is not the user; and an operation determination unit that determines whether the moving object determined by the user determination unit to be the user has performed the opening operation in the first area.
[0086] The operation detection device determines whether a moving object is a user based on the moving direction of the moving object in the second area. The operation detection device then determines whether an opening operation has been performed in the first area on a moving object determined to be a user, but does not determine whether an opening operation has been performed in the first area on a moving object determined not to be a user. Therefore, the operation detection device determines whether an opening operation has been performed on a moving object that is likely to be a user. As a result, the operation detection device can improve the accuracy of detecting an opening operation.
[0087] In the operation detection device, it is preferable that the moving object detection unit is a camera capable of capturing images of the first area and the second area, the opening operation is a gesture in which the user moves a body part of the user, the user determination unit determines that the moving object moving in the second area toward the first area is the user based on multiple images captured by the camera at different times, and determines that the moving object not moving in the second area toward the first area is not the user, and the operation determination unit determines whether the moving object determined by the user determination unit to be the user performed the gesture in the first area based on multiple images captured by the camera at different times.
[0088] The operation detection device determines both the direction of movement of the user in the second area and whether the user has made a gesture in the first area based on multiple images captured by the camera. This reduces the number of sensors required to implement the operation detection device compared to when detection results from separate sensors are used for both determinations.
[0089] In the operation detection device, it is preferable that the user determination unit calculates a first speed by dividing the amount of movement of the moving object in a first period by the first period, and determines whether the moving object is the user based on the first speed.
[0090] If the speed of the moving object deviates from the moving speed of the user when approaching the vehicle, it is highly likely that the moving object is not the user. In this regard, the operation detection device determines whether the moving object is the user based on the first speed, which is the speed of the moving object. Therefore, the operation detection device can suppress determination of whether a moving object that is highly likely not the user has made a gesture.
[0091] In the operation detection device, it is preferable that the user determination unit calculates a second speed by dividing the amount of movement of the moving object in a second period shorter than the first period by the second period, and determines whether the moving object is the user based on the second speed.
[0092] Even if a person is moving from the second area toward the first area at a reasonable speed, if the person suddenly decelerates or accelerates, it is highly likely that the person is not the user of the vehicle. Because the first speed is a speed for a relatively long first period, the first speed may be a reasonable speed even if the person's speed changes suddenly. On the other hand, because the second speed is a speed for a relatively short second period, the second speed is unlikely to be a reasonable speed if the person's speed changes suddenly. As described above, the operation detection device determines whether the moving object is a user based on the first speed as an average speed and the second speed as an instantaneous speed. Therefore, the operation detection device can further suppress determining whether a moving object that is highly likely not a user has made a gesture. [Explanation of symbols]
[0093] 10...vehicle, 20...vehicle body, 21...front opening, 22...rear opening (door opening), 30...front door, 31...side mirror, 40...rear door (vehicle door), 50...door drive unit, 60...camera (moving object detection unit), 70...door control device, 80...operation detection device, 81...storage unit, 82...image acquisition unit, 83...person extraction unit, 84...user determination unit, 85...skeleton estimation unit, 86...gesture determination unit, 87...command unit A1...first area, A2...second area, T1...first period, T2...second period, V1...first speed, V2...second speed, VHth1...first upper limit speed, VHth2...second upper limit speed, VLth1...first lower limit speed, VLth2...second lower limit speed
Claims
1. An operation detection device that detects a user's opening operation for opening a vehicle door from a fully closed position that fully closes a door opening, When an area around the door opening that is closer to the door opening is defined as a first area, and an area farther from the door opening than the first area is defined as a second area, a user determination unit that determines, based on a detection result from a moving object detection unit capable of detecting a moving object present at least in the second area, that the moving object moving through the second area toward the first area is the user, and that the moving object not moving through the second area toward the first area is not the user; an operation determination unit that determines whether the moving object determined to be the user by the user determination unit has performed the opening operation in the first area. Operation detection device.
2. the moving object detection unit is a camera capable of capturing images of the first area and the second area, the opening operation is a gesture in which the user moves a body part of the user, the user determination unit determines, based on a plurality of images captured by the camera at different times, that the moving object moving through the second area toward the first area is the user, and determines that the moving object not moving through the second area toward the first area is not the user; The operation determination unit determines whether the moving object determined to be the user by the user determination unit has performed the gesture in the first area based on a plurality of images captured by the camera at different times. The operation detection device according to claim 1 .
3. The user determination unit calculates a first speed by dividing a movement amount of the moving object in a first period by the first period, and determines whether the moving object is the user based on the first speed. The operation detection device according to claim 1 or 2.
4. The user determination unit calculates a second speed by dividing a movement amount of the moving object in a second period shorter than the first period by the second period, and determines whether the moving object is the user based on the second speed. The operation detection device according to claim 3 .
Citation Information
Patent Citations
Automatic opening / closing device for vehicle door
JP2013007171A