Vehicle control system
The vehicle control device improves posture determination by using object detection and depth estimation to assess safe postures with reduced computational load, addressing inefficiencies in existing methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2024-11-05
- Publication Date
- 2026-05-19
AI Technical Summary
Existing posture determination techniques for vehicle occupants require large amounts of learning data and high computational costs to handle variations in glove patterns and environmental characteristics, making them inefficient for determining safe gripping postures.
A vehicle control device uses object detection and depth estimation to determine if a vehicle occupant is in a safe posture by calculating the amount of movement and depth difference between the occupant's hand and posture stabilization equipment, reducing computational load through pre-stored detection areas and thresholds.
This approach allows for efficient determination of safe postures with lower computational costs compared to machine learning, enabling safer vehicle operations.
Smart Images

Figure 2026081759000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a vehicle control device.
Background Art
[0002] Conventionally, techniques for estimating the posture of an object are known. For example, Patent Document 1 discloses a technique related to a method and apparatus capable of simultaneously estimating object recognition and position and orientation by machine learning.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When using posture determination for determining permission to start a vehicle, it is desirable to determine that the posture of an occupant in the vehicle is a safe posture, for example, a gripping posture of gripping a handrail or the like. Determination of the gripping posture by image recognition is mainly performed using machine learning. However, determination by machine learning may require a large amount of learning data and high computational costs in order to cope with changes in the patterns of gloves and handrails and environmental characteristics such as the interior of the vehicle.
[0005] In view of such circumstances, an object of the present disclosure is to improve the technique for estimating the posture of an object.
Means for Solving the Problems
[0006] A vehicle control device according to an embodiment of the present disclosure includes a control unit and an imaging unit that images the interior of the vehicle. The control unit acquires an image of the interior of the vehicle from the imaging unit, detects a target from the image using object detection technology, The amount of movement of the object is calculated from the aforementioned image, If at least a portion of the hand of the subject is located within the detection area including the posture stabilization equipment, and the amount of movement of the subject is less than or equal to the movement threshold, the depth difference between the depth from the imaging unit to the hand of the subject and the depth from the imaging unit to the posture stabilization equipment is estimated. If the depth difference is less than or equal to the depth difference threshold, a flag indicating that the target's posture is in a safe posture is turned ON; if the depth difference is greater than the depth difference threshold, the flag is turned OFF. [Effects of the Invention]
[0007] According to one embodiment of the present disclosure, the technique for estimating the orientation of an object is improved. [Brief explanation of the drawing]
[0008] [Figure 1] This block diagram shows a schematic configuration of a system according to one embodiment of the present disclosure. [Figure 2] This flowchart shows the operation of the vehicle control device according to this embodiment. [Modes for carrying out the invention]
[0009] The embodiments of this disclosure will be described below with reference to the drawings.
[0010] (Summary of this embodiment) Referring to Figure 1, an overview of the vehicle control device 1 according to the embodiment of this disclosure will be described. The vehicle control device 1 is an electronic device mounted in a vehicle, such as a computer. The vehicle control device 1 detects an object from an image inside the vehicle using any object detection technology. The image may be a still image or a moving image. The object is the occupant inside the vehicle.
[0011] The vehicle is any vehicle capable of carrying one or more occupants, such as a car, bus, or shuttle bus. The vehicle may be an autonomous vehicle capable of autonomous driving at levels 1 to 5 as defined by the Society of Automotive Engineers (SAE). The vehicle may also be a manually driven vehicle at level 0. The vehicle may be remotely monitored by an observer outside the vehicle. The vehicle may be a vehicle specifically designed for MaaS (Mobility as a Service).
[0012] First, an overview of this embodiment will be described, and details will be described later. The vehicle control device 1 according to this embodiment comprises a control unit 10 and an imaging unit 12 that images the inside of the vehicle. The control unit 10 acquires an image of the inside of the vehicle from the imaging unit 12. The control unit 10 detects an object from the image using object detection technology. The control unit 10 calculates the amount of movement of the object from the image. If at least a part of the object's hand is located within the detection area including the attitude stabilization equipment, and the amount of movement of the object is less than or equal to the amount of movement threshold, the control unit 10 estimates the depth difference between the depth from the imaging unit 12 to the object's hand and the depth from the imaging unit 12 to the attitude stabilization equipment. If the depth difference is less than or equal to the depth difference threshold, the control unit 10 turns on a flag indicating that the object's attitude is in a safe attitude, and if the depth difference is greater than the depth difference threshold, it turns off the flag.
[0013] According to this embodiment, a measurement-based approach using human body detection and depth estimation makes it possible to determine whether a person is gripping a posture stabilization device (e.g., a handrail or strap) at a lower cost than a machine learning approach.
[0014] (Configuration of the vehicle control device 1) The vehicle control device 1 comprises a control unit 10, an imaging unit 12, a communication unit 14, and a storage unit 16. Each unit is connected to the others via an in-vehicle network such as CAN (Controller Area Network) or a dedicated line, enabling communication between them.
[0015] The control unit 10 includes one or more processors, one or more programmable circuits, one or more dedicated circuits, or a combination thereof. The processors are general-purpose processors such as CPUs (central processing units) or GPUs (graphics processing units), or dedicated processors specialized for specific processing. The control unit 10 controls each part of the vehicle control device 1 and executes processing related to the operation of the vehicle control device 1.
[0016] The imaging unit 12 is an arbitrary imaging module installed inside the vehicle and capable of imaging some or all of the seats and objects inside the vehicle. The imaging module includes one or more cameras. In this embodiment, the imaging unit 12 is a single camera installed on the ceiling of the vehicle. In this embodiment, the imaging unit 12 captures RGB images. The imaging unit 12 may also include a distance measuring device such as a depth sensor or stereo camera that acquires depth images.
[0017] The communication unit 14 includes at least one communication interface for connecting to an in-vehicle network. The communication interface supports, for example, mobile communication standards such as 4G (4th generation) or 5G (5th generation), V2X (vehicle-to-everything) communication standards such as DSRC (dedicated short range communications) or cellular V2X, or wireless LAN (local area network) communication standards such as IEEE 802.11 (Institute of Electrical and Electronics Engineers 802.11).
[0018] The storage unit 16 includes one or more memories. Each memory included in the storage unit 16 may function as, for example, a main memory device, an auxiliary memory device, or a cache memory. The storage unit 16 stores any information used for the operation of the vehicle control device 1. For example, the storage unit 16 stores a system program, an application program, embedded software, and any data used for object detection and pose estimation. In the present embodiment, the storage unit 16 stores in advance a detection area including a posture stabilization facility. Thereby, it is not necessary to set a detection area every time the position of the target hand is determined, and the processing load on the vehicle control device 1 is reduced. The detection area is, for example, a two-dimensional area on an image. The posture stabilization facility refers to a facility such as a handrail or a strap that a person can hold to stabilize their posture. The size of the detection area may be set according to the type of the posture stabilization facility. For example, when the posture stabilization facility is movable like a strap, the detection area may be set to include the movable range of the posture stabilization facility. In the present embodiment, the storage unit 16 stores in advance the depth from the imaging unit 12 to the posture stabilization facility. Thereby, the processing load on the vehicle control device 1 is reduced. The information stored in the storage unit 16 may be updated with information acquired from an in-vehicle network or an external network via the communication unit 14. In the present embodiment, the state (ON or OFF) of a flag indicating that the target is holding the posture stabilization facility is updated by the control unit 10 and also stored in the storage unit 16.
[0019] In the present embodiment, the storage unit 16 stores in advance an object detection AI that detects an object and the skeleton of the object included in an image, and a depth estimation AI that estimates the depth from the imaging unit 12 to the subject. The skeleton and depth can be detected from an RGB image. In order to improve the accuracy, a depth image may also be used together for the detection of the skeleton. The object detection AI may include any object detection model such as YOLO (You Only Look Once) or CNN (Convolutional Neural Network). The depth estimation AI may include any depth estimation model.
[0020] (Operation flow of the vehicle control device 1) Referring to FIG. 2, the operation of the vehicle control device 1 according to the present embodiment will be described. While the vehicle is temporarily stopped at a stop, for example, the control unit 10 executes the following S101 to S108 for each target in the vehicle to determine whether the posture of each target is a safe posture. In the following, the communication of each part of the vehicle control device 1 is performed via the communication unit 14 and the in-vehicle network.
[0021] S101: The control unit 10 of the vehicle control device 1 acquires an image captured by the imaging unit 12.
[0022] S102: The control unit 10 detects a target from the image using an object detection technique.
[0023] The control unit 10 inputs the image captured by the imaging unit 12 to the object detection AI pre-stored in the storage unit 16. The object detection AI detects the presence or absence of a target in the input image and the skeleton of the target.
[0024] S103: The control unit 10 determines whether at least a part of the target's hand is located within the detection area including the posture stabilizing equipment. If at least a part of the target's hand is located within the detection area (S103 - YES), the process proceeds to S104. If the target's hand is not located within the detection area (S103 - NO), the process proceeds to S109.
[0025] In S103, based on the position of the target's hand, it is determined whether the target is gripping the posture stabilizing equipment. If the target is gripping the posture stabilizing equipment, the posture of the target can be regarded as a safe posture. When at least a part of the target's hand is located within the detection area, there is a high possibility that the target is gripping the posture stabilizing equipment. The control unit 10 may determine that at least a part of the hand within the detection area is located when at least the joints of the hand or wrist in the target's skeleton are continuously located within the detection area for at least two frame periods.
[0026] S104: The control unit 10 calculates the amount of movement of the target from the image.
[0027] The control unit 10 may calculate the amount of movement per unit time for each feature point (e.g., joint) in the target skeleton and sum the amount of movement per unit time for all feature points to calculate the amount of movement of the target, i.e., the amount of movement of the entire target body. When summing the amount of movement for all feature points, different weighting coefficients may be assigned to each feature point.
[0028] S105: The control unit 10 determines whether the amount of movement of the object is less than or equal to the amount of movement threshold. If the amount of movement of the object is less than or equal to the amount of movement threshold (S105-YES), the process proceeds to S106. If the amount of movement of the object is greater than the amount of movement threshold (S105-NO), the process proceeds to S109.
[0029] In S105, a safety posture is determined based on the amount of movement of the object. When the object is stationary or moving only a little, it is more likely to be safe than when the object is moving.
[0030] S106: The control unit 10 estimates the depth difference between the depth from the imaging unit 12 to the target hand and the depth from the imaging unit 12 to the attitude stabilization equipment.
[0031] The control unit 10 inputs the image to the depth estimation AI pre-stored in the memory unit 16. The depth estimation AI estimates the depth difference based on the input image. The depth from the imaging unit 12 to the hand may be calculated from the hand coordinates in the image depth map, or from the skeletal coordinates of the hand or wrist joints.
[0032] S107: The control unit 10 determines whether the depth difference is less than or equal to the depth difference threshold. If the depth difference is less than or equal to the depth difference threshold (S107-YES), the process proceeds to S108. If the depth difference is greater than the depth difference threshold (S107-NO), the process proceeds to S109.
[0033] In S103, a determination is made based on the depth difference to determine whether the object is grasping the posture stabilization device. At the point of S107, it has been determined that at least a portion of the hand is located within the detection area (see S103). Therefore, if the depth difference is less than or equal to the depth difference threshold, the depth difference between the hand and the posture stabilization device can be considered small, and thus the probability that the object is grasping the posture stabilization device becomes even higher.
[0034] S108: The control unit 10 turns on a flag indicating that the target's posture is a safe posture. The process then terminates.
[0035] If the subject is gripping a posture stabilization device, or if the subject's movement is minimal, the subject's posture can be considered a safe posture.
[0036] S109: The control unit 10 turns the flag OFF. The process then returns to S101.
[0037] The control unit 10 repeats steps S101 to S109 until the flag is turned ON.
[0038] The control unit 10 executes processes S101 to S109 for each object in the vehicle. If the flag is ON for all objects, the control unit 10 permits the vehicle to start. If the flag is OFF for at least one object, the control unit 10 does not permit the vehicle to start.
[0039] This disclosure has been described based on the drawings and embodiments, but it should be noted that those skilled in the art may make various modifications and alterations based on this disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of this disclosure. For example, the functions included in each component or step can be rearranged in a logically consistent manner, and multiple components or steps can be combined into one or separated. For example, in the above-described embodiment, it is also possible to distribute the configuration and operation of the vehicle control device 1 to multiple computers or devices that can communicate with each other. Also, for example, in the above-described embodiment, it is possible to provide the imaging unit 12 in a device separate from the other parts of the vehicle control device 1.
[0040] In the embodiments described above, the determinations in S103, S105, and S107 may be performed in any order. For example, the determination of whether the object is gripping the attitude stabilization device in S103 and S107 may be followed by the determination of the safe posture based on the amount of movement of the object in S105. Furthermore, if the control unit 10 of the vehicle control device 1 permits the vehicle to start, it may start the vehicle by executing automatic driving control of the vehicle. The vehicle control device 1 may also be used to provide Mobility as a Service (MaaS), which is a mobility-based service. [Explanation of Symbols]
[0041] 1 Vehicle control device, 10 Control unit, 12 Imaging unit, 14 Communication unit, 16 Storage unit
Claims
1. A vehicle control device comprising a control unit and an imaging unit for imaging the interior of a vehicle, The control unit, From the aforementioned imaging unit, an image of the inside of the vehicle is acquired. Using object detection technology, the object is detected from the image. The amount of movement of the object is calculated from the aforementioned image, If at least a portion of the hand of the subject is located within the detection area including the posture stabilization equipment, and the amount of movement of the subject is less than or equal to the movement threshold, the depth difference between the depth from the imaging unit to the hand of the subject and the depth from the imaging unit to the posture stabilization equipment is estimated. A vehicle control device that, when the depth difference is less than or equal to a depth difference threshold, turns on a flag indicating that the attitude of the target is in a safe attitude, and when the depth difference is greater than the depth difference threshold, turns off the flag.
2. A vehicle control device according to claim 1, wherein the control unit turns off the flag if at least a part of the hand of the object is not located within the detection area, or if the amount of movement of the object is greater than the amount of movement threshold.
3. A vehicle control device according to claim 1, wherein the vehicle is an autonomous vehicle.
4. A vehicle control device according to claim 1, wherein the attitude stabilization device is a handrail or a strap.
5. A vehicle control device according to any one of claims 1 to 4, wherein the control unit is If the flag is set to ON for all objects within the vehicle, the vehicle is permitted to start. A vehicle control device that does not permit the vehicle to start if the flag is turned OFF for at least one object within the vehicle.
6. A method for providing MaaS (Mobility as a Service) using the vehicle control device described in claim 1.