Trained model, trained model generation method, training program, training data set generation method, vehicle cabin state detection method, vehicle cabin state detection program, and vehicle cabin state detection system

WO2026168119A1PCT designated stage Publication Date: 2026-08-13DENSO CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-08-13

Smart Images

  • Figure JP2026001248_13082026_PF_FP_ABST
    Figure JP2026001248_13082026_PF_FP_ABST
Patent Text Reader

Abstract

A trained model (50) receives, as input, a captured image (P) which includes one or more seats in a vehicle cabin, and outputs, as states in the vehicle cabin detected from the image, at least two from among: a seat occupancy state (S1-S4) for each seat; a seat belt wearing state (S5-S8) for each seat; a driver's steering wheel gripping state (S9); a driver's device operation state (S10); and a child seat wearing state (S11-S13) for each seat.
Need to check novelty before this filing date? Find Prior Art

Description

Trained model, method for generating trained model, learning program, method for generating learning dataset, method for detecting in-vehicle state, in-vehicle state detection program, and in-vehicle state detection system Cross-reference to related applications

[0001] This application is based on Japanese Application No. 2025-016854 filed on February 4, 2025, the contents of which are incorporated herein by reference.

[0002] The present disclosure relates to a trained model, a method for generating a trained model, a learning program, a method for generating a learning dataset, a method for detecting an in-vehicle state, an in-vehicle state detection program, and an in-vehicle state detection system used for detecting the state inside a vehicle cabin.

[0003] In recent years, with the improvement of automobile safety and the development of driving systems, the importance of monitoring the interior of a vehicle to appropriately grasp the situation inside the vehicle cabin has been increasing. For example, for ensuring the safety of drivers and passengers, technologies for accurately detecting various states such as the occupancy status of passengers, the wearing status of seat belts, and the presence or absence of steering wheel gripping are required.

[0004] Therefore, for example, in conventional technologies, there are known technologies for detecting various states using individual detectors.

[0005] Japanese Patent Application Laid-Open No. 2010-211705

[0006] However, in the method of detecting the state inside a vehicle cabin using individual detectors for each state, the computational load increases and real-time processing becomes difficult. Also, when the number of detection targets increases, the false detection rates of each detector accumulate, and the overall recognition accuracy decreases.

[0007] An object of the present disclosure is to reduce the computational load and improve the recognition accuracy when detecting the state inside a vehicle cabin.

[0008] The trained model described in this disclosure takes an image taken including one or more seats inside a vehicle as input, and outputs two or more of the following conditions for the state of the vehicle interior detected from the image: occupancy status of each seat, seat belt status of each seat, driver's grip on the steering wheel, driver's operation status of equipment, and child seat installation status of each seat.

[0009] The method for generating a trained model according to this disclosure involves generating a trained model using a training dataset that includes multiple training datasets, each containing an image of a vehicle interior including one or more seats, annotated with two or more of the following conditions for the vehicle interior included in the image: the occupancy status of each seat, the seatbelt status of each seat, the driver's grip on the steering wheel, the driver's operation status of equipment, and the child seat status of each seat.

[0010] The learning program described herein causes a computer to perform a process to generate a trained model using a training dataset that includes multiple training datasets annotated with two or more of the following conditions for the state of the vehicle interior detected from the images, for each image taken including one or more seats inside the vehicle: the occupancy status of each seat, the seatbelt status of each seat, the driver's grip on the steering wheel, the driver's operation status of the equipment, and the child seat status of each seat.

[0011] The method for generating a training dataset according to this disclosure is a method for generating a training dataset that includes multiple training datasets used to train a trained model, in which an image taken including one or more seats inside a vehicle is taken as input, and two or more of the following states inside the vehicle detected from the image are: the occupancy status of each seat, the seat belt status of each seat, the driver's grip on the steering wheel, the driver's operation status of the equipment, and the child seat status of each seat, and the training dataset includes a process of adjusting the number of training datasets with different annotations to a similar number.

[0012] The vehicle interior state detection method of this disclosure includes an image acquisition process and a detection process. The image acquisition process is the process of acquiring an image taken including one or more seats in the vehicle interior. The detection process is the process of inputting the image acquired in the image acquisition process into a trained model that takes the image taken including one or more seats in the vehicle interior as input and outputs two or more of the following as vehicle interior states detected from the image: the occupancy status of each seat, the seat belt status of each seat, the driver's grip on the steering wheel, the driver's operation status of equipment, and the child seat installation status of each seat, in order to detect the status of two or more vehicle interior states.

[0013] The vehicle interior state detection program described herein causes a computer to perform the image acquisition process and the detection process described above.

[0014] The vehicle interior state detection system disclosed herein comprises a camera and a detection device. The camera has a shooting range that includes one or more seats in the vehicle interior. The detection device takes an image taken including one or more seats in the vehicle interior as input and outputs two or more of the following as vehicle interior states detected from the image: the occupancy status of each seat, the seat belt status of each seat, the driver's grip on the steering wheel, the driver's operation status of equipment, and the child seat installation status of each seat. The detection device then takes the image taken by the camera and detects two or more vehicle interior states.

[0015] Figure 1 schematically shows an example of a learning device and an in-vehicle state detection system according to one embodiment. Figure 2 shows an example of a detection target or query, and an in-vehicle state or annotation applied to a trained model according to one embodiment. Figure 3 is a functional block diagram showing an example of a detection device according to one embodiment. Figure 4 is a functional block diagram showing an example of a trained model according to one embodiment. Figure 5 is a functional block diagram showing an example of a learning device according to one embodiment. Figure 6 is a data flow diagram showing an example of learning in the learning unit according to one embodiment. Figure 7 is a functional block diagram showing an example of a model being trained according to one embodiment. Figure 8 shows a first example in which a training image used in the training dataset has been annotated. Figure 9 shows a second example in which a training image used in the training dataset has been annotated. Figure 10 shows a third example in which a training image used in the training dataset has been annotated. Figure 11 shows a fourth example in which a training image used in the training dataset has been annotated.

[0016] The following describes one embodiment with reference to the drawings. In this description of the embodiment, the terms "first," "second," etc., are used solely to distinguish between configurations with the same or similar names, and do not imply any superiority or order of configurations.

[0017] Figure 1 shows the in-vehicle state detection system 1, the learning device 20, and the providing device 30. The in-vehicle state detection system 1 is installed in the vehicle 90 and is a system for detecting the state of the in-vehicle interior of the vehicle 90 using a trained model generated by the learning device 20. The in-vehicle state detection system 1 can also be called, for example, an in-vehicle monitoring system.

[0018] The learning device 20 and the providing device 30 are installed, for example, outside the vehicle 90. The learning device 20 is a device for generating trained models to be provided to the in-vehicle state detection system 1. The providing device 30 is composed of, for example, a database server and is communicated to the in-vehicle state detection system 1 via a telecommunications line W such as the Internet. The providing device 30 stores the trained models generated by the learning device 20 and provides the trained models to the in-vehicle state detection system 1, for example, in response to a request from the in-vehicle state detection system 1.

[0019] The vehicle interior state detection system 1 comprises a camera 2 and a detection device 10. The camera 2 is, for example, a camera installed inside the vehicle interior and captures an image including one or more seats inside the vehicle interior. When the camera 2 captures an image of one seat inside the vehicle interior, the seat to be photographed is preferably the driver's seat, where the driver is always seated when driving. In this embodiment, the camera 2 captures an image including multiple seats inside the vehicle interior. For example, if the vehicle 90 has two rows of seats, the camera 2 captures an image including the driver's seat, passenger seat, rear right seat, and rear left seat, as shown in images P1 to P4 in Figures 8 to 11.

[0020] The detection device 10 can be configured, for example, as an ECU (Electronic Control Unit) mounted on a vehicle 90. The detection device 10 takes an image captured by the imaging device 2 as input and detects two or more interior states from the image. The interior states to be detected can be arbitrarily set as long as they are captured in the image. The detection targets in the interior of the vehicle that can be detected by the detection device 10, and the interior states detected for those detection targets, are shown, for example, in Figure 2. In the example in Figure 2, the detection targets are shown as O1 to O13, and the interior states detected for detection targets O1 to O13 are shown as S1 to S13. The detection device 10 detects the interior states S1 to S13 corresponding to the pre-set detection targets O1 to O13 from the image captured by the imaging device 2.

[0021] Detected objects O1 to O4 represent the occupancy status of each seat, in this case the driver's seat, passenger seat, rear right seat, and rear left seat, and the detected states S1 to S4 for detected objects O1 to O4 are "absent" or "occupied". Detected objects O5 to O8 represent the seatbelt wearing status of each seat, in this case the driver's seat, passenger seat, rear right seat, and rear left seat, and the detected states S5 to S8 for detected objects O5 to O8 are "not worn" or "worn". Detected object O9 represents the driver's grip on the steering wheel, and the detected state S9 for detected object O9 is "both hands gripping", "right hand only gripping", "left hand only gripping", or "both hands released", meaning neither hand is gripping.

[0022] The detection target O10 is the operating status of the driver's equipment, and the detected state S10 for detection target O10 is "not operated" or "operating". The detection targets O11 to O13 are the installation status of the child seats for each seat, in this case the passenger seat, the right rear seat, and the left rear seat, and the detected states S11 to S13 for detection targets O11 to O13 are "not installed" or "installed". Furthermore, the detection targets and states in the trained model 50 can be arbitrarily added or removed.

[0023] The status of the rear seats is set appropriately depending on the vehicle model to which the in-cabin state detection system 1 is applied. For example, if there are two rear seats, the status of the rear seats will be set as follows: "Rear seat right," indicating the usage status of the right rear seat, and "Rear seat left," indicating the usage status of the left rear seat. If there are three rear seats, the status of the rear seats will be set as follows: "Rear seat right," "Rear seat left," and "Rear seat left." Furthermore, in the case of a so-called two-seater with only front seats, the status of the rear seats will be omitted.

[0024] As shown in Figure 3, the detection device 10 comprises a first control unit 11, a first interface 12, a first storage unit 13, an image acquisition processing unit 14, and a detection processing unit 15. The first control unit 11 includes a hardware processor such as a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory), and controls each component according to information processing. In Figure 3, the components of the first control unit 11 are represented as CPU 111, RAM 112, and ROM 113, respectively.

[0025] The first interface 12 is an interface for connecting the detection device 10 to an external device, and is configured appropriately depending on the external device to be connected. In this embodiment, the first interface 12 is connected to an external device such as a camera 2 or a navigation device 3 in a communicative manner. The conditions inside the vehicle detected by the detection device 10 can be sent to the navigation device 3, for example, and used for processing such as notifying the driver.

[0026] The first storage unit 13 stores the vehicle interior state detection program 131 and the learned model 50. In Figure 3, the vehicle interior state detection program 131 is simply referred to as the detection program 131. The first storage unit 13 is a tangible, non-temporary computer-readable medium, that is, a non-transitional, tangible recording medium. Examples of the first storage unit 13 include, but are not limited to, RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Drive), magnetic disk, magneto-optical disk, CD-ROM (Compact Disc Read Only Memory), DVD-ROM (Digital Versatile Disc Read Only Memory), semiconductor memory, etc.

[0027] The first storage unit 13 may be an internal media directly connected to the bus of the computer constituting the detection device 10, or it may be an external media connected via the first interface 12 or a telecommunications line. Furthermore, when the in-vehicle state detection program 131 is delivered to the detection device 10 via a communication line, the detection device 10, upon receiving the program, expands and executes it in the first storage unit 13 or the RAM 112 of the first control unit 11, thereby realizing the image acquisition processing unit 14 and the detection processing unit 15.

[0028] The vehicle interior state detection program 131 is a computer program for virtually implementing the image acquisition processing unit 14 and the detection processing unit 15 on a computer. The detection device 10 can virtually implement the image acquisition processing unit 14 and the detection processing unit 15 on a computer by, for example, the first control unit 11 reading and executing the vehicle interior state detection program 131 from the first storage unit 13. In other words, the image acquisition processing unit 14 and the detection processing unit 15 can be composed of functional units that are virtually implemented by, for example, the first control unit 11 executing the vehicle interior state detection program 131.

[0029] The trained model 50 takes an image taken including one or more seats in the vehicle interior as input, in this case an image taken by the camera 2 including multiple seats in the vehicle interior, and outputs two or more of the following vehicle interior states detected from the image: the occupancy status S1 to S4 of each seat shown in Figure 2, the seat belt wearing status S5 to S8 of each seat, the driver's grip on the steering wheel S9, the driver's equipment operation status S10, and the child seat installation status S11 to S13 of each seat. The number and content of vehicle interior states output by the trained model 50 can be arbitrarily set during the training of the trained model 50. In this embodiment, the trained model 50 outputs the occupancy status S1 to S4 of each seat, the seat belt wearing status S5 to S8 of each seat, the driver's grip on the steering wheel S9, the driver's equipment operation status S10, and the child seat installation status S11 to S13 of each seat shown in Figure 2 all at once. Furthermore, when the trained model 50 receives an input image of, for example, only the driver's seat as one seat in the vehicle interior, it outputs two or more of the following: the occupancy status of the driver's seat S1, the seatbelt status of the driver's seat S5, the driver's grip on the steering wheel S9, and the driver's operation status of the equipment S10.

[0030] As shown in Figure 4, the trained model 50 has an architecture with a transformer, and is a customized transformer-based model that applies the DETR (Detection Transformer) used in object detection. The trained model 50 has an encoder 51 and a decoder 52, and the decoder 52 holds multiple queries Q1 to Q13. All of these queries Q1 to Q13 are initialized as learnable parameters and optimized through learning. This makes it possible to automatically determine which states S1 to S13 each query Q1 to Q13 should output. In other words, each query Q1 to Q13 is processed in parallel.

[0031] Furthermore, the fact that queries Q1 to Q13 are optimized through learning is an important feature in achieving end-to-end characteristics where queries are automatically optimized through learning without requiring manual parameter adjustments, compared to conventional methods that pre-fix the position and size of the detection target, such as anchor settings. The trained model 50 outputs two or more states in a single output according to two or more queries Q1 to Q13 corresponding to each state learned by the transformer's decoder 52.

[0032] In this embodiment, a query is a learnable parameter optimized and learned within the trained model 50, used to detect specific states within the vehicle interior, such as the occupancy status of each seat and the seatbelt fastening status. The detection target refers to specific state information or attributes identified and detected from the vehicle interior image by the trained model 50. Specifically, this includes the occupancy status of each seat, the seatbelt fastening status, the driver's grip on the steering wheel, the driver's operation status of equipment, and the installation status of child seats in each seat. These detection targets serve as targets for the trained model 50 to analyze and identify specific regions or features within the image, and each is processed by an independent query. In other words, there is a dedicated query corresponding to each detection target, which enables the trained model 50 to detect multiple vehicle interior states in parallel and with high accuracy.

[0033] The learning device 20 is for generating a trained model 50. As shown in Figure 5, the learning device 20 comprises a second control unit 21, a second interface 22, a second storage unit 23, and a learning unit 24. The second control unit 21 includes a hardware processor such as a CPU, RAM, and ROM, and controls each component according to information processing. In Figure 5, the components of the second control unit 21 are shown as CPU 211, RAM 212, and ROM 213, respectively.

[0034] The second interface 22 is an interface for connecting to an external device and is configured appropriately depending on the external device to be connected. In this embodiment, the second interface 22 is connected to the providing device 30 as an external device so as to be able to communicate.

[0035] The second storage unit 23 stores the learning program 231 and the learning dataset 40, as well as the trained model 50 that was trained and generated by the learning unit 24. The second storage unit 23 is a tangible, non-temporary computer-readable medium, i.e., a non-transitional, tangible recording medium. Examples of the second storage unit 23 include, but are not limited to, RAM, ROM, HDD, SSD, magnetic disk, magneto-optical disk, CD-ROM, DVD-ROM, semiconductor memory, etc.

[0036] The second storage unit 23 may be an internal medium directly connected to the bus of the computer constituting the learning device 20, or it may be an external medium connected via the second interface 22 or a telecommunications line. Furthermore, when the learning program 231 is distributed to the learning device 20 via a communication line, the learning device 20, upon receiving the distribution, expands the program into the second storage unit 23 or the RAM 212 of the second control unit 21 and executes it, thereby realizing the learning unit 24.

[0037] The learning program 231 is a computer program for virtually realizing the learning unit 24 on a computer. The learning device 20 can virtually realize the learning unit 24 on a computer by, for example, the second control unit 21 reading the learning program 231 from the second storage unit 23 and executing it. In other words, the learning unit 24 can be composed of a functional unit that is virtually realized by, for example, the second control unit 21 executing the learning program 231.

[0038] As shown in Figures 6 and 7, the learning unit 24 generates a trained model 50 by training model 50A using the training dataset 40. In Figure 7, model 50A is a model in the process of training the trained model 50. The training dataset 40 includes multiple training data 41. The training data 41 consists of a training image P taken including one or more seats inside a vehicle, with two or more annotations indicating the state of the vehicle interior, i.e., annotations A1 to A13 as shown in Figure 2. In this embodiment, the training data 41 consists of a training image P taken including multiple seats inside a vehicle, with two or more annotations A1 to A13 attached. As a result, as shown in Figure 7, the learning unit 24 optimizes two or more queries corresponding to the annotations through training.

[0039] In this embodiment, annotation means explicitly labeling the training image with the in-car state corresponding to queries Q1 to Q13 shown in Figure 2. In this description, annotations corresponding to in-car states S1 to S13 are referred to as annotations A1 to A13. Queries corresponding to detection targets O1 to O13, that is, queries for detecting in-car states S1 to S13, are referred to as queries Q1 to Q13. In the example in Figure 6, multiple training data 41, training image P, and annotation A are each given the same reference numerals, but their contents are all different.

[0040] Furthermore, in this embodiment, the training data 41 is annotated with A1 to A13 corresponding to each query Q1 to Q13. Therefore, the learning unit 24 optimizes each query Q1 to Q13 corresponding to annotations A1 to A13 all at once through learning.

[0041] Figures 8 to 11 show examples of annotations A1 to A13 corresponding to queries Q1 to Q13 being applied to training images. Note that the training images in Figures 8 to 11 are referred to as training images P1 to P4, respectively, because their content differs. Training image P1 in Figure 8 has annotations A1 to A4, all corresponding to queries Q1 to Q4, marked "Present". Annotations A5 to A8, all corresponding to queries Q5 to Q8, are marked "Not Worn". Annotation A9, corresponding to query Q9, is marked "Right Hand Only". Annotation A10, corresponding to query Q10, is marked "Operating". And annotations A11 to A13, corresponding to queries Q11 to Q13, are marked "Not Worn". Similarly, training images P2 to P4 in Figures 9 to 11 have annotations A1 to A13 corresponding to each query Q1 to Q13 applied according to the content of training images P2 to P4.

[0042] The training dataset 40 contains multiple training data sets 41 with different annotations A, and the number of training data sets 41 with different annotations A is set to be roughly equal. Specifically, different annotations A1 to A13 corresponding to queries Q1 to Q13 are different. In other words, the training data sets 41 included in the training dataset 40 are set so that the occurrence rate of the content of each annotation A is equal.

[0043] Here, when combining the occupancy status of each seat of Q1 to Q4 and the seat belt wearing status of each seat of Q5 to Q8, the status of each seat corresponding to these combinations is three types: "absent", "occupied and seat belt worn", and "occupied and seat belt not worn". Also, the steering wheel gripping status of the driver at Q9 is four types: "both hands gripping", "only right hand gripping", "only left hand gripping", and "both hands released". Further, the operation status of the driver's device at Q10 is two types: "not operated" and "being operated". And the child seat attachment status of each seat excluding the driver's seat of Q11 to Q13 is two types: "not attached" and "attached". And the combinations of the above-mentioned various statuses are 3^4×4×2×2^3 = 5184 types. The learning dataset 40 uses images in which the 5184 types of combined states occur without bias, that is, randomly or with an equal probability.

[0044] According to the embodiment described above, the learned model 50 takes as input an image P taken including one or more seats in the vehicle interior, and outputs, as the state of the vehicle interior detected from the image P, two or more of the occupancy statuses S1 to S4 of each seat, the seat belt wearing statuses S5 to S8 of each seat, the steering wheel gripping status S9 of the driver, the operation status S10 of the driver's device, and the child seat attachment statuses S11 to S13 of each seat.

[0045] Also, the vehicle interior state detection method according to the embodiment includes an image acquisition process and a detection process. The image acquisition process is a process of acquiring an image P taken including one or more seats in the vehicle interior. The detection process takes as input an image P taken including one or more seats in the vehicle interior, and uses the learned model 50 that outputs two or more of the occupancy statuses S1 to S4 of each seat, the seat belt wearing statuses S5 to S8 of each seat, the steering wheel gripping status S9 of the driver, the operation status S10 of the driver's device, and the child seat attachment statuses S11 to S13 of each seat as the state of the vehicle interior detected from the image P, and inputs the image P acquired in the image acquisition process to detect two or more states of the vehicle interior.

[0046] In addition, the in-vehicle state detection program 131 according to the embodiment is for causing the detection device 10, which is a computer, to execute the above-described image acquisition process and the above-described detection process.

[0047] The in-vehicle state detection system 1 according to the embodiment includes a photographing device 2 and a detection device 10. The photographing device 2 has one or more seats in the vehicle interior as a photographing range. The detection device 10 takes as input an image P obtained by photographing including one or more seats in the vehicle interior, and outputs, as the state of the vehicle interior detected from the image P, two or more of the occupancy states S1 to S4 of each seat, the seat belt wearing states S5 to S8 of each seat, the steering grip state S9 of the driver, the operation state S10 of the driver's device, and the child seat mounting states S11 to S13 of each seat. The image P photographed by the photographing device 2 is input to the learned model 50 to detect two or more states of the vehicle interior.

[0048] According to this, the following operational effects can be obtained. For example, in the conventional technology, individual detectors were used for each state of the vehicle interior to be detected. In this case, since image processing was performed for each state, the computational load was high and the processing took time. On the other hand, since the learned model 50 of the present embodiment simultaneously detects two or more states, there is no need to use individual detectors for each state, and the computational load when detecting the state of the vehicle interior can be reduced.

[0049] In addition, in the process of using individual detectors for each state as in the conventional case, since the processing is performed in a so-called pipeline in which a plurality of processing stages are executed continuously and sequentially, the errors of each detector accumulate, so it was difficult to obtain high detection accuracy with short-time processing. On the other hand, since the learned model 50 of the present embodiment detects two or more states in parallel, the accumulation of errors in each detection occurring in the pipeline processing is suppressed, and the detection accuracy of each state can be improved.

[0050] As a result, by using the trained model 50 of this embodiment, it is possible to reduce the computational load and improve the recognition accuracy when detecting the state inside the vehicle. Furthermore, the reduction in computational load and improvement in detection accuracy make it possible to detect the state inside the vehicle in real time and quickly. This enables high-speed and stable performance when put into practical use as a system to improve driving safety and to immediately grasp and respond to the comfort of occupants. In addition, since multiple states inside the vehicle can be detected by a single detector, i.e., a single trained model 50, it is possible to reduce the cost and improve the functionality of the detection device 10 on which the trained model 50 is implemented.

[0051] Furthermore, the trained model 50 of the embodiment includes an architecture with a transformer. The trained model 50 outputs two or more in-cabin states in a single output according to two or more queries corresponding to each state learned by the transformer's decoder 52.

[0052] According to this, by leveraging the parallel processing capabilities of the transformer architecture, multiple states can be detected simultaneously, significantly reducing the computational load compared to conventional methods that use separate detectors for each state. This improves the overall system processing speed and enables real-time state detection.

[0053] Furthermore, the transformer's decoder 52 processes multiple queries corresponding to each state simultaneously, enabling detection that takes into account the correlations between each state. This reduces the false positive rate and improves overall recognition accuracy compared to independent detection by individual detectors.

[0054] Furthermore, because the Transformer architecture employs a query-based detection method, it is possible to easily add new detection targets and states. This enables the creation of flexible systems that can adapt to different vehicle types and interior layouts. In addition, the encoder-decoder structure of the Transformer allows for efficient feature extraction and state detection, improving the efficiency of model training and inference. This results in reduced training time and improved processing speed during inference.

[0055] As described above, the trained model 50 of this embodiment, through its transformer-based architecture and batch detection function using multiple queries, can simultaneously achieve reduced computational load and improved recognition accuracy in detecting the state inside the vehicle. This improves the performance of the vehicle interior state detection system, contributing to improved safety during driving and ensuring passenger comfort.

[0056] Furthermore, the method for generating a trained model according to the embodiment generates a trained model 50 using a training dataset 40 that includes multiple training data 41 annotated with two or more of the following states of the vehicle interior included in the image P: the occupancy status S1 to S4 of each seat, the seat belt wearing status S5 to S8 of each seat, the driver's grip on the steering wheel S9, the driver's operation status of the equipment S10, and the child seat installation status S11 to S13 of each seat.

[0057] Furthermore, the learning program 231 according to the embodiment causes the learning device 20, which is a computer, to perform a process to generate a trained model 50 by training an image P taken including one or more seats in the car interior, using a learning dataset 40 which contains multiple learning data 41 annotated with two or more of the following conditions for the state of the car interior detected from the image P: the occupancy status S1 to S4 of each seat, the seat belt wearing status S5 to S8 of each seat, the driver's grip on the steering wheel S9, the driver's operation status of the equipment S10, and the child seat installation status S11 to S13 of each seat.

[0058] In this way, by training using a training dataset 40 that includes multiple training data 41 annotated with two or more states, it is possible to generate a trained model 50 that can perform more comprehensive and accurate recognition by considering the correlation between different states. Furthermore, since each training data 41 constituting the training dataset 40 contains two or more different states simultaneously, it is possible to train multiple states in a single training session. As a result, it is not necessary to generate individual training models as in the past, and it is possible to generate a highly accurate trained model 50 while reducing training costs.

[0059] Furthermore, the method for generating a trained model according to this embodiment involves training using two or more queries Q1 to Q13 corresponding to each state. The training dataset 40 includes multiple training data 41 with different annotations A1 to A13, and the number of training data 41 with different annotations A1 to A13 is roughly equal.

[0060] Furthermore, the method for generating a training dataset according to the embodiment is a method for generating a training dataset 40 that includes multiple training data 41 used to train a trained model 50, with an image P taken including one or more seats in the vehicle interior as input, and two or more of the following states of the vehicle interior detected from the image P as the occupancy status S1 to S4 of each seat, the seat belt wearing status S5 to S8 of each seat, the driver's grip on the steering wheel S9, the driver's operation status of the equipment S10, and the child seat installation status S11 to S13 of each seat being output. The method for generating a training dataset includes a process of adjusting the training dataset 40 to have roughly the same number of training data 41 with different annotations A1 to A13.

[0061] According to this, the training dataset 40 used to train the trained model 50 has annotations evenly distributed among the training data 41, thereby reducing the bias of the trained model 50 and enabling the generation of a trained model 50 that can detect with high accuracy under various conditions.

[0062] (Other Embodiments) This disclosure is not limited to the embodiments described above and shown in the drawings, and can be modified, combined, or extended at will without departing from its essence. The numerical values ​​and other figures shown in the above embodiments are illustrative and not limiting.

[0063] If the trained model 50 generated by the learning device 20 can be provided to the detection device 10, then for example, the detection device 10 and the learning device 20 do not necessarily need to be connected via a telecommunications line W and may be configured independently. Also, the detection device 10 and the learning device 20 may be a single computer. Furthermore, the detection target, query, in-vehicle interior state, and annotation are not limited to those described above and may be set appropriately according to the purpose of detecting the in-vehicle interior state.

[0064] The control unit and its method described herein may be implemented by a dedicated computer provided by configuring a processor and memory programmed to perform one or more functions embodied by a computer program. Alternatively, the control unit and its method described herein may be implemented by a dedicated computer provided by configuring a processor by one or more dedicated hardware logic circuits. Alternatively, the control unit and its method described herein may be implemented by one or more dedicated computers configured by a combination of a processor and memory programmed to perform one or more functions and a processor configured by one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by the computer on a computer-readable non-transitional tangible recording medium.

[0065] This disclosure is described in accordance with the embodiments, but it is understood that this disclosure is not limited to such embodiments or structures. This disclosure also includes various modifications and variations within the equivalence. In addition, various combinations and forms, as well as other combinations and forms that include only one, more, or fewer of those elements, fall within the scope and concept of this disclosure.

Claims

1. A trained model that takes an image (P) taken including one or more seats inside a vehicle as input, and outputs two or more of the following conditions for the state of the vehicle interior detected from the image: the occupancy status of each seat (A1-A4), the seat belt status of each seat (A5-A8), the driver's grip on the steering wheel (A9), the driver's operation status of equipment (A10), and the installation status of a child seat for each seat (A11-A13).

2. The trained model according to claim 1, comprising an architecture having a transformer, wherein the trained model outputs two or more states inside the vehicle interior in a single batch according to two or more queries (Q1 to Q13) corresponding to each state learned by the decoder (52) of the transformer.

3. A method for generating a trained model (50) that is trained using a training dataset (40) which includes multiple training data (41) annotated with two or more of the following conditions for the state of the vehicle interior included in the image (P), namely the occupancy status of each seat (A1 to A4), the seat belt status of each seat (A5 to A8), the driver's grip on the steering wheel (A9), the driver's operation status of the equipment (A10), and the child seat status of each seat (A11 to A13).

4. A method for generating a trained model according to claim 3, wherein the training is performed using two or more queries (Q1 to Q13) corresponding to each state, and the training dataset includes multiple training data with different annotations, and the number of training data with different annotations is roughly equal.

5. A learning program (231) that causes a computer (20) to execute a process to generate a trained model by training it using a training dataset (40) which includes multiple training data (41) annotated with images (P) taken including one or more seats in the interior of a vehicle, and annotated with two or more of the following conditions for the interior of the vehicle detected from the images: the occupancy status of each seat (A1 to A4), the seat belt status of each seat (A5 to A8), the driver's grip on the steering wheel (A9), the driver's operation status of the equipment (A10), and the child seat status of each seat (A11 to A13).

6. A method for generating a training dataset (40) that includes multiple training data (41) used to train a trained model (50), wherein the input is an image (P) taken including one or more seats inside a vehicle, and the output is two or more of the following conditions for the state of the vehicle interior detected from the image: the occupancy status of each seat (A1 to A4), the seat belt status of each seat (A5 to A8), the driver's grip on the steering wheel (A9), the driver's operation status of the equipment (A10), and the child seat status of each seat (A11 to A13), the method for generating a training dataset (40) that includes multiple training data (41) used to train a trained model (50), wherein the training dataset includes a process of adjusting the number of training data with different annotations to roughly the same number.

7. A method for detecting the state of a vehicle interior, comprising: an image acquisition process that acquires an image (P) taken including one or more seats in the vehicle interior; and a detection process that takes the image taken including one or more seats in the vehicle interior as input and outputs two or more of the following states of the vehicle interior detected from the image: the occupancy status of each seat (A1 to A4), the seat belt status of each seat (A5 to A8), the driver's grip on the steering wheel (A9), the driver's operation status of equipment (A10), and the child seat installation status of each seat (A11 to A13), and inputs the image acquired in the image acquisition process to detect two or more states of the vehicle interior.

8. A vehicle interior state detection program (131) for a computer (10) to perform an image acquisition process to acquire an image (P) taken including one or more seats in the vehicle interior, and a detection process to input the image acquired by the image acquisition process to a trained model (50) which takes the image taken including one or more seats in the vehicle interior as input and outputs two or more of the following as the state of the vehicle interior detected from the image: the occupancy status of each seat (A1 to A4), the seat belt status of each seat (A5 to A8), the driver's grip on the steering wheel (A9), the driver's operation status of the equipment (A10), and the child seat installation status of each seat (A11 to A13), in order to detect two or more states of the vehicle interior.

9. A vehicle interior state detection system (1) comprising: a camera (2) whose shooting range is set to one or more seats in the vehicle interior; and a detection device (10) which takes as input an image (P) taken including one or more seats in the vehicle interior, and outputs two or more of the following as the state of the vehicle interior detected from the image: the occupancy status of each seat (A1 to A4), the seat belt status of each seat (A5 to A8), the driver's grip on the steering wheel (A9), the driver's operation status of the equipment (A10), and the child seat installation status of each seat (A11 to A13); and inputs the image taken by the camera (10) to detect two or more states of the vehicle interior.