Improved activity identification in a passenger compartment of a transport means using static or dynamic regions of interest, roi
Patent Information
- Application Number
- EP2023833032
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-13
- Publication Date
- 2025-10-22
AI Technical Summary
Existing methods for activity detection in transportation vehicle passenger compartments, such as body pose activity recognition, are prone to false-positive results due to biases in training data and incorrect skeleton recordings, making them unreliable for security-relevant functions.
The method involves specifying static or dynamic regions of interest (ROIs) using geometric shapes to improve the reliability of activity classification by validating the presence of limbs in specific areas, such as the steering wheel, and using a hierarchical classification process that combines AI-based systems with conventional algorithms for enhanced accuracy.
This approach significantly improves the reliability of activity detection, filters out false positives, and allows for adaptive safety functions by ensuring that activities are correctly recognized, enhancing the safety and comfort of transportation systems.
Smart Images

Figure 1.1
Abstract
Description
[0001] Improved activity detection in a passenger compartment of a transport vehicle using static or dynamic regions of interest, ROI
[0002] The disclosure relates to a classification of an activity of a user in a means of transport with a detection of an activity of the user in a
[0003] Passenger compartment of the vehicle using image data, by a camera unit, and a classification of the recorded activity, by a computing unit.
[0004] In the interior of a means of transport, i.e. in the passenger compartment of a means of transport, such as a passenger car or a truck, but also in a drone taxi or a gondola, it is becoming increasingly important to record the activities and possible distractions of the users, i.e. the vehicle occupants and in particular of a user responsible for driving the means of transport. This makes it possible, for example, to optimize the handover of driving responsibility from a control unit of a (partially) automated means of transport back to the driver, to adapt assistance systems of the means of transport to the respective activity of the users and, for example, to proactively offer functions, or to take more targeted safety measures, such as airbag deployment adapted to a respective detected body pose. In this way, driving functions, comfort functions and safety functions in the means of transport can be improved.
[0005] A known example is what is known as body pose activity recognition, in which one or more cameras in the passenger compartment detect people in the image data generated by the camera and a configuration of the skeleton of each detected person is recognized in three dimensions, i.e. calculated according to a predefined anatomical model. In particular, the passenger compartment can refer to a closed space intended for accommodating users, such as the passenger compartment of a limousine and / or a partially open space intended for accommodating users, such as the passenger compartment of a convertible and / or an open space intended for accommodating users, such as the passenger compartment of a motorcycle. The respective activity is then classified, i.e. recognized as belonging to a type of activity, from the movement of the limbs, the position of known objects or devices in the passenger compartment such as the steering wheel and possibly other features.For example, classes such as "opens bottle," "drinks," "uses phone," "types on laptop," "sleeps," and the like can be specified for this classification or recognition. The recognition can also include a hierarchical class model, in which classes can each have one or more subclasses. For example, the class "uses phone" can include the subclasses "uses phone with mobile phone to ear" and "uses phone with hands-free device." As an alternative to this type of activity classification, other approaches can also be used, such as end-to-end classification, in which the image data is fed directly into a classifier and activities are classified accordingly without intermediate steps such as the described analysis of the body pose.
[0006] The known methods for classifying activities are error-prone and can therefore produce false-positive results. For example, the activity "talking on the phone with a mobile phone to the ear" can be detected even though both hands are on the steering wheel. In this case, the pose of the hands is clear, they are on the steering wheel, but the classification process was deceived by unknown reasons and consequently produces an incorrect result. The reasons for this can be, for example, a bias in the training data if a learned algorithm is used, in a corrupt skeleton capture, in which, for example, only the hands are correctly recognized but the rest are incorrect, or in the classification of other features as dominant for the recognition of the activity and thus misleading the classifier. Especially when a so-called artificial intelligence (AI), i.e.If a learned algorithm such as a neural network or an algorithm generated using other machine learning methods is used, such errors cannot be ruled out with absolute reliability, since the decisions of the algorithm can never be reproduced down to the last detail, meaning that the learned algorithm generally always has the character of a "black box". Accordingly, the previous approaches are also not suitable, for example, for implementing activity recognition, which also has an impact on safety-relevant functions.
[0007] The problem of making a classifier more robust is generally not new. Typically, attempts are made to add additional, previously unconsidered features to restrict the search tree underlying the classification. For example, in body pose activity recognition, the anatomical model could assign permissible joint angles to each joint, which counteracts the incorrect detection of a body skeleton with bent limbs, since such a configuration of the skeleton would be physiologically impossible.
[0008] Accordingly, the task is to detect the activity of a user in a means of transport, in particular in a motor vehicle, with increased reliability, ideally with a reliability sufficient to adapt safety functions to a detected activity.
[0009] This problem is solved by the subject matter of the independent patent claims. Advantageous embodiments emerge from the dependent patent claims, the description, and the figure.
[0010] One aspect relates to a method for classifying a user's activity in a means of transport such as a motor vehicle, but also an aircraft or a ship. The method can thus be used in land-based, air-based, and water-based means of transport, preferably in means of transport with partially or fully autonomous control. One method step is capturing an activity of the user in a passenger compartment of the means of transport using image data, which is carried out by a camera unit. The camera unit can comprise one or more cameras. For example, the user's hands can be captured in the image data, either directly in the camera image (i.e., two-dimensionally), or in a three-dimensional scene reconstructed using the image data (i.e., three-dimensionally), and thus also the activity of the hands or the user.
[0011] A further step in the process is to specify one or more areas in the passenger compartment as regions of interest (ROI). ROIs are two-dimensional or three-dimensional areas that are marked manually or automatically in an image as a two-dimensional (image) area or in a scene as a three-dimensional (spatial) area. In the simplest case, these are rectangles or cuboids that are intended to enclose certain areas in the image or scene. Sometimes rectangles or cuboids are not sufficient to geometrically approximate objects; in these cases, other shapes such as circles, ellipses or cylinders, or combined shapes consisting, for example, of a cylinder and a cuboid, can be used.
[0012] For example, the steering wheel can be approximated as a geometric shape in order to define an associated ROI, which is then important for recognizing activities in which the hands are placed on the steering wheel or not on the steering wheel. In the case of the steering wheel, for example, a cuboid is suitable for faster calculations or a cylinder for a more precise representation. If the hands lie within the ROI used to approximate the steering wheel, i.e. the ROI of the steering wheel, activities that require the hands to be on the steering wheel or not on the steering wheel can be better recognized and recognition can be checked (validated). Accordingly, a further process step is the classification of the recorded activity based on the image data and the specified ROI by a computing unit.
[0013] This has the advantage that the ROI-based classification, for example via occupancy analysis as described below, improves the reliability of the classifier. The solution presented here is also particularly flexible and can be applied to a variety of different recognition strategies (body pose activity recognition, end-to-end activity recognition, etc.). In particular, it allows for hierarchically divided classification or recognition of the respective activity, which allows integration into safety-relevant functions and, as described below, can combine the power of AI-based systems with the reliability of conventionally predefined algorithms. The solution presented here is therefore suitable for use in conjunction with driver assistance systems, for example occupant monitoring orDriver monitoring in partially and fully automated modes of transport, in human-machine interaction with vehicle assistance systems or infotainment systems in the passenger compartment, in occupant monitoring for adaptive safety functions, and in the detection of vehicle occupants and their respective activities for entertainment media and other comfort functions. The solution described thus filters out falsely classified activities using intuitive and comprehensible heuristics: If, for example, the hands are not in the required area, this activity cannot take place. Furthermore, the ROI analysis calculations are fast due to the approximate geometric shapes and can be easily managed with known algorithms, so that they can also be performed in modes of transport with lower data processing capacity compared to stationary systems.In one embodiment, the ROIs are used for a supplementary feature that is trained into a K1-based activity classifier, i.e., a learned algorithm for classifying the recorded activity. Accordingly, the recorded activity is then classified based on image data and ROI using the image data and, for example, activation data of the respective ROI, i.e., an occupancy of ROIs, preferably resolved by limb, as entries in a corresponding input vector for the learned algorithm (activity classifier) that performs the classification. Classification then occurs in a single step in which image data and activation data are evaluated together.
[0014] In a further embodiment, it is provided that the classification, which can also be referenced as full classification or complete classification, comprises a first classification, which can be referenced as pre-classification, and a second classification, which can be referenced as final classification. A hierarchical structure can therefore be selected for the classification. In the first classification, the recorded activity is classified based on the image data and preferably without taking the predefined ROIs into account, for example using a known method. Additional data supplementing the image data, such as status data and / or sensor data of the vehicle, can also be used. For example, a capacitive sensor on the steering wheel can detect whether the hands are on the steering wheel or not, and this can be taken into account when detecting the activity.
[0015] During the second classification, a result from the first classification is processed with an occupancy analysis of the specified ROIs. For example, the results of a conventional activity classifier can be fed into another classifier instance, such as a neural network or the like, along with the ROI occupancy analysis as a new feature vector and processed in a second, hierarchically higher process step. For example, a probability distribution for activities detected in the first classification, along with the ROI occupancy analysis (particularly per limb) and an object specified and thus relevant for the respective activity detected in the first classification (e.g., the steering wheel), can be combined as a feature vector and reclassified during the second classification to prevent nonsensical hypotheses.Accordingly, the first classification can be performed by a first neural network and / or the second classification by a second, further neural network which is different from the first neural network.
[0016] In a further embodiment, it is provided that the result of the first classification comprises at least one, preferably at least two, and particularly preferably three classification result candidates. A classification result candidate can also be or comprise the probability distribution for activities recognized in the first classification mentioned in the last paragraph. The processing of the second classification comprises a compatibility check of the respective classification result candidate(s) with the occupancy analysis of the predetermined ROIs based on a stored compatibility specification. The compatibility specification can be in the form of a database and links a plurality of classification result candidates with ROIs that are each specified as compatible and / or incompatible for a classification result candidate. The classification result candidate can each be assigned a specific validity, ieProbability. Preferably, only one classification result candidate is output as the final result of the complete classification, the assigned ROI(s) of which is compatible with a result of the occupancy analysis of the specified ROIs.
[0017] For example, the (preliminary) result of the (first) classification from the assigned first activity classifier can include three classification result candidates with the highest probabilities of occurring. Accordingly, if the result of the first classification with the highest probability, i.e., for example, is "telephone with mobile phone to ear," but at the same time the occupancy analysis shows that the ROI assigned to the steering wheel is double-occupied, i.e., the hands are on the steering wheel, output or further processing of the (preliminary) result (as final result) is prevented. Accordingly, as the final result of the classification, i.e., the full classification, only one classification result candidate from the first classification is output, whose assigned ROI(s) is compatible with a result of the occupancy analysis of the specified ROIs.In this case, the first activity classifier not only outputs one result, but one or more classification result candidates, ideally the top 3 or even the top 10 classification result candidates, possibly even a (discrete) probability distribution across all activities identified as possible. Instead of selecting the activity in first place and thus the most likely activity, as in conventional approaches, the respective classification result candidates are checked for plausibility using the ROI occupancy analysis. For example, it can be checked whether hands are important for an activity and whether the hands are located in an ROI corresponding to the activity. If this is the case, the activity was probably correctly identified during the first classification. If this is not the case, the activity can be discarded and another one selected.In this way, all classification result candidates contained in the initial classification can be checked for their suitability as a possible final result, in descending order of probability. It can also be specified that only classification result candidates with a specified minimum validity or probability are checked for compatibility. This has the advantage of implementing a hierarchical procedure with independent classifiers, significantly improving reliability.
[0018] Accordingly, it can be provided that the occupancy analysis and / or the compatibility check are carried out using an algorithm explicitly specified by a human. The occupancy analysis and / or the compatibility check is then not carried out by an algorithm indirectly or directly specified by a human, but only implicitly. Such implicitly specified algorithms are, for example, algorithms generated using a learning process, such as a trained neural network. This has the advantage that the occupancy analysis or the compatibility check is not a so-called "black box," but that the structure and behavior of the algorithm can be reproduced and verified in detail. This also makes it acceptable for the algorithm to influence safety-relevant functions, for example by taking into account the detected activity in safety-relevant functions of the means of transport.Compared to the first classification, the occupancy analysis or compatibility test is clearly defined and can be solved efficiently using explicitly specified algorithms.
[0019] In a further embodiment, it is provided that the (final) result of the classification, i.e. the full classification, is provided to a control unit of the means of transport and, in particular, a safety-relevant function of the means of transport is also controlled by the control unit as a function of the provided (final) result. The safety-relevant function can, for example, comprise an acceleration and / or steering function, in particular lateral and / or longitudinal and / or vertical guidance of the means of transport, but also another safety function such as a deployment pattern of one or more airbags or a behavior of a belt pretensioner. This leads to improved safety of the means of transport because the safety-relevant function can be implemented as an adaptive function which adapts to a behavior, i.e. an activity, of the user.
[0020] In a further embodiment, at least one of the ROIs is defined as a region that can vary in size and / or shape and / or position and / or orientation. The ROIs therefore do not have to be static, but can also be dynamic. Static ROIs are suitable, for example, for the steering wheel or other objects at a fixed position in the passenger compartment, which can be described consistently and with sufficient accuracy by a once-defined ROI. However, there are also moving regions, such as the head, to which the mobile phone or the hands must be held so that the activity of "telephoning with a mobile phone to the ear" can be recorded or recognized or classified as such. For example, the ROI of the left and / or right ear can be dynamic and moving through upstream head tracking.Accordingly, at least one ROI specified as a variable in size and / or shape and / or position and / or orientation can be an ROI linked to a mobile object (such as the ear) detected and located in the passenger compartment. The dynamic ROIs can therefore be ROIs "anchored" to a mobile object. Alternatively, the object can also be detected as mobile in the environment relative to the vehicle, and an assigned ROI can, for example, be an area on a windshield that shifts according to the movement of the vehicle and which, for example, can be occupied by a pointing hand when a user points to the object outside the vehicle. The described embodiments for the ROIs have the advantage of increasing the flexibility and reliability of the classification.
[0021] A further aspect relates to a classification device for classifying an activity of a user in a vehicle, comprising a camera unit for capturing an activity of the user in a passenger compartment of the vehicle using image data, a storage unit for storing predefined areas in the passenger compartment as regions of interest (ROI), and a computing unit for classifying the captured activity based on the image data and the predefined ROI.
[0022] Advantages and advantageous embodiments of the classification device corresponding to advantages and advantageous embodiments of the method described above.
[0023] The features and feature combinations mentioned above in the description, including in the introductory part, as well as the features and feature combinations mentioned below in the description of the figures and / or shown alone in the figures can be used not only in the respective combination specified, but also in other combinations without departing from the scope of the invention. Thus, embodiments are also to be considered encompassed and disclosed by the invention that are not explicitly shown and explained in the figures, but which emerge and can be produced by separate feature combinations from the explained embodiments. Embodiments and feature combinations are also to be considered disclosed that therefore do not have all the features of an originally formulated independent claim.Furthermore, embodiments and combinations of features are to be regarded as disclosed, in particular by the embodiments set out above, which go beyond or deviate from the combinations of features set out in the reliances of the claims.
[0024] Fig. 1 shows an exemplary embodiment of a classification device 1 for classifying an activity of a user 2 in a means of transport 3, here embodied as a motor vehicle. The classification device 1 has a camera unit 4 for recording the activity of the user 2 in a passenger compartment 5 of the means of transport 3. Furthermore, the classification device 1 has a memory unit 6 for storing predefined areas 7 in the passenger compartment as regions of interest, ROI. In the example shown, a first predefined area 7 is shown, which contains a steering wheel 8 as an ROI, and a second predefined area 7', which contains an ear 13 of the user 2 as an ROI. Finally, the classification device 1 also has a computing unit 9 for classifying the recorded activity based on the image data of the camera unit 4 and the ROIs stored in the memory unit 6.In the example shown, the classification device 1 is also equipped with a control unit 10 for controlling a (here safety-relevant) function of the vehicle 3 depending on a classification (final) result provided by the classification device 1. For this purpose, the control unit 10 in the present example is coupled to a drive unit 11 and a braking unit 12.
[0025] For example, when classifying an activity of user 2, the respective activity can be recorded by the camera unit 4. Based on the image data provided by the camera unit, a pre-classification can be performed, which initially runs without the predefined ROIs. For example, "talking on the phone with the mobile phone to the ear" can be incorrectly identified as a favored classification result candidate. "steering the vehicle" can be identified as a subsequent classification result candidate. Independently of the pre-classification, an occupancy analysis of the predefined ROIs 7, 7'—in this case, the ROI "steering wheel" and the ROI "ear"—can be performed, and it can be determined that limbs, in this case hands, are located in the ROI 7 of the steering wheel 8.Accordingly, when checking the result of the first classification, it is determined that the actual ROI in which the hands are located is not compatible with the ROI containing the ear 13. Accordingly, the first classification result candidate, in this case "telephone with mobile phone to ear", is discarded and in the present case the next classification result candidate, in this case "steers the vehicle", is output as the final result of the classification and, in the present case, is sent to the control unit 10 to control a safety-relevant function such as longitudinal and / or lateral guidance of the vehicle 3. Accordingly, for example, the control unit 10 could assume that the user 2 is paying undivided attention and, for example, select a higher speed and / or a shorter distance from the vehicle in front than if the user 2 were actually using the.
[0026] mobile phone to the ear.
Claims
Claims 1. Method for classifying an activity of a user (2) in a means of transport (3), in particular in a motor vehicle, comprising the following method steps: - detecting an activity of the user (2) in a passenger compartment (5) of the vehicle (3) by means of image data, by a camera unit (4); characterized by a - specifying one or more areas (7) in the passenger compartment (5) as regions of interest, ROI; - Classifying the recorded activity based on the image data and the or the specified ROI, by a computing unit (9).
2. Method according to the preceding claim, characterized in that - the classification comprises a first classification and a second classification, wherein in the first classification the detected activity is classified based on the image data and in the second classification a result of the first classification is processed with an occupancy analysis of the predetermined ROI(s).
3. Method according to the preceding claim, characterized in that the first classification is carried out by a neural network and / or the second classification is carried out by a further neural network.
4. Method according to one of the two preceding claims, characterized in that the result of the first classification comprises at least one, preferably at least two, particularly preferably three, classification result candidates, and the processing of the second classification comprises a compatibility check of the respective classification result candidate(s) with the Occupancy analysis of the specified ROIs based on a stored compatibility specification, which links a plurality of classification result candidates with ROIs specified as compatible and / or incompatible. Method according to the preceding claim, characterized in that the occupancy analysis and / or the compatibility check is carried out by means of an algorithm explicitly specified by a human. Method according to one of the preceding claims, characterized by a - Providing a result of the classification to a control unit (10) of the means of transport (3); and - Controlling a safety-relevant function of the means of transport (3) by the control unit (10) depending on the provided result. Method according to one of the preceding claims, characterized in that three-dimensional spatial regions (7) in the passenger compartment (5) are predefined as ROIs. Method according to one of the preceding claims, characterized in that at least one of the ROIs is predefined as a region variable in size and / or shape and / or position and / or orientation. Method according to the preceding claim, characterized in that at least one ROI predefined as a region variable in size and / or shape and / or position and / or orientation is an ROI linked to a mobile object detected and located in the passenger compartment (5). Classification device (1) for classifying an activity of a user (2) in a means of transport (3), in particular in a motor vehicle, comprising: - a camera unit (4) for detecting the activity of the user (2) in a passenger compartment (5) of the vehicle (3) by means of image data; characterized by - a storage unit (6) for storing predefined areas (7) in the passenger compartment (5) as regions of interest, ROI; - a computing unit (9) for classifying the recorded activity based on the image data and the specified ROI.